A linear programming-based neural network mapping method for storage-computing integrated chip

Through the neural network mapping method based on linear programming, the arrangement of neural network weights and biases on the memory and computing integrated chip is optimized, which solves the problems of low utilization efficiency and high current noise in the existing technology, and achieves a more efficient memory and computing integrated chip design and higher computing accuracy.

CN114723024BActive Publication Date: 2025-05-02BEIJING ZHICUN (WITIN) TECH CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210227169.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-05-02
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

When the prior art maps neural networks to flash memory cell arrays of computer-integrated chips, the flash memory cell cannot be effectively utilized, resulting in an increase in the scale of the flash memory cell array, and the direct mapping of bias results in an increase in current noise, affecting the operation accuracy.

Method used

Using a neural network mapping method based on linear programming, by obtaining the weight and bias data of the neural network to be mapped and the hardware parameters of the target memory integrated chip, a pre-established linear planning solution model is input to solve the mapping scheme, optimizing the arrangement of weights and biases, reducing the bias value on a single flash memory cell, and reducing current noise.

Benefits of technology

Effectively utilize flash memory cells to reduce the scale of flash memory cell array, reduce current noise, and improve computing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114723024B_ABST
    Figure CN114723024B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a linear programming-based neural network mapping method for a storage-computing integrated chip, the method comprising: obtaining the weight arrays of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage-computing integrated chip; inputting the weight arrays of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage-computing integrated chip into a pre-established linear programming solution model to solve and obtain a mapping scheme, the mapping scheme is used to map the weight arrays of each layer of the neural network to be mapped and the corresponding bias arrays to the target storage-computing integrated chip. Among them, by converting the weight and bias data mapping process based on intuitive experience into a solution problem of a linear programming mathematical model, the calculation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor technology, and in particular to a linear programming-based neural network mapping method, device, equipment and storage medium for a storage-computing integrated chip. Background Art

[0002] In recent years, with the continuous development of algorithms, computing power and data volume, machine learning technology has continuously demonstrated its strong advantages in solving many problems. Among them, artificial neural networks have attracted widespread attention for their outstanding performance in image recognition, target detection, semantic segmentation and other fields. However, as the scale of neural networks expands, the traditional model of processing neural network algorithms with CPU+GPU architecture has gradually encountered bottlenecks in speed and power consumption. The root cause is that the separation of storage and computing under the von Neumann architecture makes the data-centric neural network algorithm bring too much data transmission overhead to the computing system, which reduces the speed and increases the power consumption.

[0003] In-memory computing technology solves the problems caused by the separation of storage and computing. It stores the weights of the neural network in the conductance of each flash memory cell in the flash memory cell array in the in-flash NPU chip, and then sends the data source represented by voltage to the flash memory cell array. According to Ohm's law, the current output by the flash memory cell array is the product of voltage and conductance, thereby completing the matrix multiplication and addition operation of the data source and weight. In essence, it is performing analog calculations rather than traditional digital calculations.

[0004] In the whole process of storage and computing integrated chip from design to production, the design of tool chain is an important link. In the design of tool chain for storage and computing integrated chip, it is a key technology to automatically map the weight parameters of a specific neural network to the flash memory cell array of the chip according to the demand; currently, when mapping the trained neural network to the flash memory cell array of the storage and computing integrated chip, the weights and biases are mapped to the storage and computing integrated chip array in sequence according to the order of each layer of the neural network; however, on the one hand, this method cannot effectively utilize the flash memory cell and increases the size of the flash memory cell array; on the other hand, since the bias is directly mapped to the storage and computing integrated chip array, the larger the bias value, the greater the conductance of the corresponding flash memory cell. Under the same voltage, the current of the flash memory cell is larger, which leads to greater noise and affects the calculation accuracy. Summary of the invention

[0005] In response to the problems in the prior art, the present invention provides a linear programming-based neural network mapping method, device, equipment and storage medium for a storage and computing integrated chip, which can at least partially solve the problems in the prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] In a first aspect, a linear programming-based neural network mapping method for a storage-computation integrated chip is provided, comprising:

[0008] Obtain the weight arrays of each layer of the neural network to be mapped, the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip;

[0009] The weight arrays of each layer of the neural network to be mapped and the corresponding bias array data, as well as the hardware parameters of the target storage-computing integrated chip, are input into a pre-established linear programming solution model to obtain a mapping scheme, which is used to map the weight arrays of each layer of the neural network to be mapped and the corresponding bias arrays to the target storage-computing integrated chip.

[0010] Further, the weight array of each layer and the corresponding bias array data include: the number of rows and columns of the weight array of each layer, and the corresponding minimum number of bias rows;

[0011] The hardware parameters of the target integrated chip include: the row maximum value and column maximum value of the flash memory cell array used to write the weight array, the row maximum value of the flash memory cell array used to write the bias, the number of rows of the maximum operation block and the number of columns of the minimum operation block.

[0012] Furthermore, the constraints of the linear programming solution model include:

[0013] The row start address of each layer weight is between 0 and the row maximum value of the flash cell array used to write the weight array;

[0014] The column start address of each layer weight is between 0 and the column maximum value of the flash cell array used to write the weight array;

[0015] The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used for writing bias;

[0016] The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight;

[0017] The weights of each layer are arranged so that they do not overlap with each other;

[0018] The layers are arranged in an offset manner so that they do not overlap with each other.

[0019] Furthermore, the input parameters also include: the first M rows, the last N rows, the first J columns, and the last K columns of the reserved space;

[0020] The constraints of the linear programming model include:

[0021] The row start address of each layer weight is between M and the difference between the row maximum value of the flash memory cell array used to write the weight array minus the number of rows of the corresponding layer weight array minus N;

[0022] The column start address of each layer weight is between J and the difference between the column maximum value of the flash memory cell array used to write the weight array minus the number of rows of the corresponding layer weight array minus K;

[0023] The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used for writing bias;

[0024] The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight;

[0025] The weights of each layer are arranged so that they do not overlap with each other;

[0026] The layers are arranged in an offset manner so that they do not overlap with each other.

[0027] Furthermore, the constraints of the linear programming solution model also include:

[0028] The flash memory cell array used to write the weight array is divided into multiple layers according to the row maximum value of the flash memory cell array used to write the weight array and the number of rows of the maximum operation block. The row start address and row end address of each layer of weight cannot cross layers.

[0029] Furthermore, the constraints of the linear programming solution model also include:

[0030] The number of rows of each layer bias arrangement is an even number.

[0031] Furthermore, the objective function of the linear programming solution model includes:

[0032] The number of rows of bias arrangement of each layer is between the minimum number of rows of bias arrangement of each layer and the maximum number of rows of the flash memory cell array used for writing bias, and the total number of rows of bias arrangement of all layers is the largest; and,

[0033] The flash memory cell array used to write the weight array is divided into Y zones vertically, and the width of each zone is X columns. After arrangement, the sum of the partitions spanned by the weights of each layer is minimized; wherein X is the number of columns of the minimum operation unit, and Y is the number of columns of the flash memory cell array divided by X, that is: X*Y=column width of the flash memory cell array.

[0034] In a second aspect, a linear programming-based neural network mapping device for a storage-computation integrated chip is provided, comprising:

[0035] The data acquisition module obtains the weight array of each layer of the neural network to be mapped, the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip;

[0036] The linear solution module inputs the weight array of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip into a pre-established linear programming solution model to obtain a mapping scheme, which is used to map the weight array of each layer of the neural network to be mapped and the corresponding bias array to the target storage and computing integrated chip.

[0037] In a third aspect, a storage-computing integrated chip is provided, comprising: a flash memory cell array for executing neural network operations, wherein a weight array and a bias array of the neural network are mapped in the flash memory cell array;

[0038] The arrangement of the weight array and the corresponding bias array is generated according to the above-mentioned linear programming-based neural network mapping method.

[0039] In a fourth aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned linear programming-based neural network mapping method when executing the program.

[0040] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned linear programming-based neural network mapping method are implemented.

[0041] The embodiment of the present invention provides a linear programming-based neural network mapping method, device, storage-computing integrated chip, electronic device and computer-readable storage medium for storage-computing integrated chip, the method comprising: obtaining the weight array of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage-computing integrated chip; inputting the weight array of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage-computing integrated chip into a pre-established linear programming solution model to solve and obtain a mapping scheme, the mapping scheme is used to map the weight array of each layer of the neural network to be mapped and the corresponding bias array to the target storage-computing integrated chip. Among them, by converting the weight and bias data mapping process based on intuitive experience into the solution problem of the mathematical model of linear programming, the calculation accuracy is improved.

[0042] In addition, in the embodiment of the present invention, the minimum number of bias rows and the limit of the bias array are used as constraints of the solver, the number of rows occupied by each bias is expanded, the numerical size of the bias stored on a single flash memory cell is reduced, the current noise is reduced, and the calculation accuracy is further improved.

[0043] In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0045] Figure 1 A flowchart of a linear programming-based neural network mapping method for a storage-computation integrated chip in an embodiment of the present invention is shown;

[0046] Figure 2 The flash memory cell array division method in the embodiment of the present invention is illustrated;

[0047] Figure 3 shows parameter details of a flash memory cell array in an embodiment of the present invention;

[0048] Figure 4 The weight matrix and the corresponding Bias arrangement results in the embodiment of the present invention are given as examples;

[0049] Figure 5 It is a structural block diagram of a linear programming-based neural network mapping device for a storage-computation integrated chip in an embodiment of the present invention;

[0050] Figure 6 4 is a structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0052] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0053] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or apparatuses.

[0054] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0055] Figure 1 FIG. 1 is a flowchart of a linear programming-based neural network mapping method for a storage-computation integrated chip in an embodiment of the present invention; Figure 1 As shown, the linear programming-based neural network mapping method for the storage-computing integrated chip may include the following contents:

[0056] Step S100: Obtain the weight arrays of each layer of the neural network to be mapped, the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip;

[0057] The weight arrays of each layer and the corresponding bias array data include: the number of rows and columns of the weight array of each layer, and the corresponding minimum number of bias rows;

[0058] The hardware parameters of the target integrated chip include: the row maximum value and column maximum value of the flash memory cell array used to write the weight array, the row maximum value of the flash memory cell array used to write the bias, the number of rows of the maximum operation block and the number of columns of the minimum operation block.

[0059] It is worth noting that the flash memory cell array of the target integrated chip has been designed during the hardware design stage, including a flash memory cell array for writing the weight array and a flash memory cell array for writing the bias; for the specific neural network that has been trained, the weight matrix parameters and Bias parameters of each layer, the Bias size and the minimum number of rows occupied on the chip, and the hardware parameters of the target storage and computing integrated chip are known.

[0060] Step S200: Input the weight arrays of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip into a pre-established linear programming solution model to obtain a mapping scheme, wherein the mapping scheme is used to map the weight arrays of each layer of the neural network to be mapped and the corresponding bias arrays to the target storage and computing integrated chip.

[0061] Among them, the solver is a pre-established linear programming model. In practical applications, the open source linear programming solver can be directly called in Python, etc. For example, Google's open source linear programming solver can be used.

[0062] The embodiment of the present invention converts the experience-based weight and bias mapping problem into a linear programming problem through appropriate mathematical modeling, provides a strict mathematical basis for the mapping of network weights and biases, and bias expansion, and finds the best arrangement result. The theoretically optimal solution can be obtained. The solution is easy to expand and convenient to increase and reduce restrictions in the mapping. It also takes into account weight and bias mapping and bias expansion, reduces current noise, and improves calculation accuracy.

[0063] It is worth mentioning that due to the particularity of the calculation of the integrated storage and computing chip, the weight matrix and its corresponding Bias need to be aligned in columns. Therefore, the arrangement position of the Bias is aligned in columns with the arrangement position of the corresponding weight matrix, and the arrangement in the row direction is arranged according to the minimum number of Bias rows.

[0064] An arrangement scheme is generated through the above steps S100 and S200, and then the parameters of each layer of the neural network are written into the flash memory cell array of the storage and computing integrated chip through the compilation tool according to the arrangement scheme. In the application reasoning stage, according to the arrangement scheme and in combination with the control requirements, when executing a certain layer of neural network operation, the row and column where the weight matrix and Bias corresponding to the neural network operation of this layer are located are selected through the row and column decoder, and the input signal of this layer of the neural network is input into the row corresponding to the weight matrix, and the matrix multiplication and addition operation is performed with the weight matrix, and then it is superimposed with the corresponding Bias to obtain the calculation result of this layer of the neural network in the corresponding column.

[0065] Figure 2 The flash memory cell array division method in the embodiment of the present invention is exemplified, such as Figure 2 As shown in (a) in the figure, the actual physical architecture of the chip consists of a main array (used to write weight arrays) and a bias array. In actual applications, since too much current in the analog calculation process will have a significant impact on the calculation results, a maximum operation block is given when designing the storage and computing integrated chip, that is, the maximum scale of a single operation. If the operation scale exceeds the maximum operation block, it may be necessary to perform the operation multiple times, such as Figure 2 As shown in (b), the main array can be divided into 2×4 blocks. The division here is based on the actual performance of the chip, and each block can be the same size or different sizes. The embodiment of the present invention does not limit this. In addition, in order to ensure work efficiency, the storage and computing integrated chip will provide a minimum operation block when designing, that is, the minimum scale of a single operation, to prevent the efficiency from being reduced due to the small amount of single operation in the application stage.

[0066] In an optional embodiment, the linear programming-based neural network mapping method for the storage-computing integrated chip may also include: writing the weight matrix and bias of each layer of the neural network to be mapped onto the storage-computing integrated chip according to the mapping scheme.

[0067] Specifically, the above-mentioned mapping method is executed on the tool chain, which can be understood as a program running on a terminal device or a server or a chip burning device. The above-mentioned mapping method generates an arrangement scheme, and then the weight matrix and Bias need to be written into the storage-computing integrated chip according to the arrangement scheme. Then the storage-computing integrated chip can be installed on the corresponding device circuit board for reasoning application and neural network operation. For example, it can be installed on a toy for voice recognition. At this time, the neural network parameters written in the chip are the parameters corresponding to the voice recognition neural network; of course, the storage-computing integrated chip can also be installed on a face recognition device, and the neural network parameters written in the chip are the parameters corresponding to the image recognition neural network. Of course, the above only cites several chip application scenarios. The embodiment of the present invention does not limit the application scenarios of the chip, which can be various devices and scenarios that require neural network operations.

[0068] Figure 3 Detailed parameters of the flash memory cell array in the embodiment of the present invention are shown; see Figure 3 , the number of rows of the i-th layer weight array is rows[i], the number of columns of the i-th layer weight array is cols[i], and the corresponding minimum number of bias rows is bias[i];

[0069] In the hardware parameters of the target integrated chip, the maximum row value of the flash memory cell array used to write the weight array is row_size, the maximum column value is col_size, and the maximum row value of the flash memory cell array used to write the bias is bias_size. The model output is: x[i], y[i], bx[i], by[i], bias[i], where x[i] is the row starting address of the i-th layer weight array in the flash memory cell array for writing the weight array, y[i] is the column starting address of the i-th layer weight array in the flash memory cell array for writing the weight array, bx[i] is the row starting address of the i-th layer bias in the flash memory cell array for writing the bias, by[i] is the column starting address of the i-th layer bias in the flash memory cell array for writing the bias, and bias[i] is the number of rows of the i-th layer bias in the flash memory cell array for writing the bias after the extended arrangement; the number of rows of the maximum operation block is P. For example, assuming that the scale of the maximum operation block set by an in-memory computing chip is 512*128, that is, the maximum scale of a single operation, at this time, P=512;

[0070] In an optional embodiment, the constraints of the linear programming solution model include:

[0071] 1. The starting address of the row of each layer weight is between 0 and the maximum row value of the flash memory cell array used to write the weight array; that is: 0≤x[i]≤rowsize-rows[i]);

[0072] 2. The column start address of each layer weight is between 0 and the column maximum value of the flash cell array used to write the weight array; that is: 0≤y[i]≤colsize-cals[i];

[0073] 3. The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used to write the bias; that is: (0≤bx[i]≤biassize-bias[i]);

[0074] 4. The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; that is: by[i] = y[i];

[0075] 5. The weights of each layer are arranged so that they do not overlap with each other; that is, x[i]≥x[j]+rows[j] or x[j]≥x[i]+rows[i] or y[i]≥y[j]+cols[j] or y[j]≥y[i]+cols[i];

[0076] 6. The biases of each layer are arranged so that they do not overlap with each other. That is, bx[i]≥bx[j]+bias[j] or bx[j]≥bx[i]+bias[i];

[0077] The objective function of the linear programming model includes:

[0078] 1. The number of rows of bias arrangement of each layer is between the minimum number of rows of bias arrangement of each layer and the maximum number of rows of the flash memory cell array used for writing bias, and the sum of the number of rows of bias arrangement of all layers is the largest, that is: (max∑bias[i]);

[0079] 2. Divide the flash memory cell array used to write the weight array into Y zones vertically, with each zone having a width of X columns. After arrangement, the sum of the zones spanned by the weights of each layer is minimized, i.e., min∑across[i]. Among them, X is the number of columns of the minimum operation block, Y = the number of columns of the current layer weight array divided by X, that is, X*Y = the column width of the flash memory cell array, and across[i] represents the number of partitions spanned after the weight matrix of the i-th layer neural network is mapped to the flash memory cell array; due to the limitation of the maximum operation scale of a single operation, when arranging, it is hoped that the number of operations of the current layer weight array is as small as possible. If the flash memory cell array is partitioned according to the maximum operation block, it is hoped that the number of partitions spanned by the current layer weight array is as small as possible, and the power consumed by the operation is lower.

[0080] By adopting the above technical solution, the constraints and solution objectives are optimized. After performing linear solution based on the above constraints and solution objectives, the optimal arrangement of the weight matrix and bias matrix of the neural network model can be obtained.

[0081] In an optional embodiment, it is also possible to support the weight array and bias array space margin, and reserve the first M rows, the last N rows, the first J columns, and the last K columns (the 0 and maximum values ​​of items 1 and 2 can be changed accordingly based on this restriction);

[0082] The constraints of the linear programming model include:

[0083] 1. The starting address of each layer of weights is between M and the difference between the maximum value of the row of the flash memory cell array used to write the weight array minus the number of rows of the corresponding layer weight array minus N; that is, M≤x[i]≤rawsize-rows[i]-N;

[0084] 2. The column start address of each layer weight is between J and the difference between the column maximum value of the flash memory cell array used to write the weight array minus the number of rows of the corresponding layer weight array minus K; that is: J≤y[i]≤colsize-cals[i]-K;

[0085] 3. The row start address of each layer bias is between 0 and the row maximum value of the flash cell array used to write the bias; that is: (0≤bx[i]≤biassize-bias[i])

[0086] 4. The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; that is: by[i] = y[i];

[0087] 5. The weights of each layer are arranged so that they do not overlap with each other; that is, x[i]≥x[j]+rows[j] or x[j]≥x[i]+rows[i] or y[i]≥y[j]+cols[j] or y[j]≥y[i]+cols[i];

[0088] 6. The biases of each layer are arranged so that they do not overlap with each other. That is, bx[i]≥bx[j]+bias[j] or bx[j]≥bx[i]+bias[i];

[0089] The objective function of the linear programming model includes:

[0090] 1. The number of rows of bias arrangement of each layer is between the minimum number of rows of bias arrangement of each layer and the maximum number of rows of the flash memory cell array used for writing bias, and the sum of the number of rows of bias arrangement of all layers is the largest, that is: (max∑bias[i]);

[0091] 2. Divide the flash memory cell array used to write the weight array into Y zones vertically, with each zone having a width of X columns. After arrangement, the sum of the zones spanned by the weights of each layer is minimized, i.e., min∑across[i]. Among them, X is the number of columns of the minimum operation block, Y = the number of columns of the current layer weight array divided by X, that is, X*Y = the column width of the flash memory cell array, and across[i] represents the number of partitions spanned after the weight matrix of the i-th layer neural network is mapped to the flash memory cell array; due to the limitation of the maximum operation scale of a single operation, when arranging, it is hoped that the number of operations of the current layer weight array is as small as possible. If the flash memory cell array is partitioned according to the maximum operation block, it is hoped that the number of partitions spanned by the current layer weight array is as small as possible, and the power consumed by the operation is lower.

[0092] In an optional embodiment, the constraints of the linear programming solution model also include: dividing the flash memory cell array used to write the weight array into multiple layers according to the row maximum value of the flash memory cell array used to write the weight array and the number of rows of the maximum operation block, and the row start address and row end address of each layer of weights cannot cross layers.

[0093] Specifically, the constraints of the linear programming model include:

[0094] 1. The starting address of the row of each layer weight is between 0 and the maximum row value of the flash memory cell array used to write the weight array; that is: 0≤x[i]≤rowsize-rows[i]);

[0095] 2. The column start address of each layer weight is between 0 and the column maximum value of the flash cell array used to write the weight array; that is: 0≤y[i]≤colsize-cols[i];

[0096] 3. The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used to write the bias; that is: (0≤bx[i]≤biassize-bias[i]);

[0097] 4. The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; that is: by[i] = y[i];

[0098] 5. The weights of each layer are arranged so that they do not overlap with each other; that is, x[i]≥x[j]+rows[j] or x[j]≥x[i]+rows[i] or y[i]≥y[j]+cols[j] or y[j]≥y[i]+cols[i];

[0099] 6. The biases of each layer are arranged so that they do not overlap with each other. That is, bx[i]≥bx[j]+bias[j] or bx[j]≥bx[i]+bias[i];

[0100] 7. The weight array is divided into Q layers according to P alignment. The row start address and row end address of each layer of weight cannot cross layers (for example, when the row is divided into Q layers, the restriction is or );

[0101] That is: the flash memory cell array used to write the weight array is evenly divided into Q layers according to P alignment; wherein P is the number of rows of the maximum operation block; and Q is the maximum value of the row of the flash memory cell array used to write the weight array divided by P and rounded. Q = row_size / P, for example, when row_size = 1024 and P = 512, Q = 1024 / 512 = 2, that is, it is evenly divided into two layers.

[0102] For example, the current maximum row values ​​of the flash memory cell array are 1792 and 2048, respectively, and are divided into two layers, with 896 and 1024 rows in each layer. When allocating weights, you cannot cross layers. The weight allocation can cross 448 / 512, but not 896 / 1024, where 448 and 512 are the number of rows of the maximum operation block, that is: divide the flash memory cell array into four layers horizontally, with 448 / 512 rows in each layer. When allocating weights, you can cross layers 1, 2, 3, and 4, but try to cross as few layers as possible, and it is forbidden to cross layers 2 and 3.

[0103] The objective function of the linear programming model includes:

[0104] 1. The number of rows of bias arrangement of each layer is between the minimum number of rows of bias arrangement of each layer and the maximum number of rows of the flash memory cell array used for writing bias, and the sum of the number of rows of bias arrangement of all layers is the largest, that is: (max∑bias[i]);

[0105] 2. Divide the flash memory cell array used to write the weight array into Y zones vertically, with each zone having a width of X columns. After arrangement, the sum of the zones spanned by the weights of each layer is minimized, i.e., min∑across[i]. Among them, X is the number of columns of the minimum operation block, Y = the number of columns of the current layer weight array divided by X, that is, X*Y = the column width of the flash memory cell array, and across[i] represents the number of partitions spanned after the weight matrix of the i-th layer neural network is mapped to the flash memory cell array; due to the limitation of the maximum operation scale of a single operation, when arranging, it is hoped that the number of operations of the current layer weight array is as small as possible. If the flash memory cell array is partitioned according to the maximum operation block, it is hoped that the number of partitions spanned by the current layer weight array is as small as possible, and the power consumed by the operation is lower.

[0106] By adopting the above technical solution, the calculation accuracy can be further improved.

[0107] In an optional embodiment, the constraints of the linear programming solution model include:

[0108] 1. The starting address of the row of each layer weight is between 0 and the maximum row value of the flash memory cell array used to write the weight array; that is: 0≤x[i]≤rowsize-rows[i]);

[0109] 2. The column start address of each layer weight is between 0 and the column maximum value of the flash cell array used to write the weight array; that is: 0≤y[i]≤colsize-cols[i];

[0110] 3. The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used to write the bias; that is: (0≤bx[i]≤biassize-bias[i]);

[0111] 4. The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; that is: by[i] = y[i];

[0112] 5. The weights of each layer are arranged so that they do not overlap with each other; that is, x[i]≥x[j]+rows[j] or x[j]≥x[i]+rows[i] or y[i]≥y[j]+cols[j] or y[j]≥y[i]+cols[i];

[0113] 6. The biases of each layer are arranged so that they do not overlap with each other. That is, bx[i]≥bx[j]+bias[j] or bx[j]≥bx[i]+bias[i];

[0114] 7. The weight array is divided into Q layers according to P alignment. The row start address and row end address of each layer of weight cannot cross layers (for example, when the row is divided into Q layers, the restriction is or );

[0115] That is: the flash memory cell array used to write the weight array is evenly divided into Q layers according to P alignment; wherein P is the number of rows of the maximum operation block; and Q is the maximum value of the row of the flash memory cell array used to write the weight array divided by P and rounded. Q = row_size / P, for example, when row_size = 1024 and P = 512, Q = 1024 / 512 = 2, that is, it is evenly divided into two layers.

[0116] 8. The number of rows of bias in each layer must be an even number; that is, the number of rows expanded by bias in each layer must be an even number (bias[i] mod 2 = 0)

[0117] The objective function of the linear programming model includes:

[0118] 1. The number of rows of bias arrangement of each layer is between the minimum number of rows of bias arrangement of each layer and the maximum number of rows of the flash memory cell array used for writing bias, and the sum of the number of rows of bias arrangement of all layers is the largest, that is: (max∑bias[i]);

[0119] 2. Divide the flash memory cell array used to write the weight array into Y zones vertically, with each zone having a width of X columns. After arrangement, the sum of the zones spanned by the weights of each layer is minimized, i.e., min∑across[i]. Among them, X is the number of columns of the minimum operation block, Y = the number of columns of the current layer weight array divided by X, that is, X*Y = the column width of the flash memory cell array, and across[i] represents the number of partitions spanned after the weight matrix of the i-th layer neural network is mapped to the flash memory cell array; due to the limitation of the maximum operation scale of a single operation, when arranging, it is hoped that the number of operations of the current layer weight array is as small as possible. If the flash memory cell array is partitioned according to the maximum operation block, it is hoped that the number of partitions spanned by the current layer weight array is as small as possible, and the power consumed by the operation is lower.

[0120] By adopting the above technical solution, the number of rows occupied by each bias is expanded, the value of the bias stored in a single flash memory unit is reduced, the current noise is reduced, and the calculation accuracy is further improved.

[0121] It is worth noting that after the neural network model is determined, the weight array and bias value of each layer are known, and the minimum number of Bias rows for each layer is calculated based on the bias value and the parameters of the target chip (the properties of each row of Bias).

[0122] The minimum number of rows of Bias can be given in advance by circuit engineers according to accuracy requirements, generally based on meeting the worst accuracy requirements, or it can be calculated according to preset rules, which will not be elaborated in the embodiments of the present invention.

[0123] In the embodiment of the present invention, one of the goals of the Bias arrangement is to minimize the number of free rows in the Bias array. In order to minimize the Bias values ​​stored in all or part of the integrated storage and computing units in the Bias array, it is necessary to use the free rows in the Bias array to expand the Bias arrangement to obtain the final arrangement scheme, and then write the neural network parameters into the integrated storage and computing chip according to the final arrangement scheme.

[0124] Since the integrated storage and computing chip essentially uses analog computing, the larger the bias value of each integrated storage and computing unit on the Bias array, the greater the noise generated by the final calculation. The excessive noise introduced by excessive Bias will have a decisive impact on the calculation accuracy. Therefore, according to the array size, the actual number of Bias array rows occupied by a logical row of Bias can be expanded as much as possible. For example, if the actual number of rows occupied is m, the Bias size stored on each row is 1 / m of the logical Bias size, thereby improving the calculation accuracy.

[0125] In order to enable those skilled in the art to better understand the present application, the following examples are used to describe an embodiment of the present invention: it is assumed that the hardware parameters of the target storage and computing integrated chip are as shown in Table 1, and the parameters of the neural network model to be mapped are as shown in Table 2.

[0126] Table 1: Hardware parameters of the target storage and computing integrated chip:

[0127]

[0128]

[0129] Among them, the single block weight array represents the largest operation block, and the column alignment width represents the column width of the smallest operation block;

[0130] Table 2: Target neural network model parameters:

[0131]

[0132] Among them, the neural network model to be mapped has 10 operation layers. For example, if the neural network to be mapped is a convolutional neural network CNN model, then there are 10 convolution layers; the number of weight rows of the first layer of the neural network is 440, the number of columns / Bias is 112, and the minimum number of Bias rows of the first layer of the neural network is 2.

[0133] Table 3: Mapping scheme

[0134]

[0135]

[0136] Table 3 shows the mapping scheme obtained based on the data in Table 1 and Table 2 by using the linear programming-based neural network mapping method for the storage-computation integrated chip in the embodiment of the present invention. The result after mapping the mapping scheme shown in Table 3 to the target neural network is as follows: Figure 4 shown.

[0137] Among them, the sequence numbers of the 1st to 10th layers of neural networks after arrangement are 0 to 9. For example, after the first layer of neural network is arranged, its row starting address x[0]=896, and its column starting address y[0]=562; the corresponding Bias row starting address is bx[0]=0, and the column starting address is by[0]=562; the minimum number of Bias rows of the first layer of neural network is 2, and the number of Bias rows after expansion is 12 rows.

[0138] An embodiment of the present invention further provides a storage-computation integrated chip, comprising: a flash memory cell array for executing neural network operations, wherein a weight array and a bias array of the neural network are mapped in the flash memory cell array;

[0139] The arrangement of the weight array and the corresponding bias array is generated according to the above-mentioned linear programming-based neural network mapping method.

[0140] It is worth noting that the integrated storage and computing chip provided in the embodiment of the present invention can be applied to various electronic devices, such as: smart phones, tablet electronic devices, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), vehicle-mounted devices, smart wearable devices, toys, smart home control devices, assembly line equipment controllers, etc. Among them, the smart wearable devices may include smart glasses, smart watches, smart bracelets, etc.

[0141] Based on the same inventive concept, the embodiments of the present application also provide a linear programming-based neural network mapping device for a storage-computing integrated chip, which can be used to implement the method described in the above embodiment, as described in the following embodiments. The principle of solving the problem by the linear programming-based neural network mapping device for the storage-computing integrated chip is similar to the above method, so the implementation of the linear programming-based neural network mapping device for the storage-computing integrated chip can refer to the implementation of the above method, and the repeated parts will not be repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0142] Figure 5 It is a structural block diagram of a linear programming-based neural network mapping device for a storage-computation integrated chip in an embodiment of the present invention. The linear programming-based neural network mapping device for a storage-computation integrated chip includes: a data acquisition module 10 and a linear solution module 20.

[0143] The data acquisition module 10 acquires the weight arrays of each layer of the neural network to be mapped, the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip;

[0144] The linear solution module 20 inputs the weight array of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip into a pre-established linear programming solution model to obtain a mapping scheme, wherein the mapping scheme is used to map the weight array of each layer of the neural network to be mapped and the corresponding bias array to the target storage and computing integrated chip.

[0145] The devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is an electronic device, and the electronic device may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0146] In a typical example, the electronic device specifically includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned neural network mapping method for a storage-computing integrated chip are implemented.

[0147] Reference below Figure 6 , which shows a structural schematic diagram of an electronic device 600 suitable for implementing an embodiment of the present application.

[0148] like Figure 6 As shown, electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In RAM 603, various programs and data required for the operation of system 600 are also stored. CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0149] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed as needed as the storage section 608.

[0150] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned neural network mapping method for a storage-computing integrated chip.

[0151] In such an embodiment, the computer program may be downloaded and installed from a network via the communication section 609 , and / or installed from the removable medium 611 .

[0152] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0153] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0154] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0155] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0157] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0158] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0159] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0160] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0161] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A neural network mapping method based on linear programming for a storage and computing integrated chip, characterized in that: include: Obtain the weight arrays of each layer of the neural network to be mapped, the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip; Input the weight arrays of each layer of the neural network to be mapped and the corresponding bias array data, and the hardware parameters of the target storage-computing integrated chip into a pre-established linear programming solution model to solve and obtain a mapping scheme, wherein the mapping scheme is used to map the weight arrays of each layer of the neural network to be mapped and the corresponding bias arrays to the target storage-computing integrated chip; The weight arrays of each layer and the corresponding bias array data include: the number of rows and columns of the weight array of each layer, and the corresponding minimum number of bias rows; The hardware parameters of the target storage-computing integrated chip include: the row maximum value and column maximum value of the flash memory cell array for writing the weight array, the row maximum value of the flash memory cell array for writing the bias, the number of rows of the maximum operation block, and the number of columns of the minimum operation block; Among them, the objective function of the linear programming solution model includes: the number of arrangement rows of each layer bias is between the minimum number of rows of each layer bias and the maximum value of the row of the flash memory cell array used to write the bias, and the total number of arrangement rows of the bias of all layers is the largest; and the flash memory cell array used to write the weight array is divided into Y zones in the longitudinal direction, the width of each zone is X columns, and the sum of the partitions spanned by the weights of each layer after arrangement is the smallest; wherein X is the number of columns of the minimum operation block, and Y is the number of columns of the flash memory cell array divided by X.

2. The linear programming-based neural network mapping method for a storage-computation integrated chip according to claim 1, characterized in that: The constraints of the linear programming solution model include: The row start address of each layer weight is between 0 and the row maximum value of the flash cell array used to write the weight array; The column start address of each layer weight is between 0 and the column maximum value of the flash cell array used to write the weight array; The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used for writing bias; The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; The weights of each layer are arranged so that they do not overlap with each other; The layers are arranged in an offset manner so that they do not overlap with each other.

3. The linear programming-based neural network mapping method for a storage-computation integrated chip according to claim 1, characterized in that: The parameters input into the linear programming solution model also include: the first M rows, the last N rows, the first J columns, and the last K columns of the reserved space; The constraints of the linear programming solution model include: The row start address of each layer weight is between M and the difference between the row maximum value of the flash memory cell array used to write the weight array minus the number of rows of the corresponding layer weight array minus N; The column start address of each layer weight is between J and the difference between the column maximum value of the flash memory cell array used to write the weight array minus the number of columns of the corresponding layer weight array minus K; The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used for writing bias; The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; The weights of each layer are arranged so that they do not overlap with each other; The layers are arranged in an offset manner so that they do not overlap with each other.

4. The linear programming-based neural network mapping method for a storage-computation integrated chip according to claim 2, characterized in that: The constraints of the linear programming solution model also include: The flash memory cell array used to write the weight array is divided into multiple layers according to the row maximum value of the flash memory cell array used to write the weight array and the number of rows of the maximum operation block. The row start address and row end address of each layer of weight cannot cross layers.

5. The linear programming-based neural network mapping method for a storage-computation integrated chip according to claim 4 is characterized in that: The constraints of the linear programming solution model also include: The number of rows of each layer bias arrangement is an even number.

6. A neural network mapping device based on linear programming for a storage and computing integrated chip, characterized in that: include: The data acquisition module obtains the weight array of each layer of the neural network to be mapped, the corresponding bias array data, and the hardware parameters of the target storage and computing integrated chip; A linear solution module, inputting the weight arrays of each layer of the neural network to be mapped and the corresponding bias array data and the hardware parameters of the target storage-computing integrated chip into a pre-established linear programming solution model to solve and obtain a mapping scheme, wherein the mapping scheme is used to map the weight arrays of each layer of the neural network to be mapped and the corresponding bias arrays to the target storage-computing integrated chip; The weight arrays of each layer and the corresponding bias array data include: the number of rows and columns of the weight array of each layer, and the corresponding minimum number of bias rows; The hardware parameters of the target storage-computing integrated chip include: the row maximum value and column maximum value of the flash memory cell array for writing the weight array, the row maximum value of the flash memory cell array for writing the bias, the number of rows of the maximum operation block, and the number of columns of the minimum operation block; Among them, the objective function of the linear programming solution model includes: the number of arrangement rows of each layer bias is between the minimum number of rows of each layer bias and the maximum value of the row of the flash memory cell array used to write the bias, and the total number of arrangement rows of the bias of all layers is the largest; and the flash memory cell array used to write the weight array is divided into Y zones in the longitudinal direction, the width of each zone is X columns, and the sum of the partitions spanned by the weights of each layer after arrangement is the smallest; wherein X is the number of columns of the minimum operation block, and Y is the number of columns of the flash memory cell array divided by X.

7. The linear programming-based neural network mapping device for a storage-computation integrated chip according to claim 6, characterized in that: The constraints of the linear programming solution model include: The row start address of each layer weight is between 0 and the row maximum value of the flash cell array used to write the weight array; The column start address of each layer weight is between 0 and the column maximum value of the flash cell array used to write the weight array; The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used for writing bias; The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; The weights of each layer are arranged so that they do not overlap with each other; The layers are arranged in an offset manner so that they do not overlap with each other.

8. The linear programming-based neural network mapping device for a storage-computation integrated chip according to claim 6, characterized in that: The parameters input into the linear programming solution model also include: the first M rows, the last N rows, the first J columns, and the last K columns of the reserved space; The constraints of the linear programming solution model include: The row start address of each layer weight is between M and the difference between the row maximum value of the flash memory cell array used to write the weight array minus the number of rows of the corresponding layer weight array minus N; The column start address of each layer weight is between J and the difference between the column maximum value of the flash memory cell array used to write the weight array minus the number of columns of the corresponding layer weight array minus K; The row start address of each layer bias is between 0 and the row maximum value of the flash memory cell array used for writing bias; The column starting address of each layer bias is equal to the column starting address of the corresponding layer weight; The weights of each layer are arranged so that they do not overlap with each other; The layers are arranged in an offset manner so that they do not overlap with each other.

9. The linear programming-based neural network mapping device for a storage-computation integrated chip according to claim 7, characterized in that: The constraints of the linear programming solution model also include: The flash memory cell array used to write the weight array is divided into multiple layers according to the row maximum value of the flash memory cell array used to write the weight array and the number of rows of the maximum operation block. The row start address and row end address of each layer of weight cannot cross layers.

10. The linear programming-based neural network mapping device for a storage-computation integrated chip according to claim 9, characterized in that: The constraints of the linear programming solution model also include: The number of rows of each layer bias arrangement is an even number.

11. A storage and computing integrated chip, comprising: A flash memory cell array for performing neural network operations, wherein a weight array and a bias array of the neural network are mapped into the flash memory cell array; The arrangement of the weight array and the corresponding bias array is generated according to the linear programming-based neural network mapping method according to any one of claims 1 to 5.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the linear programming-based neural network mapping method described in any one of claims 1 to 5 are implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the linear programming-based neural network mapping method described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method to control error floor in analog neural LDPC decoder

    CA2651256A1

  • Operating method of neural network device

    CN108537325A