Neural network reasoning method, system, device and equipment, storage medium and program product

By splitting the neural network weight matrix and quantizing different precisions, mapping to the main array and sparse computing array of the memristor array, the problem of outliers affecting inference accuracy in neural networks is solved, and more efficient neural network inference is achieved.

CN119962601APending Publication Date: 2025-05-09INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411949087.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

With the increase in the scale of neural networks, outliers in the network affect the quantization effect, resulting in low inference accuracy of neural networks, and it is difficult for the existing technology to effectively solve this problem.

Method used

By splitting the weight matrix of the target neural network, the outlier matrix and the non-outlier matrix are obtained, and they are quantized with different precisions respectively. The quantized matrix is ​​mapped to the main array and sparse computing array in the memristor array, and the inference process is performed separately.

Benefits of technology

By isolating outliers and non-outliers and performing quantization processing with different precisions, the inference accuracy and efficiency of the neural network are improved, effectively solving the on-chip inference accuracy and performance problems caused by outliers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962601A_ABST
    Figure CN119962601A_ABST
Patent Text Reader

Abstract

The invention relates to a neural network reasoning method, system, device and equipment, a storage medium and a program product, and the method comprises the steps: carrying out the splitting of a weight matrix of a target neural network, obtaining an outlier matrix and a non-outlier matrix, carrying out the quantification processing of first precision of the non-outlier matrix, obtaining a first quantification matrix, and obtaining a second quantification matrix; and performing quantization processing of second precision on the outlier matrix to obtain a second quantization matrix, then mapping the first quantization matrix to a main array in the memristor array, mapping the second quantization matrix to a sparse calculation array in the memristor array, and finally driving the mapped memristor array to perform reasoning on the neural network to obtain the memristor array. And obtaining a reasoning result. According to the method, the main array is used for processing the non-outlier matrix, the sparse calculation array is used for processing the outlier matrix, the outlier and the non-outlier are processed in an isolated and separated mode, quantization processing of different precisions is carried out, and the reasoning efficiency and the reasoning precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a neural network reasoning method, system, device, equipment, storage medium and program product. Background Art

[0002] The storage-computing integrated architecture can effectively solve the storage wall problem caused by large-scale data movement, and has the advantage of high energy efficiency in realizing neural network reasoning, and has broad prospects for edge or cloud reasoning. Memristor array is one of the important technical paths of the storage-computing integrated architecture, which can efficiently realize matrix multiplication. The core process of neural network reasoning based on memristor array is to map the weight matrix of the neural network to the memristor array, and then apply the input voltage signal to the memristor array and measure the output current to complete the matrix multiplication operation.

[0003] However, as the scale of neural networks increases, outliers in the network begin to become a new problem worthy of attention. Outliers are manifested as data with significantly larger absolute values ​​in the weight matrix. Outliers affect the quantization effect, and then affect the final reasoning effect of the neural network model, resulting in low reasoning accuracy. Therefore, how to achieve accurate reasoning of neural networks has become a problem that needs to be solved urgently. Summary of the invention

[0004] Based on this, it is necessary to provide a neural network reasoning method, system, device, equipment, storage medium and program product that can improve the reasoning accuracy of the neural network in response to the above technical problems.

[0005] In a first aspect, the present application provides a neural network reasoning method, the method comprising:

[0006] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0007] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0008] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0009] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0010] In one embodiment, the weight matrix of the target neural network is split to obtain an outlier matrix and a non-outlier matrix, including:

[0011] Determine the weight values ​​in the weight matrix that are within the preset weight range as first weight values, and determine the weight values ​​in the weight matrix that are not within the preset weight range as second weight values;

[0012] The weight matrix is ​​split according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix.

[0013] In one embodiment, the weight matrix is ​​split according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix, including:

[0014] Constructing a first initial matrix and a second initial matrix; the scale of the first initial matrix is ​​the same as the scale of the weight matrix; the scale of the second initial matrix is ​​the same as the scale of the weight matrix;

[0015] Migrating the first weight value in the weight matrix to the corresponding position of the first initial matrix, and setting the weight values ​​at other positions in the first initial matrix to the first preset values, to generate an outlier matrix;

[0016] The second weight values ​​in the weight matrix are migrated to corresponding positions in the second initial matrix, and the weight values ​​at other positions in the second initial matrix are set to second preset values ​​to generate a non-outlier matrix.

[0017] In one embodiment, the main array includes a plurality of first memristors, the sparse computing array includes a plurality of sparse computing units, each sparse computing unit includes a plurality of second memristors, mapping the first quantization matrix to the main array in the memristor array, and mapping the second quantization matrix to the sparse computing array in the memristor array, comprising:

[0018] Setting the value of each first memristor in the main array to each quantized value in the first quantization matrix;

[0019] The value of each sparse computing unit in the sparse computing array is correspondingly set to each quantized value in the second quantization matrix.

[0020] In one embodiment, the method further comprises:

[0021] Determine if the size of the non-outlier matrix exceeds the size of the main array;

[0022] If the size of the non-outlier matrix does not exceed the size of the main array, returning to the step of performing a first precision quantization process on the non-outlier matrix;

[0023] If the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and any segmented outlier matrix, return to the step of performing first precision quantization processing on the non-outlier matrix.

[0024] In one embodiment, the mapped memristor array is driven to infer the neural network to obtain an inference result, including:

[0025] Obtain input data corresponding to the target neural network and convert the input data into a voltage signal;

[0026] The voltage signal is input into the mapped memristor array, and the mapped memristor array is driven to perform reasoning to obtain the reasoning result.

[0027] In a second aspect, the present application further provides a neural network reasoning system, the reasoning system comprising a processor, at least one memristor array and a peripheral circuit; the processor is connected to an input end and an output end of the memristor array through the peripheral circuit;

[0028] The memristor array includes a main array and a sparse computing array; the main array includes a plurality of first memristors, each of which is cross-connected, the sparse computing array includes a plurality of sparse computing units, each of which includes a plurality of second memristors, and the plurality of second memristors in each sparse computing unit are connected in parallel;

[0029] A processor, configured to execute the steps of the neural network reasoning method of any one of the embodiments of the first aspect.

[0030] In a third aspect, the present application also provides a neural network reasoning device, the device comprising:

[0031] A splitting module is used to split the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix;

[0032] A quantization module, configured to perform a quantization process on the non-outlier matrix with a first precision to obtain a first quantization matrix, and to perform a quantization process on the outlier matrix with a second precision to obtain a second quantization matrix;

[0033] A mapping module, used for mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is arranged in a processor;

[0034] The inference module is used to drive the mapped memristor array to infer the neural network and obtain the inference result.

[0035] In a fourth aspect, the present application further provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:

[0036] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0037] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0038] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0039] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0040] In a fifth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0041] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0042] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0043] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0044] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0045] In a sixth aspect, the present application further provides a computer program product, the computer program product comprising a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0046] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0047] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0048] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0049] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0050] The above-mentioned neural network reasoning method, system, device, equipment, storage medium and program product, the method splits the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix, then performs a first precision quantization process on the non-outlier matrix to obtain a first quantization matrix, and performs a second precision quantization process on the outlier matrix to obtain a second quantization matrix, and then maps the first quantization matrix to the main array in the memristor array, and maps the second quantization matrix to the sparse computing array in the memristor array, and finally drives the mapped memristor array to reason the neural network to obtain a reasoning result. In the above-mentioned method, by using the main array to process the non-outlier matrix, and using the sparse computing array to process the outlier matrix, the outliers and non-outliers are isolated and processed separately in the subsequent mapping and reasoning process, and the outlier matrix and the non-outlier matrix are quantized with different precisions. For non-outliers, the low-precision quantization range can be fully utilized to improve the reasoning efficiency, and for outliers, higher-precision quantization can be performed to improve the reasoning accuracy. On the other hand, during on-chip reasoning, the sparse computing array uses multiple devices to represent a weight, which effectively improves the calculation accuracy. Therefore, this approach can effectively solve the on-chip reasoning accuracy and performance problems caused by outliers. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a structural block diagram of an inference system of a neural network in one embodiment;

[0052] Figure 2 is a flowchart of a neural network reasoning method in one embodiment;

[0053] Figure 3 is a flowchart of a neural network reasoning method in another embodiment;

[0054] Figure 4 is a flowchart of a neural network reasoning method in another embodiment;

[0055] Figure 5 is a flowchart of a neural network reasoning method in another embodiment;

[0056] Figure 6 is a schematic structural diagram of a memristor array in one embodiment;

[0057] Figure 7is a flowchart of a neural network reasoning method in another embodiment;

[0058] Figure 8 is a flowchart of a neural network reasoning method in another embodiment;

[0059] Fig. 9 is a flowchart of a neural network reasoning method in another embodiment;

[0060] Fig.10 is a flowchart of a neural network reasoning method in another embodiment;

[0061] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] The storage-computing integrated architecture can effectively solve the storage wall problem caused by large-scale data movement, and has the advantage of realizing neural network reasoning with high energy efficiency, and has broad prospects for edge or cloud reasoning. Memristor array is one of the important technical paths of storage-computing integrated architecture, which can efficiently realize matrix multiplication, so it can be used in neural network reasoning scenarios. At present, there are tasks such as image classification and target detection based on memristor array, and they have shown excellent results. The core process of neural network reasoning based on memristor array is to quantize the weight matrix of the neural network so that its quantization bit width matches the storage accuracy of the memristor, and then map the quantized neural network weight matrix to the memristor array, and finally apply the input voltage signal and measure the output current to complete a matrix multiplication operation. With the increase in the scale of neural networks, outliers in the network have begun to become a new problem worthy of attention. For example, significant outlier phenomena have appeared in large language models. Outliers refer to a small number of data points in a set of data that are significantly different from other data. In neural networks, outliers are reflected as data with significantly larger absolute values ​​in the weight matrix. Outliers will affect the quantization effect and reduce the quantization accuracy. In the case of limited storage and calculation accuracy of memristors, outliers will further affect the final reasoning effect of the model, resulting in low reasoning accuracy. Therefore, how to achieve accurate reasoning of neural networks has become an urgent problem to be solved.

[0064] The present application provides a neural network reasoning method, aiming to solve the above technical problems. The following embodiments will specifically illustrate the neural network reasoning method described in the present application.

[0065] The neural network reasoning method provided in the embodiment of the present application can be applied to Figure 1 The neural network reasoning system shown in the figure includes a processor 10, at least one memristor array 20 and a peripheral circuit 30, wherein the processor 10 is connected to the input and output ends of the memristor array 20 through the peripheral circuit 30, wherein the memristor array includes a main array and a sparse computing array. Optionally, the memristor array can be integrated into the processor or designed independently from the processor. The above-mentioned processor 10 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, portable wearable devices, etc. The processor 10 can also be implemented using an independent server or a server cluster consisting of multiple servers.

[0066] Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the neural network reasoning system to which the scheme of the present application is applied. The specific neural network reasoning system may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0067] In one embodiment, Figure 2 As shown, a neural network reasoning method is provided, which is applied to Figure 1 The processor in the example is used to illustrate, including the following steps:

[0068] S201, splitting the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix.

[0069] Among them, the target neural network can be any type of neural network. The weight matrix includes outliers and non-outliers. Outliers are weight values ​​within the preset weight range, and non-outliers are weight values ​​that are not within the weight range. The weight range includes the lower limit of the weight threshold and the upper limit of the weight threshold. The outlier is less than the lower limit of the weight threshold or greater than the upper limit of the weight threshold. The non-outlier is greater than or equal to the lower limit of the weight threshold and less than or equal to the upper limit of the weight threshold. The outlier matrix is ​​composed of outliers, and the non-outlier matrix is ​​composed of non-outliers.

[0070] In an embodiment of the present application, the processor may receive the weight matrix of the target neural network through a display interface provided thereon. Optionally, after receiving the weight matrix acquisition instruction, the processor acquires the weight matrix of the target neural network from the path indicated by the weight matrix acquisition instruction. After the processor acquires the weight matrix of the target neural network, each weight value in the weight matrix may be screened first to screen out outliers and non-outliers, and then an outlier matrix is ​​formed by the outliers, and a non-outlier matrix is ​​formed by the non-outliers. Optionally, the processor may first determine the splitting principle according to a preset weight range, specifically, the splitting principle includes that the outlier is less than the lower limit of the weight threshold or greater than the upper limit of the weight threshold, the non-outlier is greater than or equal to the lower limit of the weight threshold and less than or equal to the upper limit of the weight threshold, and then based on the splitting principle, the weight matrix of the target neural network is split using a matrix splitting algorithm to obtain an outlier matrix and a non-outlier matrix.

[0071] S202, performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix.

[0072] The first precision is less than the second precision. The first precision and the second precision are determined by the precision of the memristor device. The first quantization matrix is ​​composed of quantized non-outlier values, and the second quantization matrix is ​​composed of quantized outlier values.

[0073] In an embodiment of the present application, the processor may determine in advance the first precision of quantizing the non-outlier matrix and the second precision of quantizing the outlier matrix according to the precision of the memristor device. Specifically, the first precision and the second precision may be determined by the precision of one or more memristor devices. After the processor obtains the outlier matrix and the non-outlier matrix based on the above steps, the non-outlier matrix may be quantized with the first precision to obtain the first quantization matrix, and the outlier matrix may be quantized with the second precision to obtain the second quantization matrix. Specifically, a first quantization function may be constructed with the first precision, and a second quantization function may be constructed with the second precision, and then each non-outlier in the non-outlier matrix may be input into the first quantization function for quantization to obtain the first quantization matrix, and each outlier in the outlier matrix may be input into the second quantization function for quantization to obtain the second quantization matrix.

[0074] S203, mapping the first quantization matrix to a main array in the memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array.

[0075] Wherein, the memristor array is arranged on the processor.

[0076] In an embodiment of the present application, after the processor obtains the first quantization matrix and the second quantization matrix based on the above steps, the first quantization value in the first quantization matrix can be deployed to the main array in the memristor array, and the second quantization value in the second quantization matrix can be deployed to the sparse computing array in the memristor array, so as to obtain the mapped memristor array. Optionally, the first quantization value can be mapped to a single device in the main array, or the first quantization value can be mapped to multiple devices in the main array to obtain the mapped main array. Optionally, the second quantization value can be mapped to the memristor in a single sparse computing unit in the main array, or the second quantization value can be mapped to the memristor in multiple sparse computing units in the main array to obtain a sparse computing array.

[0077] S204, driving the mapped memristor array to infer the neural network to obtain an inference result.

[0078] The mapped memristor array includes a mapped main array and a mapped sparse computing array. The inference result is the result of inference calculation on the weight matrix of the target neural network, and the inference result is determined by the inference result of the main array and the inference result of the sparse computing array.

[0079] In an embodiment of the present application, after the processor obtains the mapped memristor array based on the above steps, it can obtain input data, then convert the input data into a drive signal, and input the drive signal to the mapped memristor array vector memristor array, and then drive the mapped memristor array to reason about the neural network to obtain a reasoning result. Specifically, the drive signal can be input to the mapped main array, reasoning is performed through the main array to obtain a reasoning result of the main array, and the drive signal can be input to the sparse computing array, reasoning is performed through the sparse computing array to obtain a reasoning result of the sparse computing array, and then the reasoning result of the memristor array is determined according to the reasoning result of the main array and the reasoning result of the sparse computing array. Specifically, the reasoning result of the main array and the reasoning result of the sparse computing array can be added to obtain the reasoning result of the memristor array.

[0080] The neural network reasoning method provided in the embodiment of the present application is to split the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix, then perform a quantization process on the non-outlier matrix with a first precision to obtain a first quantization matrix, and perform a quantization process on the outlier matrix with a second precision to obtain a second quantization matrix, and then map the first quantization matrix to the main array in the memristor array, and map the second quantization matrix to the sparse computing array in the memristor array, and finally drive the mapped memristor array to reason the neural network to obtain a reasoning result. In the above method, by using the main array to process the non-outlier matrix, and using the sparse computing array to process the outlier matrix, the outliers and non-outliers are isolated and processed separately in the subsequent mapping and reasoning process, and the outlier matrix and the non-outlier matrix are quantized with different precisions. For non-outliers, the low-precision quantization range can be fully utilized to improve the reasoning efficiency. For outliers, higher-precision quantization can be performed to improve the reasoning accuracy. On the other hand, in the process of on-chip reasoning, the sparse computing array uses multiple devices to characterize a weight, which equivalently improves the calculation accuracy. Therefore, this approach can effectively solve the on-chip inference accuracy and performance issues caused by outliers.

[0081] In one embodiment, a specific implementation method for splitting the weight matrix of the target neural network is also provided, such as Figure 3 As shown, the above step S201 of "splitting the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix" includes:

[0082] S301, determining the weight values ​​in the weight matrix that are within a preset weight range as first weight values, and determining the weight values ​​in the weight matrix that are not within the preset weight range as second weight values.

[0083] The first weight value is an outlier, and the second weight value is a non-outlier.

[0084] In an embodiment of the present application, the processor can determine a preset weight range in advance based on actual needs and conditions, and then after obtaining the weight matrix of the target neural network, it can determine whether each weight value in the weight matrix is ​​within the preset weight range, determine the weight value within the preset weight range as the first weight value, and determine the weight value in the weight matrix that is not within the preset weight range as the second weight value.

[0085] S302: Split the weight matrix according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix.

[0086] In an embodiment of the present application, after the processor obtains the first weight value and the second weight value based on the above steps, it can use a matrix splitting algorithm based on the splitting principle to split the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix.

[0087] Specifically, Figure 4 As shown, the above step S302 of "splitting the weight matrix according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix" includes:

[0088] S3021, construct a first initial matrix and a second initial matrix.

[0089] The scale of the first initial matrix is ​​the same as the scale of the weight matrix, and the scale of the second initial matrix is ​​the same as the scale of the weight matrix.

[0090] In an embodiment of the present application, after acquiring the weight matrix, the processor can construct a first initial matrix and a second initial matrix of the same scale as the weight matrix. Specifically, two blank matrices of the same scale as the weight matrix can be constructed as the first initial matrix and the second initial matrix. Optionally, two weight matrices can be copied and then used as the first initial matrix and the second initial matrix. Optionally, the two copied weight matrices can also be cleared or initialized to obtain the first initial matrix and the second initial matrix.

[0091] S3022, migrating the first weight value in the weight matrix to the corresponding position of the first initial matrix, and setting the weight values ​​at other positions in the first initial matrix to first preset values, to generate an outlier matrix.

[0092] The first preset value may be zero.

[0093] In an embodiment of the present application, after the processor obtains the first initial matrix based on the above steps, it can migrate the first weight value in the weight matrix to the corresponding position of the first initial matrix, and set the weight values ​​at other positions in the first initial matrix to the first preset value to generate an outlier matrix.

[0094] S3023, migrating the second weight values ​​in the weight matrix to corresponding positions in the second initial matrix, and setting the weight values ​​at other positions in the second initial matrix to second preset values, to generate a non-outlier matrix.

[0095] The second preset value may be zero.

[0096] In an embodiment of the present application, after the processor obtains the second initial matrix based on the above steps, the second weight value in the weight matrix can be migrated to the corresponding position of the second initial matrix, and the weight values ​​at other positions in the second initial matrix can be set to the second preset value to generate a non-outlier matrix.

[0097] In one embodiment, a specific implementation method of neural network reasoning is also provided, such as Figure 5 As shown, the above step S203 of “mapping the first quantization matrix to the main array in the memristor array, and mapping the second quantization matrix to the sparse computing array in the memristor array” includes:

[0098] S401 , setting the value of each first memristor in the main array to each quantized value in the first quantization matrix.

[0099] Among them, Figure 6 As shown, the memristor array includes a main array and a sparse computing array, the main array includes a plurality of first memristors, each of which is cross-connected, the sparse computing array includes a plurality of sparse computing units, each of which includes a plurality of second memristors, and the plurality of second memristors in each sparse computing unit are connected in parallel. The main array also includes an input terminal 1 and an output terminal 1, and the sparse computing array also includes an output terminal 2 and an output terminal 2. Each first memristor in the main array corresponds to a quantized value, and each sparse computing unit in the sparse computing array corresponds to a quantized value.

[0100] In the embodiment of the present application, after obtaining the first quantization matrix, the processor can set the value of each first memristor in the main array to each quantization value in the first quantization matrix. Optionally, the value of each first memristor can be adjusted according to each quantization value in the first quantization matrix.

[0101] S402: Set the value of each sparse computing unit in the sparse computing array to each quantized value in the second quantization matrix.

[0102] In the embodiment of the present application, after obtaining the second quantization matrix, the processor can set the value of each sparse computing unit in the sparse computing array to each quantization value in the second quantization matrix. Optionally, the value of each second memristor in each sparse computing unit can be adjusted according to each quantization value in the second quantization matrix, so that the value of the sparse computing unit corresponds to the quantization value in the second quantization matrix.

[0103] In one embodiment, Figure 7 As shown, the above neural network reasoning method also includes:

[0104] S501, determining whether the size of the non-outlier matrix exceeds the size of the main array.

[0105] In the embodiment of the present application, after obtaining the outlier matrix and the non-outlier matrix, the processor may determine whether the size of the non-outlier matrix exceeds the size of the main array.

[0106] S502: If the size of the non-outlier matrix does not exceed the size of the main array, return to the step of performing a first precision quantization process on the non-outlier matrix.

[0107] In the embodiment of the present application, if the processor determines that the size of the non-outlier matrix does not exceed the size of the main array, then a step of performing a first precision quantization process on the non-outlier matrix may be executed, that is, step S202 is executed.

[0108] S503, if the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and any segmented outlier matrix, return to the step of performing first precision quantization processing on the non-outlier matrix.

[0109] In the embodiment of the present application, if the processor determines that the scale of the non-outlier matrix exceeds the scale of the main array, the non-outlier matrix and the outlier matrix can be segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and segmented outlier matrix, the step of performing a first precision quantization process on the non-outlier matrix can be performed, that is, the step of performing S202. Specifically, the non-outliers can be segmented to obtain multiple segmented non-outlier matrices, and then the outliers in the outlier matrix can be evenly distributed according to the number of the segmented non-outlier matrices to obtain multiple segmented outlier matrices.

[0110] In one embodiment, a specific implementation method of driving the mapped memristor array to perform reasoning on a neural network is also provided, such as Figure 8 As shown, the "driving the mapped memristor array to infer the neural network to obtain an inference result" in the above step S204 includes:

[0111] S601, obtaining input data corresponding to the target neural network, and converting the input data into a voltage signal.

[0112] In an embodiment of the present application, after obtaining the mapped memristor array, the processor can obtain input data corresponding to the target neural network and then convert the input data into a voltage signal.

[0113] S602, inputting the voltage signal to the mapped memristor array, and driving the mapped memristor array to perform inference to obtain an inference result.

[0114] The main array further includes an analog-to-digital converter 1, and the sparse computing array further includes an analog-to-digital converter 2.

[0115] In an embodiment of the present application, after the processor converts the input data into a voltage signal based on the above steps, the voltage signal can be input into the mapped memristor array, and the mapped memristor array can be driven to perform reasoning to obtain a reasoning result. Specifically, the voltage signal can be input into the main array through input terminal 1 for excitation, and the main array and the voltage signal are multiplied to obtain an output signal, and the output signal is converted by analog-to-digital converter 1, and then collected through output terminal 1 to obtain the reasoning result of the main array, and the voltage signal is input into the sparse computing array through input terminal 2 for excitation, and the sparse computing array and the voltage signal are multiplied to obtain an output signal, and the output signal is converted by analog-to-digital converter 2, and then collected through output terminal 2 to obtain the reasoning result of the sparse computing array, and finally the reasoning result of the main array and the reasoning result of the sparse computing array are added to obtain the reasoning result of the memristor array. This process can be represented by the following relationship:

[0116]

[0117] in, As the inference result, is the weight matrix, is the voltage signal, is the first quantization matrix, is the second quantization matrix.

[0118] In one embodiment, a neural network reasoning system is also provided, the reasoning system includes a processor, at least one memristor array and a peripheral circuit, the processor is connected to the input and output of the memristor array through the peripheral circuit. The memristor array includes a main array and a sparse computing array, the main array includes a plurality of first memristors, each first memristor is cross-connected, the sparse computing array includes a plurality of sparse computing units, each sparse computing unit includes a plurality of second memristors, and the plurality of second memristors in each sparse computing unit are connected in parallel. The main array also includes an input terminal 1, an output terminal 1 and an analog-to-digital converter, and the sparse computing array also includes an output terminal 2, an output terminal 2 and an analog-to-digital converter. The processor is used to execute the steps of the neural network reasoning method of any of the above embodiments.

[0119] The above-mentioned main array, which is a cross array composed of several rows and columns of first memristors, is responsible for performing matrix multiplication of the non-outlier part. Each weight value in the weight matrix is ​​quantized and mapped to a single memristor at the corresponding position. Therefore, the accuracy of a single memristor is the quantization accuracy required for the weight. The accuracy of a single memristor is k bits. After the input data of the neural network is converted into a voltage signal, it is input into the array from input terminal 1. The output signal is collected by output terminal 1 after passing through the analog-to-digital converter.

[0120] The above-mentioned sparse computing array is responsible for performing matrix multiplication of the outlier part, and its characteristics are sparsity and high precision. The sparse computing array contains several rows and columns of memristor arrays, and the connection method is different from the main array. The memristors in each column can be regarded as connected in parallel, that is, the bottom electrodes of all devices in each column are connected to the same wire, and the top electrodes are connected to another wire, and the two wires correspond to the input and output ends respectively. The sparse computing array contains several sparse computing units. Each sparse computing unit contains a memristor array with N rows and 1 column, that is, it contains N memristor devices. Each sparse computing unit can perform single-digit multiplication, so multiple sparse computing units can perform sparse matrix multiplication. In the process of performing multiplication, the outlier weight is quantized and mapped to the N devices of a single sparse computing unit, that is, the N devices jointly represent a weight, so it can be equivalently regarded as having N×k bits of precision, thereby enabling high-precision calculation. After the input data is converted into a voltage signal, it is input into the array from input terminal 2, and the output signal is collected from output terminal 2 after passing through the analog-to-digital converter. The above main array and sparse computing array are the designs of a memristor array in the inference system. The inference system contains multiple arrays of the same specifications, which are connected to peripheral circuits to complete communication, control and data transmission.

[0121] The reasoning system described in the embodiment of the present application includes a main array and a sparse computing array, which realizes the separate calculation of outliers and non-outliers. On the one hand, during the deployment process, outliers and non-outliers are quantized separately. For non-outliers, the low-precision quantization range can be fully utilized, which improves the accuracy of the quantization of the non-outlier part; for outliers, higher-precision quantization can be performed. On the other hand, during the on-chip reasoning process, the sparse computing array uses multiple devices to characterize a weight, which equivalently improves the calculation accuracy. Therefore, this method can effectively solve the on-chip reasoning accuracy and performance problems caused by outliers.

[0122] Based on all the above embodiments, a neural network reasoning method is also provided. Fig. 9 As shown, the method includes:

[0123] S701, determining a weight value in the weight matrix that is smaller than a preset weight range as a first weight value, and determining a weight value in the weight matrix that is not smaller than the preset weight range as a second weight value.

[0124] S702, constructing a first initial matrix and a second initial matrix, wherein the scale of the first initial matrix is ​​the same as the scale of the weight matrix, and the scale of the second initial matrix is ​​the same as the scale of the weight matrix.

[0125] S703: Migrate the first weight value in the weight matrix to the corresponding position of the first initial matrix, and set the weight values ​​at other positions in the first initial matrix to first preset values, so as to generate an outlier matrix.

[0126] S704, migrating the second weight values ​​in the weight matrix to corresponding positions in the second initial matrix, and setting the weight values ​​at other positions in the second initial matrix to second preset values, to generate a non-outlier matrix.

[0127] S705, determine whether the size of the non-outlier matrix exceeds the size of the main array.

[0128] S706, if the size of the non-outlier matrix does not exceed the size of the main array, execute step S708.

[0129] S707, if the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and any segmented outlier matrix, execute step S708.

[0130] S708, performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix.

[0131] S709: Set the value of each first memristor in the main array to each quantized value in the first quantization matrix.

[0132] S710: Set the value of each sparse computing unit in the sparse computing array to each quantized value in the second quantization matrix.

[0133] S711, driving the mapped memristor array to infer the neural network to obtain an inference result.

[0134] In the present application embodiment, Fig.10 As shown in the figure, the weight matrix is ​​divided into two matrices: the outlier matrix and the non-outlier matrix. The non-outlier matrix is ​​quantized with low precision and deployed on the main array, and the outlier matrix is ​​quantized with high precision and deployed on the sparse computing array. After the two arrays complete the calculation, the results are merged. Specifically, the process is as follows: ① Split the weight matrix into an outlier matrix with the same rows and columns and a non-outlier matrix with the same rows and columns, that is, . The method of setting a preset threshold range can be used for screening, that is, as long as the absolute value of the weight exceeds the preset threshold range, it will be judged as an outlier and separated from the original weight matrix. The choice of threshold will affect the number of selected outliers, which can be set according to actual needs and conditions. After the separation is completed, in the non-outlier matrix, the position of the original outlier weight becomes 0; in the outlier matrix, except for the selected outliers, the remaining positions are 0, so the outlier matrix is ​​a sparse matrix. ② Quantize the outlier matrix and the non-outlier matrix with different precisions. The quantization accuracy of the non-outlier matches the accuracy of a single memristor device on the main array, so the quantization bit width is bits, which is low-precision quantization. The quantization accuracy of outliers is the same as that of coefficient calculation units. The equivalent precision of the memristor devices is matched, so the quantization bit width is bits, which is high-precision quantization. ③ Deploy the two matrices to the corresponding arrays. The non-outlier matrix is ​​deployed to the main array. When the matrix size exceeds the scale of a single main array, the matrix can be split and deployed to multiple main arrays. Here, it is assumed that a total of The inference calculation method is the same as that of the traditional memristor array. After the input vector or matrix is ​​converted into a voltage signal, it is stimulated by input terminal 1 and the calculation result is collected by output terminal 1. The outlier matrix is ​​deployed on the sparse computing array, and the outliers are evenly distributed to the above In the process of sparse matrix multiplication, the peripheral circuit sends the data of the corresponding position in the input vector or matrix to each array and stimulates it through input terminal 2. After the calculation is completed, it is collected by output terminal 2 and further sent to the peripheral circuit. In both deployment processes, peripheral circuits are required for control and data transmission. ④ After obtaining the calculation results of the outlier matrix and the non-outlier matrix, the peripheral circuit merges them, that is, , for subsequent processing and operation.

[0135] The methods described in the above steps are all described in the above embodiments. Please refer to the above description for details and will not be repeated here.

[0136] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0137] Based on the same inventive concept, the embodiment of the present application also provides a neural network reasoning device for implementing the neural network reasoning method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of one or more neural network reasoning devices provided below can refer to the limitations of the neural network reasoning method above, and will not be repeated here.

[0138] In one embodiment, a neural network reasoning device is provided, comprising:

[0139] A splitting module is used to split the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix;

[0140] A quantization module, configured to perform a quantization process on the non-outlier matrix with a first precision to obtain a first quantization matrix, and to perform a quantization process on the outlier matrix with a second precision to obtain a second quantization matrix;

[0141] A mapping module, used for mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is arranged in a processor;

[0142] The inference module is used to drive the mapped memristor array to infer the neural network and obtain the inference result.

[0143] In one embodiment, the splitting module includes:

[0144] The determining unit is used to determine the weight values ​​in the weight matrix that are within the preset weight range as the first weight values, and to determine the weight values ​​in the weight matrix that are not within the preset weight range as the second weight values.

[0145] The splitting unit is used to split the weight matrix according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix.

[0146] In one embodiment, the splitting unit comprises:

[0147] The construction subunit is used to construct a first initial matrix and a second initial matrix; the scale of the first initial matrix is ​​the same as the scale of the weight matrix; the scale of the second initial matrix is ​​the same as the scale of the weight matrix.

[0148] The first generating subunit is used to migrate the first weight value in the weight matrix to the corresponding position of the first initial matrix, and set the weight values ​​at other positions in the first initial matrix to the first preset value, so as to generate an outlier matrix.

[0149] The second generating subunit is used to migrate the second weight value in the weight matrix to the corresponding position of the second initial matrix, and set the weight value at other positions in the second initial matrix to the second preset value to generate a non-outlier matrix.

[0150] In one embodiment, the mapping module includes:

[0151] The first mapping unit is used to set the value of each first memristor in the main array to each quantization value in the first quantization matrix.

[0152] The second mapping unit is used to set the value of each sparse computing unit in the sparse computing array to each quantization value in the second quantization matrix.

[0153] In one embodiment, the neural network reasoning device further includes:

[0154] A determination module is used to determine whether the size of the non-outlier matrix exceeds the size of the main array.

[0155] The execution module is used for returning to the step of performing a first precision quantization process on the non-outlier matrix if the size of the non-outlier matrix does not exceed the size of the main array.

[0156] A splitting module is used for splitting the non-outlier matrix and the outlier matrix if the size of the non-outlier matrix exceeds the size of the main array to obtain multiple split non-outlier matrices and multiple split outlier matrices. For any of the split non-outlier matrix and the split outlier matrix, returning to the step of performing first precision quantization processing on the non-outlier matrix.

[0157] In one embodiment, the driving module includes:

[0158] The acquisition unit is used to acquire input data corresponding to the target neural network and convert the input data into a voltage signal.

[0159] The driving unit is used to input the voltage signal to the mapped memristor array, and drive the mapped memristor array to perform reasoning to obtain the reasoning result.

[0160] Each module in the above-mentioned neural network reasoning device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0161] In one embodiment, a computer device is provided. The computer device may be a terminal or a server. The internal structure diagram thereof may be as follows: Fig.11 As shown, the computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be realized through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a neural network reasoning method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0162] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0163] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0164] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0165] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0166] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0167] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0168] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0169] Determine the weight values ​​in the weight matrix that are within the preset weight range as first weight values, and determine the weight values ​​in the weight matrix that are not within the preset weight range as second weight values;

[0170] The weight matrix is ​​split according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix.

[0171] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0172] Constructing a first initial matrix and a second initial matrix; the scale of the first initial matrix is ​​the same as the scale of the weight matrix; the scale of the second initial matrix is ​​the same as the scale of the weight matrix;

[0173] Migrating the first weight value in the weight matrix to the corresponding position of the first initial matrix, and setting the weight values ​​at other positions in the first initial matrix to the first preset values, to generate an outlier matrix;

[0174] The second weight values ​​in the weight matrix are migrated to corresponding positions in the second initial matrix, and the weight values ​​at other positions in the second initial matrix are set to second preset values ​​to generate a non-outlier matrix.

[0175] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0176] Setting the value of each first memristor in the main array to each quantized value in the first quantization matrix;

[0177] The value of each sparse computing unit in the sparse computing array is correspondingly set to each quantized value in the second quantization matrix.

[0178] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0179] Determine if the size of the non-outlier matrix exceeds the size of the main array;

[0180] If the size of the non-outlier matrix does not exceed the size of the main array, returning to the step of performing a first precision quantization process on the non-outlier matrix;

[0181] If the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and any segmented outlier matrix, return to the step of performing first precision quantization processing on the non-outlier matrix.

[0182] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0183] Obtain input data corresponding to the target neural network and convert the input data into a voltage signal;

[0184] The voltage signal is input into the mapped memristor array, and the mapped memristor array is driven to perform reasoning to obtain the reasoning result.

[0185] The above embodiment provides a computer device, whose implementation principle and technical effect are similar to those of the above method embodiment, and will not be repeated here.

[0186] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0187] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0188] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0189] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0190] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0191] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0192] Determine the weight values ​​in the weight matrix that are within the preset weight range as first weight values, and determine the weight values ​​in the weight matrix that are not within the preset weight range as second weight values;

[0193] The weight matrix is ​​split according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix.

[0194] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0195] Constructing a first initial matrix and a second initial matrix; the scale of the first initial matrix is ​​the same as the scale of the weight matrix; the scale of the second initial matrix is ​​the same as the scale of the weight matrix;

[0196] Migrating the first weight value in the weight matrix to the corresponding position of the first initial matrix, and setting the weight values ​​at other positions in the first initial matrix to the first preset values, to generate an outlier matrix;

[0197] The second weight values ​​in the weight matrix are migrated to corresponding positions in the second initial matrix, and the weight values ​​at other positions in the second initial matrix are set to second preset values ​​to generate a non-outlier matrix.

[0198] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0199] Setting the value of each first memristor in the main array to each quantized value in the first quantization matrix;

[0200] The value of each sparse computing unit in the sparse computing array is correspondingly set to each quantized value in the second quantization matrix.

[0201] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0202] Determine if the size of the non-outlier matrix exceeds the size of the main array;

[0203] If the size of the non-outlier matrix does not exceed the size of the main array, returning to the step of performing a first precision quantization process on the non-outlier matrix;

[0204] If the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and any segmented outlier matrix, return to the step of performing first precision quantization processing on the non-outlier matrix.

[0205] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0206] Obtain input data corresponding to the target neural network and convert the input data into a voltage signal;

[0207] The voltage signal is input into the mapped memristor array, and the mapped memristor array is driven to perform reasoning to obtain the reasoning result.

[0208] The above embodiment provides a computer-readable storage medium, whose implementation principle and technical effect are similar to those of the above method embodiment, and will not be repeated here.

[0209] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0210] Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix;

[0211] Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix;

[0212] Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in a processor;

[0213] The mapped memristor array is driven to infer the neural network and obtain the inference result.

[0214] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0215] Determine the weight values ​​in the weight matrix that are within the preset weight range as first weight values, and determine the weight values ​​in the weight matrix that are not within the preset weight range as second weight values;

[0216] The weight matrix is ​​split according to the first weight value and the second weight value to obtain an outlier matrix and a non-outlier matrix.

[0217] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0218] Constructing a first initial matrix and a second initial matrix; the scale of the first initial matrix is ​​the same as the scale of the weight matrix; the scale of the second initial matrix is ​​the same as the scale of the weight matrix;

[0219] Migrating the first weight value in the weight matrix to the corresponding position of the first initial matrix, and setting the weight values ​​at other positions in the first initial matrix to the first preset values, to generate an outlier matrix;

[0220] The second weight values ​​in the weight matrix are migrated to corresponding positions in the second initial matrix, and the weight values ​​at other positions in the second initial matrix are set to second preset values ​​to generate a non-outlier matrix.

[0221] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0222] Setting the value of each first memristor in the main array to each quantized value in the first quantization matrix;

[0223] The value of each sparse computing unit in the sparse computing array is correspondingly set to each quantized value in the second quantization matrix.

[0224] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0225] Determine if the size of the non-outlier matrix exceeds the size of the main array;

[0226] If the size of the non-outlier matrix does not exceed the size of the main array, returning to the step of performing a first precision quantization process on the non-outlier matrix;

[0227] If the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain multiple segmented non-outlier matrices and multiple segmented outlier matrices. For any segmented non-outlier matrix and any segmented outlier matrix, return to the step of performing first precision quantization processing on the non-outlier matrix.

[0228] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0229] Obtain input data corresponding to the target neural network and convert the input data into a voltage signal;

[0230] The voltage signal is input into the mapped memristor array, and the mapped memristor array is driven to perform reasoning to obtain the reasoning result.

[0231] The above embodiment provides a computer program product, whose implementation principle and technical effect are similar to those of the above method embodiment, and will not be repeated here.

[0232] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0233] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0234] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A neural network reasoning method, characterized in that: Applied to a processor, the method comprises: Split the weight matrix of the target neural network to obtain the outlier matrix and the non-outlier matrix; Performing a quantization process with a first precision on the non-outlier matrix to obtain a first quantization matrix, and performing a quantization process with a second precision on the outlier matrix to obtain a second quantization matrix; Mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is disposed in the processor; The mapped memristor array is driven to infer the neural network and obtain the inference result.

2. The method according to claim 1, characterized in that The weight matrix of the target neural network is split to obtain an outlier matrix and a non-outlier matrix, including: Determine the weight values ​​in the weight matrix that are less than a preset weight range as first weight values, and determine the weight values ​​in the weight matrix that are not less than the preset weight range as second weight values; The weight matrix is ​​split according to the first weight value and the second weight value to obtain the outlier matrix and the non-outlier matrix.

3. The method according to claim 2, characterized in that The step of splitting the weight matrix according to the first weight value and the second weight value to obtain the outlier matrix and the non-outlier matrix includes: Constructing a first initial matrix and a second initial matrix; the scale of the first initial matrix is ​​the same as the scale of the weight matrix; the scale of the second initial matrix is ​​the same as the scale of the weight matrix; Migrating the first weight value in the weight matrix to the corresponding position of the first initial matrix, and setting the weight values ​​at other positions in the first initial matrix to first preset values, to generate the outlier matrix; The second weight values ​​in the weight matrix are migrated to corresponding positions in the second initial matrix, and the weight values ​​at other positions in the second initial matrix are set to second preset values ​​to generate the non-outlier matrix.

4. The method according to claim 1, characterized in that: The main array includes a plurality of first memristors, the sparse computing array includes a plurality of sparse computing units, each sparse computing unit includes a plurality of second memristors, and mapping the first quantization matrix to the main array in the memristor array, and mapping the second quantization matrix to the sparse computing array in the memristor array, comprises: Setting the value of each first memristor in the main array to each quantization value in the first quantization matrix; The value of each sparse computing unit in the sparse computing array is correspondingly set to each quantization value in the second quantization matrix.

5. The method according to claim 1, characterized in that The method further comprises: determining whether the size of the non-outlier matrix exceeds the size of the main array; If the size of the non-outlier matrix does not exceed the size of the main array, returning to the step of performing the quantization process of the non-outlier matrix with the first precision; If the size of the non-outlier matrix exceeds the size of the main array, the non-outlier matrix and the outlier matrix are segmented to obtain a plurality of segmented non-outlier matrices and a plurality of segmented outlier matrices. For any of the segmented non-outlier matrix and the segmented outlier matrix, the step of performing the first precision quantization process on the non-outlier matrix is ​​returned.

6. The method according to any one of claims 1 to 5, characterized in that: The memristor array after the driving mapping performs reasoning on the neural network to obtain reasoning results, including: Acquire input data corresponding to the target neural network, and convert the input data into a voltage signal; The voltage signal is input to the mapped memristor array, and the mapped memristor array is driven to perform inference to obtain an inference result.

7. A neural network reasoning system, characterized in that: The inference system comprises a processor, at least one memristor array and a peripheral circuit; the processor is connected to an input terminal and an output terminal of the memristor array through the peripheral circuit; The memristor array comprises a main array and a sparse computing array; the main array comprises a plurality of first memristors, each of which is cross-connected, the sparse computing array comprises a plurality of sparse computing units, each of which comprises a plurality of second memristors, and the plurality of second memristors in each of the sparse computing units are connected in parallel; The processor is used to execute the steps of the neural network reasoning method according to any one of claims 1 to 6.

8. A neural network reasoning device, characterized in that: The device comprises: A splitting module is used to split the weight matrix of the target neural network to obtain an outlier matrix and a non-outlier matrix; A quantization module, configured to perform a quantization process on the non-outlier matrix with a first precision to obtain a first quantization matrix, and to perform a quantization process on the outlier matrix with a second precision to obtain a second quantization matrix; A mapping module, used for mapping the first quantization matrix to a main array in a memristor array, and mapping the second quantization matrix to a sparse computing array in the memristor array; the memristor array is arranged in the processor; The inference module is used to drive the mapped memristor array to infer the neural network and obtain the inference result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.