Error correction method of neural network processor and related device
By adopting a mixed-grained fault recovery solution in NPUs and using fine-grained and coarse-grained error correction mechanisms, the problems of waste resources and lack of reconfigurability of existing NPU error correction technology are solved, and efficient fault isolation and computing reliability are achieved.
Patent Information
- Application Number
- CN202510205146.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing NPU error correction technology has wasted resources and lacks reconfigurability. The traditional three-mode redundancy technology has caused too much power consumption and area overhead in a multi-processing unit environment. The shared redundant PE solution cannot be effectively utilized in multiple failures.
The mixed-grained fault recovery scheme is adopted to expand the adjacency range through adjacent PE status query and cross-column protection mechanism in the fine-grained error correction stage; in the coarse-grained error correction stage, mark the column where the faulted PE is located as an unavailable thread, generate a routing table and start on-chip redirection to achieve error correction function.
Without adding backup PE, the error correction function is implemented to avoid resource waste, improve the success rate of fault tolerance, and take into account processing efficiency and computing reliability.
Smart Images

Figure CN120144360A_ABST
Abstract
Description
Technical Field
[0001] This application proposes an error correction method and related device for a neural network processor, belonging to the technical field of neural networks. Background Art
[0002] A neural network processor (Neuron Processing Unit, NPU) is an application specific integrated circuit (ASIC) specifically designed to accelerate the inference or training of neural network models. The purpose of the NPU fault tolerance technology is to ensure that the chip can still work normally or minimize the impact of faults on the system performance and functions in the event of hardware faults.
[0003] The NPU fault tolerance technology includes fault detection and error correction. The error correction technology corrects the detected faults after detecting hardware errors, which can ensure the normal inference operation of the NPU. Currently, the most common error correction technology for protecting functional units is to adopt an implementation scheme of adding redundant functional units, such as the direct replacement spare part technology proposed by Japanese scholars Takanami et al. Although this method can play a certain degree of protective role, it wastes hardware resources.
[0004] The disadvantages of the prior art are as follows: 1. Although the traditional triple modular redundancy technology is already very mature, due to the inherent multi-processing unit characteristics of the NPU, if this technology is directly used for NPU error correction, it will cause power consumption and area overhead three times that of the PE array, which is not worth the candle.
[0005] 2. To solve the first technical disadvantage, some scholars have proposed a scheme of sharing a redundant PE in the cross rows and columns of the PE array. However, when there are more than one faulty PE in the shared rows and columns, the chip cannot utilize the redundant PEs that have not been activated in other shared rows and columns, resulting in new resource waste.
[0006] 3. Existing NPU error correction technologies mostly adopt a single granularity implementation scheme and do not have reconfigurability. Summary of the Invention
[0007] In view of the limitations of the prior art, this application proposes an error correction method and related device for a neural network processor, adopting a hybrid granularity fault recovery scheme, which can achieve the error correction function without adding any spare PEs. The fine-grained scheme provides more flexibility and does not waste PEs, while the coarse-grained scheme ensures that the inference accuracy does not decrease as the failure rate increases.
[0008] To solve the above technical problems, the technical solutions adopted in this application are as follows: In a first aspect, the present application provides an error correction method for a neural network processor, including: According to the faulty PE table obtained by error detection, the faulty PE table includes a two-dimensional list of fault occurrence and location information, Locate the newly detected faulty PE according to the faulty PE table and perform multi-level fault tolerance decisions: i. Fine-grained error correction stage: Query the status of adjacent PEs of the faulty PE. When there are healthy adjacent PEs, generate a replication query table and perform adjacent replication operations, and extend the adjacent range through a cross-column protection mechanism; when all four adjacent PEs on the left and right are faulty, trigger coarse-grained error correction; ii. Coarse-grained error correction stage: When fine-grained error correction is not feasible, mark the column where the faulty PE is located as an unavailable thread, generate a routing table and start on-chip redirection; Select pipeline submission or a submission mechanism according to the current error correction mode, and output the calculation result.
[0009] As a further improvement of the present application, in the hidden layer of the neural network processor, in the MWMT architecture, each row of the systolic array adopts a parallel computing structure of row broadcast weights / column pipelined activation, and the PE unit independently completes the multiply-accumulate operation; in the pooling layer, it has the characteristic of adjacent activation value substitution; Add a 1-bit fault flag signal and a 2-bit replication signal inside the PE; among them, the fault flag signal includes non-faulty PEs and faulty PEs, and the replication signal supports four-level replication directions of the left / right one in the adjacent column and the left / right two across columns.
[0010] As a further improvement of the present application, the fine-grained error correction stage includes: Adjacent replication priority determination: First detect the status of the left and right two adjacent PEs in the same row of the faulty PE table; as long as one of the adjacent PEs is in a healthy state, pair these two PEs and fill the pairing information into a two-dimensional replication query table; if both the left and right adjacent PEs are faulty and isolated, then enable the cross-column protection mechanism; Cross-column protection mechanism: Query whether the status of the left and right two adjacent PEs across one column from the faulty PE is healthy. If one is healthy, form a pair and fill the information into the replication query table.
[0011] As a further improvement of the present application, the cross-column protection mechanism satisfies: The cross-column distance is limited to a single-column span, and the cross-column PE and the original faulty PE are in the same row; Cross-column protection is only triggered when all adjacent PEs are faulty, and the cross-column PE with the highest correlation with the faulty PE's calculation task is preferentially selected.
[0012] As a further improvement of the present application, the coarse-grained error correction stage includes: Mark unavailable threads so that the entire column containing the faulty PE is marked as a boolean array; Use head and tail double pointers to scan the boolean array to dynamically generate a routing table, and the routing table completes the reconstruction of the mapping relationship from the thread buffer to the target thread within a preset period; Introduce a priority selector to route the activation value from the thread buffer to the corresponding correct thread through the input routing table; when the faulty thread is isolated, the activation value in the original corresponding thread buffer will be distributed to the first healthy thread after it for calculation; As a further improvement of the present application, the use of head and tail double pointers to scan the boolean array to dynamically generate a routing table, and the routing table completes the reconstruction of the mapping relationship from the thread buffer to the target thread within a preset period, includes: Introduce a pair of head and tail double pointers, with the initial value of the head pointer pointing to the 0 address of the routing table and the tail pointer pointing to the 63 address of the routing table; Scan from the 0 address of the faulty thread registration table in each period. Whenever 0 is encountered, the address pointed to by the head pointer is filled with the current thread number, and the pointer is incremented by one; whenever 1 is encountered, the address pointed to by the tail pointer is filled with the current thread number, and it is decremented by one; After a total of preset periods, the update of the routing table is completed; the routing table is represented by a set of register files in hardware, with a depth equal to the number of threads and a width equal to the logarithm of the number of threads.
[0013] As a further improvement of the present application, the routing table generation satisfies: The sequence number of the routing table represents the thread buffer number, and its value represents the target thread number to be routed. The initial state of the routing table is that the sequence number and the target thread number correspond one by one.
[0014] As a further improvement of the present application, the selection of pipeline submission or selection submission mechanism according to the current error correction mode and the output of the calculation result include: When submitting in the fine-grained error correction stage, when the coarse-grained error correction is not triggered, add a replication period to complete the pipeline submission; when the coarse-grained error correction is triggered, use the selection submission mode; each layer of operator only performs a single submission; When submitting in the coarse-grained error correction stage, introduce a pipeline-style selection submission module. The selection submission module is added with a set of multiplexers and one-hot code thread counters for controlling the selection submission, and in the way of selecting a column of threads in each period, the pipeline submission is restored.
[0015] In a third aspect, the present application provides an error correction device for a neural network processor, including: An acquisition unit for acquiring the faulty PE table obtained by error detection, and the faulty PE table includes a two-dimensional list of fault occurrence and location information An execution unit, configured to locate a newly detected faulty PE according to a faulty PE table and perform multi-level fault tolerance decision-making: i. Fine-grained error correction phase: Query the states of adjacent PEs of the faulty PE. When there are healthy adjacent PEs, generate a replication query table and perform an adjacent replication operation, and extend the adjacent range through a cross-column protection mechanism; when all four adjacent PEs on the left and right fail, trigger coarse-grained error correction; ii. Coarse-grained error correction phase: When fine-grained error correction is not feasible, mark the column where the faulty PE is located as an unavailable thread, generate a routing table, and start on-chip redirection; A submission unit, which selects pipeline submission or a submission mechanism according to the current error correction mode and outputs a calculation result.
[0016] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the error correction method of the neural network processor is implemented.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the error correction method of the neural network processor is implemented.
[0018] In a fifth aspect, the present application provides a computer program product, which includes computer instructions that instruct a computer to execute the error correction method of the neural network processor.
[0019] The beneficial effects of the present application compared with the prior art are as follows: In the fine-grained error correction phase of the present application, local fault tolerance is achieved through adjacent PE state query, and the adjacent range is extended by using a cross-column protection mechanism, realizing dynamic migration of computing tasks at the hardware level; in the coarse-grained error correction phase, when local fault tolerance is not feasible, fault tolerance is achieved through column-level disabling and on-chip network redirection, and the data path is dynamically planned by using a routing table; a two-dimensional topology mapping is constructed through a faulty PE table to support real-time update of the processing unit state. The cross-column protection mechanism breaks through the limitations of traditional adjacent PEs and improves the fault tolerance success rate by extending the adjacent range; the hybrid mode of pipeline submission and selective submission takes into account both processing efficiency and computing reliability. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for clearly expressing some embodiments of the technical solutions in the present application for convenience. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0021] Figure 1 Flowchart of an error correction method for a neural network processor provided by the present application; Figure 2 Schematic diagram of generating a copy lookup table provided by the present application; Figure 3 Schematic diagram of a fine-grained error correction function provided by the present application; Figure 4 Flowchart of a thread-level coarse-grained error correction technology based on on-chip routing provided by the present application; Figure 5 Flowchart of generating a routing table by head and tail double pointers provided by the present application; Figure 6 Pipeline-style selection submission module provided by the present application; Figure 7 Graph of delay overhead variation and comparison table under five remaining available threads provided by the present application; Figure 8 Schematic diagram of submission accuracy under five remaining available thread counts provided by the present application; Figure 9 Error correction device for a neural network processor provided by the present application. Detailed implementation manners
[0022] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application. For the step numbers in the following embodiments, they are only set for convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0023] In the description of the present application, unless otherwise clearly defined, words such as "set", "installed", and "connected" should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.
[0024] This application implements an error correction technology for a neural network processor. Based on the method of adjacent replication or on-chip routing, damaged processing elements (PEs) or threads (i.e., containing multiple processing elements working in parallel) are isolated. Without introducing spare PEs, this application successfully implements a fault error correction technology with a hybrid granularity at the thread level and the PE level by only modifying a small amount of hardware logic.
[0025] As Figure 1 shown, the first object of this application is to provide an error correction method for a neural network processor, including: S1. According to the fault PE table obtained by error detection, the fault PE table contains a two-dimensional list of fault occurrence and location information, S2. Locate the newly detected faulty PE according to the fault PE table and execute a multi-level fault tolerance decision: i. Fine-grained error correction stage: Query the status of adjacent PEs of the faulty PE. When there are healthy adjacent PEs, generate a replication query table and perform an adjacent replication operation, and expand the adjacent range through a cross-column protection mechanism; when all four adjacent PEs on the left and right are faulty, trigger coarse-grained error correction; ii. Coarse-grained error correction stage: When fine-grained error correction is not feasible, mark the column where the faulty PE is located as an unavailable thread, generate a routing table, and start on-chip redirection; S3. Select pipeline submission or a selection submission mechanism according to the current error correction mode and output the calculation result.
[0026] According to the fault PE table, this method first locates the newly faulty PE, and then processes it in two stages: fine-grained and coarse-grained error correction. In the fine-grained stage, it checks whether the adjacent PEs are healthy. If possible, it performs a replication operation and expands the adjacent range. If all four adjacent PEs are damaged, it triggers the coarse-grained stage, marks the entire column as unavailable, and starts on-chip redirection. Finally, it selects a submission mechanism according to the error correction mode.
[0027] Among them, the fine-grained error correction stage realizes local fault tolerance through adjacent PE status query, adopts a cross-column protection mechanism to expand the adjacent range, and realizes the dynamic migration of computing tasks at the hardware level; the coarse-grained error correction stage is when local fault tolerance is not feasible, realizes fault tolerance through column-level disabling and on-chip network redirection, and uses a routing table to dynamically plan the data path; constructs a two-dimensional topology mapping through the fault PE table to support real-time update of the processing unit status. The cross-column protection mechanism breaks through the limitations of traditional adjacent PEs and improves the fault tolerance success rate by expanding the adjacent range; the hybrid mode of pipeline submission and selection submission takes into account both processing efficiency and computing reliability.
[0028] This method realizes dynamic fault isolation through hardware-software co-design while maintaining computational continuity, and is particularly suitable for high-reliability neural network acceleration scenarios. Compared with traditional ECC or dual-mode redundancy schemes, it has reduced area overhead and no performance degradation.
[0029] The following will further elaborate on this application in conjunction with the accompanying drawings and specific embodiments.
[0030] This application implements an error correction method for a neural network processor. Based on the method of adjacent replication or on-chip routing, it isolates damaged processing elements (PEs) or threads (i.e., threads containing multiple parallel working processing elements). Without introducing spare PEs, this application successfully realizes fault error correction at the thread level and PE level with a hybrid granularity by only modifying a small amount of hardware logic.
[0031] The implementation of this application is based on the fault detection technology of the MWMT AI chip architecture, and is also adapted to error detection technologies that can obtain the location information of faulty PEs, and adopts a systolic array micro-architecture that outputs a stable data stream. The technology of this application has high portability and can implement a hybrid granularity fault recovery scheme for fine-grained error correction at the PE level and coarse-grained error correction at the thread level. The basic content is as follows.
[0032] Step 1: According to the faulty PE table obtained from error detection (i.e., a two-dimensional list containing fault occurrence and location information), locate the newly detected faulty PE, and then enter the fine-grained fault error correction stage. In this stage, this application will copy the calculation results of the adjacent PEs in the same row of the faulty PE and submit them as the results of the faulty PE once during the submission cycle.
[0033] The specific steps of the PE-level fine-grained fault recovery technical solution are as follows: Since there are a large number of equal values and zero values in the adjacent activations in the hidden layer of a deep neural network, this application can utilize its own characteristics to complete error correction.
[0034] Specifically, on the one hand, in the MWMT architecture, each row of the systolic array shares weight values in a broadcast manner, each column shares activation values in a pipelined manner, and the multiply-accumulate operation is completed inside the PE and does not flow between PEs - this is called an output stable data stream. As described in the previous section, based on this characteristic, this application proposes an adjacent replication error correction scheme. On the other hand, the pooling layer of the deep neural network further enhances the substitutability of adjacent activation values.
[0035] Furthermore, when the fault tolerance state machine obtains the faulty PE table, it will complete the generation of the replication lookup table within 64 cycles, just as Figure 2As shown. When four adjacent PEs on the left and right all fail, a coarse-grained error correction scheme is triggered - however, the probability of this is actually very small. In this application, two signals are added inside the PE, namely a 1-bit fault flag signal and a 2-bit replication signal. Among them, the fault flag signal (0: non-faulty PE, 1: faulty PE), and the replication signal (00: the first on the left, 01: the first on the right, 10: the second on the left, 11: the second on the right).
[0036] In the submission stage, add one cycle to complete the replication, and then start pipeline submission (coarse-grained error correction has not been started) or select submission (coarse-grained error correction has been started). As Figure 3 shown, within the added one cycle, the main state machine will control each PE to judge that the fault flag bit is 1, then query the replication signal, and copy the value of the pointed PE to the submission value register of this PE, replacing the original result as the submission value. It should be noted that each layer of operator performs submission at most once, so the power consumption load will not be increased. Thus, the fine-grained fault correction process is completed.
[0037] In the above scheme, this process coordinates the error correction process through a fault-tolerant state machine and uses adjacent replication technology to achieve efficient recovery. Assume that the newly added faulty PE number is (2, 3). The fault-tolerant state machine queries the faulty PE table to learn whether the two adjacent PEs (1,3) and (4,3) of this faulty PE are faulty. As long as one of the adjacent PEs is in a healthy state, these two PEs are paired, and the pairing information is filled into a two-dimensional replication query table.
[0038] If by chance the adjacent PEs on the left and right also fail and are isolated, then the "cross-column protection mechanism" is enabled. That is, immediately query whether the PEs (0,3) and (5,3) across one column from the faulty PE are in a healthy state. If one of them is healthy, a pair is formed and the information is filled into the replication query table. The purpose of only crossing one column is that when adjacent PEs cannot be replicated, fine-grained error correction can continue while avoiding complex connection logic.
[0039] Step 2, if after querying the faulty PE table, it is found that neither the adjacent replication nor the cross-column protection mechanism has the execution conditions, then enter the thread-level error correction stage. In this stage, this application will mark the column thread where the faulty PE is located as an unavailable thread according to the column (i.e., thread) position of the faulty PE, and then isolate the entire column of threads.
[0040] The specific steps of the thread-level coarse-grained fault recovery technical solution are as follows: As Figure 4As shown, it is a flowchart of the thread-level coarse-grained error correction technology based on on-chip routing. In the thread-level coarse-grained error correction scheme, first, the main state machine records the column number where the faulty PE is located in the PE fault table into a one-dimensional boolean array. This array is represented by a 64-bit register in hardware, and its serial number represents the thread number (0 to 63). The value "0" in it represents a healthy thread, while "1" represents a faulty thread.
[0041] Subsequently, this application introduces a pair of head and tail double pointers to complete the update of the routing table. The process is schematically shown as Figure 5 As shown, it is a flowchart of generating a routing table by the head and tail double pointers. The initial value of the head pointer points to the 0 address of the routing table, and the tail pointer points to the 63 address of the routing table. Similarly, under the control of the fault-tolerant state machine, starting from the 0 address of the faulty thread registration table, it is scanned every cycle. Whenever a "0" is encountered, the address pointed to by the head pointer is filled with the current thread number, and the pointer is incremented by one; whenever a "1" is encountered, the address pointed to by the tail pointer is filled with the current thread number, and it is decremented by one. After 64 cycles in total, the update of the routing table can be completed. The routing table is represented by a set of register files in hardware, with a depth equal to the number of threads and a width equal to the logarithm of the number of threads.
[0042] Subsequently, this application introduces a priority selector, which routes the activation value from the thread buffer (Thread Buffers, TB) to the corresponding correct thread through the input routing table. Since the MWMT architecture uses a systolic array with a stable output stream, its columns share a section of activation in a pipelined manner, and its rows share a section of weights in a broadcast manner. The partial sum operation stays within the PE to complete the accumulation. Therefore, when the faulty thread is isolated, the activation value in the original corresponding thread buffer will be distributed to the first healthy thread behind it for calculation. Although the cost is to reduce one thread, it ensures that no errors will occur in the calculation result and the accuracy will not be affected at all.
[0043] Furthermore, as the faults occur in the PEs of threads 1, 2, and 4 in this example, the fault-tolerant state machine will set the corresponding thread addresses in the faulty thread registration table to one. Subsequently, a new routing table is generated under the coordination of the head and tail double pointers, and then the routing table is forwarded to the priority selector. Whenever the operation is executed, the priority selector will complete a process of distributing the activation value from 64 thread buffers to the correct thread according to the saved routing table. In this example, thread buffers 1 and 2 no longer send the activation to threads 1 and 2, but send it to the subsequent healthy threads 3 and 5 for calculation.
[0044] As a specific solution, this application introduces a module in the submission stage to restore the pipeline submission, as Figure 6As shown in the figure, it is a schematic diagram of a pipelined selective submission module. Before introducing this error correction technology, in the MWMT architecture, the systolic array would submit the PE submission values of an array of threads to the post-processing unit for processing in a pipelined manner in one cycle. Then, since the faulty threads would be disabled, the pipelining technology could not be directly adopted. Therefore, in this application, a set of multiplexers and a one-hot code thread counter are added to control the selective submission, and the pipelined submission is restored by selecting a column of threads in each cycle. In addition, the first-level output register is used to optimize the timing.
[0045] Furthermore, this process adopts on-chip routing technology, and the execution of error correction is controlled by a fault-tolerant state machine. This application will record the unavailable threads in a boolean array, and then scan the array with head and tail double pointers within a fixed number of cycles to generate a routing table. The sequence number of the routing table represents the thread buffer number, and its value represents the target thread number to be routed. The initial state of the routing table is that the sequence number corresponds one-to-one with the target thread number. Then, through the routing table, this application can route the activation value to the healthy and available threads to complete the computing task.
[0046] Among them, the fault-tolerant state machine of the embodiment of this application includes: Fault response module: Complete the update of the replication lookup table within 3 clock cycles after detecting a new faulty PE; Mode switching module: Automatically switch between fine-grained / coarse-grained error correction modes according to the availability evaluation result of adjacent PEs; Power consumption control module: Ensure that a single-layer operator only performs one submission operation.
[0047] Hardware signal extension unit: Add a 1-bit fault flag register and a 2-bit replication direction register for each PE; On-chip routing reconstruction unit: Includes a programmable routing table memory and a double-pointer scan controller; Submission arbitration logic: Integrate the multiplexer of the pipelined submission channel and the selective submission channel.
[0048] This application can effectively enhance the computing reliability of the chip through the above-mentioned hybrid-granularity fault recovery technology. The specific verification process is as follows: The coarse-grained fault correction of this application has passed the VCS simulation of Synopsys, and the results have been verified to be correct using the Kintex-7 410T FPGA; as of the time of editing the disclosure, the fine-grained part of this application is still in the development stage.
[0049] Specifically, this application uses the Kintex-7 410T FPGA as the prototype chip for the experiment, and the independently developed toolchain provides support for neural network model compilation, quantization, accuracy simulation, deployment, and fault simulation. The FPGA uses a customized API to interact with the host to load bitstreams, instructions, neural network models, network parameters, and feature maps.
[0050] This application runs YOLO-V3 and YOLO-Tiny to verify its protection effect. First, permanent faults are injected into multiple parts and registers of the specified PE, and the fault threshold is set to 0.5%. Then, the changes in the latency, overhead, and committed accuracy of the prototype chip are observed as the available threads gradually increase with the reduction of faults. The specific experiments are as follows: a. Explanation of the latency overhead experimental results: Figure 7 The latency curves of the prototype chip with 32 threads under five configurations are shown, as well as the latency table of the prototype chip and the baseline architecture (without fault tolerance). It is explained through the latency overhead change graph and the comparison table under five remaining available thread conditions.
[0051] It can be concluded that the latency of the prototype chip is inversely proportional to the available threads. For example, when the available threads increase from 4 to 8, the latency decreases from 1807 ms to 957 ms, which is reduced by half. This application runs the above network on the baseline architecture. The table measures the latency when using 2, 4, 8, 16, and 32 threads. In different networks, the latency overhead caused by fault recovery is slightly different from the baseline architecture, but the maximum overhead is only 2.5%.
[0052] b. Explanation of the committed accuracy experimental results: As Figure 8 shown, Figure 8 are the committed accuracies under five remaining available thread counts. This application tests the impact of the prototype chip with 64 threads on the committed accuracy of the PE (the situation of five different remaining available thread counts after injecting errors).
[0053] First, the above two networks are run on the 64-thread baseline architecture, and their commits are used as the golden reference to calculate the committed accuracy of the prototype chip. It can be observed that once there are any available threads in the prototype chip, its committed result can match the percentage of the committed value of the baseline architecture. The experimental results show that even with only a very low 1.5% (1 out of 64 threads) of remaining functional elements, the prototype chip can still complete the inference task, which is better than the NPU without a fault tolerance mechanism.
[0054] As Figure 9 shown, the second objective of this application is to provide an error correction device for a neural network processor, including: An acquisition unit 100 is used to acquire a faulty PE table obtained by error detection. The faulty PE table contains a two-dimensional list of fault occurrence and location information. An execution unit 200 is used to locate a newly detected faulty PE according to the faulty PE table and execute a multi-level fault tolerance decision: i. Fine-grained error correction stage: Query the status of adjacent PEs of the faulty PE. When there are healthy adjacent PEs, generate a replication query table and execute an adjacent replication operation, and expand the adjacent range through a cross-column protection mechanism; when all four adjacent PEs on the left and right are faulty, then trigger coarse-grained error correction. ii. Coarse-grained error correction stage: When fine-grained error correction is not feasible, mark the column where the faulty PE is located as an unavailable thread, generate a routing table and start on-chip redirection. A submission unit 300 selects pipeline submission or a submission mechanism according to the current error correction mode and outputs a calculation result.
[0055] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the error correction method of the neural network processor.
[0056] The electronic device provided by this application can implement each function in the error correction method of the neural network processor described above, and can achieve beneficial effects similar to those of the embodiments of the error correction method of the neural network processor described above. To avoid repetition, it will not be elaborated here.
[0057] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the error correction method of the neural network processor.
[0058] The computer-readable storage medium provided by this application can implement each function in the error correction method of the neural network processor described above, and can achieve beneficial effects similar to those of the embodiments of the error correction method of the neural network processor described above. To avoid repetition, it will not be elaborated here.
[0059] An embodiment of this application also provides an AI chip system, including: The fault tolerance processing device as described above; A systolic array computing unit, whose PEs are interconnected through a row broadcast weight bus and a column pipelined activation bus; A pooling layer hardware accelerator, whose output end is connected to the data replacement interface of the fault tolerance processing device.
[0060] This AI chip system adopts a heterogeneous computing architecture, mainly including three core modules: a fault-tolerant processing device, a systolic array computing unit, and a pooling layer hardware accelerator. Each module realizes data interaction through a customized bus, and the specific parameters of the embodiments can be adjusted according to the actual process node. Those skilled in the art can make appropriate modifications and variations without departing from the core idea of this application.
[0061] An embodiment of this application also provides a computer program product, which includes computer instructions that direct a computer to execute the error correction method of the above neural network processor.
[0062] The computer program product provided by this application can implement each function in the error correction method of the foregoing neural network processor of this application, and can achieve beneficial effects similar to those of the embodiments of the error correction method of the foregoing neural network processor of this application. To avoid repetition, it will not be elaborated here.
[0063] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified functions in one block or multiple blocks.
[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified functions in one block or multiple blocks.
[0065] This application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, readable storage media, optical storage, etc.) containing computer-usable program code.
[0066] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0067] Obviously, the described embodiments are only partial embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present application or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present application shall be covered by the protection scope of the claims of the present application.
Claims
1. A method for correcting an error in a neural network processor, characterized in that: include: According to the fault PE table obtained by error detection, the fault PE table contains a two-dimensional list of fault occurrence and location information, According to the faulty PE table, the newly detected faulty PE is located and multi-level fault-tolerance decisions are made: i. Fine-grained error correction stage: query the status of the adjacent PEs of the faulty PE. When there are healthy adjacent PEs, generate a replication query table and perform adjacent replication operations, and expand the adjacent range through the cross-column protection mechanism. When all four adjacent PEs on the left and right fail, coarse-grained error correction is triggered. ii. Coarse-grained error correction stage: When fine-grained error correction is not feasible, the column where the faulty PE is located is marked as an unavailable thread, a routing table is generated, and on-chip redirection is initiated; Select pipeline submission or selection submission mechanism according to the current error correction mode and output the calculation results.
2. The error correction method of a neural network processor according to claim 1, characterized in that: In the hidden layer of the neural network processor, each row of the systolic array in the MWMT architecture adopts a parallel computing structure of row broadcast weight / column pipeline activation, and the PE unit independently completes the multiplication and addition operation; in the pooling layer, it has the feature of adjacent activation value substitution; A 1-bit fault flag signal and a 2-bit replication signal are added inside the PE; the fault flag signal includes non-faulty PE and faulty PE, and the replication signal supports four-level replication directions: left one / right one in adjacent columns and left two / right two in cross-columns.
3. The hybrid granularity error correction method for a neural network accelerator according to claim 1, characterized in that: The fine-grained error correction stage includes: Adjacent replication priority determination: Prioritize the detection of the status of the two adjacent PEs on the left and right of the same row of the faulty PE table; as long as one of the adjacent PEs is in a healthy state, the two PEs are paired and the pairing information is filled into a two-dimensional replication query table; if both the left and right adjacent PEs fail and are isolated, the cross-column protection mechanism is enabled; Cross-column protection mechanism: Query whether the two adjacent PEs on the left and right of the faulty PE across a column are in a healthy state. If one of them is healthy, a pair is formed and the information is filled in the replication query table.
4. The hybrid granularity error correction method for a neural network accelerator according to claim 3, characterized in that: The cross-column protection mechanism satisfies: The span distance is limited to a single column span, and the spanning PE is located in the same row as the original faulty PE; Cross-column protection is triggered only when all adjacent PEs fail, and the cross-column PE with the highest correlation with the faulty PE computing task is selected first.
5. The method according to claim 1, characterized in that: The coarse-grained error correction stage includes: Mark the unavailable threads so that the entire column containing the faulty PE is marked as a Boolean array; The routing table is dynamically generated by scanning the Boolean array with double pointers at the beginning and the end. The routing table completes the reconstruction of the mapping relationship from the thread buffer to the target thread within a preset period. A priority selector is introduced to route the activation value from the thread buffer to the corresponding correct thread through the input routing table; when the faulty thread is isolated, the activation value in the original corresponding thread buffer will be distributed to the first healthy thread thereafter for calculation.
6. The method according to claim 5, characterized in that: The routing table is dynamically generated by scanning the Boolean array with double pointers at the beginning and the end. The routing table completes the reconstruction of the mapping relationship from the thread buffer to the target thread within a preset period, including: A pair of head and tail double pointers are introduced. The initial value of the head pointer points to address 0 of the routing table, and the tail pointer points to address 63 of the routing table. Each cycle starts scanning from address 0 of the fault thread registration table. Whenever 0 is encountered, the address pointed to by the head pointer is filled with the current thread number, and the pointer is increased by 1; whenever 1 is encountered, the address pointed to by the tail pointer is filled with the current thread number and decreased by 1; After a preset period, the updating of the routing table is completed; the routing table is represented by a group of register stacks in hardware, the depth of which is the number of threads, and the width is the logarithmic value of the number of threads.
7. The method according to claim 4, characterized in that: The routing table generation satisfies: The sequence number of the routing table represents the thread buffer number, and its value represents the target thread number to be routed. The initial state of the routing table is that the sequence number corresponds to the target thread number one by one.
8. The error correction method of a neural network processor according to claim 1, characterized in that: The step of selecting pipeline submission or selecting submission mechanism according to the current error correction mode and outputting the calculation result includes: When submitting in the fine-grained error correction stage, if the coarse-grained error correction is not triggered, an additional replication cycle is added to complete the pipeline submission; if the coarse-grained error correction is triggered, the selective submission mode is adopted; each layer of operators only performs a single submission; When submitting in the coarse-grained error correction stage, a pipelined selection submission module is introduced. The selection submission module is added with a set of multiplexers and one-hot code thread counters to control the selection submission, and restore the pipeline submission by selecting a column of threads in each cycle.
9. An error correction device for a neural network processor, characterized in that: include: an acquisition unit, configured to acquire a fault PE table obtained by error detection, wherein the fault PE table includes a two-dimensional list of fault occurrence and location information, The execution unit is used to locate the newly detected faulty PE according to the faulty PE table and execute multi-level fault-tolerance decisions: i. Fine-grained error correction stage: query the status of the adjacent PEs of the faulty PE. When there are healthy adjacent PEs, generate a replication query table and perform adjacent replication operations, and expand the adjacent range through the cross-column protection mechanism. When all four adjacent PEs on the left and right fail, coarse-grained error correction is triggered. ii. Coarse-grained error correction stage: When fine-grained error correction is not feasible, the column where the faulty PE is located is marked as an unavailable thread, a routing table is generated, and on-chip redirection is initiated; The submission unit selects pipeline submission or selection submission mechanism according to the current error correction mode and outputs the calculation results.
10. A computer-readable storage medium storing a fault-tolerant control program, characterized in that: When the program is executed by a processor, the error correction method of the neural network processor according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Integrated circuit memory fault management method and system based on NPU processor architecture
CN121743098A
An NPU array architecture with online fault detection and PE remapping capabilities and its implementation method.
CN122570243A
An npu array architecture with online fault detection and pe remapping capability and an implementation method thereof
CN122570243B