Built-in self-test repair system for PE matrix of AI accelerator chip
By designing a built-in self-test repair system in the AI accelerator chip, the problem of large proportion of circuit module area and external software for diagnosis of faults in the existing technology is solved, and the effect of reducing chip area and improving yield is achieved.
Patent Information
- Application Number
- CN202421827556.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2034-07-31
AI Technical Summary
In the circuit design of existing AI accelerator chips, the circuit module area accounts for a large proportion, resulting in a significant increase in the chip area. It is necessary to diagnose the faulty PE computing unit through external software, resulting in a decrease in yield rate and an increase in manufacturing costs.
Design a built-in self-test and repair system for AI accelerator chip PE matrix, including a test and repair controller, test excitation memory, configurable PE computing unit, output response comparator and fault PE diagnostics, which can perform self-test and self-repair without external equipment and software, reduce chip area and improve yield.
The built-in self-test circuit significantly reduces the chip area proportion, reduces the number of abandoned PE computing units, improves the computing accuracy, improves the chip yield and reduces manufacturing costs.
Smart Images

Figure CN222825904U_ABST
Abstract
Description
Technical Field
[0001] The utility model relates to the technical field of AI accelerator chip circuit structure design, and in particular to a built-in self-test and repair system of an AI accelerator chip PE matrix. Background Art
[0002] The systolic matrix composed of multipliers and accumulators can greatly improve the efficiency of deep learning algorithms and is widely used in AI accelerator chips such as tensor processors and neural network processors. The testability and fault tolerance design technology for such emerging high-performance AI accelerator chips has begun to receive widespread attention.
[0003] The structure of the systolic matrix unit composed of the multiplication and accumulation unit PE is shown in FIG. Figure 1 , PE operation unit structure diagram Figure 2 Each PE operation unit includes a weight register, a data register and an accumulation register. The weight and activation data are loaded into the weight register and data register in each row and column of PE operation units in the form of serial scanning. The weight and activation data latched by the register are multiplied by a multiplier, and the multiplication result is added to the accumulation result latched by the accumulation register in the previous PE operation unit through an adder. The addition result is stored in the accumulation register and transmitted to the next PE operation unit. The final result of the multiplication and addition operation of each column of PE operation units is saved in the accumulation result memory. During the test, the accumulation register needs to be configured as a shift register through a two-way selector, and the output response of the multiplier and adder can be serially exported by the scan chain circuit composed of the shift register for comparison.
[0004] However, since each PE operation unit contains a multi-bit register, the accumulator register is configured as a larger scan register, that is, each bit register needs to add a two-way selector to form a scan chain. The structure is shown in the figure below. Figure 2 , which will greatly increase the area of the entire pulsation matrix module, and ultimately lead to a significant increase in the area of the entire AI acceleration chip; and in the test process of the current circuit structure, the faulty PE computing unit can only be diagnosed by external software, and the diagnosed faulty unit can only be shielded and discarded. If the number of discarded PE computing units is too large or the weight is too large, the computing accuracy of the AI acceleration chip will be significantly reduced, making the AI acceleration chip a waste chip, resulting in a decrease in the chip yield and an increase in the chip manufacturing cost. Utility Model Content
[0005] In order to overcome the problem in the existing circuit structure design technology that the area of the circuit module in the circuit design accounts for a large proportion, resulting in a significant increase in the area of the entire AI acceleration chip, and the need to diagnose the faulty PE operation unit through external software. The faulty PE operation unit leads to a decrease in the chip yield and an increase in the chip manufacturing cost, the utility model proposes a built-in self-test and repair system for the PE matrix of an AI accelerator chip, which can effectively reduce the area of the circuit module, thereby reducing the area of the AI acceleration chip, and there is no need to diagnose the faulty PE operation unit through external software, thereby improving the chip yield and reducing the chip manufacturing cost.
[0006] In order to achieve the purpose of the utility model, the utility model adopts the following technical solutions:
[0007] One of the purposes of the present invention is to:
[0008] A built-in self-test system for a PE matrix of an AI accelerator chip, the system comprising: a test repair controller, a test stimulus memory, a configurable PE operation unit, an output response comparator and a fault PE diagnostic device, wherein the output end of the test stimulus memory is connected to the input end of the configurable PE operation unit, the output end of the configurable PE operation unit is connected to the input end of the output response comparator, and the output end of the output response comparator is connected to the input end of the fault PE diagnostic unit;
[0009] The output end of the test repair controller is connected to the input end of the test stimulus memory, the output response comparator and the fault PE diagnostic device respectively.
[0010] According to the above technical solution, the built-in self-test can efficiently complete the self-test and diagnosis of the PE matrix without the aid of external equipment and software, and there is no need to insert a scan chain in the PE unit. The area occupied by the chip self-test circuit is significantly lower than the area occupied by the existing test technology circuit, which can significantly reduce the chip manufacturing cost.
[0011] The second purpose of this utility model is:
[0012] A built-in self-repair system for an AI accelerator chip PE matrix, the system comprising: a fault PE address memory, an RPE redundant operation unit row, an RPE redundant operation unit column, a configurable PE operation unit, a test repair controller, an accumulation memory, an activation data bus distributor, and a weight data bus distributor;
[0013] The output end of the fault PE address memory is connected to the address input ends of the activation data bus distributor and the weight data bus distributor respectively; the output end of the activation data bus distributor is connected to the activation data input end of the RPE unit in the RPE redundant operation unit row, the accumulation output ends of the RPE units in the RPE redundant operation unit row are respectively connected to the accumulation input ends of the configurable PE operation units located in the first row of the PE matrix, and the accumulation output end of the configurable PE operation unit located in the last row of the PE matrix is connected to the input end of the accumulation memory; the output end of the weight data bus distributor is connected to the weight data input end of the RPE redundant operation unit in the RPE redundant operation unit column; the output end of the test and repair controller is respectively connected to the input ends of the fault PE address memory, the activation data bus distributor, and the weight data bus distributor.
[0014] According to the above technical solution, the built-in self-repair can efficiently complete the self-diagnosis and repair of the PE matrix without the help of external equipment and software. There is no need to insert a scan chain in the PE unit, and the faulty PE can be replaced and repaired, so that the chip yield rate will be significantly improved. The reduction in chip area and the improvement in yield rate can bring significant economic benefits.
[0015] According to the above technical scheme, through the scan test generator, configurable PE operation unit, output response comparator, fault PE diagnostic device in the built-in self-test circuit, and the activation data bus distributor and RPE redundant operation unit in the built-in self-repair circuit, combined with the test repair controller and the accumulator memory, the area of the circuit module can be effectively reduced, thereby reducing the area of the AI acceleration chip, and there is no need to diagnose the faulty PE operation unit through external software. The circuit can be self-tested through the built-in self-test circuit, and the circuit can be self-repaired through the built-in self-repair circuit, thereby reducing the number of PE operation units abandoned during the test process, improving the operation accuracy of the PE operation unit, thereby improving the chip yield and reducing the chip manufacturing cost.
[0016] Compared with the prior art, the beneficial effects of the present invention are:
[0017] The utility model proposes a built-in self-test and repair system for a PE matrix of an AI accelerator chip. Through a scan test generator, a configurable PE operation unit, an output response comparator, a faulty PE diagnostic device in a built-in self-test circuit, and an activation data bus distributor and an RPE redundant operation unit in a built-in self-repair circuit, combined with a test repair controller and an accumulator memory, the area of the circuit module can be effectively reduced, thereby reducing the area of the AI accelerator chip. Without the need to diagnose the faulty PE operation unit through external software, the circuit can be self-tested through the built-in self-test circuit, and the circuit can be self-repaired through the built-in self-repair circuit, thereby reducing the number of PE operation units discarded during the test process, improving the operation accuracy of the PE operation unit, thereby improving the chip yield rate, and reducing the chip manufacturing cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a schematic diagram of the PE pulsation matrix structure inside an existing AI accelerator chip;
[0019] Figure 2 It is a schematic diagram of the circuit structure of the PE operation unit under the existing test;
[0020] Figure 3 A schematic diagram of the structure of a built-in self-test and repair system for a PE matrix of an AI accelerator chip provided by an embodiment of the utility model;
[0021] Figure 4 A schematic diagram of the circuit structure of a configurable PE computing unit provided in an embodiment of the utility model;
[0022] Figure 5 A schematic diagram of the circuit structure of the RPE redundant computing unit provided in an embodiment of the utility model;
[0023] Figure 6 A schematic diagram of the arrangement of RPE units for a 64x64 matrix of PE computing units provided in an embodiment of the present utility model;
[0024] Figure 7 A schematic diagram of the arrangement of RPE units for a 128x128 matrix of PE computing units provided in an embodiment of the utility model;
[0025] Figure 8 A schematic diagram of the arrangement of RPE units for a 256x256 matrix of PE computing units provided in an embodiment of the present utility model. DETAILED DESCRIPTION
[0026] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. The drawings provide preferred embodiments of the present invention. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive.
[0027] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in the technical field of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0029] Embodiment 1:
[0030] This embodiment provides a built-in self-test and repair system for the PE matrix of an AI accelerator chip, see Figure 3 The system comprises: a test repair controller, a test stimulus memory, a configurable PE operation unit, an output response comparator and a fault PE diagnostic device, wherein the output end of the test stimulus memory is connected to the input end of the configurable PE operation unit, the output end of the configurable PE operation unit is connected to the input end of the output response comparator, and the output end of the output response comparator is connected to the input end of the fault PE diagnostic unit;
[0031] The output end of the test repair controller is connected to the input end of the test stimulus memory, the output response comparator and the fault PE diagnostic device respectively.
[0032] As a preferred embodiment, see Figure 4The configurable PE operation unit includes a PE weight register, a PE data register, a zero register, a PE accumulation register, a PE multiplier, a PE adder and a PE AND gate circuit. The input end of the PE multiplier is respectively connected to the output end of the PE weight register and the output end of the PE data register, and the input end of the PE AND gate circuit is respectively connected to the output end of the PE data register, the output end of the PE multiplier, and the output end of the zero register; the output end of the PE AND gate circuit is connected to the input end of the PE adder, and the output end of the PE adder is connected to the input end of the PE accumulation register.
[0033] In this embodiment, the circuit structure of the configurable PE operation unit is shown in Figure 4 , an AND gate circuit is inserted between the multiplier and the adder in the original PE operation unit. Each bit output of the multiplier is connected to an input of the AND gate circuit, each bit output of the multiplier is connected to one of the inputs of the AND gate, the output of the AND gate is connected to the input of the adder, and the other input ends of the AND gate can be controlled by a zero register or a scan test control signal.
[0034] When the zero register or scan test control signal is set to 0, the AND gate circuit output is 0 and the multiplier result is set to zero. At this time, the result of the previous PE unit output is accumulated with 0 and remains unchanged, so the result latched by the accumulation register follows the output of the previous PE unit, thereby realizing the scan shift function of the register. When the zero control register and the scan test control signal are 1, the AND gate circuit can be regarded as transparent, keeping the operation function of the PE unit unchanged.
[0035] As a preferred embodiment, see Figure 3 The system also includes a scan test generator, the output end of the test stimulus memory is connected to the input end of the scan test generator, the output end of the scan test generator is connected to the input end of the configurable PE operation unit located in the first row of the PE matrix; the output end of the test repair controller is connected to the input end of the scan test generator.
[0036] In this embodiment, the test stimulus memory stores test stimulus for testing the multiplier and adder in a single PE operation unit. The output end of the test data memory is connected to the input end of the scan test generator; the scan data generator reads the stored test stimulus data, and its output end is connected to the activation data input end and the weight data input end of the configurable PE operation unit, and the test stimulus is serially transmitted to the multiplier and adder input end of each PE operation unit through the shift register composed of the weight register and the data register in the PE.
[0037] As a preferred embodiment, see Figure 3The system also includes a faulty PE address memory, an input end of the faulty PE address memory is connected to an output end of the faulty PE diagnoser, and an output end of the test and repair controller is connected to an input end of the faulty PE address memory.
[0038] In this embodiment, the output end of the accumulator register of the configurable PE operation unit in the last row of the configurable PE operation unit matrix is connected to the input end of the output response comparator, the output end of the output response comparator is connected to the input end of the fault PE diagnostic device, and the output end of the fault PE diagnostic device is connected to the data and address input ends of the fault PE address memory.
[0039] The output response of the test stimulus after multiplication and addition operation is latched by the accumulator register of the PE operation unit, and the accumulator register is configured into a scan chain through the added AND gate circuit. The output response is serially shifted through the scan chain to the output response comparator and compared with the correct response result saved in the test stimulus memory.
[0040] The output response comparator compares the output response of each PE operation unit with the correct result stored in the on-chip memory to determine whether the PE operation unit has a fault. Each of the lowest bits of the output of the output response comparator is connected to the AND gate circuit. According to the chip PE operation matrix operation accuracy requirements, the AND gate can be opened and closed using external signals to adopt or ignore the comparison results of the lowest bits of the response output. If the operation accuracy requirements are not high, the fault that causes the lowest bits of the output response to be wrong can be ignored. At this time, the AND gate circuit will shield the output response comparison results of these bits to ensure that the PE operation matrix passes the test.
[0041] The faulty PE diagnostic device reads the comparison result data output by the detection result comparator, and adopts a clock-controlled multi-stage AND gate, XOR gate feedback circuit and decoder circuit to decode the parallel output data in one clock cycle to obtain the row address of the faulty PE operation unit, and decode the serial output data in multiple clock cycles to obtain the column address of the faulty PE operation unit, and save the address to the faulty PE address memory.
[0042] As a preferred embodiment, see Figure 3 The system further comprises a two-way selector, wherein a data input end of the two-way selector is connected to an output end of the scan test generator, a control input end of the two-way selector is connected to an output end of a test and repair controller, and an output end of the two-way selector is connected to a data input end of a configurable PE operation unit in the first column of a PE matrix. The two-way selector is configured to switch the working state of the circuit to a built-in self-test process or a built-in self-repair process according to an instruction of the test and repair controller.
[0043] Embodiment 2:
[0044] In this embodiment, a plurality of redundant RPE computing units that can be used to replace faulty PE computing units are provided, and the redundant RPE computing unit row and column structures are arranged around the configurable PE computing unit matrix to form a plurality of RPE redundant unit rows and a plurality of RPE redundant unit columns. This embodiment provides a built-in self-repair system for the PE matrix of an AI accelerator chip, see Figure 3 , the system includes: a faulty PE address memory, an RPE redundant operation unit row, an RPE redundant operation unit column, a configurable PE operation unit, a test and repair controller, an accumulation memory, an activation data bus distributor and a weight data bus distributor;
[0045] The output end of the fault PE address memory is connected to the address input ends of the activation data bus distributor and the weight data bus distributor respectively; the output end of the activation data bus distributor is connected to the activation data input end of the RPE unit in the RPE redundant operation unit row, the accumulation output ends of the RPE units in the RPE redundant operation unit row are respectively connected to the accumulation input ends of the configurable PE operation units located in the first row of the PE matrix, and the accumulation output end of the configurable PE operation unit located in the last row of the PE matrix is connected to the input end of the accumulation memory; the output end of the weight data bus distributor is connected to the weight data input end of the RPE redundant operation unit in the RPE redundant operation unit column; the output end of the test and repair controller is respectively connected to the input ends of the fault PE address memory, the activation data bus distributor, and the weight data bus distributor.
[0046] As a preferred embodiment, see Figure 3 The system also includes a weight FIFO memory, the output end of the faulty PE address memory is connected to the address input end of the weight FIFO memory, and the output end of the weight FIFO memory is respectively connected to the weight data input end of the RPE unit in the RPE redundant operation unit row and the input end of the weight data bus distributor.
[0047] As a preferred embodiment, see Figure 3 The system also includes an RPE operation result scanning accumulator, the input end of the RPE operation result scanning accumulator is respectively connected to the data output end of the RPE unit in the RPE redundant operation unit column and the output end of the test and repair controller, and the output end of the RPE operation result scanning accumulator is connected to the input end of the accumulation memory.
[0048] As a preferred embodiment, see Figure 3, the system further comprises an adder, a two-way selector and an activation data memory, the input end of the adder is respectively connected to the output end of the configurable PE operation unit and the output end of the RPE operation result scanning accumulator, and the output end of the adder is connected to the input end of the accumulator memory;
[0049] The data input end of the two-way selector is connected to the output end of the activation data memory, the control input end of the two-way selector is connected to the output end of the test and repair controller, and the output end of the two-way selector is connected to the data input end of the configurable PE operation unit in the first column of the PE matrix;
[0050] The input end of the activation data storage is connected to the output end of the test and repair controller.
[0051] As a preferred embodiment, see Figure 5 The RPE redundant operation unit includes an RPE weight counter, an RPE weight register, an RPE data register, an RPE multiplier, an RPE AND gate circuit, an RPE adder and an RPE accumulator register. The output end of the counter is connected to the input end of the RPE weight register, the input end of the RPE multiplier is respectively connected to the output end of the RPE weight register and the output end of the RPE data register, the output end of the RPE multiplier is connected to the input end of the RPE AND gate circuit, the output end of the RPE AND gate circuit is connected to the input end of the RPE adder, and the output end of the RPE adder is connected to the input end of the RPE accumulator register.
[0052] In this embodiment, in the repair mode, the output end of the faulty PE address memory is connected to the activation data bus distributor, the weight data bus distributor and the address input end of the RPE redundant operation unit, the output end of the activation data buffer is connected to the data input end of the activation data bus distributor, the output end of the weight FIFO memory is connected to the weight data input end of the RPE redundant operation unit in the RPE redundant operation unit row and the data input end of the weight data bus distributor, the output end of the activation data bus distributor is connected to the activation data input end of the RPE redundant operation unit in the RPE redundant operation unit row, the output end of the weight data bus distributor is connected to the weight data input end of the RPE redundant operation unit in the RPE redundant operation unit column, the accumulator output end of the RPE redundant operation unit in the RPE redundant operation unit column is connected to the input end of the column RPE operation result scanning accumulator, the output end of the RPE operation result scanning accumulator is connected to the adder above the accumulation memory, and the output end of the adder is connected to the data input end of the accumulation memory.
[0053] After the row and column addresses of the faulty PE computing unit are read from the faulty PE address access device, the RPE unit in the column of the RPE redundant computing unit row is called for replacement according to its column address. If this RPE unit has not been used, the data bus distributor allocates the activation data bus corresponding to the row to this RPE unit according to the row address of the current faulty PE to ensure that the activation data stream obtained by the RPE unit is the same as the activation data stream obtained by the faulty PE unit. Since the weight data input end of each RPE in the RPE unit row is respectively connected to the output end corresponding to the weight FIFO memory, there is no need to allocate the weight data bus.
[0054] If the RPE unit in the RPE redundant operation unit row has been used, the RPE unit in the RPE redundant operation unit column in the row is called for replacement. If this RPE unit has not been used, the weight data bus distributor allocates the weight data bus corresponding to the column to this RPE unit according to the column address of the current faulty PE to ensure that the weight data stream obtained by the RPE unit is the same as the weight data stream obtained by the faulty PE unit. Since the activation data input end of each RPE unit in the RPE redundant operation unit column is connected to the activation data output end of the previous PE unit, there is no need to allocate the activation data bus.
[0055] In this embodiment, multiple RPE redundant operation unit rows and RPE redundant operation unit columns can be added. If the corresponding RPE units in the current RPE redundant operation unit row and RPE redundant operation unit column have been used, the RPE units in the remaining RPE redundant operation unit rows or RPE redundant operation unit columns can be called for replacement, so that the repair probability of the chip PE matrix can be greatly improved under the premise of slightly increasing the circuit area.
[0056] After all faulty PE units have been repaired, the zero register of the faulty PE unit is set to 0. At this time, the output of the AND gate circuit is 0, the multiplier result is masked to 0, and the adder of the faulty PE unit directly transmits the calculation result of the previous PE unit connected in series to the next PE unit in series. The faulty PE unit will not participate in the calculation.
[0057] After the shielding is completed, the weight data of the RPE redundant operation unit rows and RPE redundant operation unit columns are serially preloaded. According to the column address information of the faulty PE latched by each RPE unit, the weight data corresponding to the column address is latched into the weight register of the RPE using the counter in the RPE unit, so that the weight data used by the RPE unit is the same as the weight data of the faulty PE unit.
[0058] After the replacement and repair is completed, the RPE operation units participating in the replacement and repair in the RPE redundant operation unit row and the RPE redundant operation unit column and other normal PE operation units perform multiplication operations synchronously, but the RPE units in the RPE unit column latch the operation results inside and continuously accumulate themselves with the multiplication results generated in the next clock cycle, that is, the multiplication and accumulation results in the entire operation cycle are always stored in a single RPE unit.
[0059] After the PE matrix operation is completed, multiple RPE unit column operation result scanning accumulators read the operation results of each RPE in the RPE unit column and its corresponding column address in a serial scanning manner, and continuously accumulate the RPE unit operation results with the same column address as the scanning accumulator, and distribute the final accumulated results to the accumulator at the end of the PE systolic matrix to complete the final accumulation operation.
[0060] After testing and repair, the number of faulty PEs and their row and column address data can be exported along with the scan chain to external testing software for chip design and manufacturing process maturity analysis and optimization, and to determine the optimized number of RPE units and row and column arrangement.
[0061] Embodiment three:
[0062] This embodiment provides relevant experimental data of built-in self-test and built-in self-repair based on the built-in self-test and self-repair system of the AI accelerator chip PE matrix for the first and second embodiments.
[0063] As a preferred embodiment, see Figure 4 and Figure 6 , a 64x64 matrix composed of PE computing units, two RPE redundant unit rows are arranged on the upper part of the matrix, and one RPE redundant unit column is arranged on the right side of the matrix.
[0064] See also Figure 5 ,The counter is used to latch the corresponding weight data into the weight register according to the column address information to ensure that the RPE and the replaced faulty PE have the same weight data.,There is also a two-input AND gate inside the RPE unit to support the scan chain circuit configuration of the accumulator register.
[0065] Select the RPE computing unit in the first RPE redundant computing unit row that is in the same column as the faulty PE computing unit for replacement, and use the activation data bus distributor to allocate the corresponding activation data bus to the RPE computing unit. If the RPE computing unit is already in use, select the RPE computing unit in the same column as the faulty PE computing unit. unit Replace the RPE operation unit in the same row, and use the weight data bus distributor to assign the corresponding weight data bus to the RPE unit. If the RPE unit is also used, continue to select the second RPE operation unit The RPE computing units in the row are replaced and repaired.
[0066] Monte Carlo simulation is used to verify the actual repair effect of this embodiment. Assuming that each PE computing unit has a certain probability of failure, 10,000 64x64 PE matrices composed of these PE units are randomly generated to verify the probability of the PE unit matrix being completely repaired under different failure probabilities when two rows and one column of RPE units are added. The simulation results are shown in Table 1.
[0067] It can be seen that the configuration of two rows and one column of RPE computing units has a repair rate of 100% when the failure rate is 0.5% or less, and the repair rate can still reach 97.3% when the failure rate is 1%.
[0068] Table 1:
[0069]
[0070] As a preferred embodiment, see Figure 7 For the 128x128 PE matrix, in order to improve the fault repair rate, an arrangement of two RPE operation unit rows and two RPE operation unit columns is adopted. Since an additional RPE operation unit column is added, the repair process adds an operation step of selecting an RPE operation unit from the second RPE operation unit column for replacement and repair.
[0071] Monte Carlo simulation is performed on the repair results, and the simulation results are shown in Table 2. It can be seen that the repair rate of the configuration of two rows and two columns of RPE units still reaches 100% when the failure rate is 0.4% or less, and the repair rate still reaches 93.2% at a high failure rate of 0.8%.
[0072] Table 2:
[0073]
[0074] As a preferred embodiment, see Figure 8 For a 256x256 PE operation unit matrix, four RPE operation unit rows and two RPE operation unit columns are arranged. Since two more RPE unit rows and one RPE unit column are added, the repair process adds the operation steps of selecting RPE units from the second RPE unit column, the third and the fourth RPE unit rows for replacement and repair.
[0075] The Monte Carlo simulation results using this configuration are shown in Table 3. It can be seen that even for a very large 256x256 PE computing unit matrix, the repair rate still reaches 100% when the failure rate is 0.6% or less, the repair rate is 97.4% when the failure rate is 0.8%, and the repair rate is still 61.5% at a high failure rate of 1%. If the failure rate continues to increase, the repair rate will decay exponentially. At this time, if you want to further improve the repair rate, you can continue to add another RPE unit column.
[0076] Table 3:
[0077]
[0078] The invention adds RPE redundant operation units and other logic circuits and on-chip memories for testing and repairing, which will cause a slight increase in the circuit area of the entire PE operation unit matrix.
[0079] Synopsys's synthesis tool Design Compiler is used to perform circuit area analysis on the RTL design of each circuit structure after logic synthesis. The analysis results are shown in Tables 4 and 5.
[0080] Table 4:
[0081]
[0082] Table 4 shows the increase in the area of the PE matrix of the existing test technology relative to the PE matrix area without adding test technology. It can be seen that the area increase of the existing scan chain test technology using a two-way selector accounts for 5.24%.
[0083] Table 5:
[0084]
[0085] Table 5 shows the area increment of each circuit structure of the utility model relative to the PE matrix area without adding test technology. The area increment of the configurable PE unit circuit using AND gate circuit accounts for only 0.52%.
[0086] Table 5 shows that for the above three different PE matrices, the proportion of area increase caused by adding different numbers of RPE redundant computing unit rows and columns. Assuming that the area of a single RPE computing unit is similar to that of a PE computing unit, it can be directly estimated that the area increment of the RPE redundant computing unit in Example 1 is 4.68%, the area increment of the RPE redundant computing unit in Example 2 is 2.47%, and the area increment of the RPE redundant computing unit in Example 3 is 2.34%. It can be seen that as the scale of the PE matrix increases, the proportion of the RPE redundant computing unit area increment gradually decreases.
[0087] It can also be seen from Table 5 that the incremental proportion of the built-in self-test circuit area is very small, but since the weighted data bus distributor and the activation data bus distributor in the built-in self-repair circuit occupy a large area, the incremental proportion of the built-in self-repair circuit area is relatively large. Perhaps the area proportion of the built-in self-repair circuit can be reduced by further optimizing the built-in self-repair circuit.
[0088] It should be noted that this area ratio is only relative to the PE matrix circuit area without adding test technology. Relative to the entire AI chip area, the PE matrix circuit area only accounts for about 30-40%, so the total area of the solution of the present invention generally does not exceed 5% relative to the area of the entire AI chip. However, the improvement of the AI chip yield by using the technology of the present invention is far more than 5%.
[0089] The above description is only an embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A built-in self-test system for an AI accelerator chip PE matrix, characterized in that: The system comprises: a test repair controller, a test stimulus memory, a configurable PE operation unit, an output response comparator and a fault PE diagnostic device, wherein the output end of the test stimulus memory is connected to the input end of the configurable PE operation unit, the output end of the configurable PE operation unit located in the last row of the PE matrix is connected to the input end of the output response comparator, and the output end of the output response comparator is connected to the input end of the fault PE diagnostic unit; The output end of the test repair controller is connected to the input end of the test stimulus memory, the output response comparator and the fault PE diagnostic device respectively.
2. The built-in self-test system of the AI accelerator chip PE matrix according to claim 1, characterized in that: The system also includes a scan test generator, the output end of the test stimulus memory is connected to the input end of the scan test generator, the output end of the scan test generator is connected to the input end of the configurable PE operation unit located in the first row of the PE matrix; the output end of the test repair controller is connected to the input end of the scan test generator.
3. The built-in self-test system of the AI accelerator chip PE matrix according to claim 1, characterized in that: The system further comprises a faulty PE address memory, an input end of the faulty PE address memory is connected to an output end of the faulty PE diagnoser, and an output end of the test and repair controller is connected to an input end of the faulty PE address memory.
4. The built-in self-test system of the AI accelerator chip PE matrix according to claim 2, characterized in that: The system also includes a two-way selector, wherein a data input end of the two-way selector is connected to an output end of the scan test generator, a control input end of the two-way selector is connected to an output end of a test repair controller, and an output end of the two-way selector is connected to a data input end of a configurable PE operation unit in the first column of a PE matrix.
5. The built-in self-test system of the AI accelerator chip PE matrix according to any one of claims 1 to 4, characterized in that: The configurable PE operation unit includes a PE weight register, a PE data register, a zero register, a PE accumulation register, a PE multiplier, a PE adder and a PE AND gate circuit. The input end of the PE multiplier is respectively connected to the output end of the PE weight register and the output end of the PE data register, and the input end of the PE AND gate circuit is respectively connected to the output end of the PE data register, the output end of the PE multiplier and the output end of the zero register; the output end of the PE AND gate circuit is connected to the input end of the PE adder, and the output end of the PE adder is connected to the input end of the PE accumulation register.
6. A built-in self-repair system for an AI accelerator chip PE matrix, characterized in that: The system comprises: a fault PE address memory, an RPE redundant operation unit row, an RPE redundant operation unit column, a configurable PE operation unit, a test and repair controller, an accumulation memory, an activation data bus distributor and a weight data bus distributor; The output end of the fault PE address memory is connected to the address input ends of the activation data bus distributor and the weight data bus distributor respectively; the output end of the activation data bus distributor is connected to the activation data input end of the RPE unit in the RPE redundant operation unit row, the accumulation output ends of the RPE units in the RPE redundant operation unit row are respectively connected to the accumulation input ends of the configurable PE operation units located in the first row of the PE matrix, and the accumulation output end of the configurable PE operation unit located in the last row of the PE matrix is connected to the input end of the accumulation memory; the output end of the weight data bus distributor is connected to the weight data input end of the RPE redundant operation unit in the RPE redundant operation unit column; the output end of the test and repair controller is respectively connected to the input ends of the fault PE address memory, the activation data bus distributor, and the weight data bus distributor.
7. The built-in self-repair system of the AI accelerator chip PE matrix according to claim 6, characterized in that: The system also includes a weight FIFO memory, the output end of the fault PE address memory is connected to the address input end of the weight FIFO memory, and the output end of the weight FIFO memory is respectively connected to the weight data input end of the RPE unit in the RPE redundant operation unit row and the input end of the weight data bus distributor.
8. The built-in self-repair system of the AI accelerator chip PE matrix according to claim 7, characterized in that: The system also includes an RPE operation result scanning accumulator, the input end of the RPE operation result scanning accumulator is respectively connected to the data output end of the RPE unit in the RPE redundant operation unit column and the output end of the test and repair controller, and the output end of the RPE operation result scanning accumulator is connected to the input end of the accumulation memory.
9. The built-in self-repair system of the AI accelerator chip PE matrix according to claim 8, characterized in that: The system further comprises an adder, a two-way selector and an activation data memory, wherein the input end of the adder is respectively connected to the output end of the configurable PE operation unit and the output end of the RPE operation result scanning accumulator, and the output end of the adder is connected to the input end of the accumulator memory; The data input end of the two-way selector is connected to the output end of the activation data memory, the control input end of the two-way selector is connected to the output end of the test and repair controller, and the output end of the two-way selector is connected to the data input end of the configurable PE operation unit in the first column of the PE matrix; The input end of the activation data storage is connected to the output end of the test and repair controller.
10. The built-in self-repair system of the AI accelerator chip PE matrix according to any one of claims 6 to 9, characterized in that: The RPE redundant operation unit includes an RPE weight counter, an RPE weight register, an RPE data register, an RPE multiplier, an RPE AND gate circuit, an RPE adder and an RPE accumulator register. The output end of the counter is connected to the input end of the RPE weight register, the input end of the RPE multiplier is respectively connected to the output end of the RPE weight register and the output end of the RPE data register, the output end of the RPE multiplier is connected to the input end of the RPE AND gate circuit, the output end of the RPE AND gate circuit is connected to the input end of the RPE adder, and the output end of the RPE adder is connected to the input end of the RPE accumulator register.