Apparatus and method for diagnostic coverage of neural network accelerators - Patents.com

JP2024528185A5Active Publication Date: 2025-05-09CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024506558
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-10
Filing Date
2022-07-25
Publication Date
2025-05-09
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

Existing AI accelerators face challenges in achieving high diagnostic coverage for safety-critical applications due to the high cost, area, and power consumption of redundant hardware, as well as performance inefficiencies in techniques like lockstep and BIST, and the difficulty in achieving deterministic fault detection with software-based approaches.

Method used

Incorporating safety processing engines (SPEs) within MAC arrays to perform additional operations that verify the correctness of convolution and matrix multiplication operations, reducing the need for redundant hardware and runtime overhead, while ensuring high diagnostic coverage.

Benefits of technology

This approach provides efficient, deterministic diagnostic coverage with reduced hardware resources and minimal performance impact, enabling AI accelerators to operate in safety-critical environments without significant additional cost, area, or power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems, apparatus, and methods are disclosed for implementing a safety framework for safety-critical convolutional neural network inference applications and associated convolution and matrix multiplication based systems. An exemplary system includes a safety-critical application, a hardware accelerator, and additional hardware for performing validation of the hardware accelerator. The validation hardware has a lower bandwidth than the hardware accelerator and therefore requires more machine cycles per calculation. Inconsistencies in the results indicate a faulty processing element.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Background technology An emerging technology field is machine learning or artificial intelligence (AI), and neural networks are a type of machine learning model. Artificial intelligence is widely used in various automotive applications, and more specifically, convolutional neural networks (CNNs) have been shown to provide significant accuracy improvements compared to traditional algorithms for perception and other applications. Neural networks such as CNNs have shown superior performance in tasks such as handwritten digit classification and face detection. In addition, neural networks have shown promise in performing well in other, more challenging visual classification tasks. Other applications of neural networks include speech recognition, language modeling, sentiment analysis, predictive input, etc.

[0002] However, they are often limited to non-safety-critical functions due to the lack of AI hardware and software that can meet safety requirements such as ISO26262. Most AI accelerators do not yet claim to independently meet any ASIL level, and some claim to qualify for the "strictest safety compliance standard" with plans to reach ASIL D compliance in the future. The field of safe AI acceleration can still be characterized as being in its early stages. When referring to safety in terms of hardware design, there are two aspects to be considered: Systematic Capability, i.e., coverage against deterministic failures, which is addressed by following a process standardized by, for example, ISO26262, and Diagnostic Coverage (DC), i.e., coverage against random hardware failures. The latter DC is the main challenge for any new hardware architecture to reach a high safety integrity level. Some approaches are described in the following paragraphs.

[0003] In a typical deployment of machine learning algorithms for real-time or near real-time use, a software application feeds a neural network to an inference accelerator hardware engine. Often, the accelerator is a multiply-add (MAC) array that includes multiple processing elements, each capable of performing multiplication and addition or multiply-add operations. When the inference accelerator operates in a safety-critical environment, it is desirable to monitor the inference accelerator to check for abnormal or faulty behavior. A typical implementation for monitoring the inference accelerator inserts monitoring logic into the inference accelerator processing hardware sub-block. For example, a machine check architecture is a mechanism by which monitoring logic in the processing hardware checks for abnormal behavior.

[0004] Achieving reliability requirements and regulations may require ensuring a hardware reliability level that includes continuous fault detection. To meet the diagnostic coverage required to detect and protect against random hardware faults during continuous operation, many existing common fault detection and handling techniques used in microcontrollers are also applicable and usable in AI accelerators. However, these techniques may have drawbacks.

[0005] The Lockstep Core method (US Pat. No. 5,915,082 A, WO 2011117155 A1), in which two hardware resources perform the same function and compare the results, is a technique widely used in automotive electronic control units (ECUs) to achieve high diagnostic coverage. This technique can potentially claim 100% diagnostic coverage of computational elements for single point failures, but requires duplicating hardware resources, which has a negative impact on cost, space, and power budgets.

[0006] The use of duplicated hardware resources as part of a lockstep approach introduces performance inefficiencies at added cost, area, and power consumption, and may therefore be limited to only high-end products that have sufficient performance and power margins to compensate for these drawbacks.

[0007] Similarly, using built-in self-test (BIST) with functional logic can provide high diagnostic coverage, but many implementations must safely pause and store intermediate results of the functional application to run the self-test, then restart the function, which introduces latency overhead on top of the runtime required for the self-test, which can cause further performance degradation. Retaining intermediate application results while running the self-test is another major concern. This approach requires additional circuitry to be implemented.

[0008] There are also pure software implementations that can guarantee the required diagnostic coverage of some hardware using software test libraries (STL). This approach does not require additional circuitry in the hardware since it performs safety checks at the software level. However, an undesirable side effect of the high-level approach can be that it is difficult to achieve high fault coverage of the hardware, especially for Multiply-Accumulate (MAC) arrays. For neural network accelerators, the Multiply-Accumulate (MAC) arrays may constitute a major part of the hardware. Further details on the state of the art can be derived from the ISO 26262 functional safety standard.

[0009] In addition to the above mentioned approaches to reach high ASIL levels for automotive use cases, there are several research publications proposing various design modifications to achieve higher safety requirements for neural network hardware accelerators. Typically, these safety mechanisms rely on run-time calculated checksums of the neural network weights as well as the input data or intermediate results from the layers of the network, also called activation values. This dynamic nature of the safety mechanisms, which rely on real-time calculation of the checksums and their dependency on input data received in real-time, in addition to increasing hardware performance requirements, makes it difficult to argue in favor of achieving deterministic diagnostic coverage for such hardware. Summary of the Invention [Means for solving the problem]

[0010] The inventive approach described below provides an improvement in reducing the additional cost of additional hardware resources to achieve high coverage compared to existing approaches (e.g., redundant hardware for lockstep approaches or BIST circuits for self-tests). Another improvement concerns reducing the additional runtime overhead that may result from switching between functional application and self-tests or software-based safety tests at runtime without compromising diagnostic coverage. Similarly, the inventive approach may minimize or eliminate potential interruptions in the operational flow that may be required in both BIST and STL cases. Finally, the inventive approach may provide deterministic diagnostic coverage of AI hardware accelerators, since many existing accelerators do not include the capability to provide deterministic diagnostic coverage.

[0011] Neural networks (NNs) are known to be very computationally intensive applications, in particular they involve convolutions and matrix multiplications, which are essentially multiply-accumulate (MAC) operations. A typical Convolutional Neural Network (CNN) application may include convolutions, which account for less than 90% of its computational requirements. As a result, AI hardware accelerators or inference accelerators may have multiple MAC processors or equivalents, which consume a lot of the active processor area (excluding memory), total processing power, etc. Our solution exploits this unique feature of NNs and AI accelerators, whereby applications such as CNNs, Long Short-Term Memory (LSTM), Recurrent Neural Networks (RNNs), etc., are accelerated. In fact, the concept finds application for any architecture that includes MAC arrays, whereby high safety coverage of such arrays can be guaranteed.

[0012] Convolution or matrix multiplication is an operation that includes multiple MAC operations, which use an input matrix [x] and a filter matrix [w] as inputs to generate an output matrix [y]. Existing accelerators may perform these calculations using a cluster of processing engines (PEs) as a MAC array. Each PE is responsible for performing one MAC operation. The size of the MAC array is one of the factors that determine the parallelizability of the architecture, and is constrained by factors such as the availability of internal memory or input / output (IO) or power or silicon area to store input data and intermediate results of the calculation. For example, a simple 3×3 convolution requires at most 9 PEs to enable all operations of the matrix multiplication to occur in parallel. The operation is then repeated each time a successive convolution is performed. Repeated or continuous operations can often be found in real-time data processing applications, such as those use cases found in automotive applications.

[0013] In an embodiment, a CNN typically performs a convolution of a higher dimensional input matrix with a small filter matrix. The input matrix is ​​sliced ​​into small convolution matrices, and then each small input matrix is ​​dot-producted and summed with the filter matrix to complete the corresponding convolution operation. As a result, each input data value of the input matrix is ​​multiplied with different weights of the filter matrix in successive iterations, and ends up being individually multiplied with all weights of the filter matrix.

[0014] The inventive solution proposes to include one or more additional or shared processing engines in the MAC array, which may be called Security Processing Engines (SPEs), which in an embodiment perform three operations: first, a sum of intermediate multiplications of the same input data element with different weights of the filter matrix; note that since the corresponding multiplications have already been performed in the existing different PEs of the MAC array, the SPEs only need to perform the corresponding additions, where the same input data element has a different relative position for each successive slice of the matrix, and its position may advance with each addition cycle.

[0015] The SPE then needs to perform the corresponding multiplication of the unique input data element with a value corresponding to the weight, e.g., the sum of all the weights used in a normal convolution operation. A comparison of the results makes it possible to verify whether the calculation was performed correctly.

[0016] The number of additional processing elements in the SPE is a design choice, including consideration of the number of PEs and type of convolutions, as well as the kernel size the hardware is intended to support. In an embodiment, the system requires only one processing element in the SPE per convolution, or every other convolution. In other embodiments, multiple processing elements in the SPE serve to accelerate the calculations used to verify the operations of the MAC array. There is a trade-off: as the number of individual PEs verified per cycle increases, the number of operations performed by the SPE increases. Depending on the security requirements of the system, a system in which verification is performed at widely spaced, regular intervals can be envisioned.

[0017] The above concepts are valid for checking and verifying standard convolutions, as well as other variants of convolutions, matrix multiplication operations, etc. Various embodiments also support strided convolutions, where the convolution is performed with a stride greater than 1 (the default stride is equal to 1), point-wise convolutions (convolutions with kernel size equal to 1), and matrix multiplication operations.

[0018] It is noted that the inventive approach may provide improved coverage of hardware faults with fewer hardware resources compared to other approaches, which would result in a corresponding increase in the number of cycles required to perform a check or verification. Additional advantages of the embodiments may include no or limited requirements for redundant hardware, which may mean, for example, no significantly degraded HW compared to lock-step approaches, or no need for additional circuitry to perform BIST at startup. Similarly, there may be no need for an additional system reboot typically required after a BIST operation (because BIST loses enough of the state of internal registers during testing that they need to be reset after BIST). Also, the embodiments may provide significantly higher coverage compared to software approaches, and even some BIST methods. The embodiments may be well able to save cost, area, and power compared to other approaches, without degrading speed performance, and while maintaining high and continuous coverage. The embodiments may be able to ensure coverage comparable or close to that of lock-step approaches for NN accelerators or other convolution accelerators, with significantly reduced overall additional computation. The inventive approach provides new hardware architectures with hardware safety mechanisms integrated into the design to achieve high diagnostic coverage against random hardware faults, thereby ensuring that the hardware is capable of implementing safety-critical applications in an efficient manner.

[0019] An embodiment of the inventive approach may make it possible to reduce or even eliminate the need for periodic testing for the MAC array of the accelerator, which often occupies a significant portion of the area and computational needs of a hardware accelerator. Similarly, the embodiment makes it possible to have no interruption in the operational flow, to perform any safety-related tests, and to obtain deterministic coverage of the accelerator. The embodiment also provides computation and memory requirements for ensuring safety that are minor compared to the overall computation and memory requirements of the acceleration function. In fact, it may be envisaged to use the embodiment for IC testing during the IC production flow.

[0020] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings. [Brief description of the drawings]

[0021] [Figure 1] 1 shows an example of 1D convolution in a MAC array. [Diagram 2] This shows the addition of safety checks using SPE. [Diagram 3] 1 illustrates the operation of an SPE to perform a safety check. [Figure 4] 1 provides additional details of the operation of the SPE to perform safety checks. [Diagram 5] 1 provides additional details of the operation of the SPE to perform safety checks. [Figure 6] 1 illustrates an exemplary implementation of hardware for a Conv2D layer. [Figure 7] We present safety checks for different kinds of computations. [Figure 8] Details regarding the implementation of the sum of weights are provided. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] In the following description, numerous specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, those skilled in the art should recognize that various embodiments may be practiced without these specific details. In some cases, well-known structures, components, signals, computer program instructions, and techniques have not been shown in detail so as not to obscure the methodologies described herein. It should be understood that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements.

[0023] Disclosed herein are systems, apparatus, and methods for implementing a safety processor framework for safety-critical neural network applications. In one embodiment, the system includes a safety-critical neural network application, a safety processor, and an inference accelerator engine. The safety processor receives input images, test data (e.g., test vectors), and neural network specifications (e.g., layers and weights) from the safety-critical neural network application. The disclosed techniques may be used for (but are not limited to) AI inference. A trained AI neural network (NN) is compiled and executed on a dedicated processor, also called an accelerator.

[0024] Referring now to FIG. 1a, one embodiment of a computing system 100 for a typical use case of an accelerator such as a multiply-accumulate (MAC) array 110 is shown. At runtime, the NN binary (including the model graph, weights, and parameters) is executed on the hardware as a convolution. Input data may be provided to a first layer, for example from a sensor such as a camera, and the results of the first layer are used as input to a second layer, and so on, ultimately obtaining the output of the application. This can be shown as follows: y=f(i) Here, i is the input from the sensor, f(.) is the nonlinear NN model, and y is the output.

[0025] In an embodiment, the CNN may perform a convolution using a higher dimensional input matrix (e.g., a 1920x1080x3 camera input) with a small filter matrix (e.g., 3x3, 5x5, etc.). Thus, the input matrix is ​​sliced ​​into small convolution matrices [xk] equivalent to the filter matrix [w], and then each small input matrix and filter matrix are dot-producted and summed to complete the corresponding convolution operation. yk = Σ_i(xk_i*w_i) where i ranges over the size of the filter matrix, k ranges over the number of slices of the input matrix, xk_i is the i-th input data element of [xk] (the k-th slice of the input matrix [x]), w_i is the corresponding i-th weight in the filter matrix, yk is the output value for the kth slice of the input matrix.

[0026] As a result, each input data value in the input matrix is ​​multiplied by a different weight in the filter matrix in successive iterations until it has been individually multiplied by all of the weights in the filter matrix.

[0027] A sum of intermediate multiplications of the same input data element with different weights of the filter matrix is ​​shown in hardware as 110. The system includes multiple PEs arranged as a MAC array. Successive processor elements PE121, 122, 123 receive inputs i[j], i[j+1], i[j+2]. Each PE includes weights w[0,0], w[0,1], w[0,2], which are used as multiplicands by the corresponding PE. In this example, the three inputs may be multiplied by the corresponding weights, respectively, and the results are provided to the next processing element. In an embodiment, if one multiply-add operation defines a cycle, a result of P[0,0] is generated after each cycle, a result of P[0,1] is generated each cycle starting after the second cycle, and a result Y[j+1] starts to become available after the third cycle. In this example, the results are sent to memory bank 126.

[0028] This particular example is based on, but not limited to, a weight-fixed data flow, and typically a PE consists of a local memory element 113, called a register file, as shown in FIG. 1b, which stores weight elements that are multiplied with different input elements, which are transferred via memory banks to multipliers 128 and added to partial sum elements in adders 129 to complete one MAC operation. This illustrates only one exemplary implementation among many other possible variations. The inventive approach is not limited to a certain data flow or implementation, and finds application wherever continuous data processing is required. The embodiments are described in more detail in the following paragraphs and figures.

[0029] 1c shows the corresponding graphical diagram, where the input values ​​are multiplied with their respective weights and summed according to equation 130. These individual input values, shown as i[0]...i[4] at 131, are multiplied with respective weights shown as 132 and summed to calculate a result shown as 133.

[0030] In Fig. 2a, the same embodiment 210 of the computing system 200 for the typical use case of the accelerator is shown together with the embodiment of the security processing element SPE as 211. The system includes multiple PEs arranged as a MAC array. As in the system 100, a sum of intermediate multiplications of the same input data element with different weights of the filter matrix is ​​shown in hardware as 210. Successive processor elements PE 221, 222, 223 receive inputs i[0], i[1], i[2]. Each PE includes weights w[0,0], w[0,1], w[0,2], which are used as multiplicands by the corresponding PE. In this example, the three inputs may be multiplied by the corresponding weights respectively, and the results are provided to the next processing element. In an embodiment, where one multiply-add operation defines a cycle, a result P[0,0] is produced after each cycle, a result P[0,1] is produced each cycle starting after the second cycle, and a result Y[j+1] begins to become available after the third cycle. In this example, the results are sent to memory bank 226.

[0031] The computation of the SPE 211 in this exemplary embodiment, detailed in FIG. 2b, will now be described. First, the sum of intermediate multiplications of the same input data element with different weights of the filter matrix is ​​performed, as will be described in more detail below. Note that the SPE only needs to perform the addition of these values ​​potentially generated in different machine cycles, at 212, since the corresponding multiplications have already been performed and are available in the different PEs of the MAC. In this embodiment, it should be noted that the addition of intermediate multiplications, which have the same input data element as one of the non-operators of each of these multiplications. Second, in the internal memory of the register file 213, the sum of weights shown in equation 240 is stored, which is multiplied with the input data at 214. Note that the sum of weights may also be performed at compile time with additional weight elements stored in memory, and does not need to be calculated during real-time execution. This has the advantage of allowing determinism in the approach and improving real-time performance requirements. Third, a comparison of the results of 212 and 214 is performed at 215 to check for any faults that lead to incorrect functioning of the hardware. A more detailed description of the different stages is given below.

[0032] Referring to Figure 3, this details the functionality of the SPE for performing safety checks on hardware. The figure shows a standard 1d convolution operation, as shown at 330, in which a stream of input data and weights are convolved to produce streams of outputs 331, 332, and 333, potentially at different machine cycles running on computing system 100.

[0033] Figures 4a and 4b show additional details of the embodiment. As shown in Figure 4a, for example, adding an SPE 411 to the computing system 200 provides the possibility to perform safety checks on the hardware used to perform these operations and detect possible failures. First, a summation of intermediate multiplication results 335, 336, and 337 available at different PEs in different machine cycles is performed in 412. As shown, the same input data element has a different relative position in successive slices of the matrix or multiply-accumulate cycles, denoted here by xi_i. The resulting partial sums are s1=Σ_i(xi_i*w_i) It can be expressed as: where i ranges up to the size of the filter matrix. Note that the dimension K of the input slice also varies with i, xi_i is the i-th input data element of [xi] (the i-th slice of the input matrix), w_i is the corresponding weight from the filter matrix, and s1 is the corresponding sum.

[0034] Second, the SPEs in this exemplary embodiment need to perform the following multiplications: s2=xi_i*W where xi_i denotes the corresponding data from the input matrix, and W is Σw_i, the weight sum stored in local memory (in an embodiment, typically calculated offline during compile time). A representation and graphic representation of this operation is shown as FIG. 4b. The weight sum from 440 is stored locally in register file 413. The multiplication shown at 441 in FIG. 4b is performed at 414.

[0035] In the third step of this exemplary embodiment, a comparison of s1 and s2 is performed at 415 to verify if the calculation was performed correctly. The sum result of 335, 336, and 337 should be equal to the result of 441. If there is a mismatch between the two values, this may be taken as an indication of a hardware failure. The evaluation of the mismatch may be a simple comparison of numbers, or a comparison of values ​​depending on the operation being performed.

[0036] Note that in this embodiment, K operations of an SPE may be required to verify K PE operations or K operations performed by a PE, depending on the particular configuration of the MAC array.

[0037] 5, the comparison of values ​​is shown for steps j=1 to 3. In the figure, SPE 511 receives intermediate results 335, 336, and 337 for summing from PEs 523, 522, and 521, respectively. Thus, in its comparison operation 415, SPE 511 should be able to detect any errors in the calculation caused due to single point hardware failures.

[0038] The SPE 511 itself may or may not be a separate specialized PE used to perform the safety checks and calculations necessary to verify correct operation. The inventive concept may be envisioned in a system where there are separate SPEs with dedicated connections and memory. Similarly, the inventive concept may be envisioned in a system with a pool of PEs, most of which are used for ongoing calculation of the MAC array, at least one of which is used as an SPE to check the operations of other PEs. Combinations are also possible, where multiple PEs are available and there are dedicated data connections for SPE operations, or a global data bus is used for transport, and there is a separate SPE to perform the check operations as well as a dedicated array of PEs. Similarly, the comparison of results may occur in the SPEs or in a separate processor. In an embodiment, the comparison of results occurs in a central processor or a system control processor.

[0039] The SPEs may also be envisioned as blocks implemented in software or as blocks integrated into an inference or AI model, for example at compile time.

[0040] In the exemplary embodiment described herein, for a 3x3 convolution, there are 9 multiplication operations that the SPE performs. In addition, for a comparison calculation, there is 1 multiplication and 9 additions, and a comparison step that compares the two results. As the size of the convolution increases, the number of operations for the SPE also increases, but the additional operations may be performed over a longer period of time or over more computation cycles. For example, for a 5x5 convolution, this means an additional 25 multiply-and-accumulate operations, 1 multiplication and 25 additions, and 1 result comparison to identify a fault. An exemplary implementation of such a system is shown in FIG. 6. In FIG. 6a, a typical computing system 600 is shown, including an array of PEs 610, including an SPE 611, as well as memory banks 625 and 626. A single convolution layer with input data 605, weights 606, and the resulting output 607 of the layer is shown in FIG. 6b. In a given scenario, assuming 605 is a high-resolution feature map of resolution 1024x1024x2 and a weight kernel size of 3x3, this example produces an output 607 of dimensionality 1024x1024x1 with a total of 2x3x3x1=18 parameters or weights (e.g., this may represent the last layer of a semantic segmentation model that classifies each pixel into two classes: "object" and "background"). In one implementation, the convolutional layer may perform 18 MAC operations per machine cycle in a given array 610 of 25 PEs. A single additional SPE 611 connected to all PEs in 610 should be able to successfully perform the necessary safety checks, as detailed in Figures 2-5. In another implementation, multiple blocks of the type shown as 610 may require one or more SPEs in the overall computing system. Other implementations may use different combinations of SPEs in the computing system. There may be a tradeoff between the speed of the verification or checking operation and the number of SPEs used.

[0041] FIG. 7 shows the inventive concept applied to perform safety checks on different types of computations. The SPE in the embodiment performs the same three-step operation as described above. However, possible variations and flexibility for performing safety checks on different operations are shown. One such variation is shown in FIG. 7a, where the convolution layer has a multiple filter kernel 735 that generates multiple features as output 737 according to equation 730. The weight sequence is shown as 736. In such a scenario, shown as FIG. 7b, the sum of weights used in the SPE can be defined as 740 instead of 440, so that the result 741 can be compared to the sum of elements 736 similar to the operation of the SPE described in 411. This approach or a combination thereof may be more suitable for certain types of convolutions where the same input element is not multiplied with each weight element of the kernel unlike the use cases shown in 335, 336, and 337. For example, convolutions with strides greater than 1, point-wise convolutions, dilated and atrous convolutions are some types of operations that can be checked using the inventive concepts.

[0042] Another variation of the calculation, typified by matrix multiplication, is shown in Figures 7c and 7d, where the inventive concept of safety checking can be integrated into a design with a similar SPE and its three steps, but with a modified function for the sum of weights shown in 750, so that the result 751 can be compared to the sum of elements 752.

[0043] Referring to FIG. 8, FIG. 8 gives details on the implementation of the weight sum (240, 440) shown in the previous figures. The weight sum may be calculated at compile time itself generating additional parameters, thereby not requiring the weight sum to be performed during execution. The number of these additional parameters obtained by performing the weight sum of the existing network may be implemented in various combinations depending on the use case. As shown in FIG. 8a, when there is a set of weight kernels 810, 811, 812 for an input activation value 801, the output generated after convolution results in three channels 820, 821, and 822, respectively. For such a convolution layer, as shown in FIG. 8b, the weight sum 840 may be any of the three separate parameters shown in 841, or it may be a single parameter 842, the sum of all weight kernels 810, 811, and 812, or any other different weight combination. The choice of implementation depends on a given use case, where 841 requires more memory compared to 842 due to the larger number of additional parameters, but has less latency to perform the validity checks.

[0044] In embodiments, iterative operations combined with systolic movement of data create the possibility of using a separate SPE, containing one or more processing elements, to perform a subset of the operations and thus verify the operations of the PE array. In fact, as shown in the previous examples, embodiments may address almost any form of convolution, such as depth-wise or group convolution, or dilation and atlas convolution, in addition to stride-wise and point-wise convolution.

[0045] In embodiments and configurations, the inference accelerator engine or MAC array implements one or more layers of a convolutional neural network. For example, in an embodiment, the inference accelerator engine implements one or more convolutional layers and / or one or more fully connected layers. In another embodiment, the inference accelerator engine implements one or more layers of a recurrent neural network. In general, an "inference engine" or "inference accelerator engine" is defined as hardware and / or software that, for example, receives image data and generates one or more label probabilities for the image data. In some cases, the "inference engine" or "inference accelerator engine" is referred to as a "classification engine" or "classifier." In another embodiment, the inference accelerator engine may analyze an image or video frame to generate one or more label probabilities for the frame. Potential use cases include at least eye tracking, object recognition, point cloud inference, ray tracing, light field modeling, depth tracking, and the like.

[0046] The inference accelerator engine may be used by any of a variety of different safety-critical applications, varying according to the implementation. For example, in one implementation, the inference accelerator engine may be used in an automotive application, where the inference accelerator engine may control one or more functions of a self-driving automobile (i.e., an autonomous automobile), a driver-assisted automobile, or an advanced driver assistance system. In other implementations, the inference accelerator engine may be trained and customized for other types of use cases. Depending on the implementation, the inference accelerator engine may generate probabilities of classification results for various objects detected in an input image or video frame.

[0047] The memory subsystems 125, 126 may include any number and type of memory devices, and the two memory subsystems 125 and 126 may be combined as a single memory or in any other configuration. For example, the types of memory in the memory subsystems may include high bandwidth memory (HBM), non-volatile memory (NVM), dynamic random access memory (DRAM), static random access memory (SRAM), NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), etc. The memory subsystems 125 and 126 may be accessible by the inference accelerator engine and by other processors. The I / O interface may include any type of data transport bus or channel (e.g., peripheral component interconnect (PCI) bus, PCI-Extension (PCI-X), PCIE (PCI Express) bus, Gigabit Ethernet (GBE) bus, Universal Serial Bus (USB)).

[0048] In some embodiments, the entire computing system 100, 200, 500, 600, or one or more portions thereof, are integrated into a robotic system, a self-driving car, an autonomous drone, a surgical tool, or other types of mechanical devices or systems. In fact, the inventive concepts find application in any system where hardware safety, security, and / or reliability is needed or desired. It should be noted that the number of components of the computing system 100 varies from embodiment to embodiment. For example, in other embodiments, there are more or less of each component than shown in the drawings. It should also be noted that in other embodiments, the computing system 100 includes other components not shown in FIG. 1 and other drawings. Additionally, in other embodiments, the computing system 100 is constructed in a manner other than that shown in the drawings.

Claims

1. A system (100, 600) comprising: a multiply-accumulate (MAC) array (110, 610) including a plurality of processing elements (PEs, 122); and a safety processing element (SPE, 211, 411), the multiply-accumulate array configured to perform matrix multiplication operations in parallel across a plurality of processing elements to perform convolutions on the multiply-accumulate array, perform a subset of the matrix multiplication operations on the safety processing element, compare results of the operations performed on the processing elements and the safety processing element, and determine a fault condition based on a mismatch in the results; said safety processing element performs a first operation of calculating a sum of multiplications of identical input data elements with different weights of a filter matrix, a second operation being a multiplication of said identical input data elements with a value corresponding to the sum of the different weights, and a third operation being a comparison of a result of said first operation with a result of said second operation. system.

2. 2. The system of claim 1, wherein the comparison of the results of the subset of the matrix multiplication operations enables verifying the operations of a plurality of processing elements of the multiply-accumulate array.

3. 3. The system of claim 1, wherein the comparison of the results of the subset of the matrix multiplication operations enables verifying the operations of all processing elements of the multiply-accumulate array.

4. 3. The system of claim 1 or 2, wherein the safety processing element receives the input values ​​of an input vector in sequence and multiplies each input value by a weight.

5. The system of claim 4 , wherein the safety processing element receives strided input values ​​sequentially, the strided input values ​​comprising a subset of values ​​of a given input vector.

6. 3. The system of claim 1, wherein the multiply-accumulate array and the safety processing element are part of an inference accelerator system implementing safety-critical inference applications that includes additional or shared hardware for performing inference accelerator validation, and where validation hardware has a lower processing bandwidth than the inference accelerator, thereby requiring more machine cycles per calculation to generate the result for comparison.

7. The convolutional layer (730) calculates y f,j+1 = Σi j+n *w f,n 3. The system of claim 1 or 2, wherein n=[0, k] is defined, where k is a natural number.

8. The sum of the weights (740) used by the safety processing element is W n =Σw f,n The system of claim 7, wherein:

9. The system of claim 1 or 2, wherein the calculation of the sum of weights is performed at compile time.

10. 3. The system of claim 1 or 2, wherein the sum of weights is calculated based on a full weight kernel (842) of a convolutional layer, a subset (841) of the full weight kernel (842) of the convolutional layer, or any other combination.

11. Continually executing matrix arithmetic calculations including a set of multiply-add operations on an array of processing elements (PEs, 122), separately and continuously executing a subset of the matrix arithmetic calculations as verification operations, and comparing results, the comparison of results enabling verification of whether the matrix arithmetic calculations have been performed correctly; Including, said verification operation comprising: a first operation of performing a sum of multiplications of identical input data elements with different weights of a filter matrix; a second operation of performing a multiplication of said identical input data elements with a value corresponding to the sum of the different weights; and a third operation of performing a comparison between a result of said first operation and a result of said second operation. method.

12. The method of claim 11 , wherein the separately and continuously executing a subset of the matrix operation computations is performed in a separate safety processing element (SPE, 211, 411).