Hardware accelerator device, corresponding system and operation method

By designing a memory-based hardware accelerator device, using a double-buffered memory library and a low-complexity interconnection network, the functional safety level is improved without increasing hardware resource duplication, solving the problems of excessive silicon area occupation and power consumption of hardware accelerators in on-chip systems, and meeting the ASIL-D level requirements of automotive safety-critical electronic components.

CN114594991BActive Publication Date: 2025-10-03STMICROELECTRONICS SRL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111457679.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-05
Filing Date
2021-12-02
Publication Date
2025-10-03
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

Existing hardware accelerators have problems with silicon area occupation and excessive power consumption in on-chip systems, making it difficult to meet the ASIL-D level requirements of automotive safety-critical electronic components, especially in the implementation of complex data processing algorithms.

Method used

A memory-based hardware accelerator device is designed, including processing circuits, data memory libraries, a control unit for configuration registers, and an interconnection network. Functional safety is achieved through a configurable lockstep control unit, supporting ASIL-X level safety architecture. A double-buffered memory library and a low-complexity interconnection network are used to provide redundant data paths to meet safety requirements.

Benefits of technology

This improves the functional safety level of the hardware accelerator, reduces silicon area, improves computing performance, and meets ASIL-D safety requirements without increasing hardware resource duplication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114594991B_ABST
    Figure CN114594991B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to hardware accelerator devices, corresponding systems, and operating methods. A device includes a set of processing circuits arranged in a subset, a set of data memory banks coupled to a memory controller, a control unit, and an interconnection network. The processing circuits are configurable to read first input data from the data memory bank via the interconnection network and the memory controller, process the first input data to generate output data, and write the output data to the data memory bank via the interconnection network and the memory controller. The hardware accelerator device includes a set of configurable lockstep control units that connect the processing circuit interface to the interconnection network. Each configurable lockstep control unit is coupled to a subset of the processing circuits and is selectively activated to operate in a first operating mode or in a second operating mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of Italian Patent Application No. 102020000029759, filed on December 3, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to hardware accelerators and, in particular embodiments, to automotive-grade hardware accelerators for accelerating complex data processing algorithms. Background Art

[0004] Real-time digital signal processing systems can involve processing a significant amount of data per unit time. For example, with growing demand in the automotive sector, such systems can be used to process video data, image data, radar data, wireless communication data, or a combination thereof. In various applications, this processing can be very demanding for purely core-based implementations (i.e., implementations involving general-purpose microprocessors or microcontrollers running processing software).

[0005] Therefore, the use of hardware accelerators is becoming increasingly important in certain areas of data processing, as it helps speed up the computation of certain algorithms. Compared to core-based implementations, a properly designed hardware accelerator can reduce the processing time of specific operations.

[0006] In particular, there is growing interest in the automotive field to use hardware accelerators to implement passive or active safety systems that can prevent or reduce injuries to vehicle drivers and passengers. By way of example only, such safety systems can include modern systems such as forward collision warning, blind spot monitoring, and automatic emergency braking, as well as more conventional systems such as airbags, anti-lock braking systems (ABS), and the like.

[0007] Safety-critical electronic components used in the automotive sector may be subject to certain safety requirements, for example, according to the safety standard ISO 26262. The ISO 26262 standard provides a generic means for measuring and recording the safety level of electrical and electronic (E / E) systems, which can be classified according to certain Automotive Safety Integrity Levels (ASILs), for example, from ASIL-A (which meets fewer safety requirements) to ASIL-D (which meets more safety requirements).

[0008] It would therefore be advantageous to provide automotive-grade hardware accelerators designed to accelerate certain complex data processing algorithms - such as Fast Fourier Transform (FFT), Finite Impulse Response (FIR) filters, Artificial Neural Networks (ANN), etc., which are increasingly used in modern Advanced Driver Assistance Systems (ADAS) to (for example) comply with certain safety requirements (e.g., ASIL-D requirements of the ISO26262 standard).

[0009] In the field of hardware accelerators (e.g., implemented in a system-on-chip (SoC)), functional safety can be achieved by replicating internal hardware resources according to a conventional lockstep configuration, which, especially in the case of complex hardware accelerators, may result in an increase in the silicon area occupied or the power consumption, or both, of the hardware accelerator. Summary of the Invention

[0010] It is an object of one or more embodiments to provide a hardware accelerator device that addresses one or more of the above-mentioned disadvantages.

[0011] According to one or more embodiments, this object may be achieved by a hardware accelerator device having the features set out in the appended claims.

[0012] One or more embodiments may be directed to corresponding systems (eg, system-on-chip integrated circuits including hardware accelerator devices). One or more embodiments may be directed to corresponding methods of operation.

[0013] According to one or more embodiments, a hardware accelerator device is provided, which may include: a set of processing circuits arranged in subsets (e.g., pairs) of the processing circuits, a set of data memory banks coupled to a memory controller, a control unit including configuration registers that provide storage space for configuration data of the processing circuits, and an interconnection network.

[0014] The processing circuitry may be configured according to the configuration data to read first input data from the data memory repository via the interconnect network and the memory controller, process the first input data to generate output data, and write the output data to the data memory repository via the interconnect network and the memory controller.

[0015] The hardware accelerator device may include a set of configurable lockstep control units that interface processing circuitry to an interconnect network.

[0016] Each configurable lock-step control unit in the set of configurable lock-step control units may be coupled to a subset of processing circuits in the set of processing circuits.

[0017] Each configurable lockstep control unit can be selectively activated to: operate in a first operating mode, wherein the lockstep control unit is configured to compare data read requests, data write requests, or both issued by the first processing circuit and the second processing circuit in the corresponding subset of processing circuits toward the memory controller to detect a fault; or operate in a second operating mode, wherein the lockstep control unit is configured to propagate data read requests, data write requests, or both issued by the first processing circuit and the second processing circuit in the corresponding subset of processing circuits toward the memory controller.

[0018] Thus, one or more embodiments may provide a memory-based hardware accelerator device (e.g., an enhanced data processing architecture (EDPA)) including a safety architecture that facilitates configuring the memory-based hardware accelerator device according to an ASIL-X level (e.g., ASIL-B, ASIL-C, ASIL-D) (statically, dynamically, or both).

[0019] According to one or more embodiments, a memory-based hardware accelerator device may be used to accelerate the computation of certain safety-related data processing algorithms, such as those employed in modern advanced driver assistance systems or other safety-critical applications.

[0020] For example, one or more embodiments may be applied in real-time processing systems that accelerate computationally demanding operations (e.g., vector / matrix products, convolutions, FFTs, radix-2 butterfly algorithms, multiplication of complex vectors, trigonometric, exponential, or logarithmic functions, etc.).

[0021] One or more embodiments aim to provide a certain level of functional safety (e.g., ASIL-D level) without relying on duplication of hardware resources in hardware accelerator devices. Thus, one or more embodiments can improve the trade-off between silicon area usage and performance of hardware accelerators. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which:

[0023] Figure 1 is a block diagram of an embodiment electronic system;

[0024] Figure 2 is a block diagram of an electronic device implementing an embodiment of a hardware accelerator;

[0025] Figure 3 is a block diagram of an embodiment lockstep architecture for use in a hardware accelerator;

[0026] Figure 4 is a block diagram of an embodiment phase shift generation circuit for use in a hardware accelerator;

[0027] Figure 5 is a block diagram of an embodiment memory system for use in a hardware accelerator;

[0028] Figure 6A is a block diagram of an embodiment memory watchdog architecture for use in a corresponding hardware accelerator;

[0029] Figure 6B is a block diagram of another embodiment memory watchdog architecture for use in a hardware accelerator; and

[0030] Figure 7 is a block diagram of an embodiment built-in self-test (BIST) circuit for use in a hardware accelerator. DETAILED DESCRIPTION

[0031] In the following description, one or more specific details are provided to provide a deeper understanding of the examples of the embodiments of the present description. The embodiments may be obtained without one or more of the specific details or with other methods, components, materials, etc. In other cases, well-known structures, materials, or operations are not shown or described in detail so that certain aspects of the embodiments are not obscured.

[0032] References to "an embodiment" or "one embodiment" in the context of this description are intended to indicate that a particular configuration, structure, or characteristic described with respect to that embodiment is included in at least one embodiment. Thus, phrases such as "in an embodiment" or "in one embodiment" that may appear in one or more points of this description are not necessarily referring to the same embodiment. Furthermore, in one or more embodiments, the particular configurations, structures, or characteristics may be combined in any appropriate manner.

[0033] Throughout the drawings appended hereto, the same parts or elements are indicated with the same references / numerals, and the corresponding descriptions will not be repeated for the sake of brevity.

[0034] References used herein are for convenience only and do not limit the scope or extent of the embodiments.

[0035] By way of introduction to the detailed description of exemplary embodiments, reference is made to the disclosure of the Italian patent application indicated below, filed by the same applicant (and not yet available to the public at the time of filing the present application), the content of which is incorporated herein by reference in its entirety: Italian Patent Application No. 102020000009358 filed on April 29, 2020, which briefly discloses a hardware accelerator device comprising a set of (runtime) configurable processing circuits, a set of data memory banks, and a control unit, wherein the configurable processing circuits are configured to receive data from the data memory banks via an interconnection network according to configuration data received from the control unit The library reads data and writes data to a data memory library; Italian patent application No. 102020000009364 filed on April 29, 2020, which briefly discloses a method for accessing a memory by supporting vector access with programmable stride and a memory access scheme to the memory, the method being suitable for a hardware accelerator device comprising a collection of processing circuits; and Italian patent application No. 102020000016393 filed on July 7, 2020, which briefly discloses a method for storing and extracting rotation factors in a memory of a hardware accelerator device for efficient calculation of a fast Fourier transform algorithm.

[0036] Figure 1 An electronic system 1, such as a system on a chip (SoC), according to one or more embodiments is illustrated. The electronic system 1 may include various electronic circuits, such as, for example, a central processing unit 10 (CPU, e.g., a microprocessor), a main system memory 12 (e.g., system RAM—random access memory), a direct memory access (DMA) controller 14, and a hardware accelerator device 16.

[0037] The hardware accelerator device 16 may be designed to support the execution of (basic) arithmetic functions.The electronic circuits in the electronic system 1 may be connected via a system interconnect network 18 (eg a SoC interconnect).

[0038] like Figure 1 As illustrated, the hardware accelerator device 16 may include a plurality of processing elements, including a number P of processing elements 1600, 1601, ..., 160 P-1 (also collectively referred to as reference numeral 160 in this description), and a set of local data storage libraries, optionally a number Q=2*P local data storage libraries M0, ..., M Q-1 (Also collectively referred to as reference M in this description).

[0039] The hardware accelerator device 16 may further include a local control unit 161, a local interconnect network 162, a local data memory controller 163, a local ROM controller 164 coupled to a set of local read-only memories, and optionally a number P of local read-only memories 1650, 1651, ..., 165 P-1 (also collectively referred to in this description as reference numeral 165), and a local configuration memory controller 166 coupled to a set of locally configurable coefficient memories, and optionally a number P of locally configurable coefficient memories 1670, 1671, ..., 167 P-1 (Also collectively referred to as reference numeral 167 in this description.) The memory 167 may include a volatile memory (eg, RAM memory) and / or a nonvolatile memory (eg, PCM memory).

[0040] Different embodiments may include a different number P of processing elements 160 and / or a different number Q of local data memory repositories M. For example, P may be equal to 8 and Q may be equal to 16.

[0041] Processing element 160 may support (eg, based on appropriate static configuration) different processing capabilities (eg, floating point single precision 32-bit, fixed point / integer 32-bit, or 16-bit or 8-bit with parallel computation or vector mode).

[0042] Processing element 160 may include corresponding internal direct memory access (DMA) controllers 1680, 1681, ... 168 P-1 (Also collectively referred to in this description as reference numeral 168.) The processing element 160 may be configured to retrieve input data from the local data memory repository 24 and / or from the main system memory 12 via a corresponding direct memory access controller 168. The processing element 160 may thus process the retrieved input data to generate output data. The processing element 160 may be configured to store the processed output data in the local data memory repository 24 and / or the main system memory 12 via the corresponding direct memory access controller 168.

[0043] Additionally, processing element 160 may be configured to retrieve input data from local read-only memory 165 and / or from local configurable coefficient memory 167 to perform such refinement.

[0044] Providing a set of local data storage libraries M can facilitate parallel processing of data and reduce memory access conflicts. The local data storage library M can be provided with buffering (e.g., double buffering), which can facilitate recovery of memory upload time (write operation) and / or download time (read operation).

[0045] In an embodiment, each local data memory bank may be replicated so that data can be read from one of the two memory banks and (new) data can be simultaneously stored in the other memory bank (e.g., for later processing). Thus, moving data may not negatively impact computational performance since it may be masked. A double buffering scheme for the local data memory banks M may be advantageous in combination with data processing in a streaming mode or back-to-back (e.g., as applicable to an FFT N-point processor configured to craft a continuous sequence of N data inputs).

[0046] Local control unit 161 may include a register file that includes information for setting the configuration of processing element 160. For example, local control unit 161 may set processing element 160 to execute a specific algorithm directed by a host application running on central processing unit 10. In one or more embodiments, local control unit 161 may thus (e.g., dynamically) configure each of processing elements 160 to compute a specific (basic) function and may configure each of the corresponding internal direct memory access controllers 168 with a specific memory access scheme and cycle time.

[0047] The local interconnect network 162 may include a low-complexity interconnect system, such as a bus network based on a known type (such as an AXI4-based interconnect). For example, the data parallelism of the local interconnect network 162 may be 64 bits, and the address width may be 32 bits.

[0048] Local interconnect network 162 may be configured to connect processing element 160 to local data memory bank 24 and / or main system memory 12. Additionally, local interconnect network 162 may be configured to connect local control unit 161 and local configuration memory controller 166 to system interconnect network 18.

[0049] In an embodiment, the interconnection network 162 may include: P master ports MP0, MP1, ..., MP P-1 A set of P slave ports SP0, SP1, ... SP P-1 A set of slave ports (also collectively referred to as reference numerals SP in this description), each of which is coupled to a local data memory bank M via a local data memory controller 163; another pair of ports, including a system master port MP P and the system slave port SP P, configured to be coupled to the system interconnect network 18 (eg, to receive instructions from the central processing unit 10 and / or access data stored in the system memory 12); and yet another slave port SP P+1 , coupled to the local control unit 161 and to the local configuration memory controller 166 .

[0050] In one or more embodiments, the interconnection network 162 may be fixed (ie, not reconfigurable).

[0051] In an exemplary embodiment (see, e.g., Table I provided at the end of the description, where an "X" symbol indicates an existing connection between two ports), interconnect network 162 may implement the following connections: P master ports MP0, MP1, ..., MP2, coupled to processing element 160. P-1 Each master port in the local data memory controller 163 may be connected to a corresponding slave port SP0, SP1, . . . SP1 coupled to the local data memory controller 163. P-1 and is coupled to the system master port MP of the system interconnection network 18 P may be connected to the slave port SP coupled to the local control unit 161 and the local configuration memory controller 166 P+1 .

[0052] In another exemplary embodiment (see, for example, Table II provided at the end of the description, where an "X" symbol indicates an existing connection between two ports), the interconnection network 162 may further implement the following connections: P master ports MP0, MP1, ... MP P-1 Each master port in the system may be connected to a system slave port SP coupled to the system interconnect network 18. P In this manner, connectivity may be provided between any processing element 160 and the SoC via system interconnect network 18 .

[0053] In another exemplary embodiment (see, for example, Table III provided at the end of the description, where an "X" symbol indicates an existing connection between two ports, and an "X" between brackets indicates an optional connection), the interconnection network 162 may further implement the following connections: P Can be connected to slave ports SP0, SP1, ..., SP P-1 At least one slave port (here, in the set of P slave ports SP0, SP1, ... SP P-1 In this way, the first slave port SP0 can be used at the master port MP P Provides connection between (any) slave port. According to the specific application of system 1, the master port MP PThe connection can be extended to multiple (eg all) slave ports SP0, SP1, ... SP P-1 The connection of the master port MPP to at least one of the slave ports SP0, SP1, ... SPP-1 can be used (only) to load input data to be processed into the local data memory banks M0, ... M Q-1 As long as all memory banks can be accessed via a single slave port. Only one slave port can be used to load input data, while processing data by parallel computing can utilize multiple (e.g., all) slave ports SP0, SP1, ... SP P-1 advantages.

[0054] In one or more embodiments, local data memory controller 163 may be configured to arbitrate (e.g., by processing element 160 ) access to local data memory bank M. For example, local data memory controller 163 may use a memory access scheme selected based on a signal received from local control unit 161 (e.g., for computation of a particular algorithm).

[0055] In one or more embodiments, the local data memory controller 163 may convert incoming read / write transaction bursts (e.g., AXI bursts) generated by the read / write direct memory access controller 168 into read / write memory access sequences based on a particular burst type, burst length, and memory access scheme.

[0056] In one or more embodiments, local read-only memory 165 accessible by processing element 160 via local ROM controller 164 may be configured to store digital factors and / or fixed coefficients used for implementation of a particular algorithm or operation (e.g., twiddle factors or other complex coefficients for FFT calculations). Local ROM controller 164 may implement a particular address scheme.

[0057] In one or more embodiments, local configurable coefficient memory 167, accessible by processing element 160 via local configuration memory controller 166, may be configured to store application-dependent digital factors and / or coefficients (e.g., coefficients for implementing FIR filters or beamforming operations, weights of a neural network, etc.) that may be configured by software. Local configuration memory controller 166 may implement a specific address scheme.

[0058] In one or more embodiments, local read-only memory 165 and / or local configurable coefficient memory 167 may advantageously be divided into a number P of banks equal to the number of processing elements 160 included in hardware accelerator device 16. This may help avoid conflicts during parallel computations.

[0059] Figure 2is a circuit block diagram of an embodiment processing element 160 and associated connections to a local ROM controller 164, a local configuration memory controller 166, and a local data memory vault M (wherein dashed lines schematically indicate reconfigurable connections between the processing element 160 and the local data memory vault M via the local interconnect network 162 and the local data memory controller 163).

[0060] like Figure 2 The illustrated processing element 160 can be configured to: receive a first input signal P (e.g., a digital signal indicating a binary value from a local data memory library M) via a corresponding read direct memory access 2000 and a buffer register 2020 (e.g., a FIFO register); receive a second input signal Q (e.g., a digital signal indicating a binary value from a local data memory library M) via a corresponding read direct memory access 2001 and a buffer register 2021 (e.g., a FIFO register); receive a first input coefficient W0 (e.g., a digital signal indicating a binary value from a local read-only memory 165); and receive a second input coefficient W1, a third input coefficient W2, a fourth input coefficient W3, and a fifth input coefficient W4 (e.g., digital signals indicating corresponding binary values ​​from a local configurable coefficient memory 167).

[0061] In one or more embodiments, processing element 160 may include a number of read direct memory accesses 200 equal to the number of input signals P,Q.

[0062] It should be understood that in different embodiments, the number of input signals and / or input coefficients received at processing element 160 may vary.

[0063] The processing element 160 may include a computational circuit 20 that is configurable (possibly at runtime) to process input values ​​P, Q and input coefficients W0, W1, W2, W3, W4 to generate a first output signal X0 (e.g., a digital signal indicating a binary value to be stored in a local data memory repository M via a corresponding write direct memory access 2040 and a buffer register 2060 (such as a FIFO register)) and a second output signal X1 (e.g., a digital signal indicating a binary value to be stored in a local data memory repository M via a corresponding write direct memory access 2041 and a buffer register 2061 (such as a FIFO register)).

[0064] In one or more embodiments, processing element 160 may include a number of write direct memory accesses 204 equal to the number of output signals X0 , X1 .

[0065] In one or more embodiments, programming of read and / or write direct memory access 200 , 204 (which may be included in direct memory access controller 168 ) may be performed via an interface (e.g., an AMBA interface) that may allow access to internal control registers located in local control unit 161 .

[0066] Additionally, processing element 160 may include ROM address generator circuitry 208 coupled to local ROM controller 164 and memory address generator circuitry 210 coupled to local configuration memory controller 166 to manage data retrieved therefrom.

[0067] The computing circuit 20 may include a collection of (e.g., highly parallelized) processing resources, such as four complex / real multiplier circuits, two complex addition and subtraction circuits, two accumulator circuits, and two activated nonlinear function circuits, which can be reconfigured and coupled (e.g., through multiplexers) to form different data paths, where different data paths correspond to different mathematical operations.

[0068] Reference again Figure 1 One or more embodiments of the hardware accelerator device 16 may include a plurality of read / write lockstep units, for example, a number R=P / 2 lockstep units 1690, ..., 169 R-1 (Also collectively referred to as reference numeral 169 in this description).

[0069] In one or more embodiments, each lockstep unit 169 may be configured to couple a pair of processing elements 160 to an interconnect network 162. For example, Figure 1 As illustrated, the first lockstep unit 1690 may be coupled to the DMA controllers 1680 and 1681 of the first processing element 1600 and the second processing element 1601, and the last lockstep unit 1690 may be coupled to the DMA controllers 1680 and 1681 of the first processing element 1600 and the second processing element 1601. R-1 can be coupled to the final processing element 160 P-1 and the penultimate processing element 160 P-2 DMA controller 168 P-1 and 168 P-2 .

[0070] Each lockstep unit 169 can be selectively configured (e.g., by setting registers of the local control unit 161) to pass data between the corresponding processing element and the interconnect network so that the processing elements in the corresponding pair can operate according to two different operating modes: in a first mode ("pseudo-lockstep mode"), the two processing elements operate in parallel, where the first processing element in the pair acts as a "functional" circuit and the second processing element in the pair acts as a "shadow" circuit that replicates the operations performed by the functional circuit so that safety-related algorithms can be calculated using a target level of functional safety (e.g., ASIL-D level); and in a second mode ("high-speed mode"), the two processing elements operate independently of each other as instructed by the control unit 161, so that non-safety-related algorithms can be calculated at a higher speed without meeting ASIL-D safety requirements.

[0071] Thus, when high functional safety is required, for each pair of processing elements 160 in the hardware accelerator device 16 , the “functional” computation path and the “shadow” computation path may be dynamically configurable for a specific scenario / algorithm.

[0072] Figure 3 is by reading the lockstep unit 169 0,r and write lockstep unit 169 0,w A circuit block diagram of an embodiment of a processing element pair 1600, 1601 coupled to an interconnect network 162. It will be understood that the read and write portions of the lockstep unit 1690 are shown as separate circuit blocks for clarity only, and that in one or more embodiments, the read and write portions of the lockstep unit 169 may be implemented in a single circuit. The same applies to the interconnect network 162, which is shown for clarity only. Figure 3 The interconnection network 162 is shown in FIG. 1 using two separate blocks.

[0073] like Figure 3 As illustrated, the hardware accelerator 16 according to one or more embodiments may thus include: at least one pair of processing elements 1600, 1601 configurable to support two redundant paths for data processing (a functional path and a shadow path); an interconnect network 162 configured to support data routing from / to the functional processing elements and the shadow processing elements; and a read lockstep unit 169 for each pair of processing elements (the functional processing element and the shadow processing element). 0,r and write lockstep unit 169 0,w , the lockstep unit is configured to check and / or protect data delivery from / to the local memory M and / or the system memory 12 (end-to-end) via the local interconnect network 162.

[0074] In one or more embodiments, (by Figure 3The functional path (illustrated by the solid line in Figure 3 The shadow path (illustrated by the dotted line in ) is related to the redundant path, and the interconnect access of the redundant path is gated in lockstep mode.

[0075] As an example, read the lockstep cell 169 0,r A functional read address channel RA may be provided between the interconnect network 162 and the first processing element 1600 f and a shadow read address channel RA between the interconnection network 162 and the second processing element 1601 s (Read burst request / address / acknowledge). Read lockstep unit 169 0,r A functional read response channel RR between the interconnect network 162 and the first processing element 1600 may also be provided. f and a shadow read response channel RR between the interconnection network 162 and the second processing element 1601 s (Read data).

[0076] Still as an example, write lockstep unit 169 0,w A functional write address channel Wa may be provided between the interconnect network 162 and the first processing element 1600 f and a shadow write address channel WA between the interconnection network 162 and the second processing element 1601 s (Write Burst Request / Address / Acknowledge). Write Lockstep Unit 169 0,w A functional write data channel WD may also be provided between the interconnection network 162 and the first processing element 1600. f and a shadow write data channel WD between the interconnection network 162 and the second processing element 1601 s (Write data). Write lockstep unit 169 0,w A functional write response channel WR may also be provided between the interconnection network 162 and the first processing element 1600. f and a shadow write response channel WR between the interconnection network 162 and the second processing element 1601 s (Write Response).

[0077] like Figure 3 As illustrated, the read lockstep unit 169 0,r May include: being coupled to a functional read address channel RA f The loop buffer circuit 300 is coupled to the output of the loop buffer circuit 300 and the shadow read address channel RA s The comparator and gating circuit 302 is coupled to the functional read response channel RR f and is configured as a read response channel RR in the function fA delayed copy of the read response (data channel) is generated on the delay generation circuit 308 and is coupled to the output of the delay generation circuit 308 and the shadow read response channel RR s And the lockstep enable signal LS EN A multiplexer or gating circuit 310 is controlled.

[0078] In one or more embodiments, the loop buffer circuit 300 may be configured to buffer the read address channel RA in a lockstep mode. f 1601 read request (address channel, control channel). The control logic of the circular buffer circuit 300 can be based on a simple handshake mechanism, request and confirmation, which allows support for most communication protocols (such as, AXI protocol). New read requests (address channel, control channel) can be buffered so that they can be compared later (only) when the entry is still available, otherwise the confirmation signal may be kept low waiting for the entry to become idle. The entry becomes idle (only) due to a read request from the processing element 1601. Data comparison can occur in the comparator circuit 302 between the read request from 1601 and the first entry of the buffer 300. All buffered entries (one entry per request) are shifted one position at the next clock cycle. The circular buffer circuit 300 may not be enabled in high-speed mode, and no read request is stored in the buffer.

[0079] In one or more embodiments, the comparator and gating circuit 302 can be configured to compare the functional read address channel RA stored inside the buffer 300 in lockstep mode. f and shadow read address channel RA s The read request (address channel, control channel) between the two channels is configured to read the address channel RA in the shadow s Since the lockstep mode is not activated, the comparator and gating circuit 302 can read the shadow address channel RA s The request on is propagated to the local interconnect 162 and the data comparison may not occur.

[0080] In one or more embodiments, the read lockstep unit 169 0,r It may be configured to send a fault signal to a fault collection unit (FCU) if a fault is detected by the comparator and gating circuit 302 .

[0081] In one or more embodiments, since lockstep mode is enabled (eg, LS EN =1), the multiplexer or gate circuit 310 can read the response channel RR from the function f The delayed response (output of delay generation circuit 308) is propagated to processing element 1601. Otherwise, since lockstep mode is disabled (e.g., LSEN =0), the multiplexer or gate circuit 310 can read the shadow response channel RR S The response is propagated to processing element 1601.

[0082] like Figure 3 As shown, write lockstep unit 169 0,w May include: being coupled to a functional write address channel WA f The first circular buffer circuit 312 is coupled to the functional write data channel WD f The second loop buffer circuit 314, the first comparator, and the output of the first loop buffer circuit 312 and the shadow write address channel WA are coupled to each other. s The gating circuit 316 is coupled to the output of the second loop buffer circuit 314 and the shadow write data channel WD s The second comparator and gating circuit 318 is coupled to the functional write response channel WR f and is configured as a write response channel WR in the function f The delay generation circuit 326 generates a delayed copy of the write response (data channel), and the output of the delay generation circuit 326 and the shadow write response channel WR are coupled to each other. s And the lockstep enable signal LS EN Controlled multiplexer or gating circuit 328.

[0083] In one or more embodiments, the first loop buffer circuit 312 may be configured to write the address path WA in a lockstep mode. f Buffer write requests (address channel, control channel) on the buffer. The control logic of the first circular buffer circuit 312 can be based on a simple handshake mechanism, request and confirmation, which allows support for most communication protocols (such as the AXI protocol). New write requests (address channel, control channel) can be buffered so that they can be compared later (only) when the entry is still available, otherwise the confirmation signal may remain low, waiting for the entry to become free. The entry becomes free (only) due to a write request from the processing element 1601. Data comparison can occur in the comparison circuit 316 between the write request from 1601 and the first entry of the buffer 312. All buffered entries (one entry per request) are shifted one position at the next clock cycle. The first circular buffer circuit 312 can be disabled in high-speed mode so that the write request is not stored in the buffer 312 and can be directly propagated to the local interconnect 162.

[0084] In one or more embodiments, the first comparator and gating circuit 316 can be configured to compare the functional write address channel WA stored inside the buffer 312 in lockstep mode. fand shadow write address channel WA s The first comparator and gating circuit 316 can write shadow writes to the address channel WAs because the lockstep mode is not activated. s The request on is propagated to the local interconnect 162 and the data comparison may not occur.

[0085] In one or more embodiments, the second loop buffer circuit 314 may be configured to buffer the functional write data channel WD in a lockstep mode. f 1601 ). The control logic of the second circular buffer circuit 314 can be based on a simple handshake mechanism, request and confirmation, which allows support for most communication protocols (such as, AXI protocol). New write requests (data channels) can be buffered so that they can be compared later (only) when the entry is still available, otherwise the confirmation signal may remain low, waiting for the entry to become free. The entry becomes free (only) due to a write request from the processing element 1601. Data comparison can occur in the comparison circuit 318 between the write request from 1601 and the first entry of the buffer 314. All buffered entries (one entry per request) are shifted one position at the next clock cycle. The second circular buffer circuit 314 can be disabled in high-speed mode so that the write request is not stored in the buffer 314 and can be directly propagated to the local interconnect 162.

[0086] In one or more embodiments, the second comparator and gating circuit 318 can be configured to compare the data stored in the buffer 314, the shadow write data channel WD in lockstep mode. s The internal function writes the data channel WD f Write request between (data channel), and write the shadow to the data channel WD s The request on is gated to the interconnect 162 which is not allowed to be propagated. Since the lockstep mode is not activated, the second comparator and gating circuit 318 can write the shadow into the data channel WD s The request on is propagated to the local interconnect 162 and the data comparison may not occur.

[0087] In one or more embodiments, the write lockstep unit 169 0,w The comparator and gating circuit 316 and / or 318 may be configured to send a fault signal to the fault collection unit if a fault is detected by the comparator and gating circuit 316 and / or 318. Additionally or alternatively, the comparator and gating circuit 316 and / or 318 may be configured to gate the functional path WA. f and WD f Write access on the IO pin is restricted to avoid corruption of memory contents in the event of a detected fault.

[0088] Thus, in one or more embodiments, a functional read access may be propagated immediately to the interconnect network 162 without waiting for comparison to occur in the comparison circuit 302, while a functional write access may be propagated to the interconnect network 162 only after comparison has occurred in the comparator circuits 316 and / or 318, whenever an erroneous write access may corrupt data stored in the memory.

[0089] In one or more embodiments, when the hardware accelerator device 16 is used for safety-related applications, each pair of processing elements 160 (statically defined within the hardware accelerator device 16) can be programmed to operate in a pseudo-lockstep mode by setting a corresponding configuration bit in a configuration register of the control unit 161. For example, if such a configuration bit has a first value (e.g., it is equal to 1), the read / write comparator circuits 302, 316, 318 can perform a comparison on the DMA read / write output channel. In response to detecting a mismatch, the corresponding write request may not be propagated, the flow may be stopped, and the error may be reported to the (e.g., external) logical fault collection unit (LFCU) of the system-on-chip 1. Alternatively, if such a configuration bit has a second value (e.g., it is equal to 0), the request on the bus may simply be propagated to the memory controller 163.

[0090] In one or more embodiments, when the two processing elements in a pair (i.e., the "functional" processing element 1600 and the "shadow" processing element 1601) operate in pseudo-lockstep mode, their operations may be time-shifted (e.g., out of phase). For example, such a time shift may be equal to a certain number of clock cycles, optionally two clock cycles. For example, Figure 4 As illustrated, this can be achieved through hardware architecture.

[0091] In one or more embodiments, the phase shifting mechanism can be the same for all internal DMA controllers (e.g., both the read DMA controller and the write DMA controller). This time shifting can facilitate good coverage of common cause faults. For example, a fault due to electromagnetic interference (EMI) may cause a fault to appear on both the functional path and the shadow path, and without time shifting, the lockstep unit would not be able to detect the fault.

[0092] like Figure 4 As illustrated, the DMA controller 1680 (read and / or write) of the processing element 1600 configured as a pair of functional processing elements may receive a start signal START0 from the control unit 161. The DMA controller 1681 (read and / or write) of the processing element 1601 configured as a pair of shadow processing elements may be enabled due to the pseudo lockstep mode (e.g., LS EN= 1) to receive the delayed copy START0' of the start signal START0, or because the high-speed mode is enabled (eg, LS EN =0) to receive the start signal STATR1 from the control unit 161. The delay generation circuit block 400 can receive the start signal STRAT0 and generate a delayed replica START0'. The selection circuit 402 (eg, a multiplexer or a gate circuit) can select the start signal STATR1 according to the pseudo lockstep enable signal LS EN To propagate the delayed copy START0′ or the start signal START1 to the DMA controller 1681.

[0093] In one or more embodiments, when both processing elements in a pair operate in high-speed mode, read / write lockstep unit 169 may be configured as a basic safety mechanism (without a lockstep comparator, i.e., providing two independent data paths) to (only) protect (end-to-end) data transfer to / from local memory M via local interconnect 162.

[0094] In one or more embodiments, hardware accelerator device 16 can support concurrent execution of multiple algorithms. In this case, corresponding pairs of processing elements can be configured to operate in pseudo-lockstep mode or high-speed mode, depending on the safety requirements of the specific algorithm. In an embodiment, a subset of the processing element pairs can operate in pseudo-lockstep mode, while another subset of the processing element pairs can operate in high-speed mode. Pseudo-lockstep mode or high-speed mode can be part of the configuration of hardware accelerator device 16 and can be selected based on the algorithm.

[0095] Additionally or alternatively, in one or more embodiments, similar security mechanisms may be implemented in the data path between processing element 160 and configurable coefficient memory 167 .

[0096] For example, Figure 5 FIG. 1 is a circuit block diagram illustrating implementation details of an embodiment of the configurable coefficient memory 167. Figure 5 As illustrated, each processing element 160 may include a respective memory address generator circuit 210 coupled to a memory controller 166 to which a configurable coefficient memory 167 is coupled. Figure 3 The same read lockstep unit 169 shown in FIG. r The memory address generator circuit 210 may be coupled to compare read requests in the coefficient memory 167 between the functional path and the shadow path. The same start signal START0 used for the DMA controller 168 may be used for the memory address generator circuit 210.

[0097] Additionally or alternatively, in one or more embodiments, similar security mechanisms may be implemented in the data path between processing element 160 and read-only memory 165 .

[0098] For example, each processing element 160 may include a corresponding ROM address generator circuit 208 coupled to a ROM controller 164 to which the read-only memory 165 is coupled. Figure 3 The same read lockstep unit 169 shown in FIG. r The DMA controller 168 may be coupled to the ROM address generator circuit 208 to compare read requests in the read-only memory 165 between the functional path and the shadow path. The same start signal START0 used for the DMA controller 168 may be used for the ROM address generator circuit 208.

[0099] In one or more embodiments, the hardware accelerator device 16 may comply with the requirements of the ISO 26262 security standard by protecting (all) addresses and data stored (in RAM memory and / or in ROM memory) by protection codes such as double error detection (DED), single and double error correction (SECDED), and / or PARITY codes.

[0100] In one or more embodiments, control signals (eg, burst length signal, burst type signal, etc.) of local interconnect 162 may be protected by DED or PARITY codes. Additionally, PARITY bits may be used to protect local interconnect handshake bits.

[0101] In one or more embodiments, the read / write DMA controller 168 , the local data memory controller 163 , the configurable memory controller 166 , and / or the ROM memory controller 164 may therefore implement new functionality to provide improved functional safety.

[0102] In one or more embodiments, the protection scheme may be statically configurable according to the requirements of the processing system 1 .

[0103] In one or more embodiments, the read DMA controller 200 may be configured to implement one or more of the following functions: generate a DED, SECDED, or PARITY code on a burst start address, generate a DED or PARITY code on a burst control signal, generate a PARITY bit on an output handshake signal, perform a DED, SECDED, or PARITY check on input read data and issue an error signaling to a logical fault collection unit (LFCU), and perform a PARITY check on the input handshake signal and issue an error signaling to the logical fault collection unit.

[0104] For example, Figure 3As illustrated, the read DMA controller 200 may include: a corresponding protection code generator circuit 304 configured to read the address channel RA in a functional manner; f Generates a protection code (address channel, control channel) for a read request; and a corresponding protection code checker circuit 306 configured to check the function read response channel RR f The protection code of the read response (data channel) on the shadow read response channel RR s Replication of data channels (with delay) on .

[0105] In one or more embodiments, the write DMA controller 204 may be configured to implement one or more of the following functions: generate a DED, SECDED, or PARITY code on a burst start address, generate a PARITY bit on an output handshake signal, generate a DED, SECDED, or PARITY code on write data, and perform a PARITY check on an input handshake signal, and issue an error signaling to a logic fault collection unit.

[0106] For example, Figure 3 As illustrated, the write DMA controller 204 may include: a corresponding protection code generator circuit 320 configured to write to the functional address channel WA f Generates write request on protection code (address channel, control channel), and / or is configured to write data channel WD on function f Generates a protection code for a write request (data channel); and a corresponding protection code checker circuit 324 configured to check the function write response channel WR f Write response protection code and shadow write response channel WR s Replication of the (delayed) response channel on .

[0107] In one or more embodiments, the local data memory controller 163 may be configured to implement one or more of the following functions: perform DED, SECDED, or PARITY checks on input burst start addresses and send error signaling to a logical fault collection unit, propagate WRITE DATA protection codes to the memory bank M, perform DED or PARITY checks on input control signals, propagate READ DATA ECC protection codes from the memory bank M to the local interconnect 162, perform PARITY checks on input handshake signals and send error signaling to a logical fault checking unit, and generate PARITY bits on output handshake signals.

[0108] In one or more embodiments, the configurable memory address generator circuit 210 can be configured to implement one or more of the following functions: generate DED, SECDED, or PARITY codes on addresses, generate DED or PARITY codes on burst control signals, generate PARITY bits on output handshake signals, and perform PARITY checks on input handshake signals, and send error signaling to a logic fault collection unit.

[0109] In one or more embodiments, the configurable memory controller 166 can be configured to implement one or more of the following functions: performing DED, SECDED, or PARITY checking on the input burst start address and sending error signaling to the logical fault collection unit, performing DED, SECDED, or PARITY checking on the read data value and sending error signaling to the logical fault collection unit, performing DED or PARITY checking on the input local bus control signal and sending error signaling to the logical fault collection unit, performing PARITY checking on the input handshake signal and sending error signaling to the logical fault collection unit, and generating a PARITY bit on the output handshake signal.

[0110] In one or more embodiments, ROM address generator 208 may be configured to implement generation of DED, SECDED, or PARITY codes on addresses.

[0111] In one or more embodiments, the ROM controller 164 may be configured to implement one or more of the following functions: performing DED, SECDED, or PARITY checks on input addresses and sending error signaling to a logic fault collection unit, and performing DED, SECDED, or PARITY checks on read data values ​​and sending error signaling to a logic fault collection unit.

[0112] It is important to note that error correction code (ECC) protection schemes can provide high coverage on the address path and data path, but they may not be applicable to the control path of the interface. The control signals on the target interface (e.g., memory) can be generated by logic blocks (e.g., FSM, decoder, etc.) without preserving the source information on the initiator interface (e.g., internal DMA, external AXI interface).

[0113] Thus, one or more embodiments may include (even independently of the implementation of a lockstep architecture as previously described) a memory read / write watchdog mechanism to provide an end-to-end (e.g., booting program memory (such as DMA memory)) safety mechanism on the control path that may facilitate detection of hard and / or soft faults within the control logic of the memory controller.

[0114] In one or more embodiments, a first type of memory read / write watchdog can be configured by a particular launcher (e.g., each launcher) to count the number and / or type (reads, writes, bus width) of memory operations performed during calculation of an algorithm (e.g., each algorithm).

[0115] Note that, in contrast to solutions where memory access depends on kernel implementations of policies, compilation toolchains, etc., the data flow and the number / type of operations performed in memory by the hardware accelerator device are statically defined according to the computational algorithm.

[0116] Thus, one or more embodiments may include a read / write watchdog circuit for each target interface (e.g., local memory or system memory), the read / write watchdog circuit including a set of concurrent read / write counters for each initiator device (e.g., a counter for each internal DMA controller, a counter for each external bus interface, etc.). Each read / write watchdog circuit may track all operations and store the accumulated results in a set of status registers. At the end of execution of the algorithm, the contents of the watchdog status register may be compared with the expected number of read / write operations, thereby providing a safety mechanism to prevent possible failures of the control path.

[0117] Figure 6A FIG. 1 is an example circuit block diagram of this first type of memory watchdog used in one or more embodiments. Figure 6A As illustrated, the memory watchdog circuit 60A may include a read counter circuit 62Ar and a write counter circuit 62Aw.

[0118] The read counter circuit 62Ar may be configured to receive a corresponding chip select signal CS, a read enable signal REN, and a corresponding identification signal ID, which carries information suitable for identifying a boot program (e.g., a boot program ID). The read counter circuit 62Ar may be configured to generate an output read count signal RC (e.g., to be propagated to a watchdog status register in the local control unit 161).

[0119] The write counter circuit 62Aw may be configured to receive a corresponding chip select signal CS, a write enable signal WEN, and a corresponding identification signal ID carrying information suitable for identifying a boot program (e.g., a boot program ID). The write counter circuit 62Aw may be configured to generate an output write count signal WC (e.g., to be propagated to a watchdog status register in the local control unit 161).

[0120] In one or more embodiments, the local memory read / write watchdog circuit 60A can be configured to: track the (e.g., cumulative) number of read accesses for each initiator (based on the received initiator ID) based on an algorithm, track the (e.g., cumulative) number of write accesses for each initiator (based on the received initiator ID) based on an algorithm, and optionally, issue a fault signal to a logical fault collection unit in the event of a fault.

[0121] In one or more embodiments, Figure 6A As illustrated, the watchdog status can be checked according to different strategies. Those strategies can be configurable via control registers of the local control unit. For example, the watchdog status check strategy can include a software check of the watchdog status at the end of data processing for the configured algorithm. In another example, the watchdog status check strategy can include a hardware check of an expected number of read / write accesses per memory bank and per boot program (e.g., automatically triggered at the end of data processing for the configured algorithm), optionally with error signaling to a logical fault collection unit in the event of a failure.

[0122] Purely by way of non-limiting example, reference is made to a matrix multiplication algorithm involving a single processing element 160. Figure 6A Operation of the disclosed memory watchdog mechanism. The matrix multiplication algorithm can be expressed as C[n,n]=A[n,n]*B[n,n]. The input matrices A and B can be preloaded in the local memory M by the system DMA controller 14. The output matrix C can be stored in the local memory M by the internal DMA controller 168 of the processing element 160. Therefore, the expected number of write operations performed by the system DMA controller is equal to 2*(n*n), the expected number of read operations performed by the internal DMA controller is equal to 2*(n*n*n), and the expected number of write operations performed by the internal DMA controller is equal to n*n.

[0123] Additionally or alternatively, in one or more embodiments, the second type of memory read / write watchdog can be configured to count the number of outstanding memory operations during an algorithm (e.g., each algorithm) computed by a particular initiator (e.g., each read or write internal DMA), where such number is incremented when a transaction is issued and decremented when a response is received. Thus, the second type of memory read / write watchdog may not count the absolute value of memory accesses but rather a relative value of memory accesses (e.g., by a simple up / down counter circuit), which relative value is expected to be equal to zero at the end of the algorithm.

[0124] Figure 6BFIG. 1 is an example circuit block diagram of this second type of memory watchdog used in one or more embodiments. Figure 6B As illustrated, the memory watchdog circuit 60B may include a read up-down counter circuit 62Br and / or a write up-down counter circuit 62Bw configured to count the difference between the number of words requested to be read or written by each initiator and the number of words actually read or written.

[0125] The read up / down counter circuit 62Br can be configured to receive a corresponding read enable signal REN and / or a read burst length signal RBURSTL to increase a corresponding counter value, and receive a corresponding response enable signal RRESP_EN to decrease a corresponding counter value. Thus, if the initiator interface is a read interface (e.g., a local read DMA interface), the read request and burst length signals can be used for incrementing operations, while the response enable signal (which can be used to detect valid read data) can be used to decrement the up / down counter value. In this exemplary case, all signals used by the watchdog circuit come from the same initiator interface.

[0126] The write up / down counter circuit 62Bw can be configured to receive a corresponding write enable signal WEN and / or a write burst length signal WBURSTL to increase a corresponding counter value, and receive a corresponding write enable signal W_EN and / or a corresponding identification signal INIT_ID carrying information suitable for identifying an initiator (e.g., an initiator ID) to decrease a corresponding counter value. Thus, if the initiator interface is a write interface (e.g., a local write DMA interface), the write request and burst length signals can be used to increment the number of words requested to be written, while the write enable signal (which is used to detect valid data) and the initiator ID at the target interface (e.g., a memory controller output) can be used to decrement the number of words.

[0127] In one or more embodiments, an initiator ID may be propagated from each source (eg, using AXI user signals) to a target for use in protecting write transactions inside a hardware accelerator device.

[0128] At the end of data processing (e.g., at the end of the algorithm calculation), the external host controller can read the status of the instantiated watchdog (e.g., one per boot program) and verify that the final count value is equal to zero. A mismatch between the number of requests and the number of data actually read / written can be attributed to a fault within the controller (e.g., a single point of failure, SPF).

[0129] In one or more embodiments, Figure 6BAs illustrated, the watchdog status can be checked according to different strategies. Those strategies can be configurable via control registers of the local control unit. For example, the watchdog status check strategy can include a software check of the watchdog status at the end of data processing for the configured algorithm (when the number of outstanding transactions is expected to be zero). In another example, the watchdog status check strategy can include a hardware check of the expected number of outstanding transactions for each memory bank and each boot program (for example, automatically triggered at the end of data processing for the configured algorithm), optionally with error signaling sent to a logical fault collection unit in the event of a fault. In another example, the watchdog status check strategy can include monitoring hardware or software runtime as a performance metric (wait time, peak / average number of outstanding transactions) by an interrupt asserted at a configurable threshold.

[0130] It should be noted that Figure 6B The illustrated watchdog mechanism, which relies on counting the number of outstanding transactions rather than the total number of valid transactions, may provide one or more of the following advantages: simple and immediate watchdog management at the application level (no need to configure the absolute number of transactions for each algorithm), resilience to algorithms where the number of memory accesses depends on the type / value of the data, and runtime monitoring of the peak / average number of outstanding transactions as an additional performance metric.

[0131] In one or more embodiments, the following is provided: Figure 6A or Figure 6B The illustrated memory read / write watchdog circuit can contribute to the overall safety goals (local memory bank M, configuration memory 167 and read-only memory 165) while avoiding duplication of internal memory controllers and achieving satisfactory single point failure (SPF) coverage within such blocks.

[0132] Note that in conventional devices, safety monitors are typically protected from potential faults by their replication (e.g., in the case of standard cores) or by logic built-in self-test (LBIST) routines applied to the full / partial device.

[0133] In order to reduce the area overhead and / or design complexity caused by the duplication of monitor or LBIST insertion process, one or more embodiments may include a dedicated hardware built-in self-test (BIST) for the safety monitor (even independent of the implementation of the lockstep architecture or the memory watchdog mechanism as previously described). In one or more embodiments, such a dedicated hardware BIST may also reduce the unavailability of device functions insofar as conventional devices may be unavailable during the execution of the LBIST and the next partial or full reset.

[0134] In one or more embodiments, a dedicated hardware BIST may provide one or more of the following features: latent fault (LF) detection or fault injection for the safety monitor to check the monitor at the LFCU interface, runtime checking (e.g., with rate timing defined at the application level), and fault simulation of the BIST and safety monitor to provide the required fixed coverage (e.g., equal to or higher than 90% coverage for ASIL-D safety level).

[0135] In one or more embodiments, BIST may be applied to all implemented safety monitors to support end-to-end protection schemes, lockstep comparators, memory watchdogs, and the like.

[0136] Figure 7 FIG. 1 is a circuit block diagram of a hardware safety monitor BIST according to one or more embodiments. Figure 7 The illustrated hardware BIST may include: a local control unit 161; a pattern generator 71 or a fault injector 71 based on a ROM, a lookup table, or a pseudo-random linear feedback shift register (LFSR); an adder node 72 configured to combine data received from a functional path with data generated by the pattern generator / fault injector 71; a selection circuit 73 (e.g., a multiplexer) configured to propagate output data from the adder node 72 or output data from the pattern generator / fault injector 71 to a safety monitor 74, where a subset of safety monitors (e.g., all, a single, a sequence) may be selected for execution of the BIST process; a CRCn (e.g., CRC32, CRC16, etc.) compressor 75 for a subset of safety monitor nodes (measurement points); a comparator circuit block 77 configured to compare the output of the CRCn compressor 75 with a final signature 76 (e.g., a magic value) and / or intermediate values ​​to increase coverage fault signaling; control and status registers; and an interface 78 to a logical fault collection unit.

[0137] In one or more embodiments, functionality of the hardware accelerator device may not be available at runtime during execution of the safety monitor BIST.

[0138] In one or more embodiments, a reset of the hardware accelerator device may be required after execution of the safety monitor BIST.

[0139] Because the local control unit 161 may represent a source of failure for the hardware accelerator device 16, one or more embodiments may rely on one or more of the following safety mechanisms: replication of the FSM and critical portions (status / error registers, interrupts, etc.), and protection code for control registers (parity bits or CRC32 checksums).

[0140] In one or more embodiments, interface 78 may include a (simple) dual-signal-level register interface. If at least one uncorrectable error is detected in the hardware accelerator device, a first signal (e.g., EDPA_cf) may be asserted (e.g., set to a logic level 1). If at least one correctable error is detected within the hardware accelerator device, a second signal (e.g., EDPA_ncf) may be asserted (e.g., set to a logic level 1). Thus, interface 78 may advantageously signal internally detected errors externally to allow the system (e.g., system-on-chip) to reach a safe state within an acceptable time interval (fault-tolerance time interval) required by the safety goal.

[0141] In one or more embodiments, the electronic system 1 can be implemented as an integrated circuit in a single silicon chip or die (e.g., as a system on a chip). Alternatively, the electronic system 1 can include a distributed system including multiple integrated circuits interconnected together (e.g., via a printed circuit board (PCB)).

[0142] Thus, for example, when a safety-related algorithm needs to be calculated, one or more embodiments as illustrated herein can provide a hardware accelerator device 16 that can be selectively configured (e.g., at runtime) to operate at a certain level of functional safety (e.g., at an ASIL-D level). When a non-safety-related algorithm needs to be calculated, such calculation can be accelerated, with all internal computing power available.

[0143] The functional safety architecture disclosed herein facilitates providing memory-based hardware accelerator devices, and potentially SoCs integrating hardware accelerator devices, that support ASIL-D applications with reduced area overhead.

[0144] As illustrated herein, a hardware accelerator device (e.g., 16) may include: a set of processing circuits (e.g., 160) arranged in a subset (e.g., a pair) of processing circuits; a set of data memory banks (e.g., M) coupled to a memory controller (e.g., 163); a control unit (e.g., 161) including configuration registers that provide storage space for configuration data of the processing circuits; and an interconnection network (e.g., 162). The processing circuits may be configured according to the configuration data stored in the control unit to read (e.g., 200, 202) first input data from the data memory bank via the interconnection network and the memory controller, process (e.g., 20) the first input data to generate output data, and write (e.g., 204, 206) the output data to the data memory bank via the interconnection network and the memory controller. The hardware accelerator device may include a set of configurable lockstep control units (e.g., 169) that interface processing circuits (e.g., DMA controller 168) to an interconnect network, each configurable lockstep control unit (e.g., 1690) in the set of configurable lockstep control units being coupled to a subset of processing circuits (e.g., 1600, 1601) in the set of processing circuits. Each configurable lockstep control unit may be selectively activated (e.g., LS EN ) to operate in a first operating mode (e.g., a "lockstep mode" or a "pseudo-lockstep mode"), wherein the lockstep control unit (e.g., 169 0,r 、169 0,w ) is configured to compare data read requests and / or data write requests issued by a first processing circuit (e.g., 1600) and a second processing circuit (e.g., 1601) in a corresponding subset of the processing circuits toward a memory controller to detect a fault; or operates in a second operating mode (e.g., a “high-speed mode” or a “performance mode”) in which the lockstep control unit is configured to propagate data read requests and / or data write requests issued by a first processing circuit and a second processing circuit in a corresponding subset of the processing circuits toward the memory controller.

[0145] As illustrated herein, a configurable lockstep control unit may be selectively activated to operate in a first operating mode or a second operating mode according to configuration data stored in the control unit.

[0146] As illustrated herein, a hardware accelerator device may include a clock source configured to generate a clock signal. A configurable lockstep control unit may be configured to delay (e.g., 400, 402) processing of first input data by the second processing circuit relative to the first processing circuit in response to the configurable lockstep control unit operating in a first operating mode, optionally by a period of two clock cycles of the clock signal.

[0147] As illustrated herein, the hardware accelerator device may include at least one read-only memory (e.g., 165) coupled to a ROM controller (e.g., 164). The processing circuit may be configured to read second input data from the at least one read-only memory via the ROM controller and process the second input data to generate output data. The lockstep control unit may compare data read requests issued by the first and second processing circuits in the respective subsets of the processing circuits toward the ROM controller to detect a fault in a first operating mode. In a second operating mode, the lockstep control unit may propagate the data read requests issued by the first and second processing circuits in the respective subsets of the processing circuits toward the ROM controller.

[0148] As illustrated herein, a hardware accelerator device may include at least one locally configurable memory (e.g., 167) coupled to a configuration memory controller (e.g., 166). The processing circuitry may be configured to read third input data from the at least one locally configurable memory via the configuration memory controller and process the third input data to generate output data. In a first operating mode, a lockstep control unit may compare data read requests issued by a first processing circuit and a second processing circuit in a respective subset of the processing circuits toward the configuration memory controller to detect a fault. In a second operating mode, the lockstep control unit may propagate data read requests issued by a first processing circuit and a second processing circuit in a respective subset of the processing circuits toward the configuration memory controller.

[0149] As illustrated herein, the interconnect network may include at least one control channel configured to exchange control messages.The processing circuitry and / or the memory controller may be configured to include a double error detection (DED) code or a parity check code in the control messages.

[0150] As illustrated herein, the interconnect network may include at least one address channel configured to exchange address messages and at least one data channel configured to exchange data messages. The processing circuitry and / or the memory controller may be configured to include protection codes in the address messages and in the data messages. The protection codes may include one of a double error detection (DED) code, a parity check code, or a single error correction double error detection (SECDED) code.

[0151] As illustrated herein, the ROM controller may include at least one address channel configured to exchange address messages and at least one data channel configured to exchange data messages. The processing circuit and / or the ROM controller may be configured to include a protection code in the address message and in the data message. The protection code may include a double error detection (DED) code, a parity check code, or a single error correction double error detection (SECDED) code.

[0152] As illustrated herein, the configuration memory controller may include at least one address channel configured to exchange address messages and at least one data channel configured to exchange data messages. The processing circuit and / or the configuration memory controller may be configured to include protection codes in the address messages and the data messages. The protection codes may include one of a double error detection (DED) code, a parity check code, or a single error correction double error detection (SECDED) code.

[0153] As illustrated herein, the hardware accelerator device may include an end-to-end mechanism configured to propagate protection codes from the processing circuit to the memory unit (e.g., any of the data memory bank M, the read-only memory 165, and / or the local configurable memory 167, as appropriate) and / or from the memory unit to the processing circuit via the lockstep control unit and the interconnect network as a result of the corresponding lockstep control unit operating in a first operating mode. Alternatively, the end-to-end mechanism may be configured to propagate protection codes between the processing circuit and the memory unit as a result of the corresponding lockstep control unit operating in a second operating mode.

[0154] As illustrated herein, a hardware accelerator device may include a first memory watchdog circuit coupled to a data memory bank, wherein the first memory watchdog circuit (e.g., 60A) is configured to count a first number of memory transaction requests received in the data memory bank, and the hardware accelerator device is configured to compare the first counted number of memory transactions with a first expected number of memory transactions to detect a failure. For example, the first expected number of memory transactions may include a plurality of memory transactions for executing a complete algorithm, or a plurality of memory transactions for executing a computation cycle of the algorithm. Additionally or alternatively, the first memory watchdog circuit (e.g., 60B) may be configured to count a first number of outstanding memory transaction requests received at the data memory bank, and the hardware accelerator device may be configured to check whether the first counted number of outstanding memory transactions is equal to zero to detect a failure.

[0155] Optionally, the first memory watchdog circuit may comprise a respective counter for each memory transaction initiation procedure.Optionally, the first memory watchdog circuit may be configured to store a first counted number(s) of memory transactions in a status register of the control unit.

[0156] As illustrated herein, a hardware accelerator device may include a second memory watchdog circuit coupled to at least one read-only memory, wherein the second memory watchdog circuit is configured to count a second number of memory transaction requests received at the at least one read-only memory, and the hardware accelerator device may be configured to compare the second counted number of memory transactions with a second expected number of memory transactions to detect a failure. For example, the second expected number of memory transactions may include a number of memory transactions for executing a complete algorithm, or a number of memory transactions for executing a computation cycle of the algorithm. Additionally or alternatively, the second memory watchdog circuit may be configured to count a second number of outstanding memory transaction requests received at the at least one read-only memory, and the hardware accelerator device may be configured to check whether the second counted number of outstanding memory transactions is equal to zero to detect a failure.

[0157] Optionally, the second memory watchdog circuit may comprise a respective counter for each memory transaction initiation procedure.Optionally, the second memory watchdog circuit may be configured to store a second counted number(s) of memory transactions in a status register of the control unit.

[0158] As illustrated herein, a hardware accelerator device may include a third memory watchdog circuit coupled to at least one locally configurable memory, wherein the third memory watchdog circuit may be configured to count a third number of memory transaction requests received at the at least one locally configurable memory, and the hardware accelerator device may be configured to compare the third counted number of memory transactions with a third expected number of memory transactions to detect a failure. For example, the third expected number of memory transactions may include a number of memory transactions for executing a complete algorithm, or a number of memory transactions for executing a computation cycle of the algorithm. Additionally or alternatively, the third memory watchdog circuit may be configured to count a third number of outstanding memory transaction requests received at the at least one locally configurable memory, and the hardware accelerator device may be configured to check whether the third counted number of outstanding memory transactions is equal to zero to detect a failure.

[0159] Optionally, the third memory watchdog circuit may comprise a respective counter for each memory transaction initiation procedure.Optionally, the third memory watchdog circuit may be configured to store a third counted number(s) of memory transactions in a status register of the control unit.

[0160] As illustrated herein, the hardware accelerator device may include: a built-in self-test pattern generator circuit (e.g., 71) or a fault injector circuit (e.g., 71) configured to inject a test pattern into a lock-step control unit to generate a corresponding test output signal; and a comparator circuit (e.g., 77) configured to compare the test output signal with an expected test output signal to detect a fault of the lock-step control unit.

[0161] As illustrated herein, a system (e.g., 1) may include a hardware accelerator device according to one or more embodiments and a fault collection unit, possibly coupled via a system interconnect (e.g., 18). The fault collection unit may be sensitive to faults detected by a lockstep control unit (or by any other safety monitor that may be provided in the hardware accelerator device) and may be configured to set the system to a safe operating mode in response to the detected fault.

[0162] As illustrated herein, a method of operating a hardware accelerator or system according to one or more embodiments may include: reading first input data from a data memory repository via an interconnection network and a memory controller; processing the first input data at a processing circuit to generate output data, writing the output data to the data memory repository via the interconnection network and the memory controller; and selectively activating a configurable lockstep control unit to: operate in a first operating mode, wherein the lockstep control unit is configured to compare data read requests and / or data write requests issued by the first and second processing circuits in the corresponding subset of the processing circuits toward the memory controller to detect a fault; or operate in a second operating mode, wherein the lockstep control unit is configured to propagate data read requests and / or data write requests issued by the first and second processing circuits in the corresponding subset of the processing circuits toward the memory controller.

[0163] Without prejudice to the underlying principle and without departing from the scope of protection, the details and embodiments may vary, even significantly, with respect to what has been described purely by way of example.

[0164] <![CDATA[SP0]]> <![CDATA[SP1]]> … <![CDATA[SP P-1 ]]> <![CDATA[SP P ]]> <![CDATA[SP P+1 ]]> <![CDATA[MP0]]> X <![CDATA[MP1]]> X … … <![CDATA[MP P-1 ]]> X <![CDATA[MP P ]]> X

[0165] Table 1

[0166] <![CDATA[SP0]]> <![CDATA[SP1]]> … <![CDATA[SP P-1 ]]> <![CDATA[SP P ]]> <![CDATA[SP P+1 ]]> <![CDATA[MP0]]> X X <![CDATA[MP1]]> X X … … … <![CDATA[MP P-1 ]]> X X <![CDATA[MP P ]]> X

[0167] Table 2

[0168] <![CDATA[SP0]]> <![CDATA[SP1]]> … <![CDATA[SP P-1 ]]> <![CDATA[SP P ]]> <![CDATA[SP P+1 ]]> <![CDATA[MP0]]> X X <![CDATA[MP1]]> X X … … … <![CDATA[MP P-1 ]]> X X <![CDATA[MP P ]]> X (X) (X) (X) X

[0169] Table 3

[0170] Although this description has been described in detail, it will be understood that various changes, substitutions, and modifications may be made without departing from the spirit and scope of the present disclosure as defined by the appended claims. In different figures, identical elements are represented by identical reference numerals. In addition, the scope of the present disclosure is not intended to be limited to the specific embodiments described herein, as those skilled in the art will readily appreciate from this disclosure that currently existing or later developed processes, machines, manufactures, material compositions, means, methods, or steps that may perform substantially the same functions as the corresponding embodiments described herein or achieve substantially the same results. Therefore, the appended claims are intended to include such processes, machines, manufactures, material compositions, means, methods, or steps within their scope.

[0171] Accordingly, the specification and drawings are to be regarded simply as illustrative of the present disclosure as defined by the appended claims, and are intended to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present disclosure.

Claims

1. A device comprising: a plurality of processors, including a subset of processors, each processor being configurable based on functionality of corresponding configuration data; a plurality of data storage libraries configured to store input data; Memory controller; interconnected networks; A control unit comprising a configuration register configured to store configuration data for the processors, each processor being configured to: reading first input data from the data memory bank via the interconnection network and the memory controller, generating first output data based on the first input data, and writing the first output data into the data memory repository via the interconnection network and the memory controller; as well as a plurality of configurable lockstep control units configured to connect the processor interfaces to the interconnect network, each configurable lockstep control unit coupled to a subset of the plurality of processors, each configurable lockstep control unit selectively activatable to: operating in a first operating mode, wherein the respective lockstep control units are configured to compare data read requests or data write requests issued by the first processor and the second processor toward the memory controller to detect a fault, and Operating in a second operating mode, wherein the corresponding lockstep control unit is configured to transmit the data read request or the data write request issued by the first processor and the second processor toward the memory controller. 2 . The apparatus of claim 1 , wherein the plurality of configurable lockstep control units are selectively activatable to operate in the first operating mode or the second operating mode based on configuration data of a corresponding processor.

3. The apparatus according to claim 1, further comprising: a clock source configured to generate a clock signal in response to the plurality of configurable lockstep control units operating in the first operating mode, the plurality of configurable lockstep control units configured to delay processing of the first input data by the second processor relative to the first processor by a period of two clock cycles.

4. The apparatus according to claim 1, further comprising: a read-only memory storage device coupled to the read-only memory controller, Each processor is configured as: reading second input data from the ROM storage device via the ROM controller, and generating second output data based on the second input data, and Wherein, in the first operation mode, the corresponding lockstep control unit is further configured to compare data read requests issued by the first processor and the second processor to the read-only memory controller to detect a fault, and Wherein, in the second operation mode, the corresponding lockstep control unit is further configured to transmit the data read request issued by the first processor and the second processor to the read-only memory controller.

5. The apparatus of claim 4 , wherein the read-only memory controller comprises an address channel and a data channel, the address channel being configured to exchange address messages, the data channel being configured to exchange data messages, and wherein one or more of the plurality of processors and the memory controller are configured to include a protection code in the address messages and the data messages, the protection code comprising one of a double error detection code, a parity code, or a single error correction double error detection code.

6. The apparatus of claim 4 , further comprising a first memory watchdog circuit coupled to the plurality of data memory banks, the first memory watchdog circuit being configured to: a first number of memory transaction requests received at the plurality of data storage repositories is counted, and the apparatus is configured to compare the first number of memory transactions with a first expected number of memory transactions to detect a failure; or A first number of outstanding memory transaction requests received at the plurality of data storage repositories is counted, and the apparatus is configured to check whether the first number of outstanding memory transactions is equal to zero to detect a failure.

7. The apparatus of claim 6 , further comprising a second memory watchdog circuit coupled to the read-only memory storage device, the second memory watchdog circuit being configured to: a second number of memory transaction requests received at the read-only memory storage is counted, and the apparatus is configured to compare the second number of memory transactions with a second expected number of memory transactions to detect a failure; or A second number of outstanding memory transaction requests received at the read-only memory storage is counted, and the apparatus is configured to check whether the second number of outstanding memory transactions is equal to zero to detect a failure.

8. The apparatus according to claim 7, further comprising: a local configurable memory storage device coupled to a configuration memory controller, Each processor is configured as: reading third input data from the local configurable memory storage device via the configuration memory controller, and generating third output data based on the third input data, and Wherein, in the first operating mode, the corresponding lockstep control unit is further configured to compare data read requests issued by the first processor and the second processor toward the configuration memory controller to detect a fault, and Wherein, in the second operation mode, the corresponding lockstep control unit is further configured to transmit the data read request issued by the first processor and the second processor toward the configuration memory controller.

9. The apparatus of claim 8 , wherein the configuration memory controller comprises an address channel and a data channel, the address channel being configured to exchange address messages, the data channel being configured to exchange data messages, and wherein one or more of the plurality of processors and the memory controller are configured to include a protection code in the address messages and the data messages, the protection code comprising one of a double error detection code, a parity code, or a single error correction double error detection code.

10. The apparatus of claim 9 , further comprising an end-to-end mechanism configured to transfer the protection code from the plurality of processors to one of a read-only memory storage device or a locally configurable memory storage device, or from one of a read-only memory storage device or a locally configurable memory storage device to the plurality of processors via the corresponding lockstep control unit as a result of the corresponding lockstep control unit operating in the first operating mode.

11. The apparatus of claim 10 , further comprising a third memory watchdog circuit coupled to the locally configurable memory storage device, the third memory watchdog circuit being configured to: counting a third number of memory transaction requests received at the local configurable memory storage, and the apparatus is configured to compare the third number of memory transactions with a third expected number of memory transactions to detect a failure; and A third number of outstanding memory transaction requests received at the local configurable memory storage is counted, and the apparatus is configured to check whether the third number of outstanding memory transactions is equal to zero to detect a failure.

12. The apparatus of claim 1 , wherein the interconnect network comprises a control channel configured to exchange control messages, and wherein one or more of the plurality of processors and the memory controller are configured to include a double error detection code or a parity check code in the control messages.

13. The apparatus of claim 1 , wherein the interconnect network comprises an address channel configured to exchange address messages and at least one data channel configured to exchange data messages, and wherein one or more of the plurality of processors and the memory controller are configured to include a protection code in the address messages and the data messages, the protection code comprising one of a double error detection code, a parity code, or a single error correction double error detection code.

14. The apparatus of claim 1, further comprising: circuitry configured to inject a test pattern into the plurality of configurable lockstep control units to generate corresponding test output signals; as well as A comparator circuit is configured to compare the corresponding test output signal with an expected test output signal to detect a failure of the plurality of configurable lockstep control units.

15. A system comprising: Equipment, including: a plurality of processors, including a subset of processors, each processor being configurable based on functionality of corresponding configuration data; a plurality of data storage libraries configured to store input data; Memory controller; interconnected networks; a control unit comprising configuration registers configured to store configuration data for said processors, each processor being configured to: reading first input data from the data memory bank via the interconnection network and the memory controller, generating first output data based on the first input data, and writing the first output data into the data memory repository via the interconnection network and the memory controller; a plurality of configurable lockstep control units configured to connect the processor interfaces to the interconnect network, each configurable lockstep control unit coupled to a subset of the plurality of processors, each configurable lockstep control unit selectively activatable to: operating in a first operating mode, wherein the respective lockstep control units are configured to compare data read requests or data write requests issued by the first processor and the second processor toward the memory controller to detect a fault, and operating in a second operating mode, wherein the corresponding lockstep control unit is configured to transmit the data read request or the data write request issued by the first processor and the second processor toward the memory controller; and A fault collection unit is sensitive to faults detected by the plurality of configurable lockstep control units, the fault collection unit being configured to set the system to a safe operating mode in response to the detected faults. 16 . The system of claim 15 , wherein the plurality of configurable lockstep control units are selectively activatable to operate in the first operating mode or the second operating mode based on configuration data of a corresponding processor.

17. The system of claim 15, wherein the device further comprises: a clock source configured to generate a clock signal in response to the plurality of configurable lockstep control units operating in the first operating mode, the plurality of configurable lockstep control units configured to delay processing of the first input data by the second processor relative to the first processor by a period of two clock cycles.

18. A method comprising: reading, by one or more processors, first input data from a data memory bank via an interconnection network and a memory controller; generating, by the one or more processors, first output data based on the first input data; writing, by the one or more processors via the interconnection network and the memory controller, the first output data to the data memory repository; as well as Selectively activate one or more configurable lockstep control units to: comparing data read requests or data write requests issued by the first processor and the second processor toward the memory controller in a first operating mode to detect a failure, and The data read request or the data write request issued by the first processor and the second processor is transmitted toward the memory controller in a second operation mode.

19. The method of claim 18, wherein selectively activatable comprises operating in one of the first operating mode or the second operating mode based on configuration data of a corresponding processor.

20. The method of claim 18, further comprising: generating a clock signal in response to operating in the first operating mode; and Responsive to operating in the first operating mode, processing of the first input data by the second processor is delayed relative to the first processor by a period of two clock cycles.

Citation Information

Patent Citations

  • Method and apparatus for controlling watchdog

    CN105988884A

  • Multi-Channel Network-on-a-Chip

    US20160283314A1