Enhanced Durability of System-on-Chip (SOC)
The microsequencer-driven runtime global push-to-persistence method addresses data durability challenges in complex SoCs by automatically flushing data to persistent storage, ensuring data visibility and reducing loss during unexpected events.
Patent Information
- Application Number
- CN202180076143.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-11
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-29
AI Technical Summary
The prior art is difficult to effectively ensure the durability of data during power loss or reset in computer systems, especially in complex systems on chip (SOC), and it is difficult to ensure the durability of internal data without rewriting application software or operating system.
By introducing a microsequencer into the system on chip, the cached data is automatically refreshed to external memory in response to the trigger signal, and combined with the microsequencer and the hardware circuit in the interconnect structure, the global push to persistence is achieved during runtime, ensuring the integrity of the data in the event of system failure.
It realizes automatic data saving in power loss or reset, ensuring that there is no data loss during restart, and simplifies the durability management of the data processing system.
Smart Images

Figure CN116472512B_ABST
Abstract
Description
Background Art
[0001] Computer systems are vulnerable to contingencies that cause the computer system to be temporarily shut down or powered off. For example, the power supply of the building or home in which the computer system operates may suffer a power loss due to power rationing, power outage, or natural disaster. In addition, the power supply of the computer system itself may malfunction. Another class of events that cause the computer system to shut down is an application or operating system failure that "locks" the computer system and requires the user to manually reset it. Sometimes, the conditions that require the computer to be shut down are predictable, and critical data can be saved before the shutdown. However, any data that has been modified but not yet saved in persistent storage (e.g., non-volatile memory, battery-backed memory, hard disk drive, etc.) will be lost due to power loss or reset. To prevent such accidental data loss, applications sometimes periodically save data files to persistent storage, and the operating system can intervene after detecting one of these events to save important data before the computer shuts down.
[0002] Modern data processors typically use caches (i.e., high-speed memories such as static random access memories (SRAMs) that are tightly coupled to the data processor) to allow for fast access to frequently used data and thereby improve computer system performance. When an application modifies data that has been allocated to the cache, the data processor typically keeps a copy in its cache in a modified ("dirty") form until the cache needs to make room for other data and writes the updated copy back to memory. If an event that requires a shutdown is encountered and there is enough time before the shutdown, the application or operating system can "flush" (i.e., write back) any dirty data from the cache to persistent storage, thereby allowing the updates to critical data to be saved and globally observable so that the user's work can be resumed without loss when the computer system is restarted later.
[0003] A system-on-chip (SOC) combines various data processors, caches, queues, multi-layer interconnect circuits, and input / output peripherals on a single integrated circuit chip. With the advent of deep sub-micron semiconductor manufacturing process technologies, SOCs have become increasingly complex and can contain several data processor cores, multi-layer caches, and highly buffered interconnect fabrics, making it difficult for applications and operating systems running on these SOCs to ensure that their internal data is durable without having to rewrite the application software or operating system to understand the details of the SOC. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 A data processing system known in the prior art is shown in block diagram form;
[0005] Figure 2 FIG. 13 shows a high - level abstract view of a data - processing system with runtime global push to persistence according to some embodiments in block diagram form;
[0006] Figure 3 FIG. 14 shows a data - processing system with an exemplary system - on - chip (SOC) with runtime global push to persistence according to some embodiments in block diagram form;
[0007] Figure 4 FIG. 15 shows another data - processing system with an SOC with runtime global push to persistence according to some embodiments in block diagram form;
[0008] Figure 5 FIG. 16 shows a flowchart of runtime processes available in an SOC according to some embodiments;
[0009] Figure 6 FIG. 17 shows a data - processing system according to some embodiments in block diagram form, which shows Figure 4 the SOC performing Figure 5 the runtime global push to persistence process;
[0010] Figure 7 FIG. 18 shows another data - processing system with an SOC with runtime global push to persistence according to some embodiments in block diagram form;
[0011] Figure 8 FIG. 19 shows in block diagram form Figure 4 and Figure 7 the terminal - event trigger - generation circuitry available in the SOC of FIG. 20; and
[0012] Figure 9 FIG. 20 shows in block diagram form Figure 4 and Figure 7 the non - terminal - event trigger - generation circuitry available in the SOC of FIG. 21.
[0013] In the following description, the same reference numerals are used in different figures to indicate similar or identical items. Unless otherwise specified, the word "coupled" and its associated verb forms include both direct connection and indirect electrical connection by means known in the art, and unless otherwise specified, any description of direct connection also implies an alternative embodiment using a suitable form of indirect electrical connection. DETAILED DESCRIPTION
[0014] As will be described in detail below, a system-on-chip having runtime global push to persistence includes a data processor with a cache, an external memory interface, and a microsequencer. The external memory interface is coupled to the cache and is adapted to be coupled to an external memory. The cache provides data to the external memory interface for storage in the external memory. The microsequencer is coupled to the data processor. In response to a trigger signal, the microsequencer causes the cache to flush data by sending the data to the external memory interface for transfer to the external memory.
[0015] A data processing system having runtime global push to persistence includes a system-on-chip and an external memory coupled to the system-on-chip. The system-on-chip includes a data processor with a cache, an external memory interface, and a microsequencer. The data processor selectively modifies data in the cache in response to executing instructions. The external memory interface is coupled to the cache and is adapted to be coupled to an external memory. The cache selectively provides modified data to the external memory interface for storage in the external memory. The microsequencer is coupled to the data processor and the cache, and in response to a trigger signal, the microsequencer causes the cache to flush the modified data by sending the modified data to the external memory interface for transfer to the external memory.
[0016] A method for providing runtime global push to persistence in a system-on-chip including a data processor with a cache coupled to an external memory interface through a data fabric includes receiving a trigger signal. In response to receiving the trigger signal, stopping the data processor. Flushing dirty data from the cache by sending a corresponding first write request to the data fabric. Flushing the pending write requests from the data fabric by sending all pending write requests to an external persistent memory. Providing a handshake between the data processor and the external persistent memory, thereby establishing runtime global push to persistence.
[0017] Figure 1 A prior art data processing system 100 is shown in block diagram form. The data processing system 100 generally includes a data processor 110 and an external memory 120. The data processor 110 includes an instruction fetch unit 111 labeled "IFU", an execution unit 112 labeled "EU", a cache 113, a memory controller 115 labeled "MC", and a physical interface circuit ("PHY") 116. The external memory 120 generally includes a first portion for storing an application program 121 including a FLUSH instruction 122, and a non-volatile memory 123 labeled "NVM".
[0018] The data processor 110 includes components whose operations are well known and which are not important for understanding the relevant operations of the present disclosure and will not be discussed further. The components of the data processor are connected together to exchange various signals, but Figure 1 only a set of signal flows relevant to understanding the problems of a known data processor are shown.
[0019] The cache 113 includes a set of lines that are divided into a tag portion, a data portion, and a status portion. The tag portion uses a subset of the bits of the memory address to help the cache 113 quickly index and find the accessed cache line from its cache lines. The data field stores the data corresponding to the cache line indicated by the TAG. The STATUS field stores information about the state of the lines in the cache, which allows the system to maintain data consistency in a complex data processing environment that also includes multiple processors and their associated caches in addition to different forms of main memory. There are several known cache coherence protocols, but the cache 113 implements the so-called "MOESI" protocol, which stores M, O, E, S, and I status bits that respectively indicate that the cache line is modified, owned, exclusive, shared, and / or invalid. As Figure 1 shown, the state indicating dirty data is the state where I = 0 and M = 1, as shown for Figure 1 cache line 114 in
[0020] The data processing system 100 implements known techniques for ensuring data durability. During the execution of the application 121, the instruction fetch unit 111 fetches a "FLUSH" command from the application 121. The instruction fetch unit 111 ultimately passes the FLUSH command to the execution unit 112 for execution. In response to the FLUSH command, the execution unit 112 causes the cache 113 to flush all of its dirty data to the external memory 120. It can do this by having an internal state machine that loops through the valid cache lines and writes them to the non-volatile memory 123, or the execution unit 112 itself can check all of the cache lines and write the contents of the dirty cache lines to the non-volatile memory 123. Using either technique, the cache 113 provides cache line information having an updated copy of the data to the memory controller 115, which ultimately provides the data to the external data bus to the non-volatile memory 123 through the PHY 116.
[0021] Figure 1The techniques shown have several problems or limitations. First, the technique relies on an application to initiate the FLUSH operation, and the application must know the hardware capabilities of the data processor 110. Additionally, if the data processing system 100 implements multiple processors, where distributed memory forming a memory pool is used to exchange data between the multiple processors, all processors must flush their caches and enforce serialization between different application threads to ensure reliable backup and recovery points, which can result in a substantial interruption of software operations. In larger and more complex systems such as a system-on-chip (SOC) with a complex data fabric for communication between different processing elements, it becomes difficult to predict the amount of time it takes for all writes in the system to propagate to the visible memory. Thus, such known systems do not appear sufficient to ensure data visibility and durability in the system when needed.
[0022] Figure 2 A high-level abstract diagram of a data processing system 200 with runtime global push to persistence according to some embodiments is shown in block diagram form. As Figure 2 shown, the data processing system 200 includes processors 210 and 220, attached accelerators 230, and a memory system 240. Processor 210 has an associated cache 211 and is connected to processor 220, accelerator 230, and memory system 240. The memory system 240 includes two memory levels, including a first level or "level 0" memory 241 and a second level or "level 1" memory 242. In one example, the level 0 memory 241 is a persistent memory such as non-volatile memory, and the level 1 memory 242 is a volatile memory such as high-speed dynamic random access memory (DRAM). In certain cases, it is desirable for the data processing system 200 to perform a runtime global push to persistence operation, in which dirty data in caches 211 and 222 and accelerator 230 is moved into the memory system 240 and thus becomes globally visible. As will be explained in more detail below, in response to an event indicating a need for a runtime global push to persistence, processors 210 and 220 respectively cause the dirty data in their caches 211 and 221 to be flushed into the memory system 240. As Figure 2 shown, cache 211 flushes the dirty data into memory 241 and memory 242 via paths 212 and 213 respectively, cache 221 flushes the dirty data into memory 241 and memory 242 via paths 222 and 223 respectively, and accelerator 230 will flush the dirty data into memory 241 and memory 242 indirectly through processor 210 via paths 232 and 233 respectively. Additionally, all "in-flight" memory operations in the data communication fabric, buffers, memory controllers, etc. are completed as part of the runtime global push to persistence to achieve data durability for the entire system.
[0023] Figure 3 A data processing system 300 according to some embodiments having an exemplary System-on-Chip (SOC) 310 with runtime global push to persistence is shown in block diagram form. The data processing system 300 generally includes an SOC 210 and a memory system 380. The SOC 310 includes a CPU complex 320, a fabric 330, a set of input / output (I / O) controllers 340, a Unified Memory Controller (UMC) 350, a Coherence Network Layer Interface (CNLI) 360, and a Global Memory Interface (GMI) controller 370. The CPU complex 310 includes one or more CPU cores, each having one or more dedicated internal caches, where a shared cache is shared among all CPU cores. The fabric 330 includes a coherence master block 331, an Input / Output Memory Slave block (IOMS) 333, a power / interrupt controller 334, a Coherent AMD Socket eXtender (CAKE) 335, a coherence slave block 336, an ACM 337, and a coherence slave block 338, all of which are interconnected by a fabric transport layer 332. The I / O controllers 340 include various controllers for protocols such as Peripheral Component Interconnect Express (PCIe) and their physical layer interface circuits. The UMC 350 performs command buffering, reordering, and timing qualification enforcement for efficient utilization of the bus to external memories such as Double Data Rate (DDR) memories and / or Non-Volatile Dual In-line Memory Modules with Persistent Storage (“NVDIMM-P”) memories. The CNLI 360 routes traffic to external coherence memory devices. The GMI controller 370 performs inter-chip communication to other SOCs that have their own attached storage visible to all processors in the memory map. The memory system 380 includes a DDR / NVDIMM-P memory 381 connected to the UMC 350 and a Compute Express Link (CXL) device 382 connected to the CNLI 360.
[0024] The SOC 310 is an exemplary SOC that shows the complexity of the fabric 330 for connecting various data processors, memories, and I / O components to various storage points for in-process write transactions. For example, the coherence slave blocks 336 and 338 support various memory channels and implement coherence and transaction ordering as well as runtime global push to persistence as will be described later. In an exemplary embodiment, these coherence slave blocks track coherence and address conflicts and support, for example, 256 outstanding transactions.
[0025] Figure 4Another data processing system 400 having a SOC 410 with runtime global push to persistence is shown in block diagram form. The data processing system 400 includes a SOC 410 and a memory system 490. The SOC 410 generally includes a processor layer 420, an interconnect fabric 430, a coherence network layer interface (CNLI) circuit 440, a unified memory controller (UMC) 450, a data input / output block 460, a physical interface layer 470, a microsequencer 480, and a memory system 490.
[0026] The processor layer 420 includes a CPU complex 421, a cache coherent memory 422 labeled "CCM", and a power / interrupt controller 423. The CPU complex 421 includes one or more CPU cores, each of which will typically have its own dedicated internal cache. In some embodiments, the dedicated internal cache includes both a first level 1 (L1) cache and a second level (L2) cache connected to the L1 cache. The lowest level cache in each processor core or multiple processor cores in the CPU complex 421 has an interface to the CCM 422. In some embodiments in which each CPU core has a dedicated internal L1 cache and L2 cache, the CCM 422 is a third level (L3) cache shared among all processors in the CPU complex 421. The power / interrupt controller 423 has a bi-directional connection for receiving register values and settings and signaling events such as interrupts and resets to circuits in the SOC 410, and may also be directly connected to other elements in the SOC 410 via a dedicated or special purpose bus.
[0027] The interconnect fabric 430 includes a fabric transfer layer 431, an input / output (I / O) master / slave controller 432 labeled "IOMS", an I / O hub 433 labeled "IOHUB", a peripheral component interconnect express (PCIe) controller 434, an accelerator cache coherence interconnect controller 435 labeled "ACM", and coherence slave circuits 436 and 437 each labeled "CS". The fabric transfer layer 431 includes an upstream port connected to the downstream port of the CCM 422, an upstream port connected to the power / interrupt controller 423, and four downstream ports. The IOMS 432 has an upstream port connected to the first downstream port of the fabric transfer layer 431 and a downstream port. The I / O hub 433 has an upstream port connected to the downstream port of the IOMS 432 and a downstream port. The PCIe controller 434 has an upstream port connected to the downstream port of the IOHUB 433 and a downstream port. The ACM 435 has an upstream port connected to the second downstream port of the fabric transfer layer 431 and a downstream port for conveying CXL cache transactions labeled "CXL.cache". The CS 436 has an upstream port connected to the third downstream port of the fabric transfer layer 431 and a downstream port for conveying CXL memory transactions labeled "CXL.mem". The CS 437 has an upstream port connected to the fourth downstream port of the fabric transfer layer 431 and a downstream port. The IOMS 432 is an advanced controller for input / output device access and may include an input / output memory management unit (IOMMU) for remapping memory addresses to I / O devices. The IOHUB 433 is a storage device for I / O access. The PCIe controller 434 performs I / O access according to the PCIe protocol and allows a deep hierarchy of PCIe switches, bridges, and devices in a deep PCIe fabric. The PCIe controller 434 in combination with firmware running on one or more processors in the CPU complex 421 can form a PCIe root complex. The ACM controller 435 receives and fulfills cache coherence requests from one or more external processing accelerators via a communication link. The ACM controller 435 instantiates a full CXL master agent that has the ability to make and fulfill memory access requests to memory attached to the SOC 410 or other accelerators using a full set of CXL protocol memory transaction types (see Figure 6 ). The CS 436 and 437 route other memory access requests initiated from the CPU complex 421, where the CS 436 routes CXL traffic and the CS 437 routes local memory traffic.
[0028] The CNLI circuit 440 has a first upstream port connected to the downstream port of the ACM 435, a second upstream port connected to the downstream port of the CS 436, and a downstream port. The CNLI circuit 440 performs network layer protocol activities for the CXL fabric.
[0029] The UMC 450 has an upstream port connected to the downstream port of the CS 437 and a downstream port for connecting to an external memory through a physical interface circuit. Figure 4 Not shown. The UMC 450 performs command buffering, reordering, and timing qualification enforcement to effectively utilize the downstream port of the UMC 450 and the bus between the DDR memory and / or NVDIMM-P memory.
[0030] The data input / output block 460 includes an interconnect block 461 and a set of digital I / O ("DXIO") controllers labeled 462 to 466. The DXIO controllers 462 to 466 perform data link layer protocol functions associated with PCIe or CXL transactions, as appropriate. The DXIO controller 462 is associated with the PCIe link and has an independent PCIe-compliant physical interface circuit between its output and the PCIe link ( Figure 4 not shown).
[0031] The physical interface circuit (PHY) 470 includes four separate PHY circuits 471 to 474, each PHY circuit connected between a corresponding DXIO controller and a corresponding I / O port of the SOC 410 and adapted to connect to different external CXL devices. The PHYs 471 to 474 perform physical layer interface functions according to the CXL communication protocol.
[0032] The microsequencer 480 has a first input for receiving a signal labeled "terminal event trigger", a second input for receiving a signal labeled "non-terminal event trigger", and a multi-signal output port connected to various circuits in the SOC 410 for providing control signals to be further described below. The SOC 410 includes circuits that generate the terminal event trigger signal and the non-terminal event trigger signal. These circuits are not shown in Figure 4 but will be further described below.
[0033] Memory system 490 includes memory 491 operating as a CXL MEM device 0 connected to the downstream port of PHY 471, memory 492 operating as a CXL MEM device 1 connected to the downstream port of PHY 471, a CXL accelerator coherence master controller (ACM) 493 connected to the downstream port of PHY 473, a CXL ACM 494 connected to the downstream port of PHY 474, and a storage class memory in the form of a double data rate (DDR) DRAM / NVDIMM-P memory 495 connected to the downstream port of UMC 450.
[0034] Obviously, the data interfaces and distributed memory hierarchies of contemporary SOCs such as SOC 410 are extremely complex, hierarchical, and distributed. This complex interconnect fabric poses challenges to supporting runtime global push to persistence in a data processing system, and these challenges are addressed by the techniques described herein.
[0035] Microsequencer 480 is a hardware controller that relieves application software, the operating system, or system firmware of the burden of tasks that identify and respond to runtime global push to persistence requirements. First, the microsequencer causes all caches in SOC 410 to flush their dirty data by writing the updated content to memory. Flushing can be done in two ways: by firmware running on microsequencer 480, which checks the status of each line in each cache in SOC 410 and selectively causes dirty cache lines to be written to memory; or preferably by an explicit hardware signal to each cache in the cache, which causes these caches to automatically flush the dirty data by checking all cache lines and writing the cache lines containing dirty data to main memory. Per-cache-way dirty indication can accelerate the cache flushing process. Those cache ways whose dirty indication is cleared can be skipped by the cache flushing process.
[0036] Second, the microsequencer 480 causes every in-flight memory write transaction present somewhere in the interconnect fabric 430 or other interface circuitry to complete and drain to the external persistent memory through any buffer points in the interconnect fabric. In one example, the fabric transport layer 431 may have buffers that store read and write commands to the memory system. In response to a trigger signal, the microsequencer 480 causes the fabric transport layer 431 to flush all writes to the memory system and allow those writes to overtake any reads. In another example, the UMC 450 stores DRAM writes in its internal command queue. In response to a runtime push-to-persistence trigger, the microsequencer 480 causes the UMC 450 to send all writes to the memory without acting on any pending reads while continuing to follow efficiency protocols such as preferring to combine writes to open pages rather than to closed pages.
[0037] The microsequencer 480 responds differently to two types of triggers. The first type of trigger is a terminal event trigger. A terminal event trigger is an event such as a hazard reset request, an impending power failure, a thermal overload, or a "trip" condition, or any other condition indicating a need to immediately terminate the operation of the data processing system 400. In response to a terminal event trigger condition, the microsequencer 480 performs two actions. First, the microsequencer stops the operation of all data processors. Then, the microsequencer commands all buffers in the caches and data fabrics to flush all pending memory transactions to the persistent memory. In this way, the microsequencer 480 prioritizes speed over low power consumption because the data needs to be pushed to the persistent non-volatile memory as quickly as possible.
[0038] The second type of trigger is a non-terminal event trigger. A non-terminal event trigger is a non-critical event such as encountering a certain address, detecting low processor utilization, encountering a certain time of day, detecting a certain elapsed time since the previous runtime global push-to-persistence operation, or detecting a certain level of "dirtyness" in one or more caches. A non-terminal event trigger allows the system to periodically push highly important data such as log records, shadow paging, etc. to the external persistent memory. In the case of a non-terminal event trigger, the microsequencer 480 does not stop any data processor cores, but instead causes the caches to send all dirty data in any cache to the memory interface without stopping the data processors, allows the data fabric to naturally drain the data, and restarts the operation without a reset. Thus, in response to a non-terminal trigger event, the microsequencer 480 implements a runtime global push-to-persistence in cases where only low power consumption is required.
[0039] In response to a persistent loss (which may be identified by the platform by setting a "loss" flag in non-volatile memory), the application software restarts at the last known trusted state, i.e., the application software performs checkpoint rollback and replay. For example, in some configurations, a "persistent loss" error is logged, and at startup, the system basic input-output system (BIOS) firmware identifies the persistent loss and reports it via the Advanced Configuration and Power Interface (ACPI) "NFIT" object. In other embodiments, the "persistent loss" is captured in a log so that the operating system can be directly aware of the event.
[0040] Figure 5 A flowchart of a runtime global push to persistence process 500 available in a system-on-chip (SOC) according to some embodiments is shown. The runtime global push to persistence process 500 is initiated, for example, in response to a trigger event indicated by a trigger signal received by the SOC. In action box 510, if the trigger is a terminal event trigger, the data processor is stopped by, for example, stopping each of a set of CPU cores. In action box 520, dirty data is flushed from the cache subsystem by sending a write request for the dirty data to the data fabric. Next, in action box 530, dirty data from any CXL ACM controller is flushed from the coherent memory attached to the external memory interface. This action includes reading the dirty data from the external CXL memory device into the on-chip data fabric. After this operation is completed, in action box 540, the runtime global push to persistence process flushes all pending write requests in the data fabric by sending a write request to an external persistent memory. The external persistent memory can be, for example, a CXL type 3 memory device (CXL memory without an accelerator) or an NVDIMM-P. Then, at action box 550, the system provides a handshake with the CXL type 3 memory device.
[0041] Figure 6 A data processing system 600 according to some embodiments is shown in block diagram form, which shows Figure 4 the SOC 410 executing Figure 5 the runtime global push to persistence process 500. Figure 6Reference numerals for the various blocks of data processing system 400 are not shown. In a first step, shown by dashed circle 610, all processors in CPU complex 421 are stopped. In a second step, shown by dashed arrow 620, dirty data is flushed from the cache subsystems of each processor in CPU complex 421 by sending write requests with the dirty data to the data fabric. These requests flow through fabric transport layer 431 and are stored in CS 436 (if the memory is mapped to CXL memory device 491 or 492) or CS 437 (if the memory is mapped to NVDIMM-P 495). In a third step, dirty data from the ACM cache (if present) is flushed and this data is sent through the data fabric to CS 436 or CS 437, as shown by arrow 630. In a fourth step, the data fabric is flushed by sending data to CXL memory device 491 or 492 or to NVDIMM-P 495, as shown by arrow 640. Finally, SOC 410 provides handshakes with CXL memory devices 491 and 492 according to the CXL protocol.
[0042] Figure 7 Another data processing system 700 having a SOC 710 with runtime global push to persistence is shown in block diagram form. SOC 710 is more highly integrated than Figure 4 SOC 410 and is organized into four fairly autonomous quadrants 720, 730, 740, and 750 labeled "Quadrant 0", "Quadrant 1", "Quadrant 2", and "Quadrant 3", respectively. Quadrants 720, 730, 740, and 750 have respective DDR memory interfaces 722, 732, 742, and 752, as well as interfaces to external cache coherence devices (CCDs) 760, 770, 780, and 790. However, SOC 710 has a set of shared ports labeled "P0", "P1", "P2", and "P3" for connection to external CXL devices such as CXL type 3 attached memory, which will be non-volatile and capable of storing data for durability. Additionally, because at least some triggering events, such as an impending chip-wide power loss or a critical reset, require flushing of all quadrants, a common microsequencer 760 conveniently provides control signals to enforce runtime global push to persistence by flushing dirty data from both quadrant-specific resources and shared resources such as a common chip-wide data fabric.
[0043] Figure 8 Shown in block diagram form Figure 4The terminal event trigger generation circuit 800 available in the SOC 410. The terminal event trigger generation circuit 800 includes an OR gate 810, an inverter 820, an OR gate 822, a temperature sensor 830, a comparator 832, and an AND gate 834. The OR gate 810 is a 3-input OR gate that has a first input for receiving a reset signal labeled "reset", a second input for receiving a signal labeled "power failure", a third input for receiving a signal labeled "thermal trip", and an output for providing a terminal event trigger signal. The inverter 820 has an input for receiving a signal labeled "POWER_GOOD" and an output. The OR gate 822 has a first input connected to the output of the inverter 820, a second input for receiving a signal labeled "DROOP_DETECTED", and an output connected to the second input of the OR gate 810 for providing a power failure signal thereto. The temperature sensor 830 has an output for providing a measured temperature sensing signal. The comparator 832 has a non-inverting input connected to the output of the temperature sensor 830, an inverting input for receiving a value labeled "temperature trip threshold", and an output. The AND gate 834 is a 2-input AND gate that has a first input connected to the output of the comparator 832, a second input for receiving a signal labeled "thermal trip enable", and an output connected to the third input of the OR gate 810 for providing a thermal trip signal thereto.
[0044] The terminal event trigger generation circuit 800 provides a terminal event trigger signal in response to a reset condition, a power loss condition, or a thermal trip condition. The reset condition is indicated by the activation of a reset signal, which can be generated by, for example, a software reset or a hardware reset caused by, for example, a user pressing a reset button. The power loss condition is indicated by the activation of a system power signal, as Figure 8 shown, by the deactivation of a POWER_GOOD signal from a motherboard or system board, or by an on-chip condition such as detecting a droop in a power source. The thermal trip condition is detected by the on-chip temperature sensor 830 when enabled, which indicates that the temperature of the SOC exceeds a terminal thermal trip threshold. In any case, the terminal event trigger generation circuit 800 provides a terminal event trigger signal in response to a critical system condition in which an entire system shutdown will occur or may occur. In such a case, the activation of the terminal event trigger signal will notify the microsequencer that data saving should be performed as quickly as possible to avoid loss of the system state.
[0045] Obviously, the terminal event trigger generation circuit 800 shows a representative set of conditions that constitute a terminal event, but other embodiments will detect only some of these conditions, while other conditions indicating a terminal event will be detected in other embodiments.
[0046] Figure 9 is shown in block diagram form in Figure 4 the SOC 410 and Figure 7 the non-terminal event trigger generation circuit 900 available in the SOC 700. The non-terminal event trigger generation circuit 900 generally includes an evaluation circuit 910, an address trigger circuit 920, an activity trigger circuit 930, a date and time trigger circuit 940, an elapsed time trigger circuit 950, and a cache dirtiness trigger circuit 960.
[0047] The evaluation circuit 910 includes a set of inputs for receiving trigger signals and an output for providing a non-terminal event trigger signal. The evaluation circuit 910 generally implements a logical OR operation among the inputs, where the evaluation circuit activates the non-terminal event trigger signal in response to the activation of any one of these inputs. Depending on the design, the evaluation circuit may also have a resetable clocked latch such that the non-terminal event trigger signal is activated only at a certain edge of the clock signal and is reset in response to, for example, a handshake signal indicating the completion of a runtime global push to a persistent operation.
[0048] The address trigger circuit 920 includes a trigger address register 921 and a comparator 922. The trigger address register 921 is programmable in a privileged execution state and has an output for providing the stored trigger address. The comparator 922 is a multi-bit digital comparator that has a first input for receiving an address signal labeled "address", a second input connected to the output of the trigger address register 921, and an output for providing a signal labeled "address trigger" to the first input of the evaluation circuit 910. The address trigger circuit 920 is a simple example of a trigger circuit that allows an application or an operating system to trigger a runtime global push to a persistent operation by accessing a certain address. In a data processing system having multiple CPU cores and a multi-threaded operating system, the exemplary circuits in the address trigger circuit 920 will be replicated for each CPU core.
[0049] The activity trigger circuit 930 includes a set of performance counters 931 and a logic circuit 932. The performance counters 931 respond to a set of activity signals representing the activity of the CPU cores and use corresponding counters to aggregate individual events. The performance counters 931 have an output for providing the status of the counters. The logic circuit 932 has an input connected to the output of the performance counters 931 and an output for providing a signal labeled "low utilization" to the second input of the evaluation circuit 910. In Figure 9In the example shown, the logic circuit 932 can determine which activities constitute important events. In the case of the low utilization signal, the logic circuit 932 can count the instructions executed per unit time and activate the low utilization signal in response to detecting that the instructions executed per unit time are less than a threshold. As in the case of the address trigger circuit 920, the exemplary circuit 930 will be replicated for each CPU core in the multi-core system.
[0050] The date and time trigger circuit 940 includes a real-time clock circuit 941 labeled "RTC", a date and time register 942, and a comparator 943. The RTC 941 has an output for providing a digital count value representing the current moment. The register 942 has an output for providing a selected moment, such as 4:00 am. The comparator 943 has a first input connected to the output of the real-time clock 941, a second input connected to the output of the register 942, and an output for providing a moment matching symbol labeled "TOD" to a third input of the evaluation circuit 910. The date and time trigger circuit 940 is an example of a non-terminal event that does not need to be replicated for each CPU core in the multi-CPU core system.
[0051] The elapsed time trigger circuit 950 includes a timer 951. The timer 951 has a reset input for receiving a signal labeled "last trigger", a clock input for receiving a clock signal labeled "clock", and a terminal count (TC) output for providing a signal labeled "next trigger" to a fourth input of the evaluation circuit 910. The elapsed time trigger circuit 950 is another example of a non-terminal event that will not need to be replicated for each CPU core in the multi-core system.
[0052] The cache dirtiness trigger circuit 960 includes a cache 961, an encoder 962, a cache dirty watermark 963, and a comparator 964. The cache 961 is a cache in a CPU core or a cache shared among multiple CPU cores. Figure 9In the example shown, cache 961 implements the aforementioned MOESI state protocol. Encoder 962 has an input connected to cache 961 and an output, and counts the number of dirty cache lines. This logic is slightly more complex than the logic shown in cache dirtyness trigger circuit 960, because encoder 961 will count not only the number of cache lines with the M bit set, but also the number of cache lines with the M bit set and the I bit cleared. Cache dirty watermark register 963 is programmable in the privileged execution state and has an output for providing the stored cache dirty watermark. Comparator 964 has a positive input connected to the output of encoder 962, a negative input connected to the output of cache dirty watermark register 963, and an output connected to a fifth input of evaluation circuit 910 for providing a signal labeled "cache dirty" thereto. Comparator 964 provides its output in an active logic state in response to the number of dirty lines in cache 961 exceeding the cache dirty watermark. Cache dirtyness trigger circuit 960 will be replicated for each cache in the SOC.
[0053] Obviously, non-terminal event trigger generation circuit 900 shows a representative set of conditions that constitute non-terminal events, but other embodiments will detect only some of these conditions, while other embodiments will detect other conditions indicating non-terminal events. Additionally, the evaluation circuit may implement a simple logical OR function or may implement a fuzzy logic evaluation based on a combination of factors.
[0054] Accordingly, a data processing system, SOC, and method for implementing a runtime global push to persistent operations have been disclosed. This runtime operation causes important data to be flushed from the cache hierarchies of each CPU core and then, along with other pending operations, to be flushed from the on-chip data fabric and stored in external persistent memory. The runtime global push to persistent operations allows for the protection and preservation of important data and allows the data processing system to back up to a known operating point in the event of a sudden or unexpected system failure. In various embodiments, there are two types of operations that can trigger a runtime global push to persistent operations: terminal events and non-terminal events. The specific trigger events supported by the SOC can vary between embodiments.
[0055] Although microsequencer 480 and its associated trigger generation circuits 800 and 900 have been described as hardware circuits, their functionality can be implemented using various combinations of hardware and software. Some of the software components can be stored in a computer-readable storage medium for execution by at least one processor. Additionally, Figure 5 Some or all of the methods described can also be managed by instructions stored in a computer-readable storage medium and executed by at least one processor. Figure 5Each of the operations shown can correspond to instructions stored in a non-transitory computer memory or a computer-readable storage medium. In various embodiments, the non-transitory computer-readable storage medium includes a magnetic or optical disk storage device, a solid-state storage device such as flash memory, or one or more other non-volatile memory devices. The computer-readable instructions stored on the non-transitory computer-readable storage medium can be source code, assembly language code, object code, or other instruction formats that are interpreted by or executable by one or more processors.
[0056] The SOC 410 and the microsequencer 480 or any part thereof can be described or represented by a computer-accessible data structure in the form of a database or other data structure that can be read by a program and is directly or indirectly used in the fabrication of integrated circuits. For example, the data structure can be a behavioral-level description or a register transfer level (RTL) description of the hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool, which can synthesize the description to produce a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates, which also represents the functionality of the hardware including the integrated circuit. The netlist can then be placed and routed to produce a dataset that describes the geometry to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce the integrated circuit. Alternatively, the database on the computer-accessible storage medium can be a netlist (with or without a synthesis library) or a dataset (as required) or Graphics Data System (GDS) II data.
[0057] Although specific embodiments have been described, various modifications to these embodiments will be apparent to those skilled in the art. For example, the conditions for generating a terminal event trigger signal or a non-terminal event trigger signal can vary between embodiments. Additionally, in various embodiments, the satisfaction of two or more conditions can be used to generate a trigger event signal. The interconnect protocols described herein are exemplary, and in other embodiments, other protocols can be used. The supported SOC topologies and cache hierarchies will vary in other embodiments. The status bits used to indicate dirty cache lines will also vary in other embodiments. As shown and described herein, various circuits are directly connected together, but in other embodiments, they can be indirectly connected through various intermediate circuits, and signals can be transmitted between the circuits through various electrical and optical signal technologies.
[0058] Accordingly, the appended claims are intended to cover all modifications that fall within the scope of the disclosed embodiments of the disclosed embodiments.
Claims
1. An on-chip system with runtime global push to persistence, the on-chip system comprising: A data processor having a cache; An external memory interface coupled to the cache and adapted to be coupled to an external memory; Wherein the cache provides data to the external memory interface for storage in the external memory; And A microsequencer coupled to the data processor, wherein in response to a trigger signal, the microsequencer causes the cache to flush the data by sending the data to the external memory interface for transfer to the external memory, wherein: The data processor is coupled to the external memory interface through a data fabric, the data fabric including at least one buffer that temporarily stores write requests to the external memory through the external memory interface; and The data fabric is further coupled to the microsequencer, and in response to the trigger signal, the microsequencer further causes the data fabric to flush the data by sending data associated with the write requests stored in the at least one buffer to the external memory interface for transfer to the external memory.
2. An on-chip system with runtime global push to persistence, the on-chip system comprising: A data processor having a cache; An external memory interface coupled to the cache and adapted to be coupled to an external memory; Wherein the cache provides data to the external memory interface for storage in the external memory; And A microsequencer coupled to the data processor, wherein in response to a trigger signal, the microsequencer causes the cache to flush the data by sending the data to the external memory interface for transfer to the external memory, wherein: The trigger signal includes one of a terminal event trigger signal and a non-terminal event trigger signal; In response to the terminal event trigger signal, the microsequencer stops the data processor and then sends the data to the external memory interface for transfer to the external memory; and In response to the non-terminal event trigger signal, the microsequencer sends the data to the external memory interface for transfer to the external memory without stopping the data processor.
3. The on-chip system according to claim 2, wherein: The on-chip system activates the terminal event trigger signal in response to one of detecting a power failure and detecting a thermal trip condition.
4. The on-chip system according to claim 2, wherein: The on-chip system activates the non-terminal event trigger signal in response to one of a change from a normal operating state to a reset state, elapsed time since a previous trigger event, a predetermined date and time, a state of at least one performance counter, and detecting an access to a predetermined address.
5. The on-chip system according to claim 4, wherein: The system-on-chip activates the non-terminal event trigger signal in response to the state of the at least one performance counter, and the state of the at least one performance counter indicates low utilization of the data processor.
6. The system-on-chip according to claim 2, wherein: The system-on-chip generates the non-terminal event trigger signal in response to multiple conditions.
7. The system-on-chip according to claim 6, wherein the multiple conditions include at least one of the following: the execution state of at least one software thread, and the dirty condition of the cache.
8. The system-on-chip according to claim 2, wherein in response to the trigger signal, the microsequencer causes the cache to selectively flush the data based on whether the address of the modified data corresponds to at least one address range.
9. A data processing system with runtime global push to persistence, the data processing system comprising: A system-on-chip; And An external memory coupled to the system-on-chip, Wherein the system-on-chip includes: A data processor having a cache, wherein the data processor selectively modifies data in the cache in response to executing instructions; An external memory interface coupled to the cache and the external memory; Wherein the cache selectively provides the modified data to the external memory interface for storage in the external memory; and A microsequencer coupled to the data processor, wherein in response to a trigger signal, the microsequencer causes the data processor to stop executing instructions, and then flushes the modified data from the cache by sending the modified data to the external memory interface for transmission to the external memory, wherein: The data processor is coupled to the external memory interface through a data fabric, the data fabric includes at least one buffer, and the at least one buffer temporarily stores a request to send the modified data to the external memory interface for transmission to the external memory; and The data fabric is further coupled to the microsequencer, and in response to the trigger signal, the microsequencer further causes the data fabric to flush the modified data by sending the modified data associated with the request stored in the at least one buffer to the external memory interface for transmission to the external memory.
10. The data processing system according to claim 9, wherein: The external memory includes non-volatile memory.
11. The data processing system according to claim 9, wherein: In response to the data processor accessing a data element, the cache fetches the data in the cache and sets the cache line in the cache to an unmodified state; and In response to a write access to the data element, the cache modifies the data element according to the write access and sets the data element to a modified state. The microsequencer causes the data processor to flush the cache by only flushing cache lines in the modified state.
12. The data processing system according to claim 9, wherein: In response to decoding a flush instruction, the data processor further causes the cache to flush the modified data by sending the modified data to the external memory interface for transfer to the external memory.
13. A data processing system having runtime global push to persistence, the data processing system comprising: A system-on-chip; And An external memory coupled to the system-on-chip, Wherein the system-on-chip includes: A data processor having a cache, wherein the data processor selectively modifies data in the cache in response to executing instructions; An external memory interface coupled to the cache and the external memory; Wherein the cache selectively provides modified data to the external memory interface for storage in the external memory; and A microsequencer coupled to the data processor, wherein in response to a trigger signal, the microsequencer causes the data processor to stop execution of instructions and then flushes the modified data from the cache by sending the modified data to the external memory interface for transfer to the external memory, wherein: The trigger signal includes one of a terminal event trigger signal and a non-terminal event trigger signal; In response to the terminal event trigger signal, the microsequencer stops the data processor and then sends the modified data to the external memory interface for transfer to the external memory; and In response to the non-terminal event trigger signal, the microsequencer sends the modified data to the external memory interface for transfer to the external memory without stopping the data processor.
14. The data processing system according to claim 13, wherein: The trigger signal includes the terminal event trigger signal, and the system-on-chip activates the terminal event trigger signal in response to one of detecting a power failure and detecting a thermal trip condition.
15. The data processing system according to claim 13, wherein: The trigger signal includes the non-terminal event trigger signal, and the system-on-chip activates the non-terminal event trigger signal in response to one of a change from a normal operating state to a reset state, elapsed time since a previous trigger event, a predetermined date and time, a state of at least one performance counter, and detecting an access to a predetermined address.
16. The data processing system according to claim 15, wherein the system-on-chip generates the non-terminal event trigger signal in response to the state of the at least one performance counter, and wherein the state of the at least one performance counter indicates low utilization of the data processor.
17. The data processing system according to claim 13, wherein the trigger signal includes the non-terminal event trigger signal, and the system-on-chip generates the non-terminal event trigger signal in response to a plurality of conditions.
18. The data processing system according to claim 17, wherein the plurality of conditions includes at least one of the following: the execution state of at least one software thread, and the dirty condition of the cache.
19. The data processing system according to claim 13, wherein in response to the trigger signal, the microsequencer causes the cache to selectively flush the modified data based on whether the address of the modified data corresponds to at least one address range.
20. A method for providing a runtime global push to persistence in a system-on-chip including a data processor having a cache coupled to an external memory interface through a data fabric, the method comprising: Receiving a trigger signal, and in response to receiving the trigger signal: Stopping the data processor; Flushing dirty data from the cache by sending a corresponding first write request to the data fabric; Flushing the pending write requests from the data fabric by sending all pending write requests to an external persistent memory; Providing a handshake between the data processor and the external persistent memory, Thereby establishing the runtime global push to persistence; And Using a microsequencer coupled to the data processor to control the stopping, flushing the dirty data from the cache, flushing all pending write requests from the data fabric, and providing the handshake.
21. The method according to claim 20, wherein flushing the dirty data from the cache includes: Flushing the dirty data from a coherence memory coupled to the external memory interface by sending a corresponding second write request to the data fabric, and then flushing all pending write requests from the data fabric.
22. The method according to claim 20, further comprising: Generating the trigger signal in response to at least one of a terminal event trigger signal and a non-terminal event trigger signal.
23. The method according to claim 22, further comprising: Resetting the data processor after providing the handshake when the trigger signal is generated in response to the terminal event trigger signal.
24. The method according to claim 22, further comprising: Causing the data processor to restart operations after providing the handshake when the trigger signal is generated in response to the non-terminal event trigger signal.
25. A system-on-chip having a runtime global push to persistence, the system-on-chip comprising: A plurality of data processors, each of the plurality of data processors having a cache; An external memory interface coupled to the cache and adapted to be coupled to an external memory; Wherein the cache in each of the plurality of data processors provides corresponding dirty data to the external memory interface for storage in the external memory; And A microsequencer, the microsequencer being coupled to each of the data processors in the data processor, wherein in response to a trigger signal, the microsequencer causes the cache in each of the plurality of data processors to flush the corresponding dirty data by sending the corresponding dirty data to the external memory interface for transfer to the external memory.
26. The system-on-chip according to claim 25, wherein: Each of the plurality of data processors is coupled to the external memory interface through a data fabric, the data fabric including at least one buffer that temporarily stores write requests to the external memory through the external memory interface; and The data fabric is further coupled to the microsequencer, and in response to the trigger signal, the microsequencer further causes the data fabric to flush the dirty data by sending the data associated with the write requests stored in the at least one buffer to the external memory interface for transfer to the external memory.
27. The system-on-chip according to claim 25, wherein: The trigger signal includes one of a terminal event trigger signal and a non-terminal event trigger signal; In response to the terminal event trigger signal, the microsequencer stops each of the plurality of data processors and then sends the dirty data to the external memory interface for transfer to the external memory; and In response to the non-terminal event trigger signal, the microsequencer sends the dirty data to the external memory interface for transfer to the external memory without stopping the plurality of data processors.
28. The system-on-chip according to claim 27, wherein: The system-on-chip activates the terminal event trigger signal in response to one of: detecting a power failure and detecting a thermal trip condition.
Citation Information
Patent Citations
Global persistent flush
US20200192798A1