Software configurable processor device

By configuring controllers and configurable storage elements, FIFO queues or L1 caches are implemented for the processor cores of data center processor devices, solving the problem of uneven resource utilization of processor devices and improving the computing efficiency and performance of data centers.

CN121833552APending Publication Date: 2026-04-10INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CORP
Filing Date
2025-09-05
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing data center processor devices struggle to efficiently configure and manage processor cores and storage components when handling multiple applications, leading to uneven resource utilization and performance bottlenecks.

Method used

A configurable processor device is provided that identifies and configures multiple storage elements for the processor core to implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache by a configuration controller, and dynamically adjusts the data flow configuration of the processor core by utilizing configurable interconnect structures and storage elements.

Benefits of technology

It enables flexible configuration of processor devices, improves the resource utilization and performance of processor cores, adapts to the needs of different workloads, and enhances the overall computing efficiency of data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833552A_ABST
    Figure CN121833552A_ABST
Patent Text Reader

Abstract

The invention relates to a software configurable processor device. A software configurable processor device includes a plurality of processor cores having respective storage elements configurable to implement one or more first-in first-out (FIFO) queues or first-level (L1) cache blocks for respective processor cores of the plurality of processor cores. Configuration hardware is provided to configure a first storage element associated with a first processor core of the set of processor cores based on a configuration definition of the processor device to implement a set of FIFO queues for the first processor core in the first storage element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to software-configurable processor devices. Background Technology

[0002] A data center may include one or more platforms, each platform including at least one processor and an associated memory module. Each platform in the data center can facilitate the performance of any appropriate number of processes associated with the various applications running on the platform. These processes may be executed by the platform's processor and other related logic. Each platform may also include I / O controllers, such as network adapter devices, which can be used to send and receive data over the network for use by various applications. Summary of the Invention

[0003] According to one aspect of this disclosure, an apparatus is provided, comprising: a plurality of processor cores; a plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configurable to: alternatively implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and a configuration controller for: identifying a configuration definition, wherein the configuration definition defines a configuration for a processor core among the plurality of processor cores; and configuring a first storage element among the plurality of storage elements associated with a first processor core among the plurality of processor cores based on the configuration definition to implement a set of FIFO queues for the first processor core in the first storage element.

[0004] According to another aspect of this disclosure, a method for configuring a configurable processor device is provided, the method comprising: receiving configuration definition data, wherein the configuration definition data describes a specific configuration for the configurable processor device, the configurable processor device including: a plurality of processor cores; a configurable interconnect structure for interconnecting components of the processor device; and a plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configurable to: alternatively implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and configuring at least the configurable interconnect structure and the plurality of configurable storage elements based on the specific configuration to define a data flow for the plurality of processor cores.

[0005] According to another aspect of this disclosure, a system is provided, including means for performing the above-described method.

[0006] According to another aspect of this disclosure, a system is provided, comprising: a processor device including: a plurality of processor cores; a plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configurable to: alternatively implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; a configuration controller for configuring a first storage element among the plurality of storage elements associated with a first processor core among the plurality of processor cores based on a configuration definition to implement a set of FIFO queues for the first processor core in the first storage element; and routing hardware for writing instructions and data to the set of FIFO queues based on the configuration definition, wherein the instructions are associated with a user application and the data will be consumed during the execution of the instructions by the first processor core. Attached Figure Description

[0007] Figure 1 A block diagram of components of a data center according to certain embodiments is shown.

[0008] Figure 2 This is a simplified block diagram of an example processor core with a just-in-time (JIT) first-in-first-out (FIFO) queue.

[0009] Figures 3A to 3B This is a simplified block diagram illustrating the FIFO queue and processing elements of an example processor core.

[0010] Figure 3C It is a simplified block diagram illustrating a set of interconnected processor cores.

[0011] Figure 4A This is a simplified block diagram illustrating an example processor device.

[0012] Figures 4B to 4C This is a simplified block diagram illustrating example routing hardware used with the example software-configurable processor device.

[0013] Figures 5A to 5B This is a simplified block diagram illustrating a sample reconfiguration of a sample processor device based on the configuration definition provided by the sample software.

[0014] Figure 6 This is a simplified flowchart illustrating an example technique for routing software threads to a configurable core of a processor device.

[0015] Figures 7A to 7B This is a simplified block diagram illustrating an example configuration of a configurable processor core.

[0016] Figure 8 This is a simplified flowchart illustrating an example technique for configuring processor cores based on configuration definitions.

[0017] Figure 9 This is a simplified block diagram illustrating the processing pipeline of an example processor device.

[0018] Figure 10 This is a simplified block diagram illustrating the interconnected processor cores in an example processor device.

[0019] Figure 11 This is a simplified block diagram illustrating an example configuration of the processor core in an example processor device.

[0020] Figure 12 This is a simplified block diagram illustrating the interaction between two processor cores in an example processor device based on a configuration.

[0021] Figure 13 This is a simplified block diagram illustrating the interaction between two processor cores in an example processor device based on another configuration.

[0022] Figures 14A to 14B This is a simplified block diagram illustrating an example of recursion in an example processor core of a processor device.

[0023] Figure 15 This is a block diagram illustrating a portion of a quantum computing architecture.

[0024] Figure 16A This is a block diagram illustrating an example processor pipeline.

[0025] Figure 16B It is a block diagram illustrating an example processor core.

[0026] Figure 17A The diagram illustrates both an exemplary ordered pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to an example embodiment.

[0027] Figure 17B This is a block diagram illustrating both an exemplary embodiment of an ordered architecture core to be included in a processor according to an example embodiment and an exemplary register renaming, out-of-order issue / execution architecture core.

[0028] Figure 18 It is a block diagram of a more specific, exemplary ordered core architecture, which will be one of several logical blocks in the chip (including other cores of the same type and / or different types).

[0029] Figure 19This is a block diagram of a processor according to an example embodiment, which may have more than one core, may have an integrated memory controller, and may have integrated graphics.

[0030] Figures 20 to 22 It is a block diagram of an exemplary computer architecture.

[0031] Figure 23 This is a block diagram comparing some example embodiments with the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set.

[0032] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation

[0033] Modern data centers provide the critical computing infrastructure for the massive and ever-growing volume of online services and applications upon which modern society relies. Data centers enable "warehouse computing" environments, where facilities house thousands of servers and network devices organized to ensure high performance, scalability, and reliability. Equipped with advanced cooling systems, redundant power supplies, and cutting-edge security measures, modern data centers are designed to provide scalable and uninterrupted access to a wide range of data and services. Beyond their robust physical infrastructure, modern data centers leverage sophisticated software solutions to optimize operations and improve efficiency. Virtualization, automation, and artificial intelligence are used to manage workloads, predict failures, and reduce energy consumption. With the rise of edge computing, data centers are becoming increasingly distributed, bringing processing power closer to end users to minimize latency and improve performance, as well as enabling other examples of functionalities.

[0034] Figure 1A block diagram of components of an example data center 100 system according to certain embodiments is illustrated. In the depicted embodiments, data center 100 includes multiple platforms 102, a data analytics engine 104, and a data center management platform 106 coupled together via a network 108. Platform 102 may include platform logic 110 having one or more processor devices 112 (e.g., a central processing unit (CPU) or a configurable processor device discussed herein), memory 114 (which may include any number of different modules), chipset 116, communication interface 118, and any other suitable hardware and / or software to execute a hypervisor 120 or other operating system capable of executing processes associated with applications running on platform 102. In some embodiments, platform 102 may serve as a host platform for one or more guest systems 122 that invoke these applications. The platform may be logically or physically subdivided into clusters, and these clusters may be enhanced through dedicated network accelerators and the use of Compute Express Link (CXL) memory semantics to make such clusters more efficient, as well as to implement other example enhancements.

[0035] Platform 102 may include platform logic 110. Platform logic 110 includes one or more processor devices 112, memory 114, one or more chipsets 116, and communication interfaces 118, as well as other logic that implements the functionality of platform 102. Although three platforms are illustrated, data center 100 may include any suitable number of platforms. In various embodiments, platform 102 may reside on a circuit board mounted in a chassis, rack, composable server, decomposed server, or other suitable structure, which includes multiple platforms coupled together via a network 108, which may include, for example, a rack or backplane switch.

[0036] Processor device 112 may include any suitable number of processor cores. These cores may be coupled to each other, to memory 114, to at least one chipset 116, and / or to communication interface 118 via one or more controllers residing on processor device 112 and / or chipset 116. In a particular embodiment, processor device 112 is embodied within a socket permanently or removably coupled to platform 102. Although four processor devices are shown, platform 102 may include any suitable number of processor devices. In some implementations, applications executing using the processor devices may include physical layer management applications that can implement software-based custom configurations of one or more interconnected physical layers, wherein the one or more interconnects are used to couple the processor devices (or associated processor devices) to one or more other devices in a data center system.

[0037] Memory 114 may include any form of volatile or non-volatile memory, including but not limited to magnetic media (e.g., one or more magnetic tape drives), optical media, random access memory (RAM), read-only memory (ROM), flash memory, removable media, or any other suitable one or more local or remote memory components. Memory 114 may be used for short-term, medium-term, and / or long-term storage of platform 102. Memory 114 may store any suitable data or information used by platform logic 110, including software embedded in computer-readable media and / or coded logic (e.g., firmware) contained in hardware or otherwise stored. Memory 114 may store data used by the core of processor device 112. In some embodiments, memory 114 may also include storage means for instructions that can be executed by the core of processor device 112 or other processing elements (e.g., logic residing on chipset 116) to provide functionality associated with components of platform logic 110. Additionally or alternatively, chipset 116 may include memory that may have any of the features described herein with respect to memory 114. Memory 114 may also store the results and / or intermediate results of various calculations and determinations performed by processing elements on processor device 112 or chipset 116. In various embodiments, memory 114 may include one or more modules of system memory coupled to the processor device via a memory controller (which may be external to or integrated with processor device 112). In various embodiments, one or more specific modules of memory 114 may be dedicated to a specific processor device 112 or other processor device, or may be shared among multiple processor devices 112 or other processor devices.

[0038] Platform 102 may also include one or more chipsets 116, which include any suitable logic supporting the operation of processor device 112. In various embodiments, chipset 116 may reside in the same package as processor device 112, or in one or more different packages. The chipset may support any suitable number of processor devices 112. Chipset 116 may also include one or more controllers to couple other components of platform logic 110 (e.g., communication interface 118 or memory 114) to one or more processor devices. Additionally or alternatively, processor device 112 may include an integrated controller. For example, communication interface 118 may be directly coupled to processor device 112 via an integrated I / O controller residing on the respective processor device.

[0039] Chipset 116 may include one or more communication interfaces 118. Communication interfaces 118 can be used for signaling and / or data communication between chipset 116 and one or more I / O devices, one or more networks 108, and / or one or more devices coupled to network 108 (e.g., data center management platform 106 or data analytics engine 104). For example, communication interface 118 can be used to send and receive network traffic, such as data packets. In a particular embodiment, communication interface 118 may be implemented by one or more I / O controllers, such as one or more physical network interface controllers (NICs), also known as network interface cards or network adapters. I / O controllers may include electronic circuitry to communicate using any suitable physical layer and data link layer standards, such as Ethernet (e.g., as defined by the IEEE 802.3 standard), Fibre Channel, InfiniBand, Wi-Fi, or other suitable standards. I / O controllers may include one or more physical ports that can be coupled to a cable (e.g., an Ethernet cable). I / O controllers can enable communication between any suitable component of chipset 116 (e.g., a switch) and another device coupled to network 108. In some embodiments, network 108 may include a switch with bridging and / or routing capabilities, located outside platform 102 and operable to couple various I / O controllers (e.g., NICs) distributed throughout data center 100 (e.g., on different platforms) to each other. In various embodiments, the I / O controllers may be integrated with a chipset (e.g., on the same integrated circuit or board as the rest of the chipset logic) or may be electromechanically coupled to a different integrated circuit or board of the chipset. In some embodiments, communication interface 118 may also allow I / O devices (e.g., disk drives, other NICs, etc.) integrated with or located outside the platform to communicate with the processor device core.

[0040] In some implementations, a switch may be used to couple to various ports of communication interface 118 (e.g., provided by a NIC), and data may be exchanged between these ports and various components of chipset 116 according to one or more link or interconnect protocols, such as Peripheral Component Interconnect Express (PCIe), Compute Fast Link (CXL), HyperTransport, GenZ, OpenCAPI, etc., which may apply the general principles and / or specific features discussed herein alternately or in combination. The switch and switching logic may be physical or virtual (e.g., software) switches.

[0041] Platform logic 110 may include an additional communication interface 118. Similar to communication interface 118, communication interface 118 can be used for signaling and / or data communication between platform logic 110 and one or more networks 108, as well as one or more devices coupled to network 108. For example, communication interface 118 can be used to send and receive network traffic, such as data packets. In a particular embodiment, communication interface 118 includes one or more physical I / O controllers (e.g., NICs). These NICs can enable communication between any suitable element of platform logic 110 (e.g., processor device 112) and another device coupled to network 108 (e.g., elements of other platforms or remote nodes coupled to network 108 via one or more networks). In a particular embodiment, communication interface 118 may allow devices external to the platform (e.g., disk drives, other NICs, etc.) to communicate with the processor core. In various embodiments, the NIC of communication interface 118 can be coupled to the processor device via an I / O controller (which may be external to or integrated with processor device 112). In addition, as discussed herein, the I / O controller may include a power manager 125 to implement power management functions at the I / O controller (e.g., automatically enabling power saving at one or more interfaces of the communication interface 118 (e.g., a PCIe interface that couples the NIC to another element of the system), and to implement other example functions.

[0042] Platform logic 110 can receive and execute any suitable type of processing request. A processing request may include any request that utilizes one or more resources of platform logic 110 (e.g., one or more cores or related logic). For example, a processing request may include: a processor core interrupt; a request to instantiate a software component (e.g., I / O device driver 124 or virtual machine 132); a request to process network packets received from a device external to virtual machine 132 or platform 102 (e.g., a network node coupled to network 108); a request to execute a workload (e.g., a process or thread) associated with an application running on virtual machine 132, platform 102, hypervisor 120, or other operating system executing on platform 102; or other suitable requests.

[0043] In various embodiments, processing requests may be associated with guest system 122. Guest system may include a single virtual machine (e.g., virtual machine 132a or 132b) or multiple virtual machines operating together (e.g., Virtual Network Function (VNF) 134 or Service Function Chain (SFC) 136). As shown, various embodiments may include various types of guest system 122 existing on the same platform 102.

[0044] Virtual machine 132 can use its own dedicated hardware to emulate a computer system. Virtual machine 132 can run a guest operating system on top of hypervisor 120. Components of platform logic 110 (e.g., processor device 112, memory 114, chipset 116, and communication interface 118) can be virtualized so that virtual machine 132 appears to the guest operating system to have its own dedicated components.

[0045] Virtual machine 132 may include a virtualized NIC (vNIC), which is used by the virtual machine as its network interface. Media access control (MAC) addresses can be assigned to the vNIC, thereby allowing multiple virtual machines 132 to be addressed individually within the network.

[0046] In some embodiments, virtual machine 132b may be paravirtualized. For example, virtual machine 132b may include an enhanced driver (e.g., a driver that provides higher performance or has a higher bandwidth interface to the underlying resources or functions provided by hypervisor 120). For example, an enhanced driver may have a faster interface to the underlying virtual switch 138 compared to the default driver, to achieve higher network performance.

[0047] VNF 134 may include software implementations of functional building blocks with defined interfaces and behaviors deployable within a virtualization infrastructure. In a particular embodiment, VNF 134 may include one or more virtual machines 132 that collectively provide specific functionalities (e.g., WAN optimization, VPN termination, firewall operation, load balancing operation, security functions, etc.). VNF 134 running on platform logic 110 can provide the same functionality as traditional network components implemented with dedicated hardware. For example, VNF 134 may include components for performing any suitable NFV workload, such as Virtualization Evolution Packet Core (vEPC) components, mobility management entities, 3GPP control and data plane components, etc.

[0048] SFC 136 is a set of VNF 134s that are organized into chains to perform a series of operations, such as network data packet processing. Service function chains can provide the ability to define an ordered list of network services (e.g., firewalls, load balancers) that are strung together in the network to create a service chain.

[0049] Hypervisor 120 (also known as a virtual machine monitor) may include logic for creating and running guest system 122. Hypervisor 120 may present a virtual operating platform to the guest operating system running the virtual machine (e.g., to the virtual machine it appears to be running on a separate physical node when the virtual machines are actually integrated onto a single hardware platform) and manage the execution of the guest operating system through platform logic 110. The services of hypervisor 120 may be provided through software virtualization, hardware-assisted resources requiring minimal software intervention, or both. Multiple instances of various guest operating systems may be managed by hypervisor 120. Platform 102 may have a separate instantiation of hypervisor 120.

[0050] Hypervisor 120 can be a local or bare-metal hypervisor that runs directly on platform logic 110 to control the platform logic and manage guest operating systems. Alternatively, hypervisor 120 can be a managed hypervisor that runs on the host operating system and abstracts the guest operating system from the host operating system. Various embodiments may include one or more non-virtualized platforms 102, in which case any suitable features or functionalities of hypervisor 120 described herein can be applied to the operating system of the non-virtualized platform. Further implementations, as described above, can be supported to enhance I / O virtualization. The host operating system can recognize the conditions and configuration of the system and determine which functions can be enabled or disabled (e.g., SIOV-based virtualization for SR-IOV-based devices), and can utilize appropriate application programming interfaces (APIs) to send and receive information related to such enabling or disabling, as well as implement other example functionalities.

[0051] Hypervisor 120 may include a virtual switch 138 that provides virtual switching and / or routing capabilities to virtual machines on guest system 122. Virtual switch 138 may include a logical switching structure that couples the vNICs of virtual machines 132 to each other, thereby creating a virtual network through which virtual machines can communicate with each other. Virtual switch 138 may also be coupled to one or more networks (e.g., network 108) via the physical NIC of communication interface 118 to allow communication between virtual machines 132 and one or more network nodes outside platform 102 (e.g., virtual machines running on a different platform 102 or nodes coupled to platform 102 via the Internet or other networks). Virtual switch 138 may include software elements implemented using components of platform logic 110. In various embodiments, hypervisor 120 may communicate with any suitable entity (e.g., an SDN controller) that enables hypervisor 120 to reconfigure parameters of virtual switch 138 in response to changing conditions in platform 102 (e.g., adding or removing virtual machines 132, or identifying optimizations that can be performed to improve platform performance).

[0052] The hypervisor 120 may include any suitable number of I / O device drivers 124. Each I / O device driver 124 represents one or more software components that allow the hypervisor 120 to communicate with physical I / O devices. In various embodiments, the underlying physical I / O devices may be coupled to any processor device 112 and may send and receive data to and from the processor device 112. The underlying I / O devices may utilize any suitable communication protocol, such as PCI, PCIe, Universal Serial Bus (USB), Serial Attached SCSI (SAS), Serial ATA (SATA), InfiniBand, Fibre Channel, IEEE 802.3, IEEE 802.11, or other current or future signaling protocols.

[0053] The underlying I / O device may include one or more ports operable for communicating with the core of processor device 112. In one example, the underlying I / O device is a physical NIC or physical switch. For example, in one embodiment, the underlying I / O device of I / O device driver 124 is a NIC with a communication interface 118 having multiple ports (e.g., Ethernet ports). In some implementations, I / O virtualization may be supported within the system and utilizes techniques described in more detail below. The I / O device may support I / O virtualization based on SR-IOV, SIOV, and other example methods and technologies.

[0054] In other embodiments, the underlying I / O device may include any suitable device capable of transmitting and receiving data to and from the processor device 112, such as an audio / video (A / V) device controller (e.g., a graphics accelerator or audio controller); a data storage device controller, such as a flash memory device, magnetic disk drive, or optical disk drive controller; a wireless transceiver; a network processor; or a controller for another input device (e.g., a monitor, printer, mouse, keyboard, or scanner); or other suitable devices.

[0055] In various embodiments, when a processing request is received, the I / O device driver 124 or the underlying I / O device may send an interrupt (e.g., a message semaphore interrupt) to any core of platform logic 110. For example, the I / O device driver 124 may send an interrupt to a core selected to perform an operation (e.g., a process representing virtual machine 132 or an application). Before the interrupt is delivered to the core, incoming data destined for the core (e.g., network data packets) may be cached on the underlying I / O device and / or the I / O block associated with the core's processor device 112. In some embodiments, the I / O device driver 124 may configure the underlying I / O device using instructions regarding where to send the interrupt.

[0056] In some embodiments, because the workload is distributed across cores, hypervisor 120 can route a greater number of workloads to higher-performing cores rather than lower-performing cores. In some cases, cores experiencing issues such as overheating or heavy loads may be assigned fewer tasks than other cores, or may be completely spared from being assigned tasks (at least temporarily). The workloads associated with applications, services, containers, and / or virtual machines 132 can be balanced across cores using network load and traffic patterns, not just processor device and memory utilization metrics.

[0057] The components of platform logic 110 can be coupled together in any suitable manner. For example, a bus can couple any components together. The bus can include any known interconnect, such as a multipoint bus, mesh interconnect, ring interconnect, point-to-point interconnect, serial interconnect, parallel bus, coherent (e.g., buffered coherent) bus, hierarchical protocol architecture, differential bus, or Gunning transceiver logic (GTL) bus.

[0058] The components of data center 100 can be coupled together in any suitable manner, such as through one or more networks 108. Network 108 can be any suitable network operating using one or more suitable network protocols, or a combination of one or more networks. A network can represent a series of nodes, points, and interconnected communication paths for receiving and sending packets of information propagated through a communication system. For example, a network may include one or more firewalls, routers, switches, security devices, antivirus servers, or other useful network devices. The network provides a communication interface between sources and / or hosts and may include any local area network (LAN), wireless local area network (WLAN), metropolitan area network (MAN), intranet, extranet, Internet, wide area network (WAN), virtual private network (VPN), cellular network, or any other suitable architecture or system that facilitates communication in a networked environment. A network may include any number of hardware or software components coupled (and communicating) to each other via a communication medium. In various embodiments, guest system 122 can communicate with nodes outside data center 100 via network 108.

[0059] like Figure 1 The data center 100 shown and discussed may include one or more software-configurable processor devices, as discussed herein. Go to Figure 2The illustration shows a portion of an example configurable processor device. The processor device may include a set of processor cores (e.g., 205) and storage elements for implementing cache memory (or processor memory), such as L1 cache, L2 cache, L3 cache, etc. A portion of the storage elements may be configured to implement conventional L1 cache and high-speed Just-In-Time (JIT) queues, such as First-In-First-Out (FIFO) queues or “FIFO” (e.g., 210a-b, 215a-b), to more efficiently pass data and / or instructions to one or more register files of processor core(s) 205 (e.g., when moving data to the core using a conventional L1 cache configuration, no additional processing and algorithms (e.g., replacement and sorting) are required). In other cases, cache storage elements may be programmably configured to implement queues other than FIFO queues (e.g., Last-In-First-Out (LIFO)), memory stacks, scratch pad memory structures, etc. For queues, such as FIFO queues, queues can be implemented as JIT queues by configuring the queue to be fed (e.g., written) in a continuous or predetermined manner, such that during operation, the queue is not over-fed or under-fed (within the range of available workload). For example, logic associated with the corresponding L2 cache or smart routing hardware (e.g., implemented as a smart NIC, infrastructure processing unit (IPU), data processing unit (DPU), edge processing unit (EPU), etc.) pushes data to the processor, which can cause data to be fed to the FIFO according to a time-based schedule aligned with the consumption or execution rate of the processor core retrieving information from the FIFO, among other example implementations.

[0060] In some implementations, various cores on a processor device can have corresponding cache / FIFO storage elements, and the hardware implementing these storage elements can be configured as all L1 cache, all FIFO, or a hybrid of L1 cache and FIFO. While both cache and FIFO are designed to efficiently pass data to the processor, cache includes the replacement strategy and sorting algorithm used by FIFO. Therefore, FIFO can be used as a simplified high-speed pipeline to feed data and instructions directly to the core's registers (e.g., 225a-b, 230a-b, etc.) for processing by the core's processing hardware. One or more FIFO register interfaces for the core can be implemented by providing data or instructions to the core's register file via JIT FIFO; other example implementations also exist.

[0061] exist Figure 2In the example, the memory structure may implement a set of caches (e.g., 220a-b) at the instruction register interface 225a-b, which provides instructions to core 205 for execution using the arithmetic logic unit (ALU), execution unit, and other processing elements (collectively referred to herein as "processing elements") of core 205. The data interface 230a-b may also be equipped with a set of L1 cache structures 235a-b for providing data to be operated on to core 205, which is associated with the instructions provided and executed by core 205. In this example, one or more instruction FIFOs (e.g., 210a-b, 215a-b) may also be provided to core 205 via the core's memory structure as an alternative to caches (e.g., 220a-b, 235a-b, etc.). FIFOs (e.g., by omitting the replacement and (potentially) more complex sorting algorithms used in traditional CPU caches) can allow instructions (from one FIFO (e.g., 210a)) and data (from another FIFO, e.g., 215a) to arrive at core 205 simultaneously (e.g., in the same clock cycle). Multiple FIFO-based data interfaces and / or multiple FIFO-based instruction interfaces can be provided at the respective cores in the processor device to allow the core's execution unit hardware to execute multiple instructions per clock cycle. Compared to traditional Harvard-based architectures (where the cache implements a single instruction interface and a single data interface), using a simplified FIFO structure can achieve higher bandwidth data and instruction throughput for the core, etc. Furthermore, by providing multiple JIT FIFO-based instruction and / or data interfaces for each core, a single core can be configured to operate like a Single Instruction Multiple Data (SIMD), Multiple Instruction Single Data (MISD), or Multiple Instruction Multiple Data (MIMD) machine, which can be used to implement various designs in systems comprising multiple such cores (e.g., system-on-a-chip, chiplets, or other processor devices).

[0062] In some implementations, FIFOs can be used instead of caches to accelerate and manipulate the way data and / or instructions are supplied to core 205. In multi-core systems, appropriate FIFOs can be used to customize the architecture of the multi-core system to implement dedicated processor or accelerator architectures for use within computing systems (e.g., data centers). High-speed FIFO structures can have a fixed length, and their filling speed may be faster than the speed at which the corresponding core can execute instructions in the queue or consume data in the queue. Therefore, in some implementations, FIFO overflows (e.g., 250) (e.g., in the core's L2 cache 245, L3 cache, network cache, or other caches provided on the processor device) can be provided to capture instructions and / or data destined for core 205 when the corresponding FIFO (e.g., 210a-b, 215a-b) reaches its capacity. Once the entries in the FIFO are open, the overflowing instructions and / or data can be fed back to the corresponding FIFO, and other example functionalities can be implemented.

[0063] Go to Figures 3A to 3C Simplified block diagrams 300a-c are shown, illustrating example portions of a processor device where multiple JIT FIFO structures (e.g., 320a-b, 325a-b, etc.) are provided. Using multiple JIT FIFOs allows the processor device to be configured more like a dedicated hardware accelerator than a general-purpose processor. For example, in Figure 3A In the example, multiple instruction FIFOs (e.g., 320a-b) and one or more data FIFOs (e.g., 325) are provided. The instruction FIFOs can be used to input corresponding instructions (e.g., 305a-b) to the ALUs (e.g., 310a-b) of the processor cores connected to the FIFOs, and the data FIFO 325 can be used to provide various data (e.g., 315a-b) to the ALUs 310a-b for execution in association with a corresponding instruction 305a-b. For example, instruction 305a can be input to ALU 310a to operate on data 315a, and instruction 305b can be input to ALU 310b to operate on data 315b. Figure 3A This diagram illustrates an example implementation of Multiple Instruction Multiple Data (MIMD). Go to... Figure 3B In some implementations, multiple data FIFOs (e.g., 325a-b) can be provided to the processor core. In this example, a single instruction 305 can be provided in parallel (via instruction FIFO 320) to multiple ALUs (e.g., 310a-d), and different data (e.g., 315a-d) can be provided via corresponding data FIFOs (e.g., 325a-b) for use during each instance of instruction 305 execution at each ALU 310a-d. Figure 3BThe diagram illustrates a Single Instruction Multiple Data (SIMD) implementation.

[0064] Go to Figure 3C The example illustrates an example where the outputs of various cores (e.g., 205a-d) and their corresponding execution units (e.g., 310a-h) (such as ALUs) can be directly fed into the JIT FIFO of another core of the processor device (e.g., 320a-h, 325a-h) (to be fed into execution units (in the same or different cores)). In this way, data and / or instructions can be cascaded in any intended pattern to implement logic equivalent to a dedicated hardware accelerator. In some implementations, instructions can specify the path from the output of an execution unit to the FIFO of the next execution unit. For example, instructions can pass data to the next execution unit. In some implementations, instructions can be timed to merge with data at the next execution unit. The configuration of the data and / or instruction stream from the execution unit output to the FIFO of the processor device allows the processor device to be configured with a variety of different architectures (e.g., logical architectures simulating specific neural networks, logical architectures simulating vector processing accelerators, etc.). This configuration can be executed before instructions are executed (based on the configuration definition provided to the processor device) and / or at least partially dynamically, where the routing / execution of downstream instructions or data depends on and is selected based on the results generated during the execution of other earlier upstream instructions (e.g., on the same or different cores), etc.

[0065] Refer again Figure 3C For example, in an illustrative example, the flow of data from one core (e.g., 205a) to another core (e.g., 205c-d) can be based on the result of an earlier instruction executed at a core (e.g., 205a). The instruction itself can indicate the path of the output, such as placing the instruction or data in a specific FIFO of another core (e.g., 320e-h, 325e-h), or looping the instruction or data back to the core's own FIFO. The data or instruction provided to the FIFO based on the execution of the previous instruction can be the same data or instruction, or it can be a different or modified version of the data or instruction, depending on the core's desired operation (e.g., in conjunction with configuring the core to behave as a specific processor type), and so on.

[0066] Go to Figure 4AA simplified block diagram 400a of an example processor device 405 including an array of processing cores is shown. The processing cores (e.g., 205a-e) in the processor device may include various execution units (e.g., ALUs) and have corresponding configurable cache / FIFO storage elements that can be configured to implement one or more JIT FIFOs for the core. In one example implementation, the storage elements of the core can be configured via software-defined definitions of the processor device 405 (e.g., using a software-based or on-board or package-based controller 415 that can provide configuration definitions for the processor device 405 (e.g., via interface 420)) to implement a set of JIT FIFOs or more traditional CPU caches for the core, as well as other configurations of the core and processor device hardware (e.g., on-chip networks of the processor device, etc.).

[0067] The configuration of the core's storage elements can be co-designed with a wider range of configurations to enable the core to behave as one of a variety of potential processor types, including a traditional CPU core, the core of another processor device (e.g., a GPU, TPU, etc.), or a hardware accelerator device. In this way, configuration definitions can be provided by software (e.g., 415) to processor device 405 (e.g., implemented as a system-on-a-chip (SOC), system-in-package (SIP), one or more application-specific integrated circuit (ASIC) devices, or other processor devices with multiple cores and other supporting hardware blocks) to configure the cores in the processor device to implement the corresponding processor device type. For example, configuration definitions can be processed by configuration controller hardware 430 on processor device 405 to define the corresponding FIFO / cache elements of its cores, as well as the on-chip network or interconnect structure coupling the individual cores of processor device 405, the multiplexer structure coupling the FIFO / cache elements to the execution units (or associated register files) of the cores, the interconnects or configurable streams between the execution units (e.g., ALUs) of individual cores, and other configurable components. The configurable components and processor device of each core can be configured as a whole (according to the provided configuration definition) such that a subset of the cores of the processor device is temporarily configured (e.g., combined with a specific client's workload or application) to implement a first type of processor or accelerator, while other cores in the processor device implement a different second type of processor or accelerator, and so on. With such a processor device (e.g., 405), servers and data centers can provide services and infrastructure that enable clients (e.g., 440) to define and configure custom combinations of accelerator and processor types specifically suited to the client's workload. This solution can achieve more deterministic solutions (e.g., virtually no cache misses, elimination of noisy neighbors on the cache, low-power servers, low-latency solutions, etc.), and other example advantages. In some implementations, an intelligent network controller (e.g., an intelligent NIC or infrastructure processing unit (IPU)) can be used to couple to network 445 and intelligently (e.g., via direct I / O access to cached data structures of the cores (e.g., 205a-e) direct requests and associated threads to specially configured cores on device 405 for execution, and so on.

[0068] exist Figure 4B A simplified block diagram 400b is shown, illustrating a software-configurable processor device 405 (e.g., Figure 4AThe example implementation is shown in the example device. In some implementations, an Infrastructure Processing Unit (IPU), a Smart NIC, or other external enhanced routing hardware (e.g., 460) may be provided to assist in programming and selectively bootstrapping workflows within the configurable processor device 405. When the individual cores of the processor device 405 (e.g., 205a-f) have been configured according to the configuration definitions provided by the software system 455 (e.g., by configuring the corresponding cache blocks of cores 205a-f (e.g., 450a-f)), the routing hardware (e.g., 460) can be used to precisely bootstrap software workloads to execute on the appropriately configured cores of the processor device 405. For example, the routing hardware 460 may be coupled to the processor device 405 using PCIe, CXL, or other interconnect technologies, and may be coupled to one or more software systems (e.g., 455) via one or more networks 445. Router hardware 460 may be equipped with logic and permissions to implement an interface (e.g., Data Direct I / O (DDIO) or a similar interface) that enables router hardware 460 to directly write to individual cache blocks (e.g., 450a-450f) of a specific corresponding configurable core (e.g., 205a-205f) to write specific instructions to a specific core associated with a software workflow. Furthermore, router hardware 460 may include hardware acceleration circuitry to perform packet checking or other processing on incoming data received from software system 455 (e.g., thread data from software system 455) to appropriately route the corresponding data and instructions to a specially configured core for execution on that core. In some implementations, such as Figure 4C As shown, all or part of the functionality provided by an external routing hardware device (e.g., 460) can be implemented on the processor device 405 itself, for example, via a Network Acceleration Complex (NAC) 465 and a cache controller 470. Data corresponding to a software workload can be sent from the software system 455 to the processor device 405 via network 445. This data can be examined (e.g., by the NAC 465 to identify instructions and data corresponding to a specific type of workflow or thread, and to identify cores (e.g., 205a-f) that have been configured to accelerate the execution of such workloads), and the cache controller 470 can write the data to the appropriate cache block (e.g., 450a-f) to provide instructions and data for execution by specific cores 205a-f in the processor device 405, as well as other example implementations.

[0069] In processor devices utilizing processor cores with cache interfaces (including JIT FIFOs), configuration definitions can be input into the processor device to configure various cores to implement or function as various different processor types and to configure an architecture that couples these different processor types within the system for executing workloads. For example, processor cores can be configured to implement functions such as: general-purpose processors (e.g., CPUs), tensor processor units (TPUs), graphics processors (e.g., GPUs), network processing units (e.g., NPUs), vector processing units (VCUs), compression engine units (e.g., CEUs), vision processing units (VPUs), cryptographic processing units, storage acceleration units (e.g., NVMe, NVMeoF, etc.) (SAUs), protocol accelerator units (PAUs) (e.g., for RDMA acceleration), quantum simulation accelerators (QEAs), matrix mathematics units (MMUs), and so on. Server-class processor devices including core arrays (e.g., Xeon-class processors) can implement arrays of different (or the same) processor types based on configuration definitions. In some cases, cores can be configured to operate as traditional CPUs (e.g., running Linux or Windows operating systems). In other cases, configuration definitions can allow some (or all) cores to be configured to be used as non-CPU processors (e.g., TPU, NPU, accelerators, etc.). Individual cores in the array can be configured and reconfigured over time to be used as various different processor types, including CPU cores and non-CPU processors, as well as accelerators, and so on.

[0070] Go to Figures 5A to 5B This illustrates an example configuration of an example processor device. Figure 5A In the example, all cores of processor device 405 can be configured to be used as CPU cores (e.g., by enabling or configuring the core's cache interface as a traditional L1 or CPU cache). A software controller (e.g., executing on a specific core on the processor device) can send a configuration definition 505 to the processor device to change the configuration of the various cores of the device, thereby modifying the configuration of these cores from CPU functionality to different processing functions. For example, some cores (e.g., 205a) can be configured as SAU, other cores (e.g., 205b) can be configured as NPU, and other cores (e.g., 205c) can be configured as TPU, and so on. As an example, configuration definition 505 can be provided in association with a given client or application that temporarily "owns" the processor device for a workload, and the configuration is designed to provide the client or application's workload with the optimal combination (and interconnection) of different processor components (e.g., CPU, NPU, TPU, VPU, etc.). Later at a later time or in a later session, such as... Figure 5BAs shown in the diagram, it can receive another configuration definition 510 to transfer the core functionality of the processor device 405 from configuration definition 505. Figure 5A The configuration of the driver (in the middle) is converted to a new, different configuration.

[0071] In some implementations, such software-delivered configuration definitions (e.g., 505, 510, etc.) can be controlled by a processing element in the system (e.g., to implement a configuration interface for the processor device). For example, a specific CPU core in a server or a CPU core on each chiplet of a processor device can be designated and configured to receive configuration definitions and implement the corresponding configuration on the core of the device. In some implementations, to address potential security concerns, cores configured to operate as general-purpose processing units (e.g., CPUs) can be protected and run security protocols. Cores used as accelerators or other dedicated processors may not be protected by conventional security solutions and can instead have their security managed and directed by another entity (e.g., another processing element, or external devices such as IPUs, DPUs, EPUs, etc.), and so on. Furthermore, in some implementations, cache coherence can be maintained among at least a subset of the cores of the processor device (e.g., for CPUs running conventional operating systems, cache coherence can be centrally maintained), but for other cores (e.g., implementing accelerators or other dedicated processors), conventional cache coherence can be deferred to support more efficient and simplified methods (e.g., using JITFIFO instead of a conventional cache structure or combining it with a conventional cache structure), and so on.

[0072] Applications and individual threads within applications can be directed to specific processor device cores configured to accelerate one or more functions associated with the thread or application. In some implementations, smart NICs, infrastructure processing units (IPUs), or other advanced networking or routing devices can be used to assist in directing individual applications, threads, or workloads to cores of a specific configuration within an example processor device. For example, via direct I / O protocols, advanced network devices can identify the configuration of a specific core within a processor device coupled to the network device and use direct I / O protocols to write instructions and / or data to the appropriate cache (e.g., L2 cache) of the core. These instructions and data can be pushed to a FIFO implemented in the core's L1 cache storage hardware, as well as other example implementations. Furthermore, servers, including software-configurable processor devices, can be configured to implement a specific type from a set of different processor types (as described herein) to suit the workloads, applications, or threads that will be executed using the server. For example, smart NICs, infrastructure processing units (IPUs), or other advanced networking or routing devices can be used to send configuration definitions to processor devices coupled to the advanced network device. As an example, applications involving workloads of video processing and matrix operations can be identified along with corresponding configuration definitions, and these configuration definitions can be sent to the processor device to configure the processor device's cores and network to include cores implementing tensor or vector processing units and video processing units (VPUs) to more efficiently process and accelerate functions that are expected to be called in association with the execution of the example application, as well as other examples implementing corresponding configuration definitions.

[0073] As described above, a smart NIC, IPU, or other controller can be used to ensure that threads and corresponding data (used for consumption within a thread) are routed to the appropriately configured core of the processor device. In one example, the IPU can parse incoming packets and determine the streams associated with those packets. This information helps the IPU understand the running application. The IPU can start its own threads to process the incoming data and identify the specific processing element configured to handle the incoming threads and / or data. In some implementations, the processing element (e.g., the core) can be configured by the IPU. Data to be consumed by one or more threads can similarly be routed to the processing element running the corresponding thread(s).

[0074] Go to Figure 6A simplified flowchart 600 of an example technique is shown for directing a specific software workload (e.g., identified as a specific thread by a process identifier, process address space identifier (PASID), namespace identifier, special tag data (e.g., attached to the workload), etc.) to a specific processing element that is advantageously configured for or based on the thread configuration. A thread can be identified (e.g., by routing hardware), and it can be determined 605 whether this is an instance of a familiar or previously processed thread, or whether the thread is a new thread. For a familiar thread, a lookup 610 can be performed (e.g., by software, routing hardware, or another controller) to identify one or more processing elements that have been configured to implement functionality useful or optimized for the familiar thread, and the thread can (again) pass 615 to such a configured processing element. For a new or unfamiliar thread, in some implementations, a configuration definition can be provided to the processor device to configure 620 at least a subset of the processor device's cores or other processing elements to implement processing hardware well adapted or optimized for the thread. After configuring the processing element to handle threads more efficiently (e.g., configuring the processing element to implement a hardware accelerator for the thread), the thread can be sent to the processing element. Furthermore, data to be processed associated with the execution of the thread can also be sent to the appropriately configured processing element and processed along with the thread. In one example implementation, a controller, IPU, or other device can assist in distributing and routing specific threads and data to various processing elements, and introduce configuration definitions to the processor device to initiate the configuration of one or more processing elements of the processor device. For example, a data packet or a set of data packets may arrive at the IPU coupled to the processor device. The IPU can determine whether a new thread should be started to process the incoming network data. As the first data packet, the IPU can configure the processing element connected to it to handle the thread efficiently. The data can then be sent to one or more processing elements, and the thread is processed. Subsequent data associated with that thread may need to be looked up to determine how to send the data to the processing element. The thread can then process that data.

[0075] In some implementations, to make processing more deterministic, time slots can be set to configure the processing of a given thread. In this implementation, hundreds or thousands of threads can utilize the same hardware, which is repeatedly reconfigured to best handle the current thread. When data arrives at the JIT FIFO of a processing element, the commands in the FIFO can indicate the data and / or instructions, processing element configuration, and other items to be loaded into the cache for rapid processing of the incoming data. A processing element can be equipped with multiple JIT FIFOs, where the processing element completes one FIFO before accessing the next. In this case, different JIT FIFOs can be used in conjunction with the processing of different threads, and so on.

[0076] Go to Figures 7A to 7B Simplified block diagrams 700a-b are shown, illustrating example software-configurable processor devices (such as...) Figures 4A to 5B Example configuration of the example core (also referred to herein as the "core complex") included in the example shown in the example). For example, Figure 7A An example default configuration of core complex 750a is shown, which is configured to operate as a traditional CPU core. For example, in the default configuration, cache block 450 of core 205 is configured to implement the L1 cache block of core 205. Multiplexer circuitry 730 can provide a single interface between the cache and the register file 715 of the CPU core logic, which utilizes execution units (e.g., 710) configured to implement a standard CPU core (e.g., including ALU 720). In some implementations, in response to a new thread or workload to be executed on the processor device, a configuration definition can be applied to the processor device to reconfigure one or more cores (e.g., 205) to adapt the cores to accelerate the execution of the newly incoming workload. For example, core complex 750a can be identified for reconfiguration, and the cache (e.g., 725a-h) can be flushed in response to the configuration definition (e.g., after the execution units have finished processing any remaining workload under the default CPU core configuration). For example, the memory management unit (MMU) of the processor device or core complex can be invoked to refresh those portions of the core's cache memory that will be configured to be used as a FIFO (e.g., 725e-h) (e.g., potential other portions of the cache memory (e.g., 725a-d) are reserved to continue being used as L1 cache), and so on. The processor device's configuration controller can reconfigure the cache (e.g., 725a-h), the multiplexer 730 logic, and other elements of the processor device (e.g., the configuration of the execution unit 710) according to the configuration definition.

[0077] Go to Figure 7B The diagram illustrates an example reconfiguration of core complex 750a. A configuration definition may take into account the configurable properties of core complex 750a (e.g., because some cores or processor devices may only allow reconfiguration of certain properties and elements of the core complex) and identify how to reconfigure the cores for at least one set of cores on a software-configurable processor device. In some cases, the configuration definition may identify multiple (e.g., interconnected) cores to be reconfigured, and the processor device's configuration controller may respond to the configuration definition to determine which specific cores can and / or should be reconfigured. In other implementations, the configuration definition may specify (e.g., for each configurable core on the processor device) the reconfiguration to be applied to each corresponding core, and so on.

[0078] exist Figure 7B In a specific example, by enabling cache FIFO logic (e.g., 740a-d) for cache elements (e.g., 725e-h) that will be reconfigured to operate as FIFOs (e.g., JIT FIFOs), the reconfiguration causes the cache block of core 205 to be reconfigured from the L1 cache to a set of FIFOs (e.g., after flushing the cache block), which includes FIFOs designated as data FIFOs and FIFOs designated as instruction FIFOs. To enable the respective FIFOs to be fed into the execution logic 710 of core 205 in a defined manner (e.g., to implement some form of hardware-accelerated execution), multiplexer circuitry can be reconfigured into multiplexer sub-blocks (e.g., 730a-e) to cause specific data and / or instruction FIFOs (e.g., 725e-h) to feed data (possibly also loopback data) to specific execution units (e.g., ALUs 720-720c), or even to other FIFOs, and so on. Furthermore, in some implementations, the interconnections between execution units within a core (e.g., 720a-c) can be configured according to configuration definitions to define various flows, loops, and outputs of core 205. Additionally, in some implementations, the interconnections and flows between cores (e.g., between core complex 750a and other core complexes of other cores in the processor device (e.g., 750b-c)) can be reconfigured and defined according to configuration definitions (e.g., by configuring the processor device's memory bus network or other on-chip network (NOC) 755 and / or providing output multiplexers (e.g., 745a-b) as reconfigurable blocks in the core complex) to allow multiple cores in the processor device to be combined in a defined manner, where the output of one core feeds into the input of one or more cores (or even feeds back to itself) to implement a specific accelerator or processor type, etc. When configuration is complete, an alert can be issued (e.g., in a register or via an interconnect message) to indicate to the software system or intelligent routing hardware that configuration is complete, and may also be used to identify the specific core that has been reconfigured. For example, after learning of the new configuration of a particular core, the IPU, cache controller, or other controller hardware can intelligently route and feed specific workloads (e.g., instructions and data to be consumed during instruction execution) to the corresponding core's FIFO (and L1 cache), which are now specifically configured to perform those specific workloads.

[0079] continue Figures 7A to 7BFor example, after a reconfigured processor device has had the opportunity to execute a specific workload (by reconfiguring the processor device for that specific workload), a new configuration definition can be received to reconfigure one or more cores again. In some implementations, context switching can be enabled, allowing different workloads to be executed on the same core, and some workloads can even interrupt other workloads (e.g., when a context switch occurs, allowing specific instructions and data of the interrupted workload to be cached (e.g., in L2 or L3 cache) to maintain the ability to quickly resume the interrupted workload by repopulating the corresponding data and instruction FIFO with the cached data of the interrupted workload, etc.). The new configuration definition can allow the configurable properties of each or a subset of all reconfigurable cores to be reconfigured. For example, in response to receiving a new configurable definition, it can be determined that the cache structure of a given core (e.g., 205) will be reconfigured (along with potential other elements, including the interconnect matrix of the multiplexers (e.g., 730a-e) feeding to execution units 710, the execution units themselves (e.g., 720a-c), and inter-core interconnect matrix elements (e.g., 745a-b, 755, etc.), etc.). The reconfigured cache block can be flushed as before (e.g., restoring the cache block to operate again to implement L1 cache or reconfiguring the cache block with a different FIFO arrangement), the MMU waits for the cache block to be emptied, and then flushes the cache block (and may also preload the reconfigured cache block with specific instructions and data). Once the various configurable elements of the core complex have been reconfigured according to the new configuration definition, an alarm or notification can be generated to indicate the completion of the new configuration and trigger the start of routing new workloads to specially configured cores in the processor device to accelerate these new workloads according to the new configuration, etc.

[0080] Go to Figure 8A simplified block diagram 800 illustrates an example technique for reconfiguring the core of a software-configurable processor device. For example, a configuration definition may be received describing configuration parameters for at least a subset of a configurable core complex (e.g., a core and associated cache structures) on the processor device. One of the described configurations may be identified 805 as being associated with a specific core. Based on the described configuration, it may be identified that the cache structures and / or the multiplexer interfaces of the execution units coupling the cache structures to the core will be reconfigured in a specified manner. Based on the cache structure reconfiguration, the configuration may include a refresh 810 of the cache structure to be reconfigured. Based on the configuration definition, the cache structure is reconfigured 815 to implement an L1 cache or a set of FIFOs (e.g., a defined combination of a data FIFO and an instruction FIFO as specified in the configuration definition). The reconfiguration may also include reconfiguring 820 the interconnects between the cache structure and the execution units (e.g., the core's register file) to ensure that a specific cache structure is fed to the corresponding execution circuitry of the core based on the configuration definition. Once the reconfiguration operation is complete and the cache structure has been properly refreshed and is ready to accept workloads and data for execution, an 825 indicator (e.g., via an interrupt, register write, sideband signal, interconnect message, etc.) can be generated to indicate the completion of the reconfiguration to one or more software systems, indicating that the new configuration is now in effect and that workloads can be routed to a specific core (e.g., by writing to its corresponding cache structure) for execution (at 830). This process can be restarted upon receiving a new configuration definition or a modification to the configuration definition that defines a new configuration for the core (at 805), and so on.

[0081] As described above, improved processor devices can include various processing elements (e.g., processor cores) with associated cache / FIFO structures. For example, one or more processing elements of the processor device can be provided with JIT FIFOs to load data into the processing elements for faster execution, unlike traditional Harvard architecture-based models. Multiple FIFOs can introduce multiple instruction and data streams that can be executed concurrently. Furthermore, by feeding more bandwidth into the processor, the functionality and performance of the processing elements can approach that of an accelerator more closely than that of a traditional CPU. That is, a single instruction can include multiple execution paths. For example, a single instruction can follow data through an execution unit, and the output of that execution unit can then reach multiple next-level execution units, or one or more JIT FIFOs from other processors (e.g., implementing SIMD, MIMD, and Single Instruction Multithreaded (SIMT) units). These instructions can vary depending on their destination.

[0082] At the processing element where instructions and / or data are fed by the corresponding JIT FIFO, the output of the processing element may be routed to various different elements on the processor device (e.g., using the processor device's internal on-chip network or interconnect structure). For example, the output of a processing element (e.g., a core) may be passed to one or more next processing elements, to the corresponding JIT FIFO of such processing element, to one or more recycle paths, etc. Different destinations may accept or be adapted to execute different instructions or instruction types. Therefore, in some implementations, predetermined instructions may be provided as output of a processing element or together with the output of a processing element, and different instructions may follow the output data to the next level. For example, instructions may be provided via the IPU, from a higher-level cache, or from the memory management unit (MMU) of device 405, etc. Using this approach, instructions may be executed with less bandwidth, lower latency, and lower power consumption. Furthermore, two or more instructions may be generated from a single (input) instruction and according to the data path (e.g., configured based on a configuration definition provided to the processor device), etc.

[0083] The cache memory structure associated with a single core can be configured to implement multiple FIFOs to provide data and / or instructions to the core's registers. The register interface to the FIFOs allows for virtually no latency when executing data entering the CPU from the FIFOs through that register interface. This FIFO-based interface can implement both data and / or instruction interfaces. In fact, a core's cache block can be configured to implement multiple different data FIFO interfaces and / or multiple instruction FIFO interfaces.

[0084] Go to Figure 9A simplified block diagram 900 is shown, illustrating an example implementation of processor core 205. In this example, the core may include multiple execution units that can be configured to interconnect as a staged pipeline (e.g., 910) of execution units (e.g., ALUs), such that the output of an execution unit (e.g., 915) corresponding to a first stage is directly fed to the next execution unit (e.g., 920) associated with the next stage in the pipeline (e.g., 910), and so on through the remaining execution units (e.g., 922, 930, etc.) configured to be included in pipeline 910. For example, a FIFO (e.g., 320) may feed a series of instructions and corresponding data to a given pipeline (e.g., 910). For example, in the pipeline of processing elements (e.g., 915, 920, 930, etc.), a single instruction in this series can be followed by corresponding data to set up the processing of each of the series of processing elements. For example, if there are four processing elements in the pipeline (e.g., ALUs 915, 920, 922, 930), the operation of each processing element is contained in an instruction that is passed along with data from the corresponding JIT data FIFO (e.g., 325) from the pipeline (e.g., 910). Similarly, JIT FIFO pairs (e.g., instruction FIFO 320 and corresponding data FIFO 325) can be configured and associated with each processing pipeline (e.g., 910) of core 205.

[0085] continue Figure 9For example, the output of a processing element in a pipeline (e.g., 910, 940, 945, 950, etc.) can have multiple alternative paths. These paths can have different purposes and therefore involve executing different instructions. Thus, the output of each processing unit can have registers that contain different instructions. For example, output path 910 might fetch the first instruction from a register fed by instruction FIFO 320a, output path 940 might fetch the second instruction from a register fed by instruction FIFO 320b, and so on. This would allow multiple instructions to be fed to a single core and executed in parallel (potentially on different datasets, e.g., fed by data FIFOs (e.g., 325a-b)). As an illustrative example, a set of 16 instructions (e.g., preloaded within memory elements (e.g., registers) of core 205) could be defined to be selected and used in association with any of the four processing element output paths (e.g., 910, 940, 945, 950). In this example, the 16 instructions can be indexed over time by the incoming instructions (e.g., according to the processor device's configuration definition). At any given moment, an incoming instruction may have only one output path, and the instruction's index is "1" (corresponding to selecting the first of the 16 instructions). The next incoming instruction may also have one output path, and the instruction's index is different, for example, "2". The third instruction may have two paths, one with the output instruction index "1" and the other with the output instruction index "3". The fourth instruction may have 16 paths, where each path has one of the 16 indices. The fifth instruction may have 16 paths, all using the instruction "1", and so on, such that any combination of any number of outputs can have any combination of instructions. Selecting one of the 16 instructions allows that selected instruction to then be fed into an instruction FIFO (e.g., at the same or a different core) to execute the instruction in response to the incoming instruction indexed therein, and so on.

[0086] Figure 9 The principle illustrated in the example can be extended to other architectures. For instance, instead of pipelined execution units (e.g., 915, 920, 922, 930) interconnected within a single core (e.g., 205), the processing elements can be represented as different cores, and the output of one core can be coupled to the input FIFO of another core (e.g., within the same or different processor devices) to implement a pipeline that is similar to or functionally equivalent to the pipeline of execution units, such as... Figure 9 The pipeline shown is analogous to this. Similarly, FIFO pairs can be implemented using an instruction FIFO and an associated data FIFO at the core of the pipeline. The output of one core can be used to determine the instructions and / or data to be fed into the FIFO of the next core in the pipeline, similar to... Figure 9 Examples of this approach, as well as implementations of other example functionalities, can be provided. Furthermore, more complex interconnections or pipelines can be provided by manipulating the interconnections of components (e.g., cores, execution units, etc.) to achieve more complex architectures (e.g., simplified examples beyond purely horizontal or vertical flow paths).

[0087] Go to Figure 10 A simplified block diagram 1000 illustrates cores 205a-c with corresponding instruction FIFOs (e.g., 320a-f) and data FIFOs (e.g., 325a-f). In some implementations, configuration definitions can define pipeline flows between execution units within a single core or between cores, allowing the output of executing a previous instruction (e.g., executed by core 205a) to drive the selection of the next instruction or data to be fed to the FIFO of a downstream core for execution. In this sense, the instructions or data fed to the FIFOs (e.g., 320c-f, 325c-f) can depend on the results generated when other instructions are executed upstream in the pipeline flow. As an example, in Figure 10 In this context, the output of core 205a can drive which instructions are fed to cores 205b and 205c for execution (e.g., according to configuration definitions intended to use cores 205a-c as a specific processor or accelerator type). For example, a given instruction can be selected and forwarded from core 205a to (e.g., written to) one or more instruction FIFOs (e.g., 320c-f) of other cores 205b-c, such that the selected (one or more) instructions are executed on those other cores (e.g., these cores can be configured to adapt the cores to execute such instructions in an accelerated manner). As an example, one or more instructions can be fed to core 205a (via instruction FIFOs 320a, 320b) for execution by execution units (e.g., 1020, 1025) using core 205a, consuming data provided to core 205a via their data FIFOs (e.g., 325a, 325b). Executing these instructions may produce results. Core 205a may include memory 1050 for storing a set of alternative instructions to be selected based on the result of instructions executed on core 205a. For example, the output of the execution may identify (e.g., via code) which alternative instructions to write to the instruction FIFO (e.g., 320c-f) of adjacent cores (e.g., 205b, 205c). Furthermore, the output may identify the target core or FIFO to write the selected instructions to. For example, in Figure 10In the example, the output of core 205a can allow selection to write instruction 1060a to FIFO 320f (of core 205c), to write instruction 1060b to FIFO 320d, to write instruction 1060c to FIFO 320e, and to write instruction 1060d to FIFO 320c, and so on (based on the corresponding configuration definitions used to configure the processor device and its component cores, cache storage elements, and interconnect structures). In this example, if the instruction execution results at core 205a are different, different instructions can be selected (from 1050) to write to the same or different target FIFOs. In other examples, the selected target could be the FIFO (e.g., 320a, 320b) of the core (e.g., 205a) that maintains (and pre-programs) the alternative instruction set, and so on. Additionally or alternatively, the output of an execution unit and / or core (e.g., 205a) can be used as a basis for determining which data path or location (e.g., core, execution unit, processing element, etc.) to use (e.g., proceeding to a first core configured to implement a first type of processor or a second core configured to implement a different second type of processor). In some paths, a new or next instruction may not be required; for example, the output of one or more ALUs (e.g., the ALU of core 205a or core 205b) may be the final stage in the pipeline and can be simply stored (e.g., storing data and the instruction is a null instruction, etc.), as well as other example implementations.

[0088] like Figure 10As illustrated in the examples, a configurable core (e.g., 205a) can be configured to interconnect with one or more other cores (e.g., 205b, 205c) such that the output of core 205a is fed to the input of one or more execution units of these other cores 205b, 205c. In some implementations, core 205a can be configured to conditionally route its output to the input of one or more other execution units. For example, registers, caches, or other memory elements 1050 can be provided on core 205a to load multiple alternative instructions. Core 205a can be configured to logically select one of the alternative instructions (e.g., 1060a-d) stored in memory element 1050 (e.g., based on the execution of the corresponding instructions fed to these execution units via FIFOs 320a-b) from one or more of its execution units to pass (e.g., via direct write) to one or more instruction FIFOs (e.g., 320c-f) of one or more downstream execution units (e.g., 1035). For example, this feature can be used in alternative configurations of cores 205a-c to implement hardware-accelerated neural networks on software-configurable processor devices. For instance, instructions fed into the execution unit from instruction FIFOs 320a-b of core 205a can cause output code that identifies a specific alternative instruction in memory 1050 (e.g., 0 = addition, 1 = multiplication, 2 = pass, etc.) and a target FIFO (e.g., 320c-f), where selected instructions (e.g., 1060a-d) should be pushed from core 205a to the target FIFO (e.g., 0 = target FIFO 320c, 1 = target FIFO 320d, 2 = target FIFO 320e, etc.) to implement a specific MAC flow for the neural network, as well as various other configurations that can be designed and imaged to accelerate various workloads provided to the configurable processor device for execution.

[0089] As mentioned above Figure 9 As described in the example, in a pipeline of processing elements (e.g., ALU, execution unit, or core), a single instruction passed to a processing element via a FIFO can be followed by corresponding data (from the corresponding data FIFO) to set up the processing for each processing element in a series of processing elements. Furthermore, a pipeline of processing elements executing single instructions following data can also establish interconnections between processing elements for each processing stage in a series of processing elements. This allows the output of one processing element to go to one or more processing elements in the next stage. Similarly, it can skip stages or loop back through the pipeline. Therefore, a set of processing elements arranged or grouped in the pipeline can form a matrix of processing elements based on the interconnections of stages. As an example, a matrix of four by four processing units can be configured (e.g., as shown in the example). Figure 9Examples of this (e.g., matrices of other dimensions, such as 7x3, 3x15, etc.) are also possible. It should be understood that a fairly large matrix of processing elements can be provided on a processor device and configured (e.g., 20x20, 100x100, etc.), allowing the total number of operations performed by each processing element in each cycle to grow accordingly. Furthermore, in the pipeline of processing elements, where a single instruction follows data to configure data processing or pipeline interconnects, a portion of an instruction may be dropped or decompressed. This allows an instruction of the same size to control more ALUs and interconnects. Additionally, new instructions can be passed to the JIT FIFOs of other cores, and so on.

[0090] Go to Figure 11 A simplified block diagram, shown in Simplified Block Diagram 1100, illustrates how JIT FIFOs can be used to implement instruction FIFOs (e.g., 320a-d) and data FIFOs (e.g., 325a-d). This enables multidimensional execution of instructions and data, such as providing a given instruction (e.g., via FIFO 320a) to multiple processing elements to operate on different corresponding data provided via multiple data FIFOs (e.g., 325a-d), and so on. These multiple instruction and data streams can be executed concurrently (e.g., within the same(one or more) cycles) to implement multidimensional processing units using a single general-purpose computing device. By using JIT FIFOs, this implementation enables more efficient AI and vector processing per core, achieving higher throughput than traditional Harvard architecture-based designs, as well as lower power consumption and latency per instruction, among other advantages.

[0091] continue Figure 11 For example, in one instance, data can be provided from a data FIFO (e.g., 325a-d) to processing elements for processing. Some of this data can be passed to the next level of processing elements in the system (e.g., execution units, other cores, etc.). Instructions can enter via corresponding instruction FIFOs (e.g., 320a-d) and can instruct or drive the operation of other processing elements within that column. For example, a single instruction can be executed by all ALUs in that column, or the result of a processing unit (e.g., 1105) can instruct or pass the instruction to be executed by the next processing element (e.g., 1110) in that column, causing different processing elements in the same column to execute different instructions. In some examples, one or more instructions can be bypass instructions, causing data to be passed from one execution stage to another without processing, and so on.

[0092] Processing elements can be interconnected via interconnect structures or on-chip networks, which can be at least partially configured to direct the output of a processing element to one of a potential plurality of different interconnected processing elements (e.g., a core on a single processor, an execution unit within a single core, etc.). In some implementations, the network can be configured to be fixed during execution, such that the output of a processing element is always input to a corresponding “partner” processing element. For example, the network configuration can be defined by a corresponding configuration definition, allowing the network to be programmed to implement a specific topology. For example, the processor device can initially be configured for a workload or dataset such that all data passing through a set, array, or pipeline of processing elements remains unchanged due to that data or workload. In another case (e.g., for subsequent workloads or datasets), the configuration definition can allow the data or instruction flow to be dynamic and change as the processor device is used. For example, in a dynamic case, the interconnects between processing elements can be based on instructions input to the processing elements and the results of executing those instructions. For example, the output of a processing element can be configured to alternatively flow to a plurality of alternative (or redundant) target processing elements coupled to that processing element (e.g., based on the result of executing one or more of its instructions). In some examples, the output of a processing element can be looped backward (to the same processing element or to another processing element involved in an earlier stage of the pipeline) or advanced forward to effectively skip one or more stages in the pipeline, and so on.

[0093] The configurability of processing element interconnects allows processing elements (e.g., cores) to not only be configured to implement the corresponding type of processor (e.g., TPU, GPU, CPU, hardware accelerator, etc.), but also to implement specific processing pipelines and data flows between these processors in a manner that can be used to implement or accelerate specific applications, threads, or workloads. As an example, appropriate configuration definitions can be used to implement neural networks or other machine learning or AI models to optimize the configuration of processing elements and on-chip networks for the model's structure and data / instruction flow. For example, in an example of a neural network application, the appropriate configuration definition could allow incoming data to be fed through a set of FIFOs, while weight data is fed through other data FIFOs, and corresponding instructions are fed through an instruction FIFO, potentially allowing each processing unit to receive some incoming data, weights, and instructions within a single cycle. Furthermore, the output of a processing unit can be fed (based on static or dynamic network configuration) to the next stage of the configured processing unit for processing. Impressive processing bandwidth can be achieved by transmitting data and instructions in parallel through multiple FIFO interfaces. As an example, a 128-core processor device running at 5 GHz with 8 instruction FIFOs and 8 data FIFOs per core could perform 40,960,000,000,000 operations per second (or approximately 41 trillion operations per server chip). In processor devices with a greater number of cores, more instruction FIFOs per core, or higher processing speeds, this architecture can enable peta-op-level performance per server chip, along with other examples of implementations and advantages for another part of the processing element.

[0094] As described above, the configuration definition of a software-configurable processor device can define how the output of a core (or the individual processing elements of a core) can populate the JIT FIFO of other cores coupled to that core (e.g., within a single processor device or chiplet). Directly populating instruction and / or data FIFOs can significantly reduce the power consumption and latency costs associated with data movement in traditional processor devices, because both data and instructions can be configured to arrive at a single core or ALU within the same clock cycle, allowing for efficient feeding throughout the execution of the workload. For example, data output from one core can populate one or more cache structures (e.g., configured as L1 caches or JIT FIFOs based on the configuration definition) of one or more other cores, enabling processing with low latency and low power consumption. Furthermore, once an instruction completes, it can be used to populate one or more instruction structures in one or more cache structures of another core. Instruction and routing configurations or dependencies can also be configured from one core output to another to implement a specific processor accelerator topology, and so on.

[0095] Go to Figure 12 A simplified block diagram 1200 of a core (e.g., 205a) is shown, having a set of interconnected execution units that implement one or more instruction and / or data pipelines corresponding to data FIFOs (e.g., 325a-d) and instruction FIFOs (e.g., 320a-d) fed to the pipelines. The output of executing one or more instructions can be directly passed from one core 205a to one or more other cores (e.g., 205b), where the on-chip architecture allows one core to write the output data of executing its instructions to the data FIFO of another core (e.g., 325e-h). In some implementations, multiple outputs can be generated from multiple concurrent pipelines on a core (e.g., 205a or 205b), and these multiple outputs can be fed (e.g., synchronously and in parallel) to multiple data FIFOs (e.g., 325e-h) provided on the same or multiple different cores (e.g., 205b). Similarly, the output of the instruction pipeline can be fed to one or more instruction FIFOs (e.g., 320e-h) of another core (e.g., 205b).

[0096] Depending on the range of configurable elements of a given core, a core (e.g., utilizing configurable execution units and configurable data flow paths between execution units) can be configured to implement a specific accelerator block using a single core. In other cases, desired hardware acceleration components can be implemented by configuring multiple interconnected cores using a single configuration definition. As an example, a hardware implementation of a neural network (or a portion of a neural network, such as one or more layers) can be implemented by configuring one or more cores of a software-configurable processor device, where the outputs of one or more execution units (e.g., ALUs) are configured to be coupled to the inputs of one or more other execution units, for example, to implement multiple accumulation (MAC) blocks, where a first execution unit is configured to perform multiplication and pass its output to a next execution unit, which is configured to receive the output and perform addition or accumulation. A core can be configured to scale the MAC computation accelerator by multiplying at a first-level core or execution unit (e.g., 2, 3, 4 cores, etc.) and accumulating (one or more) outputs at a second-level core or execution unit.

[0097] In addition, such as Figure 13As illustrated in the example, simplified block diagram 1300 shows an interface in which a core (e.g., 205a) can pass data and instruction outputs to a FIFO of another core (e.g., 205b) in other configurations. For example, as discussed earlier herein, a core and its FIFOs can be configured (through appropriate configuration definitions) to explicitly associate instruction FIFOs with corresponding data FIFOs, thus forming FIFO pairs (e.g., 1305). These FIFO pairs can extend to adjacent cores (e.g., 205b) beyond core 205a, to which core 205a can send its outputs. For example, data and instructions can be sent together from the output of core 205a to the FIFOs (or FIFO pairs (e.g., 1310)) of core 205b. These instructions can be the same instructions or pre-programmed instructions created or selected (e.g., from registers) based on the results of instructions executed by core 205a. In fact, in some implementations, data and instructions can be sent from core 205a to multiple cores based on their outputs, including recursively returning to their own input FIFOs, and so on. These and other data moves using JIT FIFO (as discussed in this article) can achieve significant reductions in latency and power consumption. For example, passing data from one core to another via a traditional cache might require moving the data to another cache, which could take tens or hundreds of processor cycles, including the execution and message passing required between the sending and receiving processors. However, by utilizing direct paths to adjacent cores / JIT FIFO, in some cases, data (and instructions) can move directly within a single clock cycle, making it faster and more efficient than traditional computing architectures (e.g., with lower power consumption), and enabling other example advantages.

[0098] Go to Figures 14A to 14BSimplified block diagrams 1400a-b illustrate examples of how one or more stages implemented by one or more processing elements (e.g., cores, execution units, ALUs, etc.) can be fed back or looped back to their own associated cache FIFO structure for reprocessing. Thus, instructions and / or data used in previous stages can be reloaded into the cache FIFO structure to be re-executed (e.g., in modified or unmodified form) for reprocessing. This looping data path or feedback loop can be used to implement looping in various algorithms and processing architectures (e.g., accelerators for implementing recurrent neural networks (RNNs), long short-term memory (LSTM) networks, gated recurrent units (GRUs), spike neural networks (SNNs), recurrent neural networks, and other machine learning models) while reducing overall data movement and resulting in reduced latency and power consumption, as well as other example advantages. For example, configuring a core to implement a loop from its output to its input FIFO can be used to implement accelerator processor types, such as accelerators for implementing recurrent neural networks (RNNs) or similar neural networks (whose loopback data may even loop back instructions), and other example use cases.

[0099] exist Figure 14A In this example, instructions are executed by corresponding processing elements (e.g., execution units with the same or different cores), and the same instructions are fed back into the instruction FIFO (e.g., 320a-b) for re-execution in subsequent loops. Similarly, in this example, data output by one (or both) of the processing elements (e.g., 1405, 1410) can also be fed back (e.g., transformed by executing instructions 1415, 1420) and provided again to the processing elements (e.g., 1405, 1410) for manipulation during the next execution of instructions 1415, 1420 (which can be executed in its original or modified form based on previous iterations of instruction execution). The result of executing instructions can also be used to terminate the loop execution of instructions. For example, based on the result of one or both of the executed instructions, alternative paths can be configured to direct data and / or instructions to another FIFO (e.g., another core), allowing the instruction FIFO (e.g., 320a-b) and / or data FIFO (e.g., 325) to be filled with new instructions, and so on.

[0100] As an illustrative example, an instruction executed at one core can be copied back to its own (or one or more) JIT FIFOs, or the instruction can point to another (or next) instruction to be executed at another core, which is then fed back to that core's own JIT FIFO for another round of processing. As an example, in a core configured to accelerate multiple-accumulate (MAC) operations, a single bit of the output can be used to identify whether the next instruction is an addition or a multiplication (e.g., 0 = addition; 1 = multiplication). In this case, the output can send 0 (or addition) to create an instruction for the input JIT FIFO, or send 1 (or multiplication) to create an instruction for the next ALU. This would make addition precede multiplication. For multiplication-accumulation, the reverse can be done, where the first stage performs multiplication and the second stage performs accumulation (addition), and so on. In some examples, the identifier of the next instruction (e.g., via bit 0 or 1) can indicate the path of the output, where the identifier points to the memory (e.g., in the L2 cache) with the next instruction to be executed, and other examples and implementations are also possible.

[0101] Go to Figure 14B For example, in some implementations, the configuration definition can define asymmetric loopback paths, because some data and / or instruction FIFOs (e.g., 320a-b, 325a, etc.) are fed back via loopback, while other data and / or instruction FIFOs (e.g., 325b) are fed back using new data or instruction 1450 (e.g., from registers, cache, output from another core, or another source). Figure 14B In the specific example shown, two levels of processing are configured (e.g., using processing elements 1405, 1425 and processing elements 1410, 1430), where one data FIFO (e.g., 325b) allows new incoming data, while the other data FIFO (e.g., 325a) is provided with feedback data. In this example, instructions fed to instruction FIFOs 320a-b are also fed back to implement processor configuration, where new data input to FIFO 325b is used in conjunction with the output of the previous processing stage.

[0102] In some implementations, the loopback instructions can be modified (e.g., based on the results of previous iterations) before returning to the JIT FIFO (e.g., 320a-b). For example, the instruction could be to loop back data and instructions for seven iterations, where after the seventh iteration, the instruction is modified to force the loopback to end (e.g., and each loopback instruction is modified to encode a counter value to indicate the remaining number of loopbacks (e.g., decrementing the counter by 1 each time an instruction completes)). Thus, with each loopback, the number in the instruction is decremented by 1.

[0103] The principles discussed above can be used to configure various processor designs to achieve hardware processing resources suitable for specific applications, threads, or models. For example, configuration definitions can mix and match loopbacks with output to other cores, where data and / or instructions are looped back (modified or unmodified, depending on the implementation), and implement other example functionalities. As an example, go to... Figure 15 A simplified block diagram 1500 illustrates an example model or a model of quantum computing simulation, representing a portion of a quantum computing problem. In one example, configuration definitions can be developed and applied to an improved processor device, utilizing configurable JIT FIFO data and instruction structures, as well as configurable on-chip architecture, to configure the cores of the processor device to implement a quantum computing simulation accelerator to simulate one or more quantum computing models. For example, by configuring on-chip interconnects and JIT FIFOs on a set of processor cores, a quantum computing network 1500 can be simulated by defining data and instructions passed in many different directions and paths, allowing the interconnected cores to simulate the operation of a quantum computer. For example, the configuration definition can apply recycling, processing element interconnects, and other techniques to simulate quantum computing models with multiple paths and cyclic paths. For example, cores and / or execution units can be configured to implement corresponding nodes in a quantum network (e.g., 1505, 1510, 1515, 1520, 1525, 1530, 1535, 1540, etc.), where interconnecting meshes (e.g., interconnecting cores or execution units within cores) are configured to implement paths between nodes in the network (e.g., 1545, 1550, 1555, 1560, etc.), as well as other example implementations. Therefore, core networks configured on processor devices can better simulate interactions between quantum computing elements, making the development of quantum computing algorithms and applications more efficient and accessible. Currently, quantum computing systems are still in their infancy, and most developers lack access to them. More efficient and accessible quantum computing simulation accelerators, enabled by software configuration on processor devices, can allow the simulation of more quantum computing elements and the verification of more quantum operation paths, enabling the development and verification of a wider range of quantum computing algorithms, while the industry awaits the development of more stable and commercially viable quantum computing systems, and the realization of other example advantages.

[0104] It should be understood that the examples provided in this article are only intended to illustrate more generally applicable principles, hardware implementations, and system applicability. Based on the number of cores and the configurability of the core cache storage elements and the interconnect structure of the interconnect cores (and / or individual execution units within a core), designers may have virtually unlimited potential to develop new and different configuration definitions to configure the corresponding processor devices to emulate or be used as a combination of various processor types, rather than a collection of general-purpose processor cores. Data centers, cloud service providers, and other servers can provide interfaces that allow customers to apply their configuration definitions to processor devices provided by the data center, which can enhance the services and configurability offered by the data center provider. Furthermore, application developers can leverage such configuration definitions to develop software optimized for execution on the configured processor devices, thereby achieving improved performance in their applications, enabling new applications and services that can be enabled through such processor resources, and implementing other example use cases.

[0105] Figures 16A to 23 Exemplary architectures and systems for implementing the embodiments described above (e.g., processors used in neuromorphic computing devices implementing the example SNNs described above) are detailed. In some embodiments, one or more hardware components and / or instructions described above are simulated as detailed below, or implemented as software modules. In fact, embodiments of the instructions (one or more) detailed above can be implemented using the “generic vector-friendly instruction format” detailed below. In other embodiments, this format is not used and another instruction format is used; however, the descriptions below of writing to mask registers, various data transformations (allocation, broadcasting, etc.), addressing, etc., generally apply to the descriptions of embodiments of the instructions (one or more) above. Furthermore, exemplary systems, architectures, and pipelines are detailed below. Embodiments of the instructions (one or more) above can be executed in such systems, architectures, and pipelines, but are not limited to those detailed above.

[0106] An instruction set may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, bit positions) to specify the operation to be performed (e.g., opcode) and the operand(s) to which the operation is performed, and / or other data fields(e.g., mask), etc. Some instruction formats are further decomposed through the definition of instruction templates (or subformats). For example, an instruction template for a given instruction format may be defined as having different subsets of the fields of that instruction format (the included fields are generally in the same order, but at least some may have different bit positions because fewer fields are included) and / or be defined to interpret the given fields differently. Thus, each instruction of the ISA is expressed using a given instruction format (and, if defined, by a given instruction template of that instruction format) and includes fields for specifying the operation and operand. For example, a demonstrative ADD instruction has a specific opcode and instruction format, which includes an opcode field to specify the opcode and an operand field to select the operand (source 1 / target and source 2); and the appearance of this ADD instruction in the instruction stream will have specific content in the operand field for selecting a specific operand.

[0107] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, such core implementations may include: 1) general-purpose ordered cores intended for general-purpose computing; 2) high-performance general-purpose out-of-order cores intended for general-purpose computing; and 3) dedicated cores primarily intended for graphics and / or scientific (throughput) computing. Different processor implementations may include: 1) CPUs comprising one or more general-purpose ordered cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) coprocessors comprising one or more dedicated cores primarily intended for graphics and / or scientific (throughput) computing. Such different processors lead to different computer system architectures, which may include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor on a separate die in the same package as the CPU; 3) a coprocessor on the same die as the CPU (in this case, such a coprocessor is sometimes referred to as dedicated logic, such as integrated graphics and / or scientific (throughput) logic, or a dedicated core); and 4) a system-on-a-chip that may include the described CPU (sometimes referred to as one or more application cores or one or more application processors), the coprocessors described above, and additional functionality on the same die. An exemplary core architecture is described next, followed by a description of exemplary processors and computer architectures.

[0108] Figure 16AIt is a block diagram illustrating both an exemplary ordered pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to an embodiment of the solution. Figure 16B This is a block diagram illustrating both an exemplary embodiment of an ordered architecture core to be included in a processor according to an embodiment of the solution and an exemplary register renaming, out-of-order issue / execution architecture core. Figures 16A to 16B The solid boxes in the diagram illustrate ordered pipelines and ordered cores, while the optional dashed boxes illustrate register renaming, out-of-order issue / execution pipelines, and cores. Given that the ordered aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0109] exist Figure 16A In the processor pipeline 1600, there are fetch phase 1602, length decoding phase 1604, decoding phase 1606, allocation phase 1608, renaming phase 1610, scheduling (also known as dispatching or issuing) phase 1612, register read / memory read phase 1614, execution phase 1616, write-back / memory write phase 1618, exception handling phase 1622, and commit phase 1624.

[0110] Figure 16B A processor core 1690 is shown, which includes a front-end unit 1630 coupled to an execution engine unit 1650, and both units are coupled to a memory unit 1670. The core 1690 can be a reduced instruction set computing (RISC) core (e.g., RISC-V), a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. Alternatively, the core 1690 can be a dedicated core, such as a network or communication core, a compression engine, a coprocessor core, a general-purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.

[0111] Front-end unit 1630 includes branch prediction unit 1632 coupled to instruction cache unit 1634, which is coupled to translation lookaside buffer (TLB) 1636. Instruction TLB 1636 is coupled to instruction fetch unit 1638, which is coupled to decoding unit 1640. Decoding unit 1640 (or decoder) can decode instructions and generate one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals as outputs. These micro-operations, microcode entry points, microinstructions, other instructions, or other control signals are decoded from, or otherwise reflect, or derived from, the original instructions. Various mechanisms can be used to implement decoding unit 1640. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memory (ROM), and so on. In one embodiment, core 1690 includes a microcode ROM or other medium that stores microcode for certain macro instructions (e.g., in decoding unit 1640 or otherwise within front-end unit 1630). Decoding unit 1640 is coupled to rename / allocator unit 1652 in execution engine unit 1650.

[0112] The execution engine unit 1650 includes a rename / allocator unit 1652 coupled to a retirement unit 1654 and a set of one or more scheduler units 1656. The scheduler units 1656 represent any number of different schedulers, including reservation stations, central instruction windows, etc. The scheduler units 1656 are coupled to one or more physical register file units 1658. Each of the physical register file units 1658 represents one or more physical register files, which store one or more different data types, such as scalar integers, scalar floating-point numbers, compressed integers, compressed floating-point numbers, vector integers, vector floating-point numbers, status (e.g., an instruction pointer as the address of the next instruction to be executed), etc. In one embodiment, the physical register file units 1658 include vector register units, write mask register units, and scalar register units. These register units can provide architectural vector registers, vector mask registers, and general-purpose registers. One or more physical register file units 1658 overlap with retirement units 1654 to illustrate various ways that can be used to implement register renaming and out-of-order execution (e.g., using one or more reorder buffers and one or more retirement register files; using one or more future heaps, one or more history buffers, and one or more retirement register files; using register maps and register pools; etc.). Retirement units 1654 and one or more physical register file units 1658 are coupled to one or more execution clusters 1660. Execution clusters 1660 include a set of one or more execution units 1662 and a set of one or more memory access units 1664. Execution units 1662 can perform various operations (e.g., shift, addition, subtraction, multiplication) on various types of data (e.g., scalar floating-point, compressed integer, compressed floating-point, vector integer, vector floating-point). While some embodiments may include several execution units dedicated to a particular function or set of functions, other embodiments may include only one execution unit or multiple execution units performing all functions.One or more scheduler units 1656, one or more physical register file units 1658, and one or more execution clusters 1660 are shown as potentially multiple, because some embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating-point / compact integer / compact floating-point / vector integer / vector floating-point pipelines, and / or memory access pipelines, each having its own scheduler unit, one or more physical register file units, and / or execution cluster—and in the case of separate memory access pipelines, in some embodiments only the execution cluster of that pipeline has one or more memory access units 1664). It should also be understood that, in the case of using separate pipelines, one or more of these pipelines may be issued / executed out of order, while the rest are ordered.

[0113] A set of memory access units 1664 is coupled to memory unit 1670, which includes a data TLB unit 1672. The data TLB unit 1672 is coupled to a data cache unit 1674, which is coupled to a Level 2 (L2) cache unit 1676. In one exemplary embodiment, memory access unit 1664 may include a load unit, a memory address unit, and a memory data unit, each of which is coupled to the data TLB unit 1672 in memory unit 1670. Instruction cache unit 1634 is further coupled to the Level 2 (L2) cache unit 1676 in memory unit 1670. The L2 cache unit 1676 is coupled to one or more other levels of cache and ultimately to main memory.

[0114] As an example, the exemplary register renaming, out-of-order issue / execution core architecture can implement pipeline 1600 as follows: 1) Instruction fetch 1638 executes fetch and length decoding stages 1602 and 1604; 2) Decoding unit 1640 executes decoding stage 1606; 3) Rename / allocator unit 1652 executes allocation stage 1608 and rename stage 1610; 4) (one or more) scheduler unit 1656 executes scheduling stage 1612; 5) (one or more) physical register file unit 1658 and memory unit 1670 execute register read / memory read stage 1614; (one or more) execution cluster 1660 executes execution stage 1616; 6) memory unit 1670 and (one or more) physical register file unit 1658 execute write-back / memory write stage 1618; 7) Various units may be involved in exception handling stage 1622; and 8) retirement unit 1654 and (one or more) physical register file unit 1658 execute commit stage 1624.

[0115] Core 1690 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with later versions); the MIPS instruction set of MIPS Technologies, Inc., Sunnyvale, California; the ARM instruction set of ARM Holdings, Inc., Sunnyvale, California (with optional additional extensions, such as NEON)), which include one or more of the instructions described herein. In one embodiment, Core 1690 includes logic supporting compressed data instruction set extensions (e.g., AVX1, AVX2), thereby allowing the use of compressed data to perform operations used by many multimedia applications.

[0116] It should be understood that the core can support multithreaded processing (two or more parallel sets of operations or threads), and can support multithreaded processing in various ways, including time-sliced ​​multithreaded processing, simultaneous multithreaded processing (where a single physical core provides a logical core for each thread performing simultaneous multithreaded processing on that physical core), or a combination of these (e.g., time-sliced ​​fetching and decoding, followed by simultaneous multithreaded processing, such as...). (as in Hyperthreading technology).

[0117] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in ordered architectures. While the illustrated embodiment of the processor also includes separate instruction and data cache units (e.g., 1634, 1674, etc.) and a shared L2 cache unit 1676, alternative embodiments may have a single internal cache for both instructions and data, such as a Level 1 (L1) internal cache or multiple levels of internal caches. In some embodiments, the system may include a combination of internal caches and external caches located outside the core and / or processor. Alternatively, all caches may be located outside the core and / or processor.

[0118] Figures 17A to 17B The diagram illustrates a more specific exemplary ordered core architecture, which will be one of several logic blocks (including other cores of the same and / or different types) in the chip. The logic blocks communicate with certain fixed-function logic, memory I / O interfaces, and other necessary I / O logic via a high-bandwidth interconnect network (e.g., a ring network), depending on the application.

[0119] Figure 17AThis is a block diagram of a single processor core according to an embodiment of the solution, its connection to an on-chip interconnect network 1702, and its local subset 1704 of the Level 2 (L2) cache. In one embodiment, the instruction decoder 1700 supports the x86 instruction set with a compact data instruction set extension. The L1 cache 1706 allows low-latency access to cache memory into scalar and vector units. While in one embodiment (for design simplification), scalar units 1708 and vector units 1710 use separate register sets (scalar register 1712 and vector register 1714, respectively), and data transferred between them is written to memory and then read back from the Level 1 (L1) cache 1706, alternative embodiments of the solution may use a different approach (e.g., using a single register set or including a communication path that allows data to be transferred between two register files without being written and read back).

[0120] The local subset 1704 of the L2 cache is part of the global L2 cache, which is divided into separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset 1704 of the L2 cache. Data read by a processor core is stored in its L2 cache subset 1704 and is quickly accessible, with these accesses occurring in parallel with other processor cores accessing their own local subsets of the L2 cache. Data written by a processor core is stored in its own L2 cache subset 1704 and can be flushed from other subsets as needed. The ring network ensures the consistency of shared data. The ring network is bidirectional to allow agents such as processor cores, L2 caches, and other logical blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide in each direction.

[0121] Figure 17B According to an embodiment of the solution Figure 17A An expanded view of a portion of the processor core. Figure 17B This includes the L1 data cache 1706A portion of L1 cache 1704, and further details regarding vector unit 1710 and vector register 1714. Specifically, vector unit 1710 is a 16-width vector processing unit (VPU) (see 16-width ALU 1728) that executes one or more of integer, single-precision floating-point, and double-precision floating-point instructions. The VPU supports register input allocation using allocation unit 1720, numerical conversion using value conversion units 1722A-B, and copying of memory input using copy unit 1724. Write mask register 1726 allows assertion result vector writing.

[0122] Figure 18This is a block diagram of a processor 1800 according to an embodiment of the solution. The processor 1800 may have more than one core, may have an integrated memory controller, and may have integrated graphics. Figure 18 The solid-line box in the diagram illustrates a processor 1800 having a single core 1802A, a system agent 1810, and a set of one or more bus controller units 1816, while the dashed-line box optionally illustrates an alternative processor 1800 having multiple cores 1802A-N, a set of one or more integrated memory controller units 1814 among the system agent units 1810, and dedicated logic 1808. In some implementations, the interconnect can be implemented as a bypass loop or other cache layer to enable direct core-to-core traffic, such as implementing the configurable processor device discussed above.

[0123] Therefore, different implementations of processor 1800 may include: 1) a CPU in which dedicated logic 1808 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores) and cores 1802A-N are one or more general-purpose cores (e.g., general-purpose ordered cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor in which cores 1802A-N are a large number of dedicated cores primarily intended for graphics and / or scientific (throughput); and 3) a coprocessor in which cores 1802A-N are a large number of general-purpose ordered cores. Thus, processor 1800 can be a general-purpose processor, coprocessor, or dedicated processor, such as a network or communication processor, compression engine, graphics processor, GPGPU (General-Purpose Graphics Processing Unit), high-throughput many-integrated-core (MIC) coprocessor (including 30 or more cores), embedded processor, etc. The processor may be implemented on one or more chips. Processor 1800 may be part of one or more substrates and / or may be implemented on one or more substrates using any of several process technologies, such as BiCMOS, CMOS, or NMOS.

[0124] The memory hierarchy includes one or more levels of cache within the core, a set or one or more shared cache units 1806, and external memory (not shown) coupled to the set of integrated memory controller units 1814. The set of shared cache units 1806 may include one or more intermediate level caches (e.g., Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache), the last level cache (LLC), and / or combinations thereof. While in one embodiment, ring-based interconnect units 1812 interconnect integrated graphics logic 1808, the set of shared cache units 1806, and system proxy units 1810 / (one or more) integrated memory controller units 1814, alternative embodiments may use any number of known techniques to interconnect such units. In one embodiment, consistency is maintained between one or more cache units 1806 and cores 1802A-N.

[0125] In some embodiments, one or more of the cores 1802A-N are capable of multithreaded processing. System agent 1810 includes those components that coordinate and operate the cores 1802A-N. System agent unit 1810 may include, for example, a power control unit (PCU) and a display unit. The PCU may be, or may include, the logic and components required to regulate the power state of the cores 1802A-N and the integrated graphics logic 1808. The display unit is used to drive one or more externally connected displays.

[0126] The 1802A-N cores can be homogeneous or heterogeneous in terms of their instruction sets; that is, two or more cores in the 1802A-N cores can execute the same instruction set, while other cores can execute only a subset of that instruction set or a different instruction set.

[0127] Figures 19 to 23 This is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptops, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In summary, a wide variety of systems or electronic devices capable of incorporating the processors and / or other execution logic disclosed herein are generally suitable.

[0128] Now for reference Figure 19A block diagram of a system 1900 according to one embodiment of the present disclosure is shown. System 1900 may include one or more processors 1910, 1915 coupled to a controller hub 1920. In one embodiment, controller hub 1920 includes a graphics memory controller hub (GMCH) 1990 and an input / output hub (IOH) 1950 (which may be on separate chips); GMCH 1990 includes a memory and a graphics controller coupled to a memory 1940 and a coprocessor 1945; IOH 1950 couples an input / output (I / O) device 1960 to GMCH 1990. Alternatively, one or both of the memory and the graphics controller are integrated within the processor (as described herein), memory 1940 and coprocessor 1945 are directly coupled to processor 1910, and controller hub 1920 and IOH 1950 are on a single chip.

[0129] The option of an additional processor 1915 is available. Figure 19 The numbers are indicated by dashed lines. Each processor 1910, 1915 may include one or more of the configurable processing cores described herein, and may be a version of the configurable processor device discussed herein.

[0130] The memory 1940 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of both. In at least one embodiment, the controller hub 1920 communicates with one or more processors 1910, 1915 via a multi-point branch bus (e.g., frontside bus, FSB), a point-to-point interface (e.g., QuickPath Interconnect, QPI, UltraPath Interconnect, UPI), or a similar connection 1995.

[0131] In one embodiment, the coprocessor 1945 is a dedicated processor, such as a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and so on. In one embodiment, the controller hub 1920 may include an integrated graphics accelerator.

[0132] In terms of a range of value metrics, including architectural characteristics, microarchitectural characteristics, thermal characteristics, power consumption characteristics, etc., there can be various differences between physical resources 1910 and 1915.

[0133] In one embodiment, processor 1910 executes instructions that control general types of data processing operations. Embedded within these instructions may be coprocessor instructions. Processor 1910 identifies these coprocessor instructions as types that should be executed by the attached coprocessor 1945. Therefore, processor 1910 issues these coprocessor instructions (or control signals representing coprocessor instructions) to coprocessor 1945 on a coprocessor bus or other interconnect. Coprocessor(s) 1945 receives and executes the received coprocessor instructions. In some implementations, processor 1910 and coprocessor(s) 1945 may be communicatively coupled and configured by sharing the same interconnect structure, cache structure, etc.

[0134] Now for reference Figure 20 A block diagram of a first, more specific, exemplary system 2000 according to embodiments of the present disclosure is shown. Figure 20 As shown, the multiprocessor system 2000 is a point-to-point interconnect system and includes a first processor 2070 and a second processor 2080 coupled via a point-to-point interconnect 2050. Each of processors 2070 and 2080 may be a version of processor 1800. In one embodiment of the solution, processors 2070 and 2080 are processors 1910 and 1915, respectively, and coprocessor 2038 is coprocessor 1945. In another embodiment, processors 2070 and 2080 are processor 1910 and coprocessor 1945, respectively.

[0135] Processors 2070 and 2080 are shown as including integrated memory controller (IMC) units 2072 and 2082, respectively. Processor 2070 also includes point-to-point (PP) interfaces 2076 and 2078 as part of its bus controller unit; similarly, the second processor 2080 includes PP interfaces 2086 and 2088. Processors 2070 and 2080 can exchange information via point-to-point (PP) interface 2050 using PP interface circuitry 2078 and 2088. Figure 20 As shown, IMC 2072 and 2082 couple the processor to the corresponding memory, namely memory 2032 and memory 2034, which may be portions of the main memory locally attached to the corresponding processor.

[0136] Processors 2070 and 2080 can each exchange information with chipset 2090 via separate PP interfaces 2052 and 2054 using point-to-point interface circuits 2076, 2094, 2086, and 2098, respectively. Chipset 2090 can optionally exchange information with coprocessor 2038 via high-performance interface 2039. In one embodiment, coprocessor 2038 is a dedicated processor, such as a high-throughput MIC processor, network or communication processor, compression engine, graphics processor, GPGPU, embedded processor, etc.

[0137] A shared cache (not shown) may be included in either processor or connected to both processors via a PP interconnect, such that if one processor is in a low-power mode, the local cache information of either or both processors may be stored in the shared cache.

[0138] Chipset 2090 can be coupled to first bus 2016 via interface 2096. In one embodiment, first bus 2016 may be a Peripheral Component Interconnect (PCI) bus, or a bus such as PCI Express bus or another third-generation I / O interconnect bus, but the scope of this disclosure is not limited thereto.

[0139] like Figure 20 As shown, various I / O devices 2014 can be coupled to a first bus 2016 along with a bus bridge 2018, which in turn couples the first bus 2016 to a second bus 2020. In one embodiment, one or more additional processors 2015 (e.g., coprocessors, high-throughput MIC processors, GPGPUs, accelerators (e.g., graphics accelerators or digital signal processing (DSP) units), field-programmable gate arrays, or any other processors) are coupled to the first bus 2016. In one embodiment, the second bus 2020 can be a low pin count (LPC) bus. In one embodiment, various devices can be coupled to the second bus 2020, including, for example, a keyboard and / or mouse 2022, communication devices 2027, and storage units 2028, such as disk drives or other high-capacity storage devices that may include instruction / code and data 2030. Furthermore, audio I / O 2024 can be coupled to the second bus 2020. Note that other architectures are also possible. For example, the system could implement a multi-drop bus or other such architecture, instead of Figure 20 A point-to-point architecture.

[0140] Now for reference Figure 21 A block diagram of a second, more specific exemplary system 2100 according to embodiments of the present disclosure is shown. For example, Figure 21 The illustration shows that processors 2170 and 2180 may include integrated memory and I / O control logic (“CL”) 2172 and 2182, respectively. Therefore, CL 2172 and 2182 include an integrated memory controller unit and I / O control logic. Figure 21 The diagram illustrates that not only are memories 2132 and 2134 coupled to CLs 2172 and 2182, but I / O device 2114 is also coupled to control logic 2172 and 2182. Conventional I / O device 2115 is coupled to chipset 2190.

[0141] Now for reference Figure 22 A block diagram of an SoC 2200 according to an embodiment of this disclosure is shown. Additionally, dashed boxes represent optional features on more advanced SoCs. Figure 22 In this embodiment, one or more interconnect units 2202 are coupled to: an application processor 2210, which includes a set of one or more cores 2202A-N and one or more shared cache units 2206; a system proxy unit 2212; one or more bus controller units 2216; one or more integrated memory controller units 2214; a set of one or more coprocessors 2220, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 2230; a direct memory access (DMA) unit 2232; and a display unit 2240 for coupling to one or more external displays. In one embodiment, one or more coprocessors 2220 include a dedicated processor, such as a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, an embedded processor, etc.

[0142] Embodiments of the mechanisms disclosed herein can be implemented using hardware, software, firmware, or a combination of such implementations. Embodiments of the solutions can be implemented as computer programs or program code that execute on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0143] Program code (e.g.) Figure 22The code 2230 shown in the diagram can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor, such as a digital signal processor (DSP), microcontroller, application-specific integrated circuit (ASIC), or microprocessor.

[0144] The program code can be implemented using a high-level procedural programming language or an object-oriented programming language to communicate with the processing system. Assembly or machine language can also be used if desired. In fact, the mechanisms described in this article are not limited to any particular programming language. In any case, the language can be compiled or interpreted.

[0145] Figure 23 This is a block diagram comparing an embodiment of the solution with the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set. In the illustrated embodiment, the instruction converter is a software instruction converter, although alternatively, the instruction converter can be implemented using software, firmware, hardware, or various combinations thereof. Figure 23 A program of high-level language 2302 is shown to be compiled by x86 compiler 2304 to generate x86 binary code 2306, which can be natively executed by processor 2316 having at least one x86 instruction set core. Processor 2316 having at least one x86 instruction set core means any processor that can perform substantially the same function as an Intel processor having at least one x86 instruction set core by compatiblely executing or otherwise processing (1) a substantial portion of the instruction set of an Intel x86 instruction set core or (2) a version of object code for an application or other software running on an Intel processor having at least one x86 instruction set core, in order to achieve substantially the same result as an Intel processor having at least one x86 instruction set core. x86 compiler 2304 means a compiler operable to generate x86 binary code 2306 (e.g., object code), which can be executed on processor 2316 having at least one x86 instruction set core with or without additional linking processing. Similarly, Figure 23A program in high-level language 2302 is shown to be compiled using an alternative instruction set compiler 2308 to generate alternative instruction set binary code 2310, which can be natively executed by a processor 2314 without at least one x86 instruction set core (e.g., a processor with a core executing the MIPS instruction set of MIPS Technologies, Inc., CAN, and / or the ARM instruction set of ARM Holdings, Inc., CAN,). An instruction converter 2312 is used to translate x86 binary code 2306 into code that can be natively executed by a processor 2314 without an x86 instruction set core. This translated code is unlikely to be identical to the alternative instruction set binary code 2310, as an instruction converter capable of doing so would be difficult to manufacture; however, the translated code will implement general operations and consist of instructions from the alternative instruction set. Thus, the instruction converter 2312 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device without an x86 instruction set processor or core to execute x86 binary code 2306 through emulation, simulation, or any other process.

[0146] One or more aspects of at least one embodiment may be implemented by representative instructions representing various logics within a processor, stored on a machine-readable medium, which, when read by a machine, cause the machine-made logic to perform the techniques described herein. This representation, referred to as an "IP core," may be stored on a tangible machine-readable medium and provided to various customer or manufacturing facilities for loading into the manufacturing machine that actually manufactures the logic or processor.

[0147] Such machine-readable storage media may include, but are not limited to, non-transitory tangible arrangements of articles made or formed by a machine or device, including storage media such as: hard disks, any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks), semiconductor devices (such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM)), magnetic cards or optical cards, or any other type of medium suitable for storing electronic instructions.

[0148] Therefore, embodiments of the solution also include a non-transitory tangible machine-readable medium containing instructions or design data, such as a Hardware Description Language (HDL), that defines the features of the structures, circuits, devices, processors, and / or systems described herein. Such embodiments may also be referred to as program products.

[0149] In some cases, instruction translators can be used to translate instructions from a source instruction set to a target instruction set. For example, an instruction translator can transform (e.g., using static binary translation, including dynamic binary translation with dynamic compilation), morph, emulate, or otherwise translate instructions into one or more other instructions to be processed by the kernel. Instruction translators can be implemented in software, hardware, firmware, or a combination thereof. Instruction translators can be on the processor, off-processor, or partially on-processor and partially off-processor.

[0150] While this disclosure is described in relation to certain implementations and generally associated methods, modifications and substitutions to these implementations and methods will be apparent to those skilled in the art. For example, the actions described herein may be performed in a different order than described, while still achieving the desired result. As an example, the processes depicted in the figures do not necessarily require a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous. Furthermore, other user interface layouts and functionalities may be supported. Other variations are within the scope of the appended claims.

[0151] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any solution or the scope of claims, but rather as descriptions of features specific to particular embodiments of a particular solution. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, while features may be described above as operating in certain combinations, or even as stated in the initial claims, one or more features from a claimed combination may in some cases be omitted from that combination, and a claimed combination may refer to a sub-combination or a variation of a sub-combination.

[0152] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all of the shown operations to be performed, in order to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0153] The following examples relate to embodiments according to this specification. Example 1 is an apparatus comprising: a plurality of processor cores; a plurality of storage elements associated with the plurality of processor cores, wherein the plurality of storage elements are configurable to: implement one or more Just-In-First-Out (JIT) queues or Level 1 (L1) cache blocks for a respective processor core among the plurality of processor cores; an interface for receiving configuration definitions from a software-based controller to define the configuration of the plurality of storage elements; and configuration hardware for configuring a first storage element among the plurality of storage elements associated with a first processor core among the plurality of processor cores based on the configuration definitions to implement a plurality of JIT FIFO queues for the first processor core in the first storage element.

[0154] Example 2 includes the subject of Example 1, wherein the plurality of JIT FIFO queues are used to provide corresponding information to the first processor core in the same clock cycle.

[0155] Example 3 includes the subject of any one of Examples 1-2, wherein, based on the configuration definition, the plurality of JIT FIFO queues include one or more instruction FIFO queues and one or more data FIFO queues, the one or more instruction FIFO queues being used to provide corresponding instructions for execution by the processing element of the first processor core, and the one or more data FIFO queues being used to provide data for instruction operations to be executed by the first processor core.

[0156] Example 4 includes the subject of Example 3, wherein the configuration hardware is configured to: configure a second storage element associated with a second processor core among the plurality of processor cores based on the configuration definition, so as to implement an L1 cache block for the second processor core in the second storage element.

[0157] Example 5 includes the subject of any one of Examples 3-4, wherein the one or more instruction FIFO queues include at least a first instruction FIFO queue and a second instruction FIFO queue, and the one or more data FIFO queues include at least a first data FIFO queue and a second data FIFO queue.

[0158] Example 6 includes the subject of Example 5, wherein the first instruction FIFO queue is associated with the first data FIFO queue to pass data from the first data FIFO queue for use during the execution of instructions provided by the first instruction FIFO queue, and the second instruction FIFO queue is associated with the second data FIFO queue to pass data from the second data FIFO queue for use during the execution of instructions provided by the second instruction FIFO queue, wherein instructions from the first instruction FIFO queue are executed in parallel at the first processor core with instructions from the second instruction FIFO queue.

[0159] Example 7 includes the subject of any one of Examples 1-6, and further includes: a configurable interconnect structure for interconnecting the plurality of processor cores, wherein the configuration definition defines the configuration of the interconnect structure.

[0160] Example 8 includes the subject matter of Example 7, wherein the configuration of the plurality of storage elements and the interconnect structure is for implementing at least one first type of processor and at least one different second type of processor device through the plurality of processor cores.

[0161] Example 9 includes the subject of Example 8, wherein the first type includes a general-purpose processor, and the second type includes one of the following: a graphics processing unit (GPU), a network processing unit, a tensor processing unit (TPU), a vector processing unit (VPU), a compression engine unit (CEU), a cryptographic processing unit, a storage acceleration unit (SAU), or a machine learning accelerator.

[0162] Example 10 includes the subject of any one of Examples 8-9, wherein a default configuration for the plurality of storage elements is used to implement a corresponding L1 cache for the plurality of processor cores, and the plurality of processor cores are used to implement a general-purpose processor core in the default configuration.

[0163] Example 11 includes the subject of any one of Examples 8-10, wherein the configuration definition includes a first configuration definition for a first operating window, wherein the configuration hardware is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operating window, and the configuration hardware is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operating window.

[0164] Example 12 includes the subject of Example 11, wherein a first user application is executed in the first operation window based on the first plurality of processors of different types, and a different second user application is executed in the second operation window based on the second plurality of processors of different types.

[0165] Example 13 includes the subject of any one of Examples 7-12, wherein the first processor core is coupled to a second processor core among the plurality of processor cores based on the configuration of the interconnect structure, so as to feed output from the first processor core to a set of JIT FIFO queues configured in a second storage element associated with the second processor core among the plurality of storage elements based on the configuration definition.

[0166] Example 14 includes the subject of Example 13, wherein the interconnect structure is configured based on the configuration definition for the first processor core to alternatively feed the output of the first processor core to the set of JIT FIFO queues of the second processor core or the FIFO queue of another processor core.

[0167] Example 15 includes the subject of Example 14, wherein the other processor core includes the first processor core to feed back the output of the first processor core to at least one of the plurality of JIT FIFO queues of the first processor core.

[0168] Example 16 includes the subject of any one of Examples 13-15, wherein the output includes instructions to be fed into the instruction FIFO in the set of JIT FIFO queues for execution by the second processor core.

[0169] Example 17 includes the subject of any one of Examples 13-16, wherein the output includes data of instruction operations to be executed by the second processor core.

[0170] Example 18 is a non-transitory machine-readable storage medium storing instructions executable by a machine to cause the machine to perform the following actions: generating a configuration definition for a processor device in a server system, wherein the processor device includes a plurality of processor cores having associated storage elements interconnected via an interconnect structure on the processor device, and the storage elements are configurable to: implement one or more Just-In-Time (JIT) First-In-First-Out (FIFO) queues for a corresponding processor core among the plurality of processor cores, wherein the configuration definition defines the configuration of the storage elements of the plurality of processor cores and the configuration of the interconnect structure for enabling the plurality of processor cores to implement a plurality of different processor types, wherein the configuration of the storage elements is configured to enable at least one storage element of a given processor core among the plurality of processor cores to implement a set of JIT FIFO queues, instead of a Level 1 (L1) cache, to deliver at least one of instructions or data to the given processor core; and sending the configuration definition to the server system to enable the processor device to implement the configuration of the storage elements of the plurality of processor cores and the configuration of the interconnect structure.

[0171] Example 19 includes the subject matter of Example 18, wherein the configuration of the storage elements of the plurality of processor cores and the configuration of the interconnect structure modify the default configuration of the processor device, wherein the plurality of processor cores implement general-purpose processor cores in the default configuration, and the plurality of different processor types include at least one dedicated hardware accelerator processor type.

[0172] Example 20 includes the subject of any one of Examples 18-19, wherein the interconnect structure is configured to enable a loopback from the output of a particular processor core among the plurality of processor cores to a JIT FIFO queue of that particular processor core.

[0173] Example 21 includes the subject of any one of Examples 18-19, wherein the interconnect structure is configured to route the output of a first processor core among the plurality of processor cores to a JIT FIFO queue of a second processor core among the plurality of processor cores.

[0174] Example 22 includes the subject of Example 21, wherein the interconnect structure is configured to route the output of the first processor core to a JIT FIFO queue of two or more of the plurality of processor cores.

[0175] Example 23 includes the subject of Example 22, wherein the output of the first processor core is routed to one of the two or more processor cores based on the result at the first processor core.

[0176] Example 24 is a method comprising: generating a configuration definition for a processor device in a server system, wherein the processor device includes a plurality of processor cores having associated storage elements, the plurality of processor cores being interconnected via an interconnect structure on the processor device, and the storage elements being configurable to: implement one or more Just-In-Time (JIT) First-In-First-Out (FIFO) queues for a corresponding processor core among the plurality of processor cores, wherein the configuration definition defines the configuration of the storage elements of the plurality of processor cores and the configuration of the interconnect structure for enabling the plurality of processor cores to implement a plurality of different processor types, wherein the configuration of the storage elements is configured to enable at least one storage element of a given processor core among the plurality of processor cores to implement a set of JIT FIFO queues, instead of a Level 1 (L1) cache, to deliver at least one of instructions or data to the given processor core; and sending the configuration definition to the server system to enable the processor device to implement the configuration of the storage elements of the plurality of processor cores and the configuration of the interconnect structure.

[0177] Example 25 includes the subject matter of Example 24, wherein the configuration of the storage elements of the plurality of processor cores and the configuration of the interconnect structure modify the default configuration of the processor device, wherein the plurality of processor cores implement general-purpose processor cores in the default configuration, and the plurality of different processor types include at least one dedicated hardware accelerator processor type.

[0178] Example 26 includes the subject of any one of Examples 24-25, wherein the interconnect structure is configured to enable a loopback from the output of a particular processor core among the plurality of processor cores to a JIT FIFO queue of that particular processor core.

[0179] Example 27 includes the subject of any one of Examples 24-25, wherein the interconnect structure is configured to route the output of a first processor core among the plurality of processor cores to a JIT FIFO queue of a second processor core among the plurality of processor cores.

[0180] Example 28 includes the subject of Example 27, wherein the interconnect structure is configured to route the output of the first processor core to a JIT FIFO queue of two or more of the plurality of processor cores.

[0181] Example 29 includes the subject of Example 28, wherein the output of the first processor core is routed to one of the two or more processor cores based on the result at the first processor core.

[0182] Example 30 is a system including means for performing the methods of any one of Examples 24-29.

[0183] Example 31 is a system comprising: a server system including: a processor device including: a plurality of processor cores; a plurality of storage elements associated with the plurality of processor cores, wherein the plurality of storage elements are configurable to: implement one or more Just-In-First-Out (JIT) queues or Level 1 (L1) cache blocks for a respective processor core among the plurality of processor cores; an interface for receiving configuration definitions from a software-based controller to define the configuration of the plurality of storage elements of the plurality of processor cores; and configuration hardware for configuring a first storage element among the plurality of storage elements associated with a first processor core among the plurality of processor cores based on the configuration definitions to implement a plurality of JIT FIFO queues for the first processor core in the first storage element.

[0184] Example 32 includes the subject of Example 31, and further includes: a software controller for generating the configuration definition and sending the configuration definition to the interface.

[0185] Example 33 includes the subject of any one of Examples 31-32, wherein the configuration hardware resides on the processor device.

[0186] Example 34 includes the subject of any one of Examples 31-33, wherein the plurality of JIT FIFO queues are used to provide corresponding information to the first processor core in the same clock cycle.

[0187] Example 35 includes the subject of any one of Examples 31-34, wherein, based on the configuration definition, the plurality of JITFIFO queues include one or more instruction FIFO queues and one or more data FIFO queues, the one or more instruction FIFO queues being used to provide corresponding instructions for execution by the processing element of the first processor core, and the one or more data FIFO queues being used to provide data for instruction operations to be executed by the first processor core.

[0188] Example 36 includes the subject of Example 35, wherein the configuration hardware is configured to: configure a second storage element associated with a second processor core among the plurality of processor cores based on the configuration definition, to implement an L1 cache block for the second processor core in the second storage element.

[0189] Example 37 includes the subject matter of any one of Examples 35-36, wherein the one or more instruction FIFO queues include at least a first instruction FIFO queue and a second instruction FIFO queue, and the one or more data FIFO queues include at least a first data FIFO queue and a second data FIFO queue.

[0190] Example 38 includes the subject of Example 37, wherein the first instruction FIFO queue is associated with the first data FIFO queue to pass data from the first data FIFO queue for use during the execution of instructions provided by the first instruction FIFO queue, and the second instruction FIFO queue is associated with the second data FIFO queue to pass data from the second data FIFO queue for use during the execution of instructions provided by the second instruction FIFO queue, wherein instructions from the first instruction FIFO queue are executed in parallel with instructions from the second instruction FIFO queue at the first processor core.

[0191] Example 39 includes the subject matter of any one of Examples 31-38, and further includes: a configurable interconnect structure for interconnecting the plurality of processor cores, wherein the configuration definition defines the configuration of the interconnect structure.

[0192] Example 40 includes the subject of Example 39, wherein the plurality of storage elements and the interconnect structure are configured to implement at least one first type of processor and at least one different second type of processor device through the plurality of processor cores.

[0193] Example 41 includes the subject of Example 40, wherein the first type includes a general-purpose processor, and the second type includes one of the following: a graphics processing unit (GPU), a network processing unit, a tensor processing unit (TPU), a vector processing unit (VPU), a compression engine unit (CEU), a cryptographic processing unit, a storage acceleration unit (SAU), or a machine learning accelerator.

[0194] Example 42 includes the subject of any one of Examples 40-41, wherein a default configuration for the plurality of storage elements is used to implement a corresponding L1 cache for the plurality of processor cores, and the plurality of processor cores are used to implement a general-purpose processor core in the default configuration.

[0195] Example 43 includes the subject of any one of Examples 40-42, wherein the configuration definition includes a first configuration definition for a first operating window, wherein the configuration hardware is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operating window, and the configuration hardware is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operating window.

[0196] Example 44 includes the subject of Example 43, wherein a first user application is executed in the first operation window based on the first plurality of processors of different types, while a different second user application is executed in the second operation window based on the second plurality of processors of different types.

[0197] Example 45 includes the subject of any one of Examples 39-44, wherein the first processor core is coupled to a second processor core among the plurality of processor cores based on the configuration of the interconnect structure, so as to feed output from the first processor core to a set of JIT FIFO queues configured in a second storage element associated with the second processor core among the plurality of storage elements based on the configuration definition.

[0198] Example 46 includes the subject of Example 45, wherein the interconnect structure is configured based on the configuration definition for the first processor core to alternatively feed the output of the first processor core to the set of JIT FIFO queues of the second processor core or the FIFO queue of another processor core.

[0199] Example 47 includes the subject of Example 46, wherein the other processor core includes the first processor core to feed back the output of the first processor core to at least one of the plurality of JIT FIFO queues of the first processor core.

[0200] Example 48 includes the subject of any one of Examples 45-47, wherein the output includes instructions to be fed into the instruction FIFO in the set of JIT FIFO queues for execution by the second processor core.

[0201] Example 49 includes the subject of any one of Examples 45-48, wherein the output includes data to be operated by instructions to be executed by the second processor core.

[0202] Example 50 is an apparatus comprising: a plurality of processor cores; a plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configurable to: alternatively implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and a configuration controller for: identifying a software-defined configuration definition, wherein the configuration definition defines a configuration for a processor core among the plurality of processor cores; and configuring a first storage element among the plurality of storage elements associated with a first processor core among the plurality of processor cores based on the configuration definition to implement a set of FIFO queues for the first processor core in the first storage element.

[0203] Example 51 includes the subject of Example 50, wherein the set of FIFO queues comprises multiple FIFO queues.

[0204] Example 52 includes the subject of Example 51, wherein the plurality of FIFO queues include a plurality of data FIFO queues and a plurality of instruction FIFO queues, a first instruction FIFO queue of the plurality of instruction FIFO queues is used to provide a first instruction for execution by the execution unit of the first processor core, and a first data FIFO queue of the plurality of data FIFO queues is used to provide data for the first instruction operation to be executed by the first processor core.

[0205] Example 53 includes the subject of Example 52, wherein a second instruction FIFO queue in the plurality of instruction FIFO queues is used to provide a second instruction for execution by the execution unit of the first processor core, and a second data FIFO queue in the plurality of data FIFO queues is used to provide data for the second instruction operation to be executed by the first processor core, wherein the first instruction from the first instruction FIFO queue will be executed by the first processor core at least partially in parallel with the second instruction.

[0206] Example 54 includes the subject matter of any one of Examples 51-53, and further includes: a configurable multiplexer circuit for directing information from the first storage element to a corresponding execution unit among a plurality of execution units of the first processor core, wherein the configuration controller is configured based on the configuration definition to configure the configurable multiplexer circuit to direct information from the plurality of FIFO queues to the corresponding execution units among the plurality of execution units.

[0207] Example 55 includes the subject of any one of Examples 51-54, wherein the plurality of FIFO queues are used to provide corresponding information to the first processor core in the same clock cycle.

[0208] Example 56 includes the subject of any one of Examples 51-55, wherein the configuration controller is configured to: configure a second storage element associated with a second processor core among the plurality of processor cores based on the configuration definition, so as to implement an L1 cache for the second processor core.

[0209] Example 57 includes the subject matter of any one of Examples 50-56, and further includes: a configurable interconnect structure for interconnecting the plurality of processor cores, wherein the configuration controller is configured to configure the configurable interconnect structure based on the configuration definition to define the data flow between the plurality of processor cores.

[0210] Example 58 includes the subject of Example 57, wherein the configurable interconnect structure is configured to feed the output of the first processor core to a FIFO queue implemented in one of the plurality of memory elements, wherein the output includes either executable instructions or data to be operated by executable instructions.

[0211] Example 59 includes the subject of Example 57, wherein the FIFO queue is implemented in a second storage element associated with a second processor core among the plurality of processor cores.

[0212] Example 60 includes the subject of Example 57, wherein the FIFO queue is implemented in the first storage element to define feedback information for use by the first processor core based on the configuration.

[0213] Example 61 includes the subject of any one of Examples 50-60, wherein the configuration for processor cores described in the configuration definition is for implementing at least one first type of processor and at least one different second type of processor device through the plurality of processor cores.

[0214] Example 62 includes the subject of Example 61, wherein the first type includes a general-purpose processor, and the second type includes one of the following: a graphics processing unit (GPU), a network processing unit, a tensor processing unit (TPU), a vector processing unit (VPU), a compression engine unit (CEU), a cryptographic processing unit, a storage acceleration unit (SAU), or a machine learning accelerator.

[0215] Example 63 includes the subject of Example 61, wherein the configuration definition includes a first configuration definition for a first operation window, wherein the configuration controller is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operation window, and the configuration controller is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operation window.

[0216] Example 64 includes the subject of Example 63, wherein a first user application will execute in the first operation window on the plurality of processor cores configured based on the first configuration definition, while a different second user application will execute in the second operation window on the plurality of processor cores configured based on the second configuration definition.

[0217] Example 65 includes the subject of Example 64, wherein, for the first processor core, a specific workload of the first user application is routed to the set of FIFO queues based on the first configuration definition.

[0218] Example 66 is a method comprising: receiving configuration definition data, wherein the configuration definition data describes a specific configuration for a configurable processor device, the configurable processor device comprising: a plurality of processor cores; a configurable interconnect structure for interconnecting the plurality of processor cores; and a plurality of configurable storage elements associated with the plurality of processor cores, wherein the storage elements of the plurality of storage elements are respectively configurable to: alternatively implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and configuring at least the configurable interconnect structure and the plurality of configurable storage elements based on the specific configuration to define a data flow for the plurality of processor cores.

[0219] Example 67 includes the subject of Example 66, wherein the particular configuration uses the plurality of processor cores to implement a collection of multiple different types of processors.

[0220] Example 68 includes the subject of Example 67, wherein the first type includes a general-purpose processor and the second type includes one of the following: a graphics processing unit (GPU), a network processing unit, a tensor processing unit (TPU), a vector processing unit (VPU), a compression engine unit (CEU), a cryptographic processing unit, a storage acceleration unit (SAU), or a machine learning accelerator.

[0221] Example 69 includes the subject matter of any one of Examples 67-68, wherein the configuration definition includes a first configuration definition for a first operating window, wherein the configuration hardware is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operating window, and the configuration hardware is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operating window.

[0222] Example 70 includes the subject of Example 69, wherein a first user application will execute in the first operation window on the plurality of processor cores configured based on the first configuration definition, while a different second user application will execute in the second operation window on the plurality of processor cores configured based on the second configuration definition.

[0223] Example 71 includes the subject of Example 70, wherein, for a first processor core, a specific workload of the first user application is routed to a set of FIFO queues based on the first configuration definition.

[0224] Example 72 includes the subject of any one of Examples 66-71, wherein the set of FIFO queues comprises multiple FIFO queues.

[0225] Example 73 includes the subject matter of Example 72, wherein the plurality of FIFO queues include a plurality of data FIFO queues and a plurality of instruction FIFO queues, a first instruction FIFO queue in the plurality of instruction FIFO queues is used to provide a first instruction for execution by an execution unit of a first processor core, and a first data FIFO queue in the plurality of data FIFO queues is used to provide data for the first instruction operation to be executed by the first processor core.

[0226] Example 74 includes the subject matter of Example 73, wherein a second instruction FIFO queue in the plurality of instruction FIFO queues is used to provide a second instruction for execution by the execution unit of the first processor core, and a second data FIFO queue in the plurality of data FIFO queues is used to provide data for the second instruction operation to be executed by the first processor core, wherein the first instruction from the first instruction FIFO queue will be executed by the first processor core at least partially in parallel with the second instruction.

[0227] Example 75 includes the subject matter of any one of Examples 72-74, and further includes: a configurable multiplexer circuit for directing information from a first storage element to a corresponding execution unit among a plurality of execution units of the first processor core, wherein a configuration controller is configured, based on the configuration definition, to configure the configurable multiplexer circuit to direct information from the plurality of FIFO queues to the corresponding execution units among the plurality of execution units.

[0228] Example 76 includes the subject of any one of Examples 72-75, wherein the plurality of FIFO queues are used to provide corresponding information to the first processor core in the same clock cycle.

[0229] Example 77 includes the subject of any one of Examples 72-76, wherein the configuration controller is configured to: configure a second storage element associated with a second processor core among the plurality of storage elements based on the configuration definition, so as to implement an L1 cache for the second processor core.

[0230] Example 78 includes the subject matter of any one of Examples 66-77, and further includes: a configurable interconnect structure for interconnecting the plurality of processor cores, wherein the configuration controller is configured to configure the configurable interconnect structure based on the configuration definition to define the data flow between the plurality of processor cores.

[0231] Example 79 includes the subject matter of Example 78, wherein the configurable interconnect structure is configured to feed the output of the first processor core to a FIFO queue implemented in one of the plurality of memory elements, wherein the output includes either executable instructions or data to be operated by executable instructions.

[0232] Example 80 includes the subject of any one of Examples 78-79, wherein the FIFO queue is implemented in a second storage element associated with a second processor core among the plurality of processor cores.

[0233] Example 81 includes the subject of any of Examples 78-79, wherein the FIFO queue is implemented in the first storage element to define feedback information for use by the first processor core based on the configuration.

[0234] Example 82 is a system including means for performing the method of any one of Examples 66-81.

[0235] Example 83 is a system comprising: a processor device including: a plurality of processor cores; a plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configurable to: alternatively implement a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; a configuration controller for configuring a first storage element among the plurality of storage elements associated with a first processor core among the plurality of processor cores based on a configuration definition to implement a set of FIFO queues for the first processor core in the first storage element; and routing hardware for writing instructions and data to the set of FIFO queues based on the configuration definition, wherein the instructions are associated with a user application and the data will be consumed during the execution of the instructions by the first processor core.

[0236] Example 84 includes the subject matter of Example 83, wherein the routing hardware is configured to: receive the instructions and data from the network; and determine that the configuration of the first processor core is suitable for executing the instructions, wherein, based on the determination that the configuration of the first processor core is suitable for executing the instructions, the instructions and data are written into the set of FIFO queues.

[0237] Example 85 includes the subject matter of Example 84, wherein the first processor core is configured to implement at least a portion of a particular type of processor device based on the configuration definition, and the particular type of processor device is adapted to execute the instructions.

[0238] Example 86 includes the subject of Example 85, wherein the particular type of processor device includes a hardware accelerator.

[0239] Example 87 includes the subject of Example 85, wherein the configuration for processor cores described in the configuration definition is used to implement at least one processor of the particular type and at least one different second type of processor device through the plurality of processor cores.

[0240] Example 88 includes the subject of Example 87, wherein the specific type and the second type are selected from the group consisting of: graphics processing unit (GPU), network processing unit, tensor processing unit (TPU), vector processing unit (VPU), compression engine unit (CEU), encryption processing unit, storage acceleration unit (SAU), or machine learning accelerator.

[0241] Example 89 includes the subject of any one of Examples 83-88, wherein the routing hardware includes an Infrastructure Processing Unit (IPU).

[0242] Example 90 includes the subject of any one of Examples 83-89, wherein the routing hardware is used to write instructions and data directly into the set of FIFO queues.

[0243] Example 91 includes the subject of any one of Examples 83-90, wherein the set of FIFO queues comprises multiple FIFO queues.

[0244] Example 92 includes the subject matter of Example 91, wherein the plurality of FIFO queues include a plurality of data FIFO queues and a plurality of instruction FIFO queues, a first instruction FIFO queue of the plurality of instruction FIFO queues is used to provide a first instruction for execution by the execution unit of the first processor core, and a first data FIFO queue of the plurality of data FIFO queues is used to provide data for the first instruction operation to be executed by the first processor core.

[0245] Example 93 includes the subject matter of Example 92, wherein a second instruction FIFO queue in the plurality of instruction FIFO queues is used to provide a second instruction for execution by the execution unit of the first processor core, and a second data FIFO queue in the plurality of data FIFO queues is used to provide data for the second instruction operation to be executed by the first processor core, wherein the first instruction from the first instruction FIFO queue will be executed by the first processor core at least partially in parallel with the second instruction.

[0246] Example 94 includes the subject matter of Example 91, and further includes: a configurable multiplexer circuit for directing information from the first storage element to a corresponding execution unit among a plurality of execution units of the first processor core, wherein the configuration controller is configured based on the configuration definition to configure the configurable multiplexer circuit to direct information from the plurality of FIFO queues to the corresponding execution units among the plurality of execution units.

[0247] Example 95 includes the subject of Example 91, wherein the plurality of FIFO queues are used to provide corresponding information to the first processor core in the same clock cycle.

[0248] Example 96 includes the subject of Example 91, wherein the configuration controller is configured to: configure a second storage element associated with a second processor core among the plurality of processor cores based on the configuration definition, so as to implement an L1 cache for the second processor core.

[0249] Example 97 includes the subject matter of any one of Examples 83-96, and further includes: a configurable interconnect structure for interconnecting the plurality of processor cores, wherein the configuration controller is configured to configure the configurable interconnect structure based on the configuration definition to define the data flow between the plurality of processor cores.

[0250] Example 98 includes the subject matter of Example 97, wherein the configurable interconnect structure is configured to feed the output of the first processor core to a FIFO queue implemented in one of the plurality of memory elements, wherein the output includes either executable instructions or data to be operated by executable instructions.

[0251] Example 99 includes the subject of Example 97, wherein the FIFO queue is implemented in a second storage element associated with a second processor core among the plurality of processor cores.

[0252] Example 100 includes the subject of Example 97, wherein the FIFO queue is implemented in the first storage element to define feedback information for use by the first processor core based on the configuration.

[0253] Example 101 includes the subject matter of any one of Examples 83-100, wherein the configuration for processor cores described in the configuration definition is for implementing at least one first type of processor and at least one different second type of processor device through the plurality of processor cores.

[0254] Example 102 includes the subject of Example 101, wherein the first type includes a general-purpose processor, and the second type includes one of the following: a graphics processing unit (GPU), a network processing unit, a tensor processing unit (TPU), a vector processing unit (VPU), a compression engine unit (CEU), a cryptographic processing unit, a storage acceleration unit (SAU), or a machine learning accelerator.

[0255] Example 103 includes the subject of Example 101, wherein the configuration definition includes a first configuration definition for a first operation window, wherein the configuration controller is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operation window, and the configuration controller is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operation window.

[0256] Example 104 includes the subject of Example 103, wherein a first user application will execute in the first operation window on the plurality of processor cores configured based on the first configuration definition, while a different second user application will execute in the second operation window on the plurality of processor cores configured based on the second configuration definition.

[0257] Example 105 includes the subject of Example 104, wherein, for the first processor core, a specific workload of the first user application is routed to the set of FIFO queues based on the first configuration definition.

[0258] Example 106 is an apparatus that includes a set of processing elements that can be reorganized or reconfigured for different processing loads.

[0259] Example 107 includes the subject of Example 106, wherein the execution unit in the set of processing elements may have a configuration based on software definition applied.

[0260] Example 108 includes the subject of Example 107, wherein the configuration of the execution unit can change over time.

[0261] Example 109 includes the subject of any one of Examples 107-108, wherein the execution unit is configured to engage with a FIFO queue.

[0262] Example 110 includes the subject matter of any one of Examples 107-109, wherein one of the set of processing elements can be configured to engage with another of the set of processing elements.

[0263] Example 111 includes the subject of any one of Examples 107-110, wherein the processing element is configured to perform two or more additions.

[0264] Example 112 includes the subject of any one of Examples 107-110, wherein the processing element is configured to perform two or more multiplications.

[0265] Example 113 includes the subject of any one of Examples 107-110, wherein the processing element is configured to perform at least one addition and at least one multiplication.

[0266] Example 114 includes the subject matter of any one of Examples 106-110, wherein the processing element includes a configurable memory interface.

[0267] Example 115 includes the subject matter of Example 114, wherein the memory interface includes configurable multiplexer circuitry.

[0268] Example 116 includes the subject of any one of Examples 114-115, wherein the memory interface can be configured to change over time.

[0269] Example 117 includes the subject of any one of Examples 114-116, wherein the memory interface couples a FIFO queue to the execution unit of the processing element.

[0270] Example 118 includes the subject of any one of Examples 106-117, wherein the set of processing elements is configured to implement an AI or machine learning accelerator.

[0271] Example 119 includes the subject of any one of Examples 106-118, wherein the configuration of the set of processing elements defines a coupling that is configured to couple adjacent processing elements to each other to transfer data between the adjacent processing elements.

[0272] Example 120 includes the subject of Example 119, wherein the data is transferred between the adjacent processing elements with lower latency, lower power consumption, and / or reduced cache misses.

[0273] Example 121 includes the subject of any of Examples 106-120, wherein the set of processing elements can be configured to be configured or reconfigured based on precise time.

[0274] Example 122 includes the subject of Example 121, wherein the precise time is based on CPU time.

[0275] Example 123 includes the subject of Example 121, wherein the precise time is based on network time.

[0276] Example 124 includes the subject of Example 121, wherein the precise time is based on IEEE 1588.

[0277] Example 125 includes the subject of Example 121, wherein the precise time is based on PCIe Precision Time Measurement (PTM).

[0278] Example 126 includes the subject of any of Examples 121-125, wherein the precise time is less than 1 µs between adjacent devices.

[0279] Example 127 includes the subject of any one of Examples 121-125, wherein the precise time is less than 10 ns between adjacent processing elements.

[0280] Example 128 includes the subject of any of 121-127, where reconfiguration or reorganization occurs independently of the SOC clock.

[0281] Example 129 includes the subject of Example 128, wherein the SOC clock includes an always-running timer (ART).

[0282] Example 130 is an apparatus including intelligent routing hardware for routing data from a software thread to a FIFO queue of a specific processor core among a plurality of processor cores in a configurable processor device, based on the configuration of that specific processor core.

[0283] Example 131 includes the subject of Example 130, wherein the configuration causes the particular processor core to perform a particular function when the configuration is applied.

[0284] Example 132 includes the subject of any one of Examples 130-131, wherein the configuration defines the configuration of the cache structure for the particular processor core.

[0285] Example 133 includes the subject of any one of Examples 130-132, wherein the configuration defines the configuration of the execution unit of the particular processor core.

[0286] Example 134 includes the subject of any one of Examples 130-133, wherein the data points to a precise time.

[0287] Example 135 includes the subject of Example 134, wherein the precise time is based on CPU time.

[0288] Example 136 includes the subject of Example 134, wherein the precise time is based on network time.

[0289] Example 137 includes the subject of Example 134, wherein the precise time is based on IEEE 1588.

[0290] Example 138 includes the subject of Example 134, wherein the precise time is based on PCIe PTM.

[0291] Example 139 includes the subject of any of Examples 134-138, wherein the precise time is less than 1 µs between adjacent devices.

[0292] Example 140 includes the subject of any of Examples 134-138, wherein the precise time is less than 10 ns between adjacent processing elements.

[0293] Example 141 includes the subject of any of Examples 130-140, wherein the data is directed away from the SOC clock.

[0294] Example 142 includes the subject of Example 141, wherein the SOC clock includes an Always-Running Timer (ART).

[0295] Example 143 includes the subject of Example 141, wherein the SOC clock is also used by an adjacent SOC.

[0296] Example 144 includes the subject of Example 141, wherein the SOC clock is used by one or more chiplets in the SOC.

[0297] Example 145 is an apparatus that includes a configurable processor device, wherein the configurable processor device includes a set of processing elements that can be configured for different functions.

[0298] Example 146 includes the subject of Example 145, wherein the configuration allows for single-dimensional processing.

[0299] Example 147 includes the subject of Example 145, wherein the configuration allows for multidimensional processing.

[0300] Example 148 includes the subject of Example 145, wherein the configuration causes an input of a processing element to be based on the result of a previous processing element.

[0301] Example 149 includes the subject of any one of Examples 145-148, where different features include multiple accumulation features.

[0302] Example 150 includes the subject matter of any one of Examples 145-149, wherein a processing element in the set of processing elements can be configured to configure a set of execution units of that processing element.

[0303] Example 151 includes the subject of Example 150, wherein the interface cached to the execution unit changes according to the configuration based on multiple data inputs for the processing element.

[0304] Example 152 includes the subject of any one of Examples 150-151, wherein the interface cached to the execution unit changes according to the configuration based on multiple instruction inputs for the processing element.

[0305] Example 153 includes the subject of any one of Examples 150-152, wherein the interface cached to the execution unit is changed according to the configuration based on multiple data outputs for the processing element.

[0306] Example 154 includes the subject of any one of Examples 150-153, wherein the result of one of the execution units can be looped back to the previous execution unit according to the configuration.

[0307] Example 155 includes the subject of any one of Examples 150-153, wherein instructions used by one of the execution units can be reused by at least one other execution unit according to the configuration.

[0308] Example 156 is a method that includes reconfiguring one or more processing elements in a configurable processor device to generate future inputs to another processing element.

[0309] Example 157 includes the subject of Example 156, wherein the reconfiguration is a path reconfiguration.

[0310] Example 158 includes the subject of Example 156, wherein the reconfiguration allows for recirculation.

[0311] Example 159 includes the subject of Example 156, wherein the reconfiguration is based on precise time.

[0312] Example 160 includes the subject of Example 156, wherein the reconfiguration is based on an SOC or a shared chiplet clock.

[0313] Example 161 includes the subject of Example 156, wherein the reconfiguration is based on the functionality to be performed by the configurable processor device.

[0314] Example 162 includes the subject of Example 161, wherein the features include the ability to accelerate AI or machine learning workloads.

[0315] Example 163 includes the subject of Example 161, wherein the functionality includes quantum computing capabilities.

[0316] Example 164 includes the subject of Example 161, wherein the functionality includes quantum computing simulation.

[0317] Example 165 includes the subject of any one of Examples 156-164, wherein the future input is the result of a mathematical operation.

[0318] Example 166 includes the subject of any one of Examples 156-165, wherein the future input is an instruction.

[0319] Example 167 includes the subject of Example 166, wherein the instruction is a copy of the current instruction.

[0320] Example 168 includes the subject of Example 166, wherein the instructions are generated based on the current instructions.

[0321] Example 169 includes the subject matter of Example 166, wherein the instructions contain information for other future instructions.

[0322] Example 170 includes the subject of any one of Examples 156-169, wherein the reconfiguration enables the configurable processor device to perform multidimensional processing.

[0323] Example 171 includes the subject of any one of Examples 156-169, wherein the reconfiguration causes the configurable processor device to perform one-dimensional processing.

[0324] Example 172 includes the subject of any one of Examples 156-169, wherein the reconfiguration enables the configurable processor device to perform two-dimensional processing.

[0325] Example 173 includes the subject of any one of Examples 156-169, wherein the reconfiguration allows the configurable processor device to perform multidimensional processing.

[0326] Example 174 is a system including means for performing the methods of any one of Examples 156-173.

[0327] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order while still achieving the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result.

Claims

1. An apparatus comprising: Multiple processor cores; A plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configured to: alternatively implement either a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and Configure the controller for: Identify configuration definitions, wherein the configuration definitions define configurations for processor cores among the plurality of processor cores; and Based on the configuration definition, configure a first storage element associated with a first processor core among the plurality of processor cores to implement a set of FIFO queues for the first processor core in the first storage element.

2. The apparatus of claim 1, wherein, The set of FIFO queues includes multiple FIFO queues.

3. The apparatus of claim 2, wherein, The plurality of FIFO queues include a plurality of data FIFO queues and a plurality of instruction FIFO queues. The first instruction FIFO queue in the plurality of instruction FIFO queues is used to provide a first instruction for execution by the execution unit of the first processor core, and the first data FIFO queue in the plurality of data FIFO queues is used to provide data for the first instruction operation to be executed by the first processor core.

4. The apparatus of claim 3, wherein, The second instruction FIFO queue in the plurality of instruction FIFO queues is used to provide a second instruction for execution by the execution unit of the first processor core, and the second data FIFO queue in the plurality of data FIFO queues is used to provide data for the second instruction operation to be executed by the first processor core, wherein the first instruction from the first instruction FIFO queue will be executed in at least partially parallel with the second instruction by the first processor core.

5. The apparatus according to any one of claims 2-4, further comprising: A configurable multiplexer circuit is provided for directing information from the first storage element to a corresponding execution unit among a plurality of execution units of the first processor core, wherein the configuration controller is configured based on the configuration definition to configure the configurable multiplexer circuit to direct information from the plurality of FIFO queues to the corresponding execution units among the plurality of execution units.

6. The apparatus according to any one of claims 2-5, wherein, The multiple FIFO queues are used to provide corresponding information to the first processor core in the same clock cycle.

7. The apparatus according to any one of claims 2-6, wherein, The configuration controller is configured to: configure a second storage element associated with a second processor core among the plurality of storage elements based on the configuration definition, so as to implement an L1 cache for the second processor core.

8. The apparatus of any one of claims 1-7, further comprising: A configurable interconnect structure is provided for interconnecting the plurality of processor cores, wherein the configuration controller is configured to configure the configurable interconnect structure based on the configuration definition to define the data flow between the plurality of processor cores.

9. The apparatus of claim 8, wherein, The configurable interconnect structure is configured to feed the output of the first processor core to a FIFO queue implemented in one of the plurality of storage elements, wherein the output includes either executable instructions or data to be operated by executable instructions.

10. The apparatus of claim 9, wherein, The FIFO queue is implemented in a second storage element associated with a second processor core among the plurality of processor cores.

11. The apparatus of claim 9, wherein, The FIFO queue is implemented in the first storage element to send back information based on the configuration definition for use by the first processor core.

12. The apparatus according to any one of claims 1-11, wherein, The configuration for processor cores described in the configuration definition is used to implement at least one first type of processor and at least one different second type of processor device through the plurality of processor cores.

13. The apparatus of claim 12, wherein, The first type includes a general-purpose processor, and the second type includes one of the following: a graphics processing unit (GPU), a network processing unit, a tensor processing unit (TPU), a vector processing unit (VPU), a compression engine unit (CEU), a cryptographic processing unit, a storage acceleration unit (SAU), or a machine learning accelerator.

14. The apparatus of claim 12, wherein, The configuration definition includes a first configuration definition for a first operation window, wherein the configuration controller is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operation window, and the configuration controller is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operation window.

15. The apparatus of claim 14, wherein, A first user application will execute in the first operation window on the plurality of processor cores configured based on the first configuration definition, while a different second user application will execute in the second operation window on the plurality of processor cores configured based on the second configuration definition.

16. The apparatus of claim 15, wherein, For the first processor core, the specific workload of the first user application is routed to the set of FIFO queues based on the first configuration definition.

17. A method for configuring a configurable processor device, the method comprising: Receive configuration definition data, wherein the configuration definition data describes a specific configuration for the configurable processor device, the configurable processor device comprising: Multiple processor cores; Configurable interconnect structures for interconnecting components of the processor device; and A plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configured to: alternatively implement either a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and Based on the specific configuration, at least the configurable interconnect structure and the configurable multiple storage elements are configured to define the data flow of the multiple processor cores.

18. The method of claim 17, wherein, The specific configuration uses the multiple processor cores to implement a collection of multiple different types of processors.

19. The method of claim 18, wherein, The first type includes general-purpose processors, and the second type includes one of the following: graphics processing unit (GPU), network processing unit, tensor processing unit (TPU), vector processing unit (VPU), compression engine unit (CEU), encryption processing unit, storage acceleration unit (SAU), or machine learning accelerator.

20. The method of any one of claims 18-19, wherein, The configuration definition includes a first configuration definition for a first operation window, wherein the configuration hardware is configured to implement a first plurality of processors of different types through the plurality of processor cores based on the first configuration definition during the first operation window, and the configuration hardware is configured to implement a second plurality of processors of different types through the plurality of processor cores based on a second configuration definition during a subsequent second operation window.

21. The method of claim 20, wherein, A first user application will execute in the first operation window on the plurality of processor cores configured based on the first configuration definition, while a different second user application will execute in the second operation window on the plurality of processor cores configured based on the second configuration definition.

22. The method of claim 21, wherein, For the first processor core, the specific workload of the first user application is routed to a set of FIFO queues based on the first configuration definition.

23. A system comprising means for performing the method of any one of claims 17-22.

24. A system comprising: Processor device, including: Multiple processor cores; A plurality of configurable storage elements associated with the plurality of processor cores, wherein each of the plurality of storage elements is configured to: alternatively implement either a first-in-first-out (FIFO) queue or a level 1 (L1) cache for a corresponding associated processor core among the plurality of processor cores; and A configuration controller is configured, based on a configuration definition, to configure a first storage element associated with a first processor core among the plurality of storage elements and the plurality of processor cores, to implement a set of FIFO queues for the first processor core within the first storage element; and Routing hardware for writing instructions and data into the set of FIFO queues based on the configuration definition, wherein the instructions are associated with a user application and the data will be consumed during the execution of the instructions by the first processor core.

25. The system of claim 24, wherein, The routing hardware is used for: Receive the instructions and data from the network; and The configuration of the first processor core is determined to be suitable for executing the instructions, wherein, based on the determination that the configuration of the first processor core is suitable for executing the instructions, the instructions and data are written into the set of FIFO queues.