Multiprotocol header processing
By using CPI network and protocol processing in chiplet systems, the problem of high packaging overhead in multi-protocol networks is solved, achieving efficient multi-protocol processing, improving system performance and reducing power consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICRON TECHNOLOGY INC
- Filing Date
- 2021-06-23
- Publication Date
- 2026-05-12
AI Technical Summary
Existing network protocols require significant encapsulation overhead when handling multiple protocols, and it is difficult to support new protocols in existing networks. This leads to increased transmission and decapsulation processing loops, increased power consumption, and decreased system performance.
By employing the CPI network in chiplet systems, multiple chiplets are integrated on the intermediary layer, and header processing is performed using the CPI protocol. This enables flexible bridging and interpretation of multiple protocols, reduces packaging overhead, and improves system performance.
It reduces power consumption during transmission, reception, and decapsulation processing, improves system performance, supports existing networks to carry new protocols without modification, and reduces design and production costs.
Smart Images

Figure CN116235487B_ABST
Abstract
Description
[0001] Priority application
[0002] This application claims the benefit of U.S. Application No. 17 / 007,748, filed on August 31, 2020, the entire contents of which are incorporated herein by reference.
[0003] Statement regarding government support
[0004] This invention was developed with the support of the U.S. government under Agreement No. HR00111830003, and was granted by the Defense Advanced Research Projects Agency (DARPA). The U.S. government holds certain rights to this invention. Technical Field
[0005] Embodiments of this disclosure generally relate to network protocols, and more specifically, to networking that uses headers to distinguish between multiple protocols supported by a network. Background Technology
[0006] The network supports specific protocols and interprets packets according to a packet format defined for those protocols. Packet fields are located in predefined positions within the packet and are interpreted consistently across all packets. For example, the Internet Protocol (IP) uses an IP header with a 4-bit version field, a 4-bit Internet Header Length field, a 6-bit Differential Service Code Point field, a 2-bit Explicit Congestion Notification field, and a 16-bit Total Length field, which appear in the order described above within the first 32 bits of each IP packet.
[0007] Chiplets are an emerging technology for integrating various processing functionalities. Generally, a chiplet system consists of discrete chips (e.g., integrated circuits (ICs) on different substrates or dies) integrated on an interposer and packaged together. This arrangement differs from a single chip (e.g., an IC) containing dissimilar device blocks (e.g., intellectual property blocks) on a substrate (e.g., a single die) or a discrete packaged device integrated on a board. Generally, chiplets offer better performance (e.g., lower power consumption, reduced latency, etc.) compared to discretely packaged devices, and offer greater manufacturing benefits than single-die chips. These manufacturing benefits may include higher yields or reduced development costs and time.
[0008] Chiplet systems typically consist of one or more application chiplets and support chiplets. Here, the distinction between application and support chiplets pertains only to possible design scenarios for chiplet systems. Thus, for example, a synthetic vision chiplet system may include application chiplets for generating synthetic vision output, as well as support chiplets, such as memory controller chiplets, sensor interface chiplets, or communication chiplets. In typical use cases, synthetic vision designers may design the application chiplet and obtain the support chiplet from other sources. Therefore, design costs (e.g., in terms of time or complexity) are reduced by avoiding the design and manufacturing of functionality embodied in the support chiplet. Chiplets also support the tight integration of IP blocks that might otherwise be difficult, such as using IP blocks with different feature sizes. Therefore, devices with larger feature sizes designed during previous manufacturing, or those where feature sizes are optimized for power, speed, or heat generation—as might be the case for sensors—can be integrated with devices of different feature sizes, which is easier than attempting to do so on a single die. Furthermore, by reducing the overall die size, chiplet yields tend to be higher than those of more complex single-die devices. Attached Figure Description
[0009] This disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments thereof. However, the drawings should not be construed as limiting this disclosure to the particular embodiments, but are for explanation and understanding only.
[0010] Figure 1A and 1B An example of a chiplet system according to an embodiment is described.
[0011] Figure 2 This describes the components of an example of a memory controller chiplet according to an embodiment.
[0012] Figure 3 This describes an example of routing between chiplets using a chiplet protocol interface (CPI) network according to an embodiment.
[0013] Figure 4 This is a block diagram of data grouping including multiple flow control units (flits) according to some embodiments of the present disclosure.
[0014] Figure 5 This is a flowchart illustrating the operation of a method performed by a circuit when processing a packet having a header that supports multiple protocols, according to some embodiments of the present disclosure.
[0015] Figure 6 This is a flowchart illustrating the operation of a method performed by a circuit when processing a packet having a header that supports multiple protocols, according to some embodiments of the present disclosure.
[0016] Figure 7 This is a flowchart illustrating the operation of a method performed by a circuit in generating a header for a packet supporting multiple protocols, according to some embodiments of the present disclosure.
[0017] Figure 8 This is a block diagram of an example computer system in which embodiments of the present disclosure may be operated. Detailed Implementation
[0018] This disclosure relates to systems and methods for processing headers that support multiple protocols. The header of a packet includes a Bridge Type (BTYPE) field indicating the protocol of the packet. Based on the value of the BTYPE field, the command field of the packet is interpreted differently.
[0019] One of the benefits of the embodiments of this disclosure is that a single network can be used to carry packets of different protocols without the overhead of encapsulation. This allows existing networks supporting existing protocols to support new protocols without modifying existing network clients. It reduces the processing loops spent transmitting, receiving, and decapsulating encapsulated data. Additionally, power consumption during processing is reduced. Due to the reduced network overhead, the performance of the system, including the communication devices, is also improved. Other benefits will be apparent to those skilled in the art who benefit from this disclosure.
[0020] Figure 1A and 1B An example of a chiplet system 110 according to an embodiment is described. Figure 1A This is an illustration of a chiplet system 110 mounted on a peripheral board 105, which can be connected to a wider range of computer systems, for example, via peripheral component interconnect (PCIe). The chiplet system 110 includes a package substrate 115, an interposer 120, and four chips: an application chiplet 125, a host interface chiplet 135, a memory controller chiplet 140, and a memory device chiplet 150. Other systems may include numerous additional chipsets to provide additional functionality, as will be understood from the following discussion. The package of the chiplet system 110 is described as having a cap or cover 165, but other packaging technologies and structures used for chiplet systems may be used. Figure 1B This is a block diagram that labels the components in a chiplet system for clarity.
[0021] Application chip 125 is described as including a network on chip (NOC) 130 to support a chiplet network 155 for inter-chiplet communication. In an example embodiment, NOC 130 may be included on application chip 125. In an example, NOC 130 may be defined in response to the selected supporting chiplets (e.g., chiplets 135, 140, and 150), allowing the designer to select an appropriate number of chiplet network connections or switches for NOC 130. In an example, NOC 130 may reside on a single chiplet or even within interposer 120. In the example discussed herein, NOC 130 implements a CPI network.
[0022] CPI is a packet-based network that supports virtual channels to enable flexible and high-speed interaction between chiplets. CPI bridges the chiplet-internal network to the chiplet network 155. For example, the Advanced Extensible Interface (AXI) is a widely used specification for designing intra-chip communication. However, the AXI specification covers various physical design options, such as the number of physical channels, signal timing, power, etc. In a single chip, these options are typically selected to meet design goals, such as power consumption and speed. However, to achieve flexibility in chiplet systems, adapters (such as CPI) are used to intersect between various AXI design options that can be implemented in various chiplets. By enabling physical-to-virtual channel mapping and encapsulating time-based signaling with packet-based protocols, CPI bridges the chiplet-internal network 155.
[0023] CPI can use various physical layers to transmit packets. A physical layer may contain simple conductive connections or drivers to increase voltage or otherwise facilitate signal transmission over longer distances. An instance of this physical layer may contain an Advanced Interface Bus (AIB), which in various instances may be implemented in intermediate layer 120. The AIB uses source-synchronous data transfer with a forwarding clock to transmit and receive data. Regarding the transmitted clock, packets are transmitted across the AIB at Single Data Rate (SDR) or Double Data Rate (DDR). The AIB supports various channel widths. In SDR mode, the AIB channel width is a multiple of 20 bits (20, 40, 60, ...), and for DDR mode, it is a multiple of 40 bits (40, 80, 120, ...). The AIB channel width includes transmit (TX) and receive (RX) signals. Channels can be configured with a symmetrical number of TX and RX inputs / outputs (I / O), or an asymmetrical number of transmitters and receivers (e.g., all transmitters or all receivers). Depending on the chiplet providing the master clock, the channel can act as either an AIB master channel or a slave channel. The AIB I / O unit supports three timing modes: asynchronous (i.e., non-timing), SDR, and DDR. In various instances, the non-timing mode is used for the clock and some control signals. SDR mode can use a dedicated SDR-only I / O unit or a dual-purpose SDR / DDR I / O unit.
[0024] In this example, the CPI packet protocol (e.g., point-to-point or routable) can use symmetrical receive and transmit I / O units within an AIB channel. The CPI streaming protocol allows for more flexible use of AIB I / O units. In this example, an AIB channel for streaming mode can be configured with I / O units as all TX, all RX, or half TX and half RX. The CPI packet protocol can use AIB channels in SDR or DDR operating modes. In this example, AIB channels are configured in increments of 80 I / O units (i.e., 40 TX and 40 RX) for SDR mode and in increments of 40 I / O units for DDR mode. The CPI streaming protocol can use AIB channels in SDR or DDR operating modes. Here, in this example, for both SDR and DDR modes, the AIB channel increments by 40 I / O units. In this example, a unique interface identifier is assigned to each AIB channel. This identifier is used during CPI reset and initialization to identify paired AIB channels across adjacent chiplets. In this example, the interface identifier is a 20-bit value, comprising a 7-bit chiplet identifier, a 7-bit column identifier, and a 6-bit link identifier. The AIB physical layer uses an AIB out-of-band shift register to transmit the interface identifier. Bits 32-51 of the shift register are used to transmit the 20-bit interface identifier in both directions across the AIB interface.
[0025] AIB defines a stacked set of AIB channels as an AIB channel column. An AIB channel column has a certain number of AIB channels plus an auxiliary (AUX) channel. The AUX channel contains signals used for AIB initialization. All AIB channels within a column (except the AUX channel) have the same configuration (e.g., all TX, all RX, or half TX and half RX, and the same number of data I / O signals). In this example, AIB channels are numbered sequentially, starting with the AIB channel adjacent to the AUX channel. The AIB channel adjacent to the AUX channel is defined as AIB channel 0.
[0026] Generally, the CPI interface on an individual chiplet may include serialization-deserialization (SERDES) hardware. SERDES interconnects work well for situations where high-speed signaling with low signal counts is desired. However, SERDES can lead to additional power consumption and longer latency for multiplexing and demultiplexing, error detection or correction (e.g., using block-level cyclic redundancy check (CRC)), link-level retries, or forward error correction. However, when low latency or power consumption is the primary concern for very short-range chiplet-to-chiplet interconnects, parallel interfaces with clock rates that allow data transmission with minimal latency can be utilized. CPIs include elements to minimize both latency and power consumption in these very short-range chiplet interconnects.
[0027] For flow control, CPI employs a credit-based technique. The receiver (e.g., application chip 125) provides the sender (e.g., memory controller chip 140) with credits representing available buffers. In this example, the CPI receiver contains buffers for each virtual channel within a given transmission time unit. Therefore, if the CPI receiver supports five messages and a single virtual channel in time, the receiver has five buffers arranged in five rows (e.g., one row per unit time). If four virtual channels are supported, the receiver has twenty buffers arranged in five rows. Each buffer holds the payload of one CPI packet.
[0028] When a sender transmits data to a receiver, the sender decrements its available credits based on the transmission. Once all of the receiver's credits are exhausted, the sender will stop sending packets to the receiver. This ensures that the receiver always has an available buffer to store transmissions.
[0029] When the receiver processes the received packet and releases the buffer, it communicates the available buffer space back to the sender. This credit can then be used by the sender to allow the transmission of additional information.
[0030] It also describes a chiplet mesh network 160 that uses direct chiplet-to-chiplet technology without requiring NOC 130. The chiplet mesh network 160 can be implemented in CPI or another chiplet-to-chiplet protocol. The chiplet mesh network 160 typically enables a chiplet pipeline, where one chiplet acts as an interface to the pipeline, while other chips in the pipeline only interface with themselves.
[0031] Additionally, dedicated device interfaces, such as one or more industry-standard memory interfaces 145 (e.g., synchronous memory interfaces, such as DDR5, DDR6), can also be used to interconnect chiplets. Chiplet systems or individual chiplets can be connected to external devices (e.g., larger systems) via a desired interface (e.g., a PCIe interface). In this example, this external interface is implemented via a host interface chiplet 135, which provides a PCIe interface external to the chiplet system 110 in the depicted example. This dedicated interface 145 is typically used when industry practice or standards have focused on such an interface. The illustrated example of connecting the memory controller chiplet 140 to the DDR interface 145 of the dynamic random access memory (DRAM) memory device chiplet 150 is merely one such industry practice.
[0032] Among the various possible supporting chiplets, the memory controller chiplet 140 is likely to be present in the chiplet system 110, due to the near-ubiquitous use of memory in computer processing and the complex, advanced technology used for memory devices. Therefore, using the memory device chiplet 150 and memory controller chiplet 140 manufactured by other vendors allows chiplet system designers to leverage robust products from advanced manufacturers. Generally, the memory controller chiplet 140 provides a memory device-specific interface for reading, writing, or erasing data. Typically, the memory controller chiplet 140 can provide additional features such as error detection, error correction, maintenance operations, or atomic operation execution. For some types of memory, maintenance operations are often specific to the memory device chiplet 150, such as garbage collection in NAND flash or memory-class memory, and temperature regulation (e.g., cross-temperature management) in NAND flash memory. In instances, maintenance operations may involve logic-to-physical (L2P) mapping or management to provide an indirection level between the physical and logical representations of data. In other types of memory, such as DRAM, some memory operations (e.g., refresh) may be controlled by the host processor or memory controller at some times, and by the DRAM memory device or by logic (e.g., interface chip (in this example, buffer)) associated with one or more DRAM devices at other times.
[0033] Atomic operations are data manipulations that can be performed, for example, by the memory controller chiplet 140. In other chiplet systems, atomic operations can be performed by other chipsets. For example, an "incremental" atomic operation can be specified in a command by the application chiplet 125, wherein the command includes a memory address and a possible increment value. Upon receiving the command, the memory controller chiplet 140 retrieves a number from the specified memory address, increments the number by the amount specified in the command, and stores the result. Upon successful completion, the memory controller chiplet 140 provides the application chiplet 125 with an indication that the command was successful. Atomic operations avoid transferring data across the chiplet mesh network 160, thereby enabling lower latency execution of such commands.
[0034] Atomic operations can be classified as built-in atoms or programmable (e.g., custom) atoms. Built-in atoms are a finite set of operations that are immutably implemented in the hardware. Programmable atoms are small programs that can run on programmable atomic units (PAUs) (e.g., custom atomic units (CAUs)) of the memory controller chiplet 140. Figure 1A and 1B This section describes an example of a memory controller chiplet including a PAU.
[0035] Memory device chip 150 may be or include any combination of volatile memory devices or non-volatile memory. Examples of volatile memory devices include (but are not limited to) random access memory (RAM)—such as DRAM, synchronous DRAM (SDRAM), and graphics dual data rate type 6 SDRAM (GDDR6 SDRAM). Examples of non-volatile memory devices include (but are not limited to) NAND flash memory, memory-class memory (e.g., phase-change memory or memristor-based technology), and ferroelectric RAM (FeRAM). The illustrated examples include memory devices as memory device chip 150; however, the memory devices may reside elsewhere, such as in different packages on board 105. For many applications, multiple memory device chips may be provided. In examples, these memory device chips may each implement one or more memory technologies. In examples, the memory chip may include multiple stacked memory dies of different technologies (e.g., one or more SRAM devices stacked with one or more DRAM devices or otherwise communicating with one or more DRAM devices). The memory controller chiplet 140 can also be used to coordinate the operation between multiple memory chipsets in the chiplet system 110 (e.g., utilizing one or more memory chipsets in one or more levels of cache storage, and using one or more additional memory chipsets as main memory). The chiplet system 110 may also include multiple memory controller chipsets 140, which can be used to provide memory control functionality for individual processors, sensors, networks, etc. For example, the chiplet architecture of the chiplet system 110 provides the advantage of allowing adaptation to different memory storage technologies and different memory interfaces through updated chiplet configurations without requiring redesign of the rest of the system architecture.
[0036] Figure 2 The components of an example of a memory controller chiplet 205 according to an embodiment are described below. The memory controller chiplet 205 includes a cache 210, a cache controller 215, an off-die memory controller 220 (e.g., for communicating with off-die memory 175), a network communication interface 225 (e.g., for interfacing with chiplet network 180 and communicating with other chiplets), and a set of atom and merge operation units 250. Members of this set may include, for example, a write merge unit 255, a danger clearing unit 260, a built-in atom unit 265, or a PAU 270. The components are described logically and not necessarily as to whether they will be implemented. For example, a built-in atom unit 265 may include different means along the path to off-die memory. For example, a built-in atom unit may be in an interface means / buffer on the memory chiplet, as discussed above. Conversely, a PAU 270 may be implemented in a separate processor on the memory controller chiplet 205 (but in various instances, it may be implemented elsewhere, such as on the memory chiplet).
[0037] The off-die memory controller 220 is directly coupled to off-die memory 275 (e.g., via a bus or other communication connection) to provide write operations to and read operations from one or more off-die memories, such as off-die memory 275 and off-die memory 280. In the depicted example, the off-die memory controller 220 is also coupled for outputs to the atom and merge operation unit 250 and for inputs to the cache controller 215 (e.g., a memory-side cache controller).
[0038] In the instance configuration, the cache controller 215 is directly coupled to the cache 210 and can be coupled to the network communication interface 225 for input (e.g., incoming read or write requests) and coupled to the off-chip memory controller 220 for output.
[0039] Network communication interface 225 includes packet decoder 230, network input queue 235, packet encoder 240, and network output queue 245 to support packet-based chiplet network 285, such as CPI. Chiplet network 285 can provide packet routing between and within processors, memory controllers, mixed-thread processors, configurable processing circuitry, or communication interfaces. In this packet-based communication system, each packet typically contains destination and source addressing, as well as any data payload or instructions. In examples, chiplet network 285 may be implemented as a collection of crossbar switches with a folded Clos configuration or a mesh network providing additional connectivity, depending on the configuration.
[0040] The packet decoder 230 can convert received packets into memory commands (e.g., read commands, write commands, burst reads or writes, atomic operations, or any suitable combination thereof) for the memory controller 205. The commands can be selected based on the packet's protocol. For example, a specific value in the command field can be interpreted as a first command for a first protocol and a second command for a second protocol.
[0041] In various instances, the chiplet network 285 may be part of an asynchronous switching configuration. Here, data packets may be routed along any of various paths, such that, depending on the route, any selected data packet may arrive at its addressed destination at any of multiple different times. Alternatively, the chiplet network 285 may be implemented at least partially as a synchronous communication network, such as a synchronous mesh communication network. Both configurations of the communication network are considered for examples according to this disclosure.
[0042] The memory controller chip 205 can receive packets having, for example, a source address, a read request, and a physical address. In response, the off-die memory controller 220 or the cache controller 215 reads data from the specified physical address (which may be in off-die memory 275 or cache 210) and assembles a response packet containing the requested data to the source address. Similarly, the memory controller chip 205 can receive packets having a source address, a write request, and a physical address. In response, the memory controller chip 205 writes data to the specified physical address (which may be in cache 210 or off-die memory 275 or 280) and assembles a response packet containing confirmation that the data has been stored in memory to the source address.
[0043] Therefore, if possible, the memory controller chiplet 205 can receive read and write requests via chiplet network 285 and process the requests using cache controller 215, which interfaces with cache 210. If the request cannot be processed by cache controller 215, then off-chip memory controller 220 processes the request by communicating with off-chip memory 275 or 280, atomic and merge operation unit 250, or both. As described above, one or more levels of cache can also be implemented in off-chip memory 275 or 280, and in some such instances, can be directly accessed via cache controller 215. Data read by off-chip memory controller 220 can be cached in cache 210 by cache controller 215 for later use.
[0044] Atom and merge operation unit 250 is coupled to receive (as input) the output of off-die memory controller 220 and provides output to cache 210, network communication interface 225, or directly to chiplet network 285. Memory danger clear (reset) unit 260, write merge unit 255, and built-in (e.g., predetermined) atom operation unit 265 can each be implemented as a state machine with other combinational logic circuitry (e.g., adders, shifters, comparators, AND gates, OR gates, XOR gates, or any suitable combination thereof) or other logic circuitry. These components may also include one or more registers or buffers to store operands or other data. PAU 270 can be implemented as one or more processor cores or control circuitry and various state machines with other combinational logic circuitry or other logic circuitry, and may also include one or more registers, buffers, or memories to store addresses, executable instructions, operands, and other data, or may be implemented as a processor.
[0045] Write merging unit 255 receives read data and request data, and merges the request data and read data to create a single unit containing the read data and source address to be used in response or return data packets. Write merging unit 255 provides the merged data to the write port of cache 210 (or equivalently, to cache controller 215 for writing to cache 210). Optionally, write merging unit 255 provides the merged data to network communication interface 225 to encode and prepare response or return data packets for transmission on chiplet network 285.
[0046] When requested data is used in a built-in atomic operation, the built-in atomic operation unit 265 receives the request and reads the data from the write merging unit 255 or directly from the off-chip memory controller 220. The atomic operation is performed, and using the write merging unit 255, the resulting data is written to the cache 210 or provided to the network communication interface 225 to encode and prepare response or return data packets for transmission on the chiplet network 285.
[0047] Built-in atomic operation unit 265 handles predefined atomic operations, such as fetch and increment or compare and swap. In examples, these operations perform simple read-modify-write operations on a single memory location of 32 bytes or less. Atomic memory operations are initiated from request packets transmitted via chiplet network 285. The request packet has a physical address, atomic operator type, operand size, and optionally up to 32 bytes of data. The atomic operation performs a read-modify-write operation on the cache line of cache 210, filling the cache if necessary. The atomic operator response can be a simple complete response or a response with up to 32 bytes of data. Example atomic memory operators include fetch and AND, fetch and OR, fetch and XOR, fetch and addition, fetch and subtraction, fetch and increment, fetch and decrement, fetch and minimum, fetch and maximum, fetch and swap, and compare and swap. In various example embodiments, 32-bit and 64-bit operations are supported, as well as operations on 16 or 32 bytes of data. The methods disclosed herein are also compatible with hardware that supports larger or smaller operations and more or less data.
[0048] Built-in atomic operations may also involve requests for "standard" atomic operations on the requested data, such as relatively simple single-cycle integer atoms, such as fetch and increment or compare and swap, which will occur with the same amount of processing as regular memory read or write operations that do not involve atomic operations. For these operations, cache controller 215 typically reserves cache lines in cache 210 by setting a danger bit (in hardware) so that the cache line cannot be read by another process while it is in transition. Data is obtained from off-die memory 275 or cache 210 and is provided to built-in atomic operation unit 265 to perform the requested atomic operation. After the atomic operation, in addition to providing the obtained data to data packet encoder 240 to encode outgoing data packets for transmission on chiplet network 285, built-in atomic operation unit 265 also provides the obtained data to write merging unit 255, which also writes the obtained data to cache 210. After the obtained data is written to cache 210, any corresponding danger bits that were set are cleared by memory danger clearing unit 260.
[0049] The PAU 270 achieves high performance (high throughput and low latency) for programmable atomic operations (also known as "custom atomic operations") comparable to built-in atomic operations. Instead of performing multiple memory accesses, in response to an atomic operation request specifying a programmable atomic operation and a memory address, the circuitry in the memory controller chiplet 205 transmits the atomic operation request to the PAU 270 and sets a danger bit in a memory danger register stored at the memory address corresponding to the memory line used in the atomic operation. This ensures that no other operation (read, write, or atomic operation) is performed on that memory line, and the danger bit is then cleared when the atomic operation is complete. The additional direct data path provided to the PAU 270 for performing programmable atomic operations allows for additional write operations without being limited by the bandwidth of the communication network or adding any congestion to the communication network.
[0050] The PAU 270 includes a multi-threaded processor, such as a RISC-VIS-based multi-threaded processor with one or more processor cores and further having an extended instruction set for performing programmable atomic operations. When equipped with the extended instruction set for performing programmable atomic operations, the PAU 270 can be embodied as one or more hybrid-threaded processors. In some example implementations, the PAU 270 provides bucket-loop instantaneous thread switching to maintain a high instruction-per-clock rate.
[0051] Programmable atomic operations can be executed by PAU 270, involving requests for programmable atomic operations on requested data. The user can prepare programming code to provide this programmable atomic operation. For example, the programmable atomic operation can be a relatively simple multi-loop operation (e.g., floating-point addition) or a relatively complex multi-instruction operation, such as Bloom filter insertion. The programmable atomic operation can be the same as or different from a predetermined atomic operation, as long as it is defined by the user, not the system vendor. For these operations, cache controller 215 can reserve cache lines in cache 210 by setting a danger bit (in hardware) so that the cache lines cannot be read by another process while they are in transition. Data is obtained from cache 210 or off-chip memory 275 or 280 and provided to PAU 270 to execute the requested programmable atomic operation. After the atomic operation, PAU 270 provides the obtained data to network communication interface 225 to directly encode outgoing data packets containing the obtained data for transmission on chiplet network 285. Additionally, PAU 270 provides the obtained data to cache controller 215, which in turn writes the obtained data to cache 210. After writing the obtained data to cache 210, cache controller 215 clears any corresponding dangerous bits that were set.
[0052] In the selected example, the approach taken for programmable atomic operations is to provide multiple custom atomic request types, which can be sent from the original source, such as a processor or other system component, to the memory controller chiplet 205 via chiplet network 285. Cache controller 215 or off-die memory controller 220 recognizes the request as a custom atom and forwards the request to PAU 270. In a representative embodiment, PAU 270: (1) is a programmable processing element capable of efficiently performing user-defined atomic operations; (2) can perform load and store operations to memory, arithmetic and logical operations, and control flow decisions; and (3) utilizes a RISC-V ISA with a new set of specialized instructions to facilitate interaction with such controllers 215, 220 to perform user-defined operations atomically. In the desired example, the RISC-V ISA contains a complete set of instructions that support high-level language operators and data types. PAU 270 may utilize a RISC-V ISA, but when included within memory controller chiplet 205, it will typically support a more limited instruction set and a limited register file size to reduce the die size of the unit.
[0053] As mentioned above, before writing read data to cache 210, the memory danger clearing unit 260 clears the set danger bits for reserved cache lines. Therefore, when a request and read data are received by write merging unit 255, a reset or clear signal can be transmitted from memory danger clearing unit 260 to cache 210 to reset the set memory danger bits for reserved cache lines. Furthermore, resetting this danger bit releases pending read or write requests involving specified (or reserved) cache lines, providing pending read / write requests to the inbound request multiplexer for selection and processing.
[0054] Figure 3 This describes an example of routing between chiplets in chiplet layout 300 using a CPI network according to an embodiment. Chiplet layout 300 includes chiplets 310A, 310B, 310C, 310D, 310E, 310F, 310G, and 310H. Chipslets 310A to 310H are interconnected via a network including nodes 330A, 330B, 330C, 330D, 330E, 330F, 330G, and 330H. Each of chiplets 310A to 310H includes hardware transceivers labeled 320A to 320H.
[0055] AIBs can be used to transmit CPI packets between chiplets 310. AIBs provide physical layer functionality. The physical layer uses source-synchronous data transfer with a forwarding clock to transmit and receive data. Packets are transmitted across AIBs at SDR or DDR relative to the transmitted clock. Various channel widths are supported by AIBs. When operating in SDR mode, AIB channel widths are multiples of 20 bits (20, 40, 60, ...), and for DDR mode, AIB channel widths are multiples of 40 bits (40, 80, 120, ...). AIB channel widths include both TX and RX signals. Channels can be configured with a symmetrical number of TX and RX (I / O) or an asymmetrical number of transmitters and receivers (e.g., all transmitters or all receivers). Depending on which chiplet provides the master clock, a channel can act as an AIB master channel or a slave channel.
[0056] The AIB adapter provides interfaces to the AIB link layer and to the AIB physical layer (PHY). The AIB adapter provides a data hierarchy register, a power-on reset sequencer, and a control signal shift register.
[0057] The AIB physical layer consists of AIB I / O units. AIB I / O units (implemented by hardware transceiver 320 in some embodiments) can be input-only, output-only, or bidirectional. An AIB channel consists of a set of AIB I / O units, and the number of units depends on the configuration of the AIB channel. A receive signal on a chiplet is connected to a transmit signal on a pair of chipslets. In some embodiments, each column includes an AUX channel numbered 0 to N and a data channel.
[0058] AIB channels are typically configured as half TX data and half RX data, all TX data, or all RX data with associated clock and miscellaneous control. In some implementations, the number of TX and RX data signals is determined at design time and cannot be configured as part of the system initialization.
[0059] The CPI packet protocol (point-to-point and routable) uses symmetrical receive and transmit I / O units within an AIB channel. The CPI streaming protocol allows for more flexible use of AIB I / O units. In some example implementations, the AIB channel for streaming mode can be configured with I / O units as all TX, all RX, or half TX and half RX.
[0060] Network node 330 routes data packets within chip 310. Node 330 can determine the next node 330 to forward a received data packet based on one or more data fields of the data packet. For example, source or destination address, source or destination port, virtual channel, or any suitable combination thereof can be hashed to select consecutive network nodes or available network paths. This method of path selection can be used to balance network traffic.
[0061] Therefore, in Figure 3 The diagram illustrates the data path from chiplet 310A to chiplet 310D. Data packets are sent from hardware transceiver 320A to network node 330A; forwarded by network node 330A to network node 330C; forwarded by network node 330C to network node 330D; and transmitted by network node 330D to hardware transceiver 320D of chiplet 310D.
[0062] Figure 3The diagram also illustrates a second data path from chiplet 310A to chiplet 310G. Data packets are sent from hardware transceiver 320A to network node 330A; forwarded by network node 330A to network node 330B; forwarded by network node 330B to network node 330D; forwarded by network node 330D to network node 330C; forwarded by network node 330C to network node 330E; forwarded by network node 330E to network node 330F; forwarded by network node 330F to network node 330H; forwarded by network node 330H to network node 330G; and delivered by network node 330G to hardware transceiver 320G of chiplet 310G. (The last sentence appears to be incomplete and possibly refers to a different data path.) Figure 3 As can be clearly seen, multiple paths through the network can be used to transfer data between any pair of small chips.
[0063] The AIB I / O unit supports three timing modes: asynchronous (i.e., non-timing), SDR, and DDR. Non-timing mode is used for clock and some control signals. SDR mode can use dedicated SDR-only I / O units or dual-purpose SDR / DDR I / O units.
[0064] The CPI packet protocol (point-to-point and routable) can use AIB channels in either SDR or DDR operating modes. In some implementations, for SDR mode, AIB channels increment by 80 I / O units (i.e., 40 TX and 40 RX), and for DDR mode, increments by 40 I / O units.
[0065] The CPI streaming protocol can use AIB channels in either SDR or DDR operating mode. In some implementations, AIB channels are incremented by 40 I / O units for both modes (SDR and DDR).
[0066] A unique interface identifier is assigned to each AIB channel. This identifier is used during CPI reset and initialization to identify paired AIB channels spanning adjacent chiplets. In some implementations, the interface identifier is a 20-bit value comprising a seven-bit chiplet identifier, a seven-bit column identifier, and a six-bit link identifier. The AIB physical layer uses an AIB out-of-band shift register to transmit the interface identifier. Bits 32-51 of the shift register are used to transmit the 20-bit interface identifier in both directions across the AIB interface.
[0067] In some example implementations, AIB channels are numbered sequentially, starting with the AIB channel adjacent to the AUX channel. The AIB channel adjacent to the AUX channel is defined as AIB channel 0.
[0068] Figure 3The example demonstrates eight chiplets 310 connected by a network comprising eight nodes 330. More or fewer chiplets 310 and more or fewer nodes 330 can be included in the chiplet network, thus allowing the creation of networks of chiplets of any size.
[0069] Figure 4 This is a block diagram of a data packet 400 comprising multiple micro-slices according to some embodiments of the present disclosure. The data packet 400 is divided into micro-slices, each of which consists of 36 bits. The first micro-slice of the data packet 400 includes a control path field 405, a path field 410, a destination identifier (DID) field 415, a sequence continuation (SC) field 420, a length field 425, and a command field 430. The second micro-slice includes address fields 435 and 445, a service ID (TID) field 440, and a reserved (RSV) field 450. The third micro-slice includes a credit return (CR) / RSV field 455, an address field 460, a source identifier (SID) 465, a bridge type (BTYPE) 470, and an extended command (EXCMD) 475. Each remaining micro-slice includes CR / RSV fields (e.g., CR / RSV fields 480 and 490) and data fields (e.g., data fields 485 and 495).
[0070] The control path field 405 is a two-bit field that indicates whether the CR / RSV fields of later micro-pieces in the group contain CR data, RSV data, or should be ignored, and whether the path field 410 is applied to control the sorting of the group. In some example embodiments, a value of 0 or 1 in the control path field 405 indicates that CR / RSV fields 455, 480, and 490 contain CR data; a value of 2 or 3 in the control path field 405 indicates that CR / RSV fields 455, 480, and 490 contain RSV data; a value of 0 indicates that path field 410 is ignored; a value of 1 or 3 indicates that path field 410 is used to determine the path of data group 400; and a value of 2 indicates that a single-path sorting will be used. In some example embodiments, a one-bit field is used. Alternatively, the high-order bit of the control path field 405 may be considered a one-bit field controlling whether CR / RSV fields 440 and 450 contain CR data or RSV data.
[0071] The path field 410 is an eight-bit field. When the control path field 405 indicates that the path field 410 is used to determine the path of data packets 400, it ensures that all data packets with the same value for the path field 410 take the same path through the network. Therefore, the order of the data packets will not change between the sender and receiver. If the control path field 405 indicates that a single-path ordering will be used, then a path is determined for each packet as if the path field 410 were set to zero. Therefore, all packets take the same path, and the order will not change, regardless of the actual value of the path field 410 for each data packet. If the control path field 405 indicates that the path field 410 will be ignored, then data packets are routed without considering the value of the path field 410, and the receiver can receive the data packets in a different order than the order in which the sender sent them. This avoids congestion in the network and allows for greater throughput in the device.
[0072] The DID field 415 stores a twelve-bit DID. The DID uniquely identifies the destination (e.g., the destination chip) within the network. It is guaranteed that the sequence of data packets with the SC field 420 set will be delivered sequentially. The length field 425 is a five-bit field indicating the number of chips comprising data packet 400. The interpretation of the length field 425 can be non-linear. For example, values 0 to 22 can be interpreted as 0 to 22 chips in data packet 400, and values 23 to 27 can be interpreted as 33 to 37 chips in data packet 400 (i.e., 10 more than the indicated value). Other values for the length field 425 can be vendor-defined rather than protocol-defined.
[0073] Commands for data packet 400 are stored in command field 430, which is a seven-bit field. The command can be a write command, a read command, a predefined atomic operation command, a custom atomic operation command, a read response, an acknowledgment response, or a vendor-specific command. Additionally, the command can indicate a virtual channel for data packet 400. For example, different commands can be used for different virtual channels, or bits 1, 2, 3, or 4 of the seven-bit command field 430 can be used to indicate a virtual channel, with the remaining bits used to indicate the command. The following table illustrates virtual channels based on protocols and commands, according to some example embodiments.
[0074]
[0075] The address of the command can be indicated in the path field 410, address fields 435, 445, and 460, or any suitable combination thereof. For example, the 38 most significant bits of a 41-bit address aligned to 4 bytes can be indicated by concatenating address fields 460, 435, 410, and 445 in sequence (most significant bits first). The TID field 440 is used to match the response to the request. For example, if the first packet 400 is a read request identifying a memory location to be read, then the responsive second packet 400 containing the read data will contain the same value in the TID field 440.
[0076] The SID field 465 identifies the source of data packet 400. Therefore, the receiver of packet 400 can send a responsive packet by copying the value in the SID field 465 to the DID field 415 of the responsive packet. The 4-bit BTYPE field 470 specifies the set of commands used for packet 400. A BTYPE of 0 indicates a first method for determining the commands of packet 400 (e.g., commands determined based on the CPI protocol and command field 430). A BTYPE of 1 indicates a second method for determining the commands of packet 400 (e.g., commands based on the AXI protocol and EXCMD field 475). Other BTYPE values indicate other methods for determining the commands of packet 400. For example, a BTYPE of 2 may indicate a command based on the PCIe protocol, wherein the command is determined from the format and type fields in the following microchip. By adding (or appending) zeros before each 32-bit PCIe microchip, PCIe packets can be transmitted over a 36-bit CPI network without further modification.
[0077] In some implementations, if the command field 430 indicates the presence of an extended header, then the EXCMD field 475 contains the command regardless of the value of the BTYPE field 470. In these implementations, the command indicated in the EXCMD field 475 is interpreted based on the value of the BTYPE field 470.
[0078] Therefore, the much lower overhead of adding one or two micro-pieces of the protocol that identify packet 400 allows the network to support multiple protocols (in this example, CPI and AXI or PCIe) instead of completely encapsulating packets of the second protocol within the data field of data packet 400 for transmission over the network using the first protocol.
[0079] Memory access commands can identify the number of bytes to be written or accessed, the memory space to be accessed (e.g., off-die memory 275 or instruction memory for custom atomic operations), or any suitable combination thereof. In some example embodiments, the command can indicate additional bits identifying the command for a later microchip. For example, a multi-byte command can be sent by using a vendor-specific command in the seven-bit command field 430 and including a portion or all of the seven-bit Extended Command (EXCMD) field 475. Thus, for certain values of the command field 430, the packet 400 contains only a header microchip (e.g., ...). Figure 4 The first header micro-piece shown contains fields 405 to 430. For other values of command field 430, group 400 contains a predetermined additional number of header micro-pieces (e.g., as shown in...). Figure 4 The two additional header micro-pieces shown in the figure contain fields 435 to 475) or a predetermined total number of header micro-pieces (e.g., as in Figure 4 The three main header micro-pieces shown in the document contain fields 405 to 475.
[0080] If CR is enabled, two bits of CR / RSV fields 445, 480, and 490 identify whether CR is used for virtual channels 0, 1, 2, or 3, and the other two bits of CR / RSV fields 445, 480, and 490 indicate whether the number of credits to be returned is 0, 1, 2, or 3.
[0081] Figure 5 This is a flowchart illustrating the operation of method 500 performed by circuitry when processing packets having headers supporting multiple protocols, according to some embodiments of the present disclosure. Method 500 includes operations 510, 520, and 530. By way of example, but not limitation, method 500 is described as being performed by... Figure 1A , 1B Devices 1, 2, and 3 are used Figure 4 Data is grouped and executed.
[0082] In operation 510, the logic (e.g., implemented in the CPI network as) Figure 3 The small chip 310D Figure 2 The memory controller chip 205 receives a packet including a first field and a second field from a source (e.g., chip 310A). In some example embodiments, the packet is a data packet 400, the first field is a BTYPE field 470, and the second field is an EXCMD field 475. For clarity, it should be understood that, as from... Figure 4 It is clear that the above identifiers for "First Field" and "Second Field" are only used for identification, but to distinguish these two fields from each other; and do not indicate the location of the fields within data group 400.
[0083] In operation 520, the logic determines the protocol of the packet based on the first field. In some example implementations, if the BTYPE stored in the BTYPE field 470 is 0, then the protocol of the packet is CPI, and if the BTYPE is 1, then the protocol of the packet is AXI. Other BTYPE values may be reserved or defined by the vendor.
[0084] In operation 530, the command indicated by the group is logically determined based on the protocol and the second field. For example, if the protocol is CPI, the command is determined by comparing the value in the EXCMD field 475 with the table of CPI commands. As another example, if the protocol is AXI, the command is determined by comparing the value in the EXCMD field 475 with the table of AXI commands. In some example embodiments, the commands of one supported protocol are a strict subset of the commands of another supported protocol. For example, one protocol may support unpublished writes, published writes, configuration writes, reads, and atomic operations of different data sizes (e.g., 1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64, and 128 bytes), while another protocol supports read and write operations but not atomic operations.
[0085] Additional operations can be performed on packets based on protocols. For example, a first protocol (e.g., AXI) may include fields in a header indicating the packet's quality of service (e.g., using 4-bit values where higher values have higher priority) or protection type (e.g., secure or insecure), while a second protocol (e.g., CPI) may not include these fields, or may have one or more fields in different locations in the header. Thus, in some example embodiments, the packet-based protocol is the first protocol and the value of a third field indicating the storage quality of service, allowing the receiving device to determine the packet's quality of service. The second packet-based protocol is the second protocol, ignoring the third field and not determining the quality of service. As another example, the packet-based protocol is the first protocol and the value of a third field indicating the storage protection type, allowing the receiving device to determine the packet's protection type. The second packet-based protocol is the second protocol, ignoring the third field and not determining the protection type.
[0086] Packets with higher service quality can be processed before those with lower service quality. Packets with equal service quality can be processed in the order they were received, in round-robin service of a virtual channel, or any suitable combination thereof. Packets without service quality (e.g., packets received using a protocol that does not provide a service quality field) can be treated as if the service quality were a predetermined value (e.g., minimum, maximum, or median).
[0087] Therefore, by using method 500, a single packet header can support command sets of multiple protocols. This enables existing devices supporting different protocols to communicate without needing to completely encapsulate data packets in a single format to conform to the network protocol. This saves computational, network, and memory resources, thereby reducing latency and power consumption.
[0088] Figure 6 This is a flowchart illustrating the operation of method 600 performed by circuitry when processing packets having headers supporting multiple protocols, according to some embodiments of the present disclosure. Method 600 includes operations 610, 620, 630, 640, and 650. By way of example, but not limitation, method 600 is described as being performed by... Figure 1A , 1B Devices 1, 2, and 3 are used Figure 4 Data is grouped and executed.
[0089] In operation 610, the circuit (e.g., implemented in a CPI network as) Figure 3 The small chip 310D Figure 2 The memory controller chip 205 receives the first microchip of the packet, the first microchip including a third field (e.g., the CMD field 430 of the first microchip of packet 400). For clarity, it should be understood that, as from Figure 4 It is clear that the above identifiers for "First Field", "Second Field" and "Third Field" are for identification purposes only, but distinguish these three fields from each other; and do not indicate the location of the fields within data group 400. Similarly, the identifiers for "First Micro Piece" and "Second Micro Piece" are for identification purposes only, but distinguish the micro pieces from each other, and do not indicate the location of the micro pieces within data group 400.
[0090] In operation 620, the circuit receives a second micro-piece of the packet, the second micro-piece including a first field (e.g., the BTYPE field 470 of another micro-piece of packet 400) and a second field (EXCMD field 475).
[0091] The circuit determines the protocol of the packet indicated by the first field based on the third field (operation 630). For example, in a non-extended packet format, the CMD field 430 may indicate a command in the default protocol, but certain values of the CMD field 430 indicate that the packet contains an extended header including the BTYPE field 470 and the EXCMD field 475. When the extended header is interpreted, the BYTPE field 470 indicates the protocol of the packet, and the EXCMD field 475 indicates a command. In some example embodiments, a bit of the CMD field 430 indicates whether the header is extended.
[0092] In operation 640, the circuit determines the packet's protocol based on the first field indicating the packet's protocol and the determination of the first field. For example, if the CMD field 430 indicates the presence of an extended header and the BTYPE field 470 is 0, then the packet's protocol can be determined to be CPI. Alternatively, if the CMD field 430 indicates the presence of an extended header and the BTYPE field 470 is 1, then the packet's protocol can be determined to be AXI.
[0093] Based on the protocol and the second field, the circuit determines the command indicated by the packet (operation 650). For example, a specific command value in the EXCMD field 475 can be interpreted as a read command using one protocol and a write command using another protocol, but by interpreting the command according to the protocol indicated in the BTYPE field 470, the intent of the packet's source (e.g., chiplet 310A) can be correctly identified.
[0094] Therefore, by using method 600, a single packet header can support command sets of multiple protocols. This enables existing devices supporting different protocols to communicate without needing to fully encapsulate data packets in one format to conform to the network protocol. This saves computational, network, and memory resources, thereby reducing latency and power consumption. Furthermore, compared to method 500, enhanced support for the default protocol reduces header overhead for packets sent with non-extended headers, further reducing computational, network, and memory resource consumption.
[0095] Figure 7 This is a flowchart illustrating the operations of a method performed by circuitry when generating headers for packets supporting multiple protocols, according to some embodiments of the present disclosure. Method 700 includes operations 710, 720, and 730. By way of example and not limitation, method 700 is described as being performed by... Figure 1A , 1B Devices 1, 2, and 3 are used Figure 4 The data is processed in groups. Method 700 can be executed by the transmitting device in conjunction with the receiving device that executes method 500 or method 600.
[0096] In operation 710, the circuit (e.g., Figure 3 The hardware transceiver of the chip 310A uses a first protocol to receive commands. For example, commands can be received from an AXI device connected to the circuit using a dedicated signaling protocol or an AXI protocol within the chip 310A.
[0097] Based on the first protocol, in operation 720, the circuit generates a header for the packet. The header includes a first field indicating the first protocol and a second field indicating a command. For example, it may generate a header including... Figure 4 The header fields 405 to 475, where BYTPE field 470 indicates the first protocol and command field 430 or extended command field 475 indicates a command.
[0098] In operation 730, the circuit transmits a header for the packet. The header is followed by the rest of the packet, if present. For example, a write command header indicates the address to be written, and the body of the packet indicates the data to be written, but a read command header indicates the address to be read and is not followed by additional data.
[0099] Therefore, by using method 700, data packets from multiple protocols can be transmitted across a network with minor modifications to the packet header. This enables existing devices supporting different protocols to communicate without needing to completely encapsulate data packets in a single format to conform to the network protocol. Consequently, computational, network, and memory resources are saved, thereby reducing latency and power consumption.
[0100] Figure 8 This block diagram illustrates an instance machine 800, which may be used to implement any or more of the techniques (e.g., methods) discussed herein, in or through the instance machine 800. As described herein, an instance may contain logic or components or mechanisms in or operable by said logic or components or mechanisms. A circuit system (e.g., a processing circuit system) is a collection of circuits implemented in a tangible entity of machine 800 containing hardware (e.g., simple circuits, gates, logic, etc.). The membership of a circuit system may be flexible over time. A circuit system contains members that can perform specified operations individually or in combination during operation. In an instance, the hardware of the circuit system may be designed immutably to perform a specific operation (e.g., hardwired). In an instance, the hardware of the circuit system may contain variably connected physical components (e.g., execution units, transistors, simple circuits, etc.) containing machine-readable media that have been physically modified (e.g., magnetically, electrically, or movably placed particles of invariant mass, etc.) to encode instructions for a specific operation. For example, when connecting physical components, the fundamental electrical properties of the hardware composition change from insulator to conductor, and vice versa. Instructions enable embedded hardware (e.g., execution unit or loading mechanism) to create members of a circuit system within the hardware via variable connections, portions of which perform specific operations during operation. Thus, in an example, a machine-readable media element is part of the circuit system, or communicatively coupled to other components of the circuit system during device operation. In an example, any of the physical components can be used in more than one member of more than one circuit system. For example, during operation, an execution unit can be used at one point in time in a first circuit of a first circuit system and can be reused at different times by a second circuit in the first circuit system or a third circuit in the second circuit system. Additional examples of these components with respect to machine 800 are as follows.
[0101] In alternative embodiments, machine 800 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, machine 800 may operate as a server machine, a client machine, or both in a server-client network environment. In an example, machine 800 may act as a peer-to-peer (P2P) (or other distributed) network environment. Machine 800 may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, network device, network router, switch, or bridge, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by said machine. Furthermore, while only a single machine is described, the term "machine" should also be considered to include any collection of machines that individually or collectively execute a set (or more) of instructions to perform one or more of the methods discussed herein, such as cloud computing, Software as a Service (SaaS), or other computer cluster configurations.
[0102] Machine (e.g., computer system) 800 may include a hardware processor 802 (e.g., a central processing unit (CPU), graphics processing unit (GPU), hardware processor core, or any combination thereof), main memory 804, static memory (e.g., memory or storage device for firmware, microcode, basic input / output (BIOS), unified extensible firmware interface (UEFI), etc.) 806, and mass storage device 808 (e.g., hard disk drive, tape drive, flash memory, or other block device), some or all of which may communicate with each other via interconnect (e.g., bus) 830. Machine 800 may further include a display unit 810, an alphanumeric input device 812 (e.g., keyboard), and a user interface (UI) navigation device 814 (e.g., mouse). In an example, the display unit 810, input device 812, or UI navigation device 814 may be a touchscreen display. Machine 800 may additionally include a signal generating device 818 (e.g., speaker), a network interface device 820, and one or more sensors 816, such as a global positioning system (GPS), a sensor, a compass, an accelerometer, or other sensors. Machine 800 may include an output controller 828, such as a serial (e.g., Universal Serial Bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC) etc.) connection, to communicate or control one or more peripheral devices (e.g., printer, card reader, etc.).
[0103] The registers of processor 802, main memory 804, static memory 806, or mass storage device 808 may be or contain machine-readable medium 822, on which one or more sets of data structures or instructions 824 (e.g., software) embody or be utilized by any or more of the techniques or functions described herein. Instructions 824 may also reside wholly or at least partially in any of the registers of processor 802, main memory 804, static memory 806, or mass storage device 808 during execution by machine 800. In an example, one or any combination of hardware processor 802, main memory 804, static memory 806, or mass storage device 808 may constitute machine-readable medium 822. Although machine-readable storage medium 822 is described as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database or associated cache and server) configured to store one or more instructions 824.
[0104] The term "machine-readable medium" may include any medium capable of storing, encoding, or carrying instructions executed by machine 800 and causing machine 800 to perform any or more of the technologies disclosed herein, or any medium capable of storing, encoding, or carrying data structures used by or associated with such instructions. Examples of non-limiting machine-readable media may include solid-state memory, optical media, magnetic media, and signals (e.g., radio frequency signals, other photon-based signals, sound signals, etc.). In examples, non-transitory machine-readable media includes machine-readable media containing a plurality of particles having invariant (e.g., rest) mass, and thus being composed of matter. Therefore, non-transitory machine-readable media is machine-readable media that does not contain transiently propagating signals. Specific examples of non-transitory machine-readable media may include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disc read-only memory (CD-ROM) and digital versatile disc read-only memory (DVD-ROM).
[0105] In this example, information stored or otherwise provided on machine-readable medium 822 may represent instructions 824, such as instructions 824 themselves or a format from which instructions 824 may be derived. This format from which instructions 824 may be derived may contain source code, encoded instructions (e.g., in compressed or encrypted form), packaged instructions (e.g., split into multiple packages), or the like. Information representing instructions 824 in machine-readable medium 822 may be processed by a processing circuitry system into instructions for carrying out any of the operations discussed herein. For example, deriving instructions 824 from information (e.g., processed by processing circuitry) may include: compiling (e.g., from source code, object code, etc.), interpreting, loading, organizing (e.g., dynamic or static linking), encoding, decoding, encrypting, decrypting, packaging, depackaging, or otherwise manipulating the information into instructions 824.
[0106] In an example, the derivation of instruction 824 may include the assembly, compilation, or interpretation (e.g., by processing a circuit system) of information to create instruction 824 from some intermediate or preprocessed format provided by machine-readable media 822. When provided in multiple parts, information may be combined, unpacked, and modified to create instruction 824. For example, the information may be in multiple compressed source code packages (or object code or binary executable code, etc.) on one or more remote servers. The source code packages may be encrypted during transmission over a network and, if necessary, decrypted, decompressed, assembled (e.g., linked), and compiled or interpreted at the local machine (e.g., compiled or interpreted into libraries, standalone executables, etc.) and executed by the local machine.
[0107] Instruction 824 may further utilize any of several transmission protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.) to transmit or receive via the communication network 826 using the transmission medium through the network interface device 820. Example communication networks may include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile phone networks (e.g., cellular networks), conventional telephone (POTS) networks, and wireless data networks (e.g., IEEE 802.11 series standards, known as Wi-Fi®, IEEE 802.16 series standards, known as WiMax®), IEEE 802.15.4 series standards, peer-to-peer (P2P) networks, etc. In examples, network interface device 820 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas for connection to the communication network 826. In an example, network interface device 820 may include multiple antennas to conduct wireless communication using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) technologies. The term "transmission medium" should be considered to include any intangible medium capable of storing, encoding, or carrying instructions executable by machine 800, and includes digital or analog communication signals or other intangible media to facilitate communication of this software. The transmission medium is a machine-readable medium.
[0108] In the foregoing description, some exemplary embodiments of this disclosure have been described. It will be apparent that various modifications can be made to the embodiments of this disclosure without departing from the broader scope and spirit set forth in the appended claims. Therefore, the description and drawings should be considered illustrative rather than restrictive. The following is a non-exhaustive list of examples of embodiments of this disclosure.
[0109] Example 1 is a system comprising: a memory device; a memory controller connected to the memory device; and logic configured to perform operations including: receiving a header of a packet, the header including a first field and a second field; determining a protocol of the packet based on the first field; and determining a command indicated by the packet based on the protocol and the second field.
[0110] In Example 2, the object of Example 1 is included, wherein the operation further includes: determining the quality of service of the group based on the protocol and the third field.
[0111] In Example 3, the objects of Examples 1 to 2 are included, wherein the operation further includes: determining the protection type of the packet based on the protocol and the third field.
[0112] In Example 4, the objects of Examples 1 to 3 include: the group is a first group; the command is a first command; the protocol is a first protocol; and the operation further includes: processing the first command; receiving a second header of a second group, the second header including a third field and a fourth field; determining a second protocol of the second group based on the third field, the second protocol being different from the first protocol; determining a second command indicated by the second group based on the second protocol and the fourth field; and processing the second command.
[0113] In Example 5, the object of Example 4 includes: the second field is equal to the fourth field; and the first command is different from the second command.
[0114] In Example 6, the subject matter of Examples 4 to 5 includes, wherein the operation further includes: determining the quality of service of the first packet based on the first protocol and the fifth field of the first packet; and determining the quality of service of the second packet based on the second protocol.
[0115] In Example 7, the objects of Examples 4 to 6 are included, wherein the operation further includes: determining the protection type of the first packet based on the first protocol and the fifth field of the first packet; and determining the protection type of the second packet based on the second protocol.
[0116] In Example 8, the objects of Examples 1 to 7 are included, wherein the operation further includes: identifying a virtual channel of the group based on the protocol and the command.
[0117] In Example 9, the object of Examples 1 to 8 includes, wherein the operation further includes: determining, based on the protocol, several flow control units (chips) in the header of the packet.
[0118] In Example 10, the objects of Examples 1 to 9 include, wherein receiving the header of the packet includes: receiving a first flow control unit (chip) including a third field; receiving a second chip including the first field and the second field; and determining, based on the third field, to perform the protocol of the packet based on the first field.
[0119] Example 11 is a method comprising: receiving a header of a packet by a memory controller, the header including a first field and a second field; determining a protocol of the packet based on the first field; and determining a command indicated by the packet based on the protocol and the second field.
[0120] In Example 12, the subject of Example 11 includes determining the quality of service of the group based on the protocol and the third field.
[0121] In Example 13, the subject matter of Examples 11 to 12 includes determining the protection type of the packet based on the protocol and the third field.
[0122] In Example 14, the object of Examples 11 to 13 includes, wherein: the packet is a first packet; the command is a first command; the protocol is a first protocol; and the method further includes: processing the first command; receiving a second header of a second packet, the second header including a third field and a fourth field; determining a second protocol of the second packet based on the third field, the second protocol being different from the first protocol; determining a second command indicated by the second packet based on the second protocol and the fourth field; and processing the second command.
[0123] In Example 15, the object of Example 14 includes: the second field is equal to the fourth field; and the first command is different from the second command.
[0124] In Example 16, the subject matter of Examples 14 to 15 includes determining the quality of service of the first packet based on the first protocol and the fifth field of the first packet; and determining the quality of service of the second packet based on the second protocol.
[0125] In Example 17, the subject matter of Examples 14 to 16 includes, wherein the operation further includes: determining the protection type of the first packet based on the first protocol and the fifth field of the first packet; and determining the protection type of the second packet based on the second protocol.
[0126] Example 18 is a system comprising: logic of a first chiplet configured to perform operations including: receiving a command using a first protocol; generating a header of a packet including a first field and a second field based on the first protocol, the first field indicating the first protocol and the second field indicating the command; and transmitting the header of the packet to a memory controller; and logic of the memory controller configured to perform operations including: receiving the header of the packet; determining, based on the first field, that the protocol of the packet is the first protocol; and determining, based on the determined protocol and the second field, the indicated command.
[0127] In Example 19, the object of Example 18 includes, wherein the operation further includes: determining the quality of service of the packet based on the protocol of the packet being the first protocol and the value of the third field.
[0128] In Example 20, the objects of Examples 18 to 19 are included, wherein the operation further includes: determining the protection type of the packet based on the protocol of the packet being the first protocol and the value of the third field.
[0129] In Example 21, the object of Examples 18 to 20 includes: the packet is a first packet; the command is a first command; and the operation further includes: processing the first command; receiving a second header of a second packet, the second header including a third field and a fourth field; determining a second protocol of the second packet based on the third field, the second protocol being different from the first protocol; determining a second command indicated by the second packet based on the second protocol and the fourth field; and processing the second command.
[0130] In Example 22, the object of Example 21 includes: the second field is equal to the fourth field; and the first command is different from the second command.
[0131] In Example 23, the subject matter of Examples 21 to 22 includes, wherein the operation further includes: determining the quality of service of the first packet based on the first protocol and the fifth field of the first packet; and determining the quality of service of the second packet based on the second protocol.
[0132] In Example 24, the objects of Examples 21 to 23 include, wherein the operation further includes: determining the protection type of the first packet based on the first protocol and the fifth field of the first packet; and determining the protection type of the second packet based on the second protocol.
[0133] Example 25 is at least one machine-readable medium containing instructions that, when executed by a processing circuitry system, cause the processing circuitry system to perform operations for implementing any of Examples 1 to 24.
[0134] Example 26 is a device that includes components for implementing any of Examples 1 to 24.
[0135] Example 27 is a system for implementing any of Examples 1 through 24.
[0136] Example 28 is a method for implementing any of Examples 1 through 24.
Claims
1. A system for processing headers supporting multiple protocols, comprising: Memory devices; A memory controller connected to the memory device; as well as Logic, configured to perform operations including: The first fragment of the header of the received packet includes a first field; The second micro-fragment of the header of the packet is received, the second micro-fragment including a second field and a third field; Based on the first field, determine that the second field is used to determine the protocol of the packet based on the second field; Based on the second field, the protocol of the group is determined; Based on the protocol and the third field, determine the command indicated by the group; as well as Based on the protocol and the fourth field, the protection type of the packet is determined.
2. The system according to claim 1, wherein the operation further comprises: Based on the protocol and the fifth field, the quality of service of the group is determined.
3. The system according to claim 1, wherein: The group is the first group; The command mentioned is the first command; The protocol is the first protocol; and The operation further includes: Process the first command; Receive the second header of the second packet, which includes the fifth and sixth fields; Based on the fifth field, a second protocol for the second group is determined, which is different from the first protocol; Based on the second protocol and the sixth field, determine the second command indicated by the second group; and Process the second command.
4. The system according to claim 3, wherein: The third field is equal to the sixth field; and The first command is different from the second command.
5. The system according to claim 3, wherein the operation further comprises: Based on the second protocol, the service quality of the second group is uncertain.
6. The system of claim 3, wherein the operation further comprises: Based on the second protocol, the protection type of the second packet is uncertain.
7. The system of claim 1, wherein the operation further comprises: Based on the protocol and the command, identify the virtual channel of the group.
8. The system of claim 1, wherein the operation further comprises: Based on the protocol, several micro-pieces in the header of the group are determined.
9. The system of claim 1, wherein the command includes a memory address.
10. The system of claim 1, wherein the command is a memory command.
11. The system of claim 1, wherein the command is a memory read command.
12. A method for processing headers supporting multiple protocols, comprising: The first fragment of the header of the received packet includes a first field; The second micro-fragment of the header of the packet is received, the second micro-fragment including a second field and a third field; Based on the first field, determine that the second field is used to determine the protocol of the packet based on the second field; Based on the protocol and the third field, determine the command indicated by the group; as well as Based on the protocol and the fourth field, the protection type of the packet is determined.
13. The method of claim 12, further comprising: Based on the protocol and the fifth field, the quality of service of the group is determined.
14. The method according to claim 12, wherein: The group is the first group; The command mentioned is the first command; The protocol is the first protocol; and The method further includes: Process the first command; Receive the second header of the second packet, which includes the fifth and sixth fields; Based on the fifth field, a second protocol for the second group is determined, which is different from the first protocol; Based on the second protocol and the sixth field, determine the second command indicated by the second group; and Process the second command.
15. The method according to claim 14, wherein: The third field is equal to the sixth field; and The first command is different from the second command.
16. The method of claim 14, further comprising: Based on the second protocol, the service quality of the second group is uncertain.
17. The method of claim 14, further comprising: Based on the first protocol and the seventh field of the first packet, the protection type of the first packet is determined; as well as Based on the second protocol, the protection type of the second packet is uncertain.
18. A system for processing headers supporting multiple protocols, comprising: The logic of the first chiplet, configured to perform operations including the following: Receive commands using the first protocol; Based on the first protocol, a header is generated for a group including a first field, a second field, and a third field, wherein the second field indicates the first protocol and the third field indicates the command; as well as Transmit the header of the packet to the memory controller; as well as The logic of the memory controller is configured to perform operations including: The first micro-fragment of the header of the packet is received, the first micro-fragment including the first field; The second micro-fragment of the header of the packet is received, the second micro-fragment including the second field and the third field; Based on the first field, determine that the second field is used to determine the protocol of the packet based on the second field; Based on the second field, it is determined that the protocol of the group is the first protocol; The indicated command is determined based on the determined protocol and the third field. as well as The protection type of the packet is determined based on the protocol of the packet, which is the first protocol and the value of the fourth field.
19. The system of claim 18, wherein the operation further comprises: The quality of service of the packet is determined based on the protocol of the packet, which is the first protocol and the value of the fifth field.
20. The system according to claim 18, wherein: The group is the first group; The command is the first command; and The operation further includes: Process the first command; Receive the second header of the second packet, which includes the fifth and sixth fields; Based on the fifth field, a second protocol for the second group is determined, which is different from the first protocol; Based on the second protocol and the sixth field, determine the second command indicated by the second group; and Process the second command.
21. The system according to claim 20, wherein: The third field is equal to the sixth field; and The first command is different from the second command.
22. The system of claim 20, wherein the operation further comprises: Based on the first protocol and the seventh field of the first packet, the service quality of the first packet is determined; as well as Based on the second protocol, the service quality of the second group is uncertain.
23. The system of claim 20, wherein the operation further comprises: Based on the second protocol, the protection type of the second packet is uncertain.
24. The system of claim 18, wherein the command includes a memory address.