On-package die-to-die (d2D) interconnect for memory using universal die interconnect fast (UCIE) PHY
By adopting UCIe PHY technology within the semiconductor package, efficient communication between SoC and memory chips is achieved, which solves the high bandwidth and low latency requirements of the interconnection architecture in high-performance computing devices and provides improvements in cost-effectiveness and power efficiency.
Patent Information
- Application Number
- CN202380093229.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-18
- Filing Date
- 2023-06-02
- Publication Date
- 2025-09-16
AI Technical Summary
Existing semiconductor interconnect architectures have difficulty meeting the requirements of high bandwidth, low latency, and power efficiency in high-performance computing devices, especially due to bottlenecks in communication between multiple processors and memories.
It adopts Universal Chip Interconnect Express (UCIe) PHY technology, realizes efficient communication between SoC and memory chip by implementing UCIe interface and interface logic in semiconductor package, and uses UCIe interface and associated interface logic for signal mapping and transmission, supporting various scenarios from handheld devices to high-performance computing applications.
It achieves reduced cost and latency at the same bandwidth density, provides higher memory bandwidth density, is suitable for the needs of different market segments, and supports seamless interoperability and power efficiency of multiple protocols.
Smart Images

Figure CN120660080A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is an international application and claims the benefit of and priority to the previously filed Indian patent application serial number 202341018509, filed on March 18, 2023, entitled “ON-PACKAGE DIE-TO-DIE (D2D) INTERCONNECT FOR MEMORY USING UNIVERSAL CHIPLET INTERCONNECT EXPRESS (UCIE) PHY”, which is incorporated herein by reference in its entirety. Background Art
[0003] Advances in semiconductor processing and logic design have increased the amount of logic that can reside on integrated circuit devices. Consequently, computer system configurations have evolved from a single or multiple integrated circuits within a system to multiple cores, multiple hardware threads, and multiple logical processors on a single integrated circuit, along with other interfaces integrated within such processors. A processor or integrated circuit typically includes a single physical processor die, which can include any number of cores, hardware threads, logical processors, interfaces, memory, controller hubs, and the like.
[0004] Smaller computing devices are becoming increasingly popular due to the increased ability to pack more processing power into smaller packages. Smartphones, tablets, ultra-thin laptops, and other consumer devices are growing exponentially. However, these smaller devices rely on servers for data storage and complex processing beyond their form factor. As a result, the demands of the high-performance computing market (e.g., the server space) have increased. For example, in modern servers, there is often not only a single processor with multiple cores, but also multiple physical processors (also known as multiple sockets) to increase computing power. But as processing power grows along with the number of devices in a computing system, communication between the sockets and other devices becomes more important.
[0005] In fact, interconnects have evolved from more traditional multi-drop buses that primarily handled electrical communications to mature interconnect architectures that facilitate fast communications. Unfortunately, the demand for future processors to consume at higher speeds has placed corresponding demands on the capabilities of existing interconnect architectures. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which the element is first introduced.
[0007] Figure 1 A first system according to one embodiment is shown.
[0008] Figure 2 An interconnect stack according to one embodiment is shown.
[0009] Figure 3 A second system according to one embodiment is shown.
[0010] Figure 4 A third system according to one embodiment is shown.
[0011] Figure 5 A fourth system according to one embodiment is shown.
[0012] Figure 6 A fifth system according to one embodiment is shown.
[0013] Figure 7 A sixth system according to one embodiment is shown.
[0014] Figure 8 A seventh system is shown according to one embodiment.
[0015] Figure 9 An eighth system is shown according to one embodiment.
[0016] Figure 10 A ninth system is shown according to one embodiment.
[0017] Figure 11 A top view of a semiconductor package according to one embodiment is shown.
[0018] Figure 12 A standard package according to one embodiment is shown.
[0019] Figure 13 A first advanced package according to one embodiment is shown.
[0020] Figure 14 A second high-level package according to one embodiment is shown.
[0021] Figure 15 A third advanced package according to one embodiment is shown.
[0022] Figure 16 A first logic flow according to one embodiment is shown.
[0023] Figure 17 A second logic flow according to one embodiment is shown.
[0024] Figure 18 A computer-readable storage medium according to one embodiment is shown. DETAILED DESCRIPTION
[0025] Embodiments described herein may include apparatus, systems, techniques, or processes for on-package die-to-die (D2D) interconnects. Specifically, embodiments herein may relate to on-package D2D interconnects for memories that utilize or involve a Universal Chip Interconnect Express (UCIe) adapter or physical layer (PHY).
[0026] Embodiments are generally directed to improvements in interconnects suitable for transferring information between semiconductor dies arranged in an on-package memory architecture. Some embodiments are particularly directed to implementing interconnects such as the Universal Chip Interconnect Express (UCIe Express) standard promulgated by the UCIe Consortium. TM Universal Chip Interconnect Express (UCIe) TM ) interconnect, such as the UCIe specification version 1.0 released on February 17, 2022, and any descendants, revisions, and variants thereof (collectively, the "UCIe specification"). Although some embodiments implement a UCIe interconnect, it will be appreciated that other semiconductor interconnects defined by other semiconductor specifications may also be used. The embodiments are not limited in this context.
[0027] In general, UCIe is an open industry standard interconnect that provides high bandwidth, low latency, power efficient, and cost-effective on-package connections between chiplets that are application specific integrated circuits (ICs). Although the UCIe interconnect is primarily designed to handle communications between chiplets, embodiments implement techniques to extend the UCI interconnect to manage communications between any type of semiconductor die (e.g., a system on a chip (SoC)) and one or more memory ICs (referred to as memory chips). In various architectures, the SoC and the memory chips are implemented as different semiconductor dies on the same semiconductor package. The SoC and the memory chips each instantiate a UCIe interface and associated interface logic to allow the SoC and the memory chips to transmit standard memory signals as UCIe signals transmitted over the UCIe interconnect. This allows for a unified solution to scale on-package memory to support a variety of applications ranging from applications suitable for handheld computers to high performance computing (HPC) applications.
[0028] For example, in one embodiment, it is assumed that the SoC implements application logic and a memory controller that controls memory operations on behalf of the application logic. It is further assumed that the on-package memory chip implements one or more memory units. The memory controller can use a standard memory interface and associated memory signals to send read or write commands and associated data to the on-package memory chip. Similarly, the memory units of the on-package memory chip can receive memory signals and send responses to the memory controller using a standard memory interface and associated memory signals. In the meantime, the UCIe interface and associated interface logic implement mapping between the UCIe interface and the memory interface. The UCIe interface uses mapping to map memory signals to UCIe signals for transmission on the UCIe interconnect, and maps UCIe signals back to memory signals for transmission to the memory controller or memory unit.
[0029] For example, in one embodiment, it is assumed that the SoC implements the application logic and the SoC structure. It is further assumed that the on-package memory chip implements the memory controller and one or more memory units. The SoC structure can send read or write commands and associated data to the on-package memory chip using a standard memory interface and associated memory signals on behalf of the application logic. Similarly, the memory controller of the memory unit of the on-package memory chip can receive memory signals and send responses to the SoC structure using a standard memory interface and associated memory signals. In between, the UCIe interface and the associated interface logic implement the mapping between the UCIe interface and the memory interface. The UCIe interface uses mapping to map memory signals to UCIe signals for transmission on the UCIe interconnect, and maps the UCIe signals back to memory signals for transmission to the memory controller or memory unit.
[0030] Traditional solutions for on-package memory can be viewed as falling into two broad categories. The first category includes direct double data rate (DDR) (e.g., low-power DDR (LPDDR) memory) package-on-package memory, which can be used for area-constrained client applications (e.g., handheld devices, laptops, etc.). The second category includes high-bandwidth memory (HBM), which runs at low speeds over very wide links on advanced packages and supports multiple independent channels (e.g., 16 channels of 64 bits each, 6.4 gigatransfers per second (GT / s) in HBM3), which is used for high-performance computing (HPC) type applications, but at a significantly higher cost.
[0031] Embodiments are generally directed to a unified solution that is universally applicable to the handheld or client market segment and applied to the server and HPC market segment. In some embodiments, the unified solution uses LPDDR memory on the same package as another semiconductor device. For example, in one embodiment, the semiconductor device can be implemented with the following features and specifications: Universal Chip Interconnect Express (UCIe) issued by the UCIe Consortium TM ) specifications, such as the UCIe Specification Version 1.0, released on February 17, 2022, and any descendants, revisions, and variants thereof (collectively, the "UCIe Specifications"), and other semiconductor specifications. Thus, even with standard packaging processing, such packages can achieve cost, power, and latency advantages at nearly the same bandwidth density. For example, a DDR pin can be approximately 5-10pJ / b, while UCIe in a standard package is approximately 0.5pJ / b. In another example, based on the latency delta between the UCIe physical layer (PHY) and the DDR PHY, an estimated round-trip latency savings of approximately 10 nanoseconds (ns) can be achieved. For higher memory bandwidth applications, embodiments provide multiple channels using a similar approach to that for the client, while being able to connect to multiple dies (if needed) via logic die fan-out. Compared to HBM, embodiments provide significantly higher bandwidth density than UCIe, and provide the ability to provide a lower cost solution when using standard packaging with similar latency and power efficiency profiles.
[0032] Embodiments may include one or more of the following independent applications to meet the needs of different market segments. Some examples of independent applications are described below. Embodiments are not limited to these examples.
[0033] In some embodiments, such as for mobile or client use, mapping LPDDR timing onto the UCIe PHY for in-package memory provides latency, power, and bandwidth density advantages over out-of-package dynamic random access memory (DRAM). The mapping can be achieved by porting a subset of DDR PHY interface (DFI) signals onto UCIe to minimize changes to existing memory controller designs while retaining all the advantages of UCIe. In some examples, DFI signals can refer to signals processed by the DFI. DFI is used in a variety of consumer electronic devices, including smartphones. DFI is an interface protocol that defines the signals, timing, and programmable parameters required to transmit control information and data to and from DRAM devices and between microcontrollers and PHYs. DFI is applicable to all DRAM protocols, including DDR4, DDR3, DDR2, DDR, LPDDR4, LPDDR3, LPDDR2, and LPDDR.
[0034] In some embodiments, for example, to scale to HBM-like densities or usage, scaling to 32GT / s provides the appropriate scale-up for standard and advanced packaging applications of UCIe. UCIe 3D can unlock further exponential scaling of bandwidth density for these applications.
[0035] In some embodiments, for example, to fully decouple the memory subsystem from the system-on-chip (SoC) structure at a higher level of abstraction while maintaining the advantages of in-package disaggregation, embodiments provide example mappings of Compute Express Link (CXL) memory protocol (CXL.mem) signals, which can reduce area and power overhead relative to traditional CXL stack construction.
[0036] In some embodiments, the optimization techniques are also applicable to traditional serializer / deserializer (SERDES) applications, some of which may be added as options to the CXL specification.
[0037] In some embodiments, techniques are used to asymmetrically extend the UCIe PHY to enable a memory controller to share the same PHY with Peripheral Component Interconnect Express (PCIe) or CXL and utilize the rest of the UCIe infrastructure.
[0038] In some embodiments, the UCIe PHY and die-to-die (D2D) adapter can be used "as is," with extensions for additional lanes as required for each of the solutions described below. This can include natively adding additional width in an asymmetric manner, as well as using multiple UCIe clusters where some lanes can be "reserved" and turned off to save power while utilizing existing bump-outs.
[0039] Figure 1 An example of a system 100 is shown. System 100 can be an electronic system, such as a computing system, that implements one or more semiconductor devices on a semiconductor die in a system-on-chip (SoC) implementation. For example, the semiconductor die can implement one or more interconnects and associated link protocols defined by a semiconductor specification, such as the UCIe specification. UCIe is an open industry standard interconnect that provides high-bandwidth, low-latency, power-efficient, and cost-effective on-package connections between die.
[0040] System 100 illustrates an example of multiple semiconductor devices in SoC 118. System 100 illustrates die-level integration to provide power-efficient and cost-effective performance. SoC 118 can be integrated with other semiconductor dies at the package level, suitable for applications ranging from handheld devices to high-end servers, with dies from multiple sources even connected on the same package with different packaging options.
[0041] like Figure 1 As depicted, SoC 118 includes one or more processors 102 coupled to SoC fabric 104 and one or more memories 106. Processor 102 may include any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or other processor. For example, each processor 102 is coupled to SoC fabric 104 via a link 120 (e.g., a front-side bus (FSB)). In one embodiment, link 120 is a serial point-to-point interconnect as described below. In another embodiment, link 120 includes a serial differential interconnect architecture that conforms to a different interconnect standard. The interconnect protocols and features discussed below may be used to implement link 120, which is coupled to the SoC fabric 104 described herein. Figure 1 A collection of components introduced in .
[0042] The memory 106 includes any memory device, such as random access memory (RAM), non-volatile (NV) memory, or other memory accessible to devices in the system 100. For example, the memory 106 is coupled to the SoC fabric 104 via a link 120, such as a memory interface. Examples of a memory interface include a double data rate (DDR) memory interface, a dual-channel DDR memory interface, a low power DDR (LPDDR), a dynamic RAM (DRAM) memory interface, or other types of memory interfaces.
[0043] SoC fabric 104 is the interconnect infrastructure within the SoC design. It contains wires, buses, and interconnected networks that facilitate communication and data transfer between the various components and subsystems integrated on a single chip. SoC fabric 104 acts as a backbone for connecting different functional blocks, such as central processing unit (CPU) cores, memory controllers, input / output interfaces, graphics processors, accelerators, and other intellectual property (IP) blocks. It enables these components to communicate, share data, and work together to efficiently perform desired tasks. SoC fabric 104 is responsible for managing data flow, routing signals, and maintaining the necessary data paths between different components. It ensures that data is transmitted accurately and with low latency, while also optimizing power consumption and overall performance.
[0044] Although not shown, the SoC fabric 104 can be coupled to off-die devices. Input / output (I / O) modules (also referred to as interfaces or ports) can implement a layered protocol stack to provide communication between the SoC fabric 104 and one or more semiconductor devices implemented on other semiconductor dies in the semiconductor package.
[0045] The SoC 118 may also include one or more accelerators 108. An accelerator is a separate architectural substructure that is architected with a different set of goals than the base processor, where these goals are derived from the requirements of a specific class of applications. Examples of accelerators include graphics accelerators for enhancing graphics rendering capabilities, cryptographic accelerators for assisting with encryption or decryption, web accelerators for improving web applications, hypertext preprocessor (PHP) accelerators for assisting with web development, and the like.
[0046] SoC 118 may include other semiconductor devices or components to implement various computing or communication functions, such as RF circuit 110 , modem 112 , optical device 114 , analog device 116 , etc. Embodiments do not limit the type or number of semiconductor devices implemented for SoC 118 .
[0047] Figure 2 A UCIe interconnect stack 200 is shown. In various embodiments, the SoC 118 may implement a UCIe interconnect 212 as defined by the UCIe specification. The SoC 118 may use the UCIe interconnect 212 to communicate with other semiconductor dies within the same semiconductor package, as described in the referenced embodiments. Figure 3 As stated.
[0048] In general, UCIe is an open industry standard interconnect that provides high-bandwidth, low-latency, power-efficient, and cost-effective on-package connectivity between chiplets. The UCIe specification defines ubiquitous interconnect at the package level and covers the die-to-die (D2D) input / output (I / O) physical layer, D2D protocol, and software stack, which leverages the mature Peripheral Component Interconnect Express (PCI) or ) and Compute Express Link TM , CXL TM ) industry standards.
[0049] The UCIe specification details a complete standardized D2D interconnect, including the physical layer, protocol stack, software model, and compliance testing, which will enable end users to easily mix and match core components from a multi-vendor ecosystem for system-on-chip (SoC) builds, including custom SoCs. The physical layer supports up to 32GT / s, has 16 to 64 channels, and uses a 256-byte stream control unit (FCU) for data, similar to PCIe 6.0. The protocol layer is based on a compute fast link with CXL.io (PCIe), CXL.mem, and CXL.cache protocols. Various on-die interconnect technologies are defined, such as organic substrates for "standard" 2D packaging, or embedded silicon bridges (EMIBs), silicon interposers, and fan-out embedded bridges for "advanced" 2.5D / 3D packaging. In some embodiments, the physical specification is based on the Advanced Interface Bus (AIB) defined by Intel Corporation, headquartered in Santa Clara, California. However, other embodiments may implement the physical specification based on other bus specifications. The embodiments are not limited to this context.
[0050] UCIe supports different data rates, widths, bump spacing, and channel ranges to ensure the widest possible interoperability. It defines a sideband interface for ease of design and verification. The building block of the interconnect is a cluster, which includes: N single-ended, unidirectional, full-duplex data channels (N=16 for standard packages and N=64 for advanced packages); one single-ended channel for valid; one channel for tracking; differential forwarding clocks in each direction; and 2 channels for sidebands in each direction (e.g., single-ended, one for 1200MHz clock and one for data). Advanced packages support spare channels to handle failed channels (e.g., including clock, valid, sideband, etc.), while standard packages support width degradation to handle failures. Multiple clusters can be aggregated to provide higher performance for each link.
[0051] UCIe is a layered protocol represented by a UCIe interconnect stack 200. The UCIe interconnect stack 200 may include a protocol layer 202, a die-to-die adapter 204, and a physical layer 206.
[0052] The physical layer 206 may be coupled to a UCIe interconnect 212, for example, Figure 3 The UCIe interconnect 304 is depicted. The physical layer 206 is responsible for electrical signaling, clocking, link training, sideband, and other physical layer operations. Information can be passed between the physical layer 206 and the die-to-die adapter 204 via the raw D2D interface (RDI) 210.
[0053] The die-to-die adapter 204 provides link state management and parameter negotiation between the core and the semiconductor die (e.g., SoC 118). It optionally ensures reliable data delivery through its cyclic redundancy check (CRC) and link-level retry mechanisms. It defines the underlying arbitration mechanism when supporting multiple protocols. The 256-byte Stream Control Unit (FCU) Level Interface Transport (FLIT) defines the underlying transport mechanism when the adapter is responsible for reliable transmission. Information can be passed between the die-to-die adapter 204 and the protocol layer 202 via the FLIT-aware D2D interface (FDI) 208.
[0054] UCIe natively maps PCIe and CXL protocols because these protocols are widely deployed at the board level across all computing components. This is done to ensure seamless interoperability by leveraging the existing ecosystem. Already deployed SoC builds, link management, and security solutions using PCIe and CXL can be used with UCIe. The usage models involved are also comprehensive: data transfers using direct memory access, software discovery, error handling, etc. are addressed through PCIe / CXL.io; memory use cases are handled through CXL.mem; and cache requirements for applications such as accelerators are addressed through CXL.cache. UCIe also defines a "streaming protocol" that can be used to map any other protocol. In addition, as usage models evolve in the future, the UCIe Consortium can innovate on future protocols optimized for chiplets and semiconductor dies.
[0055] Although the UCIe interconnect 212 is designed to handle communications between cores that are application-specific integrated circuits (ICs), it can be extended to manage communications between any semiconductor die (e.g., SoC 118) and one or more memory ICs (referred to as memory chips). The SoC 118 and the memory chips are implemented as different semiconductor dies on the same semiconductor package. The SoC 118 and the memory chips each instantiate a UCIe interface and associated interface logic to allow the SoC 118 and the memory chips to transmit standard memory signals as UCIe signals transmitted over the UCIe interconnect 212. For example, the memory controller 122 of the SoC 118 can use the standard memory interface and associated memory signals to send read or write commands and data to the on-package memory units. Similarly, the on-package memory units can receive memory signals and send responses to the memory controller 122 using the standard memory interface and associated memory signals. Meanwhile, the UCIe interface and associated interface logic maps memory signals to UCIe signals for transmission on the UCIe interconnect 212, and maps the UCIe signals back to memory signals for transmission to the memory controller 122 or the memory unit. This allows a unified solution to scale on-package memory to support a variety of applications ranging from applications suitable for handheld computers to HPC-type applications.
[0056] Figure 3 System 300 is shown. System 300 isolates a pair of semiconductor dies that may be implemented by system 100. System 300 may include SoC 118 and memory chip 302 integrated on a single semiconductor package 306. SoC 118 may exchange signals with memory chip 302 over UCIe interconnect 304. SoC 118 and memory chip 302 may implement UCIe interface 308 and UCIe interface 310, respectively. SoC 118 and memory chip 302 may also implement interface logic 322 and interface logic 324, respectively. In one embodiment, UCIe interface 308, UCIe interface 310, interface logic 322, interface logic 324, and UCIe interconnect 304 may conform to the UCIe specification.
[0057] In various embodiments, the operation of the UCIe interface 308 and the interface logic 322 mirrors the operation of the UCIe interface 310 and the interface logic 324, and vice versa. Thus, the operations discussed with respect to the UCIe interface 308 and the interface logic 322 of the SoC 118 are the same as or similar to the operations of the UCIe interface 310 and the interface logic 324 of the memory chip 302, and vice versa.
[0058] In various embodiments, the memory controller 122 of the SoC 118 may communicate with the memory unit 312 of the memory chip 302 via the UCIe interconnect 304. In one embodiment, for example, the UCIe interface 308 and the UCIe interface 310 may implement interface logic 322 and interface logic 324, respectively, to manage communications over the UCIe interconnect 304. Each of the interface logic 322 and the interface logic 324 manages data transmission over the UCIe interconnect 304 in a transmit mode and a receive mode.
[0059] From the perspective of SoC 118, in transmit mode, UCIe interface 308 receives memory signals from memory controller 122. The memory signals may include, for example, write or read commands represented by DFI signal 314. Interface logic 322 decodes DFI signal 314 and maps the memory signals to corresponding UCIe signals. Interface logic 322 encodes the UCIe signals for transmission on transmit channel 318 of UCIe interconnect 304.
[0060] From the perspective of the memory chip 302, in receive mode, the UCIe interface 310 receives UCIe signals from the transmit channel 318. The interface logic 324 decodes the UCIe signals and maps them to corresponding memory signals (e.g., DFI signals 316). The UCIe interface 310 sends the DFI signals 316 to the memory unit 312. The memory unit 312 receives the DFI signals 316 and generates a response. For example, if the DFI signal 316 represents a read request, the memory unit 312 retrieves the requested data and sends it back to the UCIe interface 310 as the DFI signals 316. In transmit mode, the interface logic 324 decodes the DFI signals 316 and maps the memory signals to corresponding UCIe signals. The UCIe interface 310 sends the UCIe signals to the SoC 118 on the receive channel 320 of the UCIe interconnect 304.
[0061] Returning to the perspective of SoC 118, in receive mode, UCIe interface 308 receives UCIe signals from receive channel 320 of UCIe interconnect 304. Interface logic 322 decodes the UCIe signals, maps the UCIe signals to memory signals, and encodes the memory signals into DFI signals 314 for transmission to memory controller 122.
[0062] In various embodiments, the memory signal is a double data rate (DDR) memory signal or a high bandwidth memory (HBM) memory signal. For example, the memory signal may include a signal associated with a DDR memory interface, a dual-channel DDR memory interface, an LPDDR memory interface, a dynamic RAM (DRAM) memory interface, or other types of memory interfaces.
[0063] In one embodiment, the memory signals include double data rate (DDR) physical layer (PHY) interface (DFI) signals. In transmit mode, the interface logic 322 maps the DFI signals 314 (e.g., DFI command and data timing signals) for the DDR PHY interface to UCIe command and data timing signals for the UCIe PHY interface. In receive mode, the interface logic 322 maps the UCIe command and data timing signals for the UCIe PHY interface to the DFI signals 314, e.g., DFI command and data timing signals for the DDR PHY interface.
[0064] In one embodiment, interface logic 322 encodes UCIe signals for transmission over transmit channel 318 of UCIe interconnect 304 to memory chip 302 located in the same semiconductor package 306 as SoC 118. Similarly, interface logic 322 decodes UCIe signals from receive channel 320 of UCIe interconnect 304 from memory chip 302 located in the same semiconductor package 306 as SoC 118.
[0065] In one embodiment, the interface logic 322 encodes a cyclic redundancy check (CRC) or error checking and correction (ECC) signal of the UCIe signal for transmission on the command channel 326 of the UCIe interconnect 304. The interface logic 322 decodes the CRC signal or the ECC signal of the UCIe signal from the command channel 326 of the UCIe interconnect 304.
[0066] As an example and not a limitation, assume that memory cell 312 is an LPDDR5 (x8) memory cell with a burst length of 8 to illustrate the benefits of moving memory devices within semiconductor package 306. It is worth noting that, given how the signal mapping and timing are performed, no buffering or logic overhead may be required on memory chip 302. Aside from the deserialization of command signals, the remaining timing may match the native LPDDR5 timing performed by memory controller 122. In some embodiments, the UCIe timing may be the same or similar to the timing that the memory chip 302 would need to take when sending packets over the DFI for an LPDDR connection. An example mapping is shown in Table 1.
[0067] Table 1
[0068]
[0069] Table 1 is an example of a low-power DDR5 (LP5x) to UCIe frequency mapping with an 8-gigabit-per-second (Gbps) data rate. As depicted in Table 1, the LP5x write clock (WCK) can be considered to have the same frequency as the UCIe clock. Typically, the command interface is simplified in UCIe by mapping the command interface to the same clock as the data. Each valid frame in UCIe is 1 nanosecond (ns) at this data rate, which is the same as the clock cycle time (tCK) in LP5. This allows all command and data timing from the DFI to be passed without additional design overhead.
[0070] Table 2 shows an example signal mapping between one LP5x(x8) channel and a UCIe signal.
[0071] Table 2
[0072]
[0073]
[0074]
[0075] Table 2 shows an example of LP5x to UCIe signal mapping. LPDDR5(x) has seven CMD pins (CA[6:0]) and one chip select (CS). In this example, the command (CMD) and CS buses are mapped to the same clock as the data on UCIe. For example, on UCIe, CA[3:0] is mapped to the lower command channel, and CA[7:4] and CS are mapped to the upper command channel. The eight (bidirectional) date / read / write data (DQ) lanes on LPDDR5(x) are assigned to eight lanes in each direction on UCIe for x8 mapping, as shown in Table 2, and x16 mapping can have 16 data lanes in each direction. The Direct Media Interface (DMI) signals, which perform various functions (data masking (1 bit per byte) for writes and carrying ECC for reads on LPDDR), are mapped to one lane in each direction on UCIe for x8 and two lanes in each direction for x16. To support ECC for writes, an extra lane is allocated in UCIe for x8 (two extra lanes for x16). LPDDR can repurpose the read data strobe (RDQS) for this purpose, and it can carry 6 bits of ECC for DMI and 9 bits of ECC for 128-bit data; these ECCs are mapped to bits [15:1] of the byte sent on the ECC lane on UCIe. WCK and RDQS are mapped to the clocks on the UCIe transmit (Tx) and receive (Rx) blocks, respectively.
[0076] The total lane count for UCIe can be shown in either or both of Table 2, discussed above, or Table 3, discussed further below. For x8 mapping, one module of the standard package can have redundant lanes for any additional features. For x16 mapping, 1.5 modules of the UCIe standard package may be sufficient.
[0077] Regarding error handling, for UCIe link speeds with a bit error rate (BER) of 1e-27 or better (e.g., the standard package example above running at 8GT / s), additional cyclic redundancy check (CRC) or error checking and correction (ECC) protection may not be needed. This is because it can provide a better BER than traditional double data rate I / O (DDRIO) package traces. When scaling to higher speeds with a UCIe BER of 1e-15, additional ECC or CRC protection can be added on the command line as a separate UCIe channel, and the error signal is returned from the memory chip 302 to the SoC 118. In some cases, there may be an upper limit on the time it takes to issue a signal indicating an error. Any pulsed error indication may require the memory controller to take appropriate action, for example, the memory controller 122 replaying a command from a previous time.
[0078] Figure 4 System 400 is shown. System 400 is similar to system 300. However, system 400 is an example of a first option for extending UCIe interconnect 212 to include multiple channels. System 400 shows an example of scaling to 2 channels of memory devices on the same memory chip 302.
[0079] For situations where higher density or bandwidth is required, the DFI mapping on UCIe is extended to support multiple channels on the memory chip 302. Embodiments may implement three example options related to this extension, where Figure 4 Depicts the first option, Figure 5 depicts the second option, and Figure 6 A third option is depicted. It will be appreciated that other embodiments may include Figure 4 、 Figure 5 ,or Figure 6 More / less / different options depicted in the examples shown.
[0080] System 400 depicts a first option for supporting multiple channels. In various embodiments, interface logic 322 of SoC 118 encodes UCIe signals for transmission on multiple transmit channels of UCIe interconnect 304. Interface logic 322 decodes UCIe signals from multiple receive channels of UCIe interconnect 304.
[0081] As an example and not a limitation, the command signal is scaled and an additional set of data returns is added to support 2 reads 1 write (2rlw) optimized scaling. The first command channel sends the CMD and CS signals for CH0 402. The second command channel sends the CMD and CS signals for CH1 404. The transmit channel sends a data set for writing, which is shared between CH0 402 and CH1 404. The two receive channels are used for: a data set for reading from CH0 402 and a data set for reading from CH1 404. In some embodiments, it may be desirable to ensure that the memory chip 302 routes the shared write data bus to both CH0 402 and CH1 404. In addition, the memory controller 122 on the SOC 118 may also be responsible for ensuring that there are no data conflicts on the shared data bus.
[0082] Figure 5System 500 is shown. System 500 is an example of a second option for extending the DFI mapping on UCIe to support multiple channels in a semiconductor package. In this second option, the single-channel assignment of system 300 is treated as a unit module. This second option instantiates the unit module of system 300 multiple times to scale to multiple channels. System 500 is an example of scaling memory device channels, where four semiconductor dies include two SoCs and two memory chips on a single semiconductor package 520.
[0083] like Figure 5 As depicted, system 500 includes four dies, including Die 0 524, Die 1 514, Die 2 502, and Die 3 522. Die 0 524 and Die 1 514 each represent a SoC, such as SoC 118. Die 2 502 and Die 3 522 each represent a memory unit, such as memory unit 312. Die 0 524 and Die 3 522 operate in a manner similar to system 300 to transmit memory signals mapped to UCIe signals over UCIe interconnect 532. Transmission occurs in both directions, as previously described with reference to systems 300 and 400.
[0084] Similar to systems 300 and 400, system 500 includes SoC Die 0 524, which further includes a memory controller 526 and a UCIe interconnect 532. The UCIe interconnect 532 includes interface logic (e.g., interface logic 322) for managing data transmission on the UCIe interconnect 532 in transmit mode and receive mode. The interface logic encodes UCIe signals for transmission on transmit channels 530 of the UCIe interconnect 532 to memory chip Die 3 522 located in the same package as SoC Die 0 524. The interface logic decodes UCIe signals from memory chip Die 3 522 located in the same package as SoC Die 0 524, which are received on receive channels 534 of the UCIe interconnect 532.
[0085] System 500 also includes a second SoC die 1 514, which further includes a second memory controller 510 and a second UCIe interconnect 504. The second UCIe interconnect 504 includes second interface logic (e.g., interface logic 322) for managing data transmission on the second UCIe interconnect 504 in transmit mode and receive mode. The second interface logic encodes UCIe signals for transmission on a transmit channel 516 of the second UCIe interconnect 504 to a second memory chip die 2 502 located in the same package as the second SoC die 0 524. The second interface logic decodes UCIe signals from the second memory chip die 2 502 located in the same package as the SoC die 0 524, which are received from the receive channel 518 of the second UCIe interconnect 504.
[0086] Figure 6 System 600 is shown. System 600 is an example of a third option for extending DFI mapping over UCIe to support multiple channels in a semiconductor package. System 600 shows the case of a daisy-chained UCIe link between memory dies.
[0087] System 600 assumes that the UCIe link to the memory controller is configured to support an aggregate bandwidth of 2 channels. However, in some embodiments, the UCIe link to the memory controller is configured to be asymmetric for 2 reads and 1 write (2rlw) flows. This option provides a cost advantage over the second option shown in system 500, but at the expense of memory latency. UCIe-to-UCIe conversion for memory channel CH1 device traffic can be accomplished by "buffer" logic on the die carrying the memory channel CH0 device.
[0088] System 600 illustrates a memory architecture that includes a daisy-chain of UCIe interconnects 620 between SoC Die 0 612, a first memory chip Die 2 610, and a second memory chip Die 3 602. Similar to systems 300 and 400, system 600 includes SoC Die 0 612, which further includes a memory controller 614 and a UCIe interconnect 620. The UCIe interconnect 620 includes interface logic (e.g., interface logic 322) for managing data transmission over the UCIe interconnect 620 in both transmit and receive modes. The interface logic encodes UCIe signals for transmission over a transmit channel 618 of the UCIe interconnect 620 to a memory chip Die 2 610 located in the same package as the SoC Die 0 612. The interface logic decodes the UCIe signal from the receive channel 622 of the UCIe interconnect 620 from the memory chip Die 2 610 that is in the same package as the SoC Die 0 612 .
[0089] In addition, the interface logic encodes the UCIe signal for transmission over the transmit channel 618 of the UCIe interconnect 620 via the buffer logic 630 of the first memory chip die 2 610, which is located in the same package as the SoC die 0 612, for the second memory chip die 3 602, which is located in the same package as the SoC die 0 612 and the first memory chip die 2 610. The interface logic decodes the UCIe signal from the second memory chip die 3 602, which is from the receive channel 622 of the UCIe interconnect 620, via the buffer logic 630 of the first memory chip die 2 610.
[0090] Some embodiments implement techniques for scaling the density or bandwidth of system 300, system 400, system 500, or system 600. To expand bandwidth density, UCIe advanced packaging can provide grouping of x64 lanes. When operating at 32 GT / s, four modules in such a cluster can help achieve nearly 1 TB / s using the same DFI mapping as provided above.
[0091] In some embodiments, UCIe can be extended to three-dimensional (3-D) interconnects and can use the same DFI mapping with a wider data bus to expand bandwidth density by about 25 times or more. This scaling can be achieved because it can be targeted at at least 9 micron (um) bump pitch (relative to 45um bump pitch), which can increase the overall density by 25 times. Table 3 provides an example comparison of bandwidth density in different cases.
[0092] Table 3
[0093]
[0094]
[0095] Table 3 shows a bandwidth density comparison. UCIe-A may refer to the advanced package of UCIe. UCIe-S may refer to the standard package of UCIe. The rows corresponding to transmit only (Tx) or receive only (Rx) may refer to unidirectional bandwidth, and the rows corresponding to (Tx+Rx) may refer to bidirectional bandwidth. When aggregating multiple devices behind a given set of UCIe links, additional "buffer memory" logic may be required to split the command and data streams between different devices or memory channels. Note that the cluster size can be different for different technologies, and one metric that may be considered useful is BW density.
[0096] Figure 7A system 700 is shown. The system 700 is an example of a memory subsystem decomposition technique for transmitting memory signals over a UCIe interconnect. The system 700 shows the decoupling of a SoC fabric (eg, SoC fabric 104) from an on-package memory subsystem.
[0097] In some applications, it would be useful to provide a higher level of abstraction that provides die disaggregation. This could allow independent scaling of in-package memory relative to the SoC structure, allowing for rapid repackaging for different stock keeping units (SKUs) or applications. It could also allow mixing and matching different memory technologies with the same SoC structure. The growing ecosystem around the CXL.mem IP provides a good abstraction layer to enable the protocol for these applications.
[0098] like Figure 7 As depicted, the SoC fabric 104 includes a UCIe interface 706 having interface logic 722 for transmitting information over a UCIe interconnect 720 that implements an asymmetric UCIe link 704. The interface logic 722 manages data transmission over the UCIe interconnect 720 to the memory chip 702 in transmit mode. The interface logic 722 decodes memory signals from the memory interface 714 of the SoC fabric 104. For example, in one embodiment, the memory signals are compute express link (CXL) memory signals. The interface logic 722 maps the memory signals to UCIe signals and encodes the UCIe signals for transmission over the asymmetric UCIe link 704 of the UCIe interconnect 720 via the UCIe interface 706.
[0099] The SoC fabric 104 includes interface logic 722 for managing data transmission from the memory chip 702 over the UCIe interconnect 720 in receive mode. The interface logic 722 decodes UCIe signals from the asymmetric UCIe link 704 of the UCIe interconnect 720. The UCIe signals may be encoded UCIe signals representing memory signals from the memory unit CHO 712 via the memory controller 710, the memory interface 716, and the UCIe interface 708. The UCIe interface 708 transmits the UCIe signals over the asymmetric UCIe link 704 of the UCIe interconnect 720. The UCIe interface 706 maps the received UCIe signals to memory signals and encodes the memory signals for transmission over the SoC fabric 104 via the memory interface 714. The SoC fabric 104 routes the memory signals to components of the SoC 118, such as the processor 102 or the application logic 718.
[0100] The UCIe interface 708, memory interface 716, and interface logic 724 operate in a manner similar to the corresponding UCIe interface 706, memory interface 714, and interface logic 722. The UCIe interface 708, memory interface 716, and interface logic 724 encode UCIe signals representing memory signals from the memory controller 710 and the memory unit CHO 712, and decode UCIe signals representing memory signals from the SoC fabric 104.
[0101] Figure 8 Shown is a system 800. System 800 is an example of a memory subsystem decomposition for memory signals (eg, CXL memory signals defined by the CXL.mem protocol).
[0102] In a traditional implementation, a traditional CXL stack may need to fully support the CXL.io protocol and CXL.mem for each link. In addition, CXL.io may need to support at least 50% of the link's maximum bandwidth. For die disaggregation applications, the primary protocol is CXL.mem, while the CXL.io protocol may only be needed for configuration input / output (CFG / IO) transactions, messages, and message signaled interrupts (MSIs). These may be considered sufficient for device discovery and enumeration, as well as error reporting capabilities.
[0103] To achieve the above capabilities without incurring the overhead of the entire CXL.io stack for each link, embodiments may be configured as follows: Figure 8 The system shown. CXL.io transactions can be packetized by sending their transaction layer packets (TLPs) within the data payload of a UCIe Vendor Defined Message (VDM). Each UCIe sideband packet can carry 64b or 2 data words (DWords or DW) of data. However, some embodiments may introduce an operation code (opcode) that can carry up to 8 DWs of data, allowing a 4DW TLP header and 4DW TLP data to be sent using a single UCIe VDM. Only update flow control (FC) (UpdateFC) data link layer packets (DLLPs) (UpdateFC DLLPs) may be allowed and they may be tunneled using a similar mechanism as TLPs. The DLLP initialization protocol may not be required. For virtual channel (VC) 0 (VC0), the initial credit of each type of header credit (P, NP, C) is implicitly assumed to be 1, and the initial credit of each type of data credit is implicitly assumed to be 1. Support for only a single VC is allowed. Device discovery and enumeration can follow CXL's "ganged link" topology, allowing multiple CXL.mem UCIe links to be associated with a single CXL.io path.
[0104] For CXL.mem, the main band data flow can use the latency optimized flow control unit (FCU) level interface transactions (FLITs) defined in the traditional UCIe specification. The number of links can be asymmetric. For example, the memory structure 840 to the SoC structure 804 can have more UCIe links than the SoC structure 804 to the memory structure 840. The asymmetric links can be optimized for 2 reads and 1 writes (2rlw) workloads. The SoC structure 804 can perform an address-based hash to determine which UCIe link the transaction should be sent to. A unique tag encoding / bit can be used to determine how the completion is returned to the initiator, regardless of which UCIe link the completion is returned on.
[0105] In some embodiments, this extension may also be useful for legacy CXL SERDES (out-of-package) stacks. This will be enabled in the CXL specification. It may be negotiated during link training as part of a modified Transaction Sequence (TS) bit 1 (TS1) and / or TS bit 2 (TS2) to identify whether this feature is supported and whether the link carries CXL.io. Since out-of-package SERDES may not have sidebands, embodiments may still use CXL.io FLITs to send TLPs, however, the aforementioned simplifications in terms of limited flow control and feature set, and the fact that only one link must carry it, may still apply.
[0106] System 800 is an example of memory subsystem decomposition for memory signals (e.g., CXL memory signals defined by the CXL.mem protocol). Figure 8 As depicted, system 800 includes a SoC fabric 804 that communicates with a memory fabric 840 via a set of UCIe links 850 of a UCIe interconnect 852. The SoC fabric 804 transmits memory signals to the memory fabric 840 and vice versa.
[0107] like Figure 8As depicted, the UCIe interconnect 852 may include multiple UCIe links 850. For example, the system 800 includes four UCIe links 850, by way of example and not limitation. Each of the four UCIe links 850 may comprise a channel or pathway between the SoC fabric 804 and the memory fabric 840. For example, a first UCIe link may include CXL.mem 806, UCIe 814, UCIe 816, and CXL.mem 830. A second UCIe link may include CXL.mem 808, UCIe 816, UCIe 824, and CXL.mem 832. A third UCIe link may include CXL.mem 810, UCIe 818, UCIe 826, and CXL.mem 834. A fourth UCIe link may include CXL.mem 812, UCIe 820, UCIe 828, and CXL.mem 836. CXL.mem 806, CXL.mem 808, CXL.mem 810, and CXL.mem 812 constitute a first CXL layer 844. CXL.mem 830, CXL.mem 832, CXL.mem 834, and CXL.mem 836 constitute a second CXL layer 848. UCIe 814, UCIe 816, UCIe 818, and UCIe 820 constitute a first UCIe layer 842. UCIe 822, UCIe 824, UCIe 826, and UCIe 828 constitute a second UCIe layer 846.
[0108] In one embodiment, SoC fabric 804 and memory fabric 840 communicate memory signals including CXL signals (e.g., CXL.mem signals according to the CXL.mem protocol). CXL signals may include, for example, CXL command and data timing signals for a CXL memory interface (implemented as CXL layer 844 and CXL layer 848). Interface logic 854 maps the CXL command and data timing signals for the CXL memory interface to UCIe command and data timing signals for a UCIe adapter and physical layer (PHY) interface (implemented as UCIe layer 842 and UCIe layer 846). Conversely, interface logic 856 maps the UCIe command and data timing signals for the UCIe adapter and PHY interface to CXL command and data timing signals for the CXL memory interface.
[0109] CXL.io 802 and CXL.io 838 execute CXL.io transactions. CXL.io transactions can be packetized by sending their transaction layer packets (TLPs) within the data payload of a UCIe Vendor Defined Message (VDM). Each UCIe sideband packet can carry 64 bytes or 2 data words (DWords or DWs) of data.
[0110] Figure 9 System 900 is shown. System 900 may be similar to system 800. However, system 900 replaces the CXL layer 844 and the CXL.mem block in CXL layer 848 of system 800 with a gearbox.
[0111] For die-disaggregation applications, embodiments improve the stack in terms of load latency and area by taking one more step and eliminating the packing / unpacking overhead of sending information as CXL.mem FLITs. Alternatively, embodiments can provide a dedicated channel for sending CXL.mem commands (e.g., as 16B slots) and a dedicated channel for sending data. Given the increased density and power efficiency of advanced packaging (UCIe), this trade-off may be worthwhile for high-performance computing (HPC) applications. An example of this is shown in system 900.
[0112] Using a SoC fabric 904 interface, such as a compute plug-in (CPI) or similar, commands and data can typically be carried independently. By using a gearbox to serialize / deserialize commands and data independently, embodiments can significantly reduce packetization / unpacking delays. For example, in some embodiments, these delays are reduced by approximately 15ns.
[0113] like Figure 9 As depicted, for example, system 900 may include a gearbox in each UCIe link 950. For example, a first UCIe link includes gearbox 906, UCIe 914, UCIe 922, and gearbox 930. A second UCIe link includes gearbox 908, UCIe 916, UCIe 924, and gearbox 932. A third UCIe link includes gearbox 910, UCIe 918, UCIe 926, and gearbox 934. A fourth UCIe link includes gearbox 912, UCIe 920, UCIe 928, and gearbox 936. Gearbox 906, gearbox 908, gearbox 910, and gearbox 912 constitute a first gearbox layer 944. Gearbox 930, gearbox 932, gearbox 934, and gearbox 936 constitute a second gearbox layer 948. UCIe 914 , UCIe 916 , UCIe 918 , and UCIe 920 constitute a first UCIe layer 942 . UCIe 922 , UCIe 924 , UCIe 926 , and UCIe 928 constitute a second UCIe layer 946 .
[0114] In one embodiment, SoC fabric 904 and memory fabric 940 communicate memory signals including CXL signals (e.g., CXL.mem signals according to the CXL.mem protocol). CXL signals may include, for example, CXL command and data timing signals for a CXL memory interface (implemented as gearbox layer 944 and gearbox layer 948). Interface logic 954 maps the CXL command and data timing signals for the CXL memory interface to UCIe command and data timing signals for a UCIe adapter and physical layer (PHY) interface (implemented as UCIe layer 942 and UCIe layer 946). Conversely, interface logic 956 maps the UCIe command and data timing signals for the UCIe adapter and PHY interface to CXL command and data timing signals for the CXL memory interface.
[0115] CXL.io 902 and CXL.io 938 execute CXL.io transactions. CXL.io transactions can be packetized by sending their transaction layer packets (TLPs) within the data payload of a UCIe Vendor Defined Message (VDM). Each UCIe sideband packet can carry 64 bytes or 2 data words (DWords or DWs) of data.
[0116] Figure 10 A system 1000 is shown. System 1000 illustrates an example of an asymmetric UCIe link 1010 for memory subsystem disaggregation within a semiconductor package. System 1000 may be similar to system 700. However, system 1000 implements a UCIe interconnect 1008 that implements an asymmetric UCIe link 1010 having multiple UCIe links or lanes.
[0117] Flow control can go directly into the SoC or memory fabric bridge, saving transaction layer queues, etc. Assuming UCIe runs at 16GT / s, the UCIe channel map can be arranged to carry the equivalent bandwidth of 4 channels of LPDDR5. Figure 10 An example diagram of this arrangement is shown.
[0118] System 1000 may include a SoC fabric 104 and a memory chip 702, as described with reference to system 700. SoC fabric 104 includes a memory interface 714 and a UCIe interface 706, along with associated interface logic 724. Memory chip 702 includes a memory controller 710 that communicates with memory unit CH0 1012, memory unit CH1 1002, memory unit CH2 1004, and memory unit CH3 1006. Memory controller 710 (or alternatively, a memory fabric) may communicate with the memory units via a memory protocol bus (e.g., a DFI bus for carrying DFI signals). SoC fabric 104 and memory chip 702 may transmit memory signals as UCIe signals over a UCIe interconnect 1008 that implements an asymmetric UCIe link 1010 using multiple lanes.
[0119] For example, from the SoC structure 104 to the memory chip 702, the asymmetric UCIe link 1010 may include or implement 32 data channels, where 2 valid frames are used to send 64B of data. The asymmetric UCIe link 1010 implements 18 command channels, where 3 commands are in parallel and each command has 6 channels. Therefore, it can carry 2 read 1 write (2rlw) commands in parallel. The asymmetric UCIe link 1010 implements 1 channel for header (HDR) and credit return, where the header carries acknowledgment (ack) and no acknowledgment (nak) as well as sequence number information for retry. The asymmetric UCIe link 1010 implements 1 channel for CRC, where 2BCRC is performed on commands and data every 2 valid frames, and 1 valid frame is delayed to save latency.
[0120] For example, from the memory chip 702 to the SoC fabric 104, the asymmetric UCIe link 1010 may include or implement 64 data lanes, with two read responses in parallel. The asymmetric UCIe link 1010 implements nine command lanes, with two data response headers and one write completion header in parallel. The asymmetric UCIe link 1010 implements one lane for headers and credit returns and one lane for CRC.
[0121] For example, in one embodiment, the system 1000 includes interface logic 724 for encoding UCIe signals for transmission over the asymmetric UCIe link 1010 of the UCIe interconnect 1008 to the memory chip 702 in the same package as the SoC fabric 104. The interface logic 722 decodes the UCIe signals from the asymmetric UCIe link 1010 of the UCIe interconnect 1008 from the memory chip 702 in the same package as the SoC fabric 104. The interface logic 724 performs similar operations on behalf of the memory chip 702.
[0122] In one embodiment, for example, the interface logic 722 may further encode a CRC signal or an ECC signal of a UCIe signal for transmission on the asymmetric UCIe link 1010 of the UCIe interconnect 1008 , and decode a CRC signal or an ECC signal of a UCIe signal from the asymmetric UCIe link 1010 of the UCIe interconnect 1008 .
[0123] In one embodiment, for example, the interface logic 722 encodes UCIe signals for transmission on multiple transmit channels of the asymmetric UCIe link 1010 of the UCIe interconnect 1008 and decodes UCIe signals from multiple receive channels of the asymmetric UCIe link 1010 of the UCIe interconnect 1008 .
[0124] In one embodiment, for example, the interface logic 722 encodes CXL.io signals of UCIe signals for transmission on the asymmetric UCIe link 1010 of the UCIe interconnect 1008 and decodes CXL.io signals of UCIe signals from the asymmetric UCIe link 1010 of the UCIe interconnect 1008 .
[0125] Figure 11 A semiconductor package 1100 is shown. The semiconductor package 1100 includes a SoC 118 and a memory chip 302 assembled on a substrate 1102. The SoC 118 and the memory chip 302 can communicate via a set of reference channels 1104. In one embodiment, the reference channels 1104 are embedded in the substrate 1102. The reference channels 1104 can be implemented as wire bonds, traces, silicon bridges, interposers, and other communication media.
[0126] Semiconductor package 1100 may have different physical configurations for semiconductor testing, where the different physical configurations are defined using different reference packages (e.g., a rigid reference package and form factor, a flexible reference package and form factor, and a custom reference package for OEMs).
[0127] Figure 12 A standard package 1200 is shown. The UCIe specification defines two types of packages. Standard package 1200 is an example of a standard package (2D) for cost-effective performance. There are multiple commercially available options, some of which are shown in the figure. The UCIe specification covers all types of package options within these categories.
[0128] In one embodiment, semiconductor package 1100 may be instantiated as a standard package 1200. This allows semiconductor package 1100 to be tested as it would be deployed in a commercially available option.
[0129] Standard package 1200 may show SoC 118 and memory chip 302 assembled on substrate 1102. The substrate may embed a set of reference channels 1104. Reference channels 1104 may be conductive paths between SoC 118 and memory chip 302. Standard package 1200 may optionally include memory chip 1202. Memory chip 1202 may include another memory chip for parallel configuration or daisy-chain configuration as previously discussed. The embodiments are not limited in this context.
[0130] Figure 13 Advanced package 1300 is shown. As previously discussed, the UCIe specification defines two types of packaging processes. Standard package 1200 is an example of a standard package (2D) for cost-effective performance. Advanced package 1300 is an example of a more advanced package for power-efficient performance. There are multiple commercially available options for advanced packaging, and advanced package 1300 is one commercially available option. The UCIe specification covers all types of packaging options within these categories.
[0131] In one embodiment, semiconductor package 1100 may be instantiated as advanced package 1300. This allows semiconductor package 1100 to be deployed in commercially available options.
[0132] Advanced package 1300 may show SoC 118 and memory chip 302 assembled on substrate 1102. However, instead of substrate 1102 having an embedded set of reference channels 1104, substrate 1102 may implement silicon bridge 1302 and / or silicon bridge 1304.
[0133] Silicon bridge 1302 and / or silicon bridge 1304 can include reference channel 1104 as a conductive path between SoC 118 and memory chip 302. Generally speaking, a silicon bridge in a substrate refers to a structure in which a layer of silicon material is used to connect two or more isolation regions or components on a substrate. Silicon bridges are typically formed using semiconductor process technologies such as photolithography, etching, and deposition. Silicon bridges 1302, 1304 can be used to establish electrical connections between isolation components (e.g., SoC 118 and memory chip 302) on a substrate, which can be used to integrate multiple functions or devices on a single chip. Silicon bridges can also be used to isolate or separate different regions on a substrate. For example, a silicon bridge can be used to create a barrier between different types of materials (e.g., a metal layer and a silicon layer) to prevent unwanted interaction or contamination.
[0134] Advanced package 1300 may optionally include memory chip 1202. Memory chip 1202 may include another memory chip for use in a parallel configuration or a daisy-chain configuration as previously discussed. The embodiments are not limited in this context.
[0135] Figure 14 An advanced package 1400 is shown. The UCIe specification defines two types of packaging processes. Advanced package 1400 is an advanced package (2.5D) for power efficient performance. There are multiple commercially available options for advanced packaging.
[0136] Advanced package 1400 is another example of a more advanced package for power efficient performance. In one embodiment, semiconductor package 1100 can be instantiated as advanced package 1400. This allows semiconductor package 1100 to be deployed in commercially available options.
[0137] High-level package 1400 may show SoC 118 and memory chip 302 assembled on interposer 1402. Interposer 1402 may have an embedded set of reference channels 1104. Interposer 1402 may be assembled on substrate 1102.
[0138] Interposer 1402 is an electronic component that acts as an interface between a chip or integrated circuit and its package or substrate. The interposer provides a connection between the chip and the package by routing electrical signals between the chip and the package. The interposer is typically a thin sheet of material, such as silicon or an organic substrate, which contains a network of electrical traces or vias. These traces are used to route signals between the chip and the package and may also include power and ground connections. The interposer can be used for advanced packaging technologies such as 2.5D and 3D packaging processes, in which multiple chips are stacked on top of each other to improve performance and reduce the overall size of the package. By using interposer 1402, chips can be connected to each other and to the package substrate 1102 without the need for wire bonding or flip-chip packaging. The interposer can also achieve heterogeneous integration, in which chips with different technologies (e.g., CPU and memory) can be integrated into a single package. Compared to traditional packaging technology, this allows for improved performance, improved power efficiency, and reduced costs.
[0139] Advanced package 1400 may optionally include memory chip 1202. Memory chip 1202 may include another memory chip for use in a parallel configuration or a daisy-chain configuration as previously discussed. The embodiments are not limited in this context.
[0140] Figure 15 Advanced package 1500 is shown. The UCIe specification defines two types of packaging processes. Advanced package 1500 is an advanced package (2.5D or 3D) for power efficient performance. There are multiple commercially available options for advanced packaging.
[0141] Advanced package 1500 is yet another example of a more advanced package for power efficient performance. In one embodiment, semiconductor package 1100 can be instantiated as advanced package 1500. This allows semiconductor package 1100 to be tested as it would be deployed in a commercially available option.
[0142] Advanced package 1500 may show SoC 118 and memory chip 302 assembled on interposer 1502. Interposer 1502 may embed silicon bridge 1504 and / or silicon bridge 1506 with reference channel 1104. Silicon bridge 1504 and silicon bridge 1506 may be similar to reference channel 1104. Figure 13 Silicon bridge 1302 and silicon bridge 1304 are depicted. Interposer 1502 can be assembled on substrate 1102.
[0143] Advanced package 1500 may optionally include memory chip 1202. Memory chip 1202 may include another memory chip for use in a parallel configuration or a daisy-chain configuration as previously discussed. The embodiments are not limited in this context.
[0144] The operation of the disclosed embodiments may be further described with reference to the following figures. Some of the figures may include logic flows. Although the figures presented herein may include specific logic flows, it will be understood that the logic flows merely provide examples of how the general functionality described herein may be implemented. Furthermore, unless otherwise indicated, a given logic flow does not necessarily have to be executed in the order presented. Furthermore, in some embodiments, not all actions shown in the logic flow may be required. Furthermore, a given logic flow may be implemented by hardware elements, software elements executed by a processor, or any combination thereof. The embodiments are not limited in this context.
[0145] Figure 16 An embodiment of a logic flow 1600 is shown. Logic flow 1600 may represent some or all of the operations performed by one or more embodiments described herein. For example, logic flow 1600 may include some or all of the operations performed by a device or entity described herein. More specifically, logic flow 1600 illustrates an example of SoC 118 and memory chip 302 communicating memory signals over UCIe interconnect 304.
[0146] In block 1602, logic flow 1600 decodes memory signals from a memory controller. In block 1604, logic flow 1600 maps the memory signals to UCIe signals. In block 1606, logic flow 1600 encodes the UCIe signals for transmission on a transmit channel of a UCIe interconnect. In block 1608, logic flow 1600 decodes UCIe signals from a receive channel of the UCIe interconnect. In block 1610, logic flow 1600 maps the UCIe signals to memory signals. In block 1612, logic flow 1600 encodes the memory signals for transmission to the memory controller.
[0147] As an example, referring to the SoC 118 and memory chip 302 of the system 300, the SoC 118 includes a memory controller 122 and a UCIe interconnect 304. The UCIe interconnect 304 includes interface logic 322 for managing data transmission on the UCIe interconnect 304 in a transmit mode and a receive mode. In the transmit mode, the interface logic 322 decodes memory signals from the memory controller 122, maps the memory signals to UCIe signals, and encodes the UCIe signals for transmission on a transmit channel 318 of the UCIe interconnect 304 via the UCIe interface 308. In the receive mode, the interface logic 322 decodes UCIe signals from a receive channel 320 of the UCIe interconnect 304, maps the UCIe signals to memory signals, and encodes the memory signals for transmission to the memory controller 122.
[0148] Figure 17 An embodiment of a logic flow 1700 is shown. The logic flow 1700 may represent some or all of the operations performed by one or more embodiments described herein. For example, the logic flow 1700 may include some or all of the operations performed by a device or entity described herein. More specifically, the logic flow 1700 illustrates an example of the SoC fabric 104 and the memory chip 702 communicating memory signals over the UCIe interconnect 720 implementing an asymmetric UCIe link 704.
[0149] In block 1702, logic flow 1700 decodes memory signals from the SoC fabric. In block 1704, logic flow 1700 maps the memory signals to UCIe signals. In block 1706, logic flow 1700 encodes the UCIe signals for transmission over the asymmetric link of the UCIe interconnect. In block 1708, logic flow 1700 decodes the UCIe signals from the asymmetric link of the UCIe interconnect. In block 1710, logic flow 1700 maps the UCIe signals to memory signals. In block 1712, logic flow 1700 encodes the memory signals for transmission over the SoC fabric.
[0150] As an example, referring to the SoC fabric 104 and the UCIe interconnect 720 of the system 700, the UCIe interconnect 720 includes interface logic 722 for managing data transmission on the UCIe interconnect 720 in a transmit mode and a receive mode. In the transmit mode, the interface logic 722 decodes memory signals from the SoC fabric 104, maps the memory signals to UCIe signals, and encodes the UCIe signals for transmission on the asymmetric UCIe link 704 of the UCIe interconnect 720. In the receive mode, the interface logic 722 decodes UCIe signals from the asymmetric UCIe link 704 of the UCIe interconnect 720, maps the UCIe signals to memory signals, and encodes the memory signals for transmission on the SoC fabric 104.
[0151] Figure 18 Device 1800 is shown. Device 1800 may include any non-transitory computer-readable storage medium 1802 or machine-readable storage medium, such as optical, magnetic, or semiconductor storage media. In various embodiments, device 1800 may include an article of manufacture or product. In some embodiments, computer-readable storage medium 1802 may store computer-executable instructions that are executable by circuits. For example, computer-executable instructions 1804 may include instructions for implementing the operations described with respect to any logic flow described herein. Examples of computer-readable storage medium 1802 or machine-readable storage medium may include any tangible medium capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. Examples of computer-executable instructions 1804 may include any appropriate type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, object-oriented code, visual code, etc.
[0152] The components and features of the above-described devices may be implemented using any combination of discrete circuits, application-specific integrated circuits (ASICs), logic gates, and / or single-chip architectures. Additionally, the features of the devices may be implemented using a microcontroller, a programmable logic array, and / or a microprocessor, or any suitable combination of the foregoing. It should be noted that hardware, firmware, and / or software elements may be collectively or individually referred to herein as "logic" or "circuitry."
[0153] It should be understood that the exemplary devices shown in the above block diagrams can represent one functional description example of many potential implementations. Therefore, the division, omission, or inclusion of block functions depicted in the accompanying drawings does not mean that the hardware components, circuits, software, and / or elements used to implement these functions must be divided, omitted, or included in the embodiments.
[0154] At least one computer-readable storage medium may include instructions that, when executed, cause a system to perform any of the computer-implemented methods described herein.
[0155] Some embodiments may be described using the expression "one embodiment" or "an embodiment" and their derivatives. These terms mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment. Furthermore, unless otherwise indicated, the above-described features are intended to be used together in any combination. Therefore, any features discussed individually may be used in combination with each other unless otherwise indicated that the features are incompatible with each other.
[0156] The detailed description herein may be presented in terms of program processes executed on a computer or computer network, with general reference to the notations and terminology used herein. These process descriptions and representations are used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art.
[0157] A process is generally considered to be a self-consistent sequence of operations producing a desired result. These operations are those requiring physical manipulation of physical quantities. Usually, although not necessarily, these quantities take the form of electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc., primarily for reasons of common usage. It should be noted, however, that all of these terms and similar terms are associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.
[0158] Furthermore, the manipulations performed are often referred to in terms such as addition or comparison, which are often associated with mental operations performed by a human operator. In most cases, no such capabilities of a human operator are required or expected in any of the operations described herein (which form part of one or more embodiments). Instead, these operations are machine operations. Useful machines for performing the operations of the various embodiments include general-purpose digital computers or similar devices.
[0159] Some embodiments may be described using the expressions "coupled" and "connected" and their derivatives. These terms are not necessarily intended to be synonyms for each other. For example, some embodiments may be described using the terms "connected" and / or "coupled" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0160] Various embodiments also relate to apparatus or systems for performing these operations. The apparatus may be specially constructed for the desired purpose, or it may comprise a general-purpose computer that can be selectively activated or reconfigured by a computer program stored in the computer. The processes presented herein are not inherently related to a particular computer or other apparatus. Various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure of a variety of these machines will be apparent from the description given.
[0161] What has been described above includes examples of the disclosed architecture. It is, of course, not possible to describe every conceivable combination of components and / or methodologies, but one skilled in the art will recognize that many further combinations and permutations are possible. Accordingly, the novel architecture is intended to embrace all such changes, modifications, and variations that fall within the spirit and scope of the appended claims.
[0162] As previously referenced Figures 1 to 18 The various elements of the described devices may include various hardware elements, software elements, or a combination thereof. Examples of hardware elements may include devices, logic devices, components, processors, microprocessors, circuits, processors, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), memory cells, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software elements may include software components, programs, applications, computer programs, application programs, system programs, software development programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. However, determining whether an embodiment is implemented using hardware elements and / or software elements can vary depending on any number of factors, such as desired computing rate, power levels, thermal tolerances, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints (as needed for a given implementation).
[0163] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium, which represent various logic within a processor, which, when read by a machine, causes the machine to manufacture logic to perform the techniques described herein. Such representations, referred to as “IP cores,” may be stored on tangible machine-readable media and provided to various customers or manufacturing facilities to be loaded into manufacturing machines that manufacture the logic or processor. Some embodiments may be implemented, for example, using a machine-readable medium or article that may store instructions or instruction sets that, if executed by a machine, may cause the machine to perform methods and / or operations according to the embodiments. Such a machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware and / or software. The machine-readable medium or article may include, for example, any suitable type of memory unit, memory device, memory article, memory medium, storage device, storage article, storage medium, and / or storage unit, such as memory, removable or non-removable media, erasable or non-erasable media, writable or rewritable media, digital or analog media, hard disk, floppy disk, compact disk read-only memory (CD-ROM), compact disk recordable (CD-R), compact disk rewritable (CD-RW), optical disk, magnetic media, magneto-optical media, removable memory cards or diskettes, various types of digital versatile disks (DVDs), magnetic tape, tape cassettes, etc. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, etc., implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0164] It should be understood that the exemplary devices shown in the above block diagrams can represent one functional description example of many potential implementations. Therefore, the division, omission, or inclusion of block functions depicted in the accompanying drawings does not mean that the hardware components, circuits, software, and / or elements used to implement these functions must be divided, omitted, or included in the embodiments.
[0165] At least one computer-readable storage medium may include instructions that, when executed, cause a system to perform any of the computer-implemented methods described herein.
[0166] Some embodiments may be described using the expression "one embodiment" or "an embodiment" and their derivatives. These terms mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment. Furthermore, unless otherwise indicated, the above-described features are intended to be used together in any combination. Therefore, any features discussed individually may be used in combination with each other unless otherwise indicated that the features are incompatible with each other.
[0167] The following examples relate to further embodiments in which various arrangements and configurations will be apparent.
[0168] A first example apparatus includes a system on a chip (SoC) including a memory controller and a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect including interface logic to manage data transfer on the UCIe interconnect in a transmit mode, the interface logic to: decode memory signals from the memory controller; map the memory signals to UCIe signals; and encode the UCIe signals for transmission on a transmit channel of the UCIe interconnect.
[0169] The first example apparatus also includes interface logic to manage data transmission on the UCIe interconnect in a receive mode, the interface logic to: decode UCIe signals from a receive channel of the UCIe interconnect; map the UCIe signals to memory signals; and encode the memory signals for transmission to a memory controller.
[0170] The first example apparatus also includes any of the previous examples, including where the memory signal is a double data rate (DDR) memory signal or a high bandwidth memory (HBM) memory signal.
[0171] The first example apparatus also includes any of the previous examples, including wherein the memory signals include double data rate (DDR) physical layer (PHY) interface (DFI) signals, and the interface logic is to: map the DFI command and data timing signals to UCIe command and data timing signals for a UCIe PHY interface; and map the UCIe command and data timing signals for the UCIe PHY interface to the DFI command and data timing signals.
[0172] The first example apparatus also includes any of the previous examples, including interface logic for: encoding UCIe signals for transmission on a transmit channel of the UCIe interconnect to a memory chip in the same package as the SoC; and decoding UCIe signals from a receive channel of the UCIe interconnect from the memory chip in the same package as the SoC.
[0173] The first example apparatus also includes any of the previous examples, including interface logic for: encoding a cyclic redundancy check (CRC) or error check and correction (ECC) signal of a UCIe signal for transmission on a command channel of a UCIe interconnect; and decoding a CRC signal or an ECC signal of a UCIe signal from the command channel of the UCIe interconnect.
[0174] The first example apparatus also includes any of the previous examples, including interface logic to: encode UCIe signals for transmission on a plurality of transmit channels of a UCIe interconnect; and decode UCIe signals from a plurality of receive channels of the UCIe interconnect.
[0175] The first example apparatus also includes any of the previous examples, including a second SoC, the second SoC including a memory controller and a second UCIe interconnect, the second UCIe interconnect including second interface logic to manage data transmission on the second UCIe interconnect in a transmit mode and a receive mode, the second interface logic to: encode UCIe signals for transmission on a transmit channel of the second UCIe interconnect to a second memory chip located in the same package as the second SoC; and decode UCIe signals from a receive channel of the second UCIe interconnect from the second memory chip located in the same package as the second SoC.
[0176] The first example apparatus also includes any of the previous examples, including interface logic for: encoding a UCIe signal for transmission on a transmit channel of the UCIe interconnect through buffer logic of a first memory chip in the same package as the SoC, for a second memory chip in the same package as the SoC and the first memory chip; and decoding a UCIe signal from the second memory chip from a receive channel of the UCIe interconnect through the buffer logic of the first memory chip.
[0177] The first example apparatus also includes any of the previous examples, comprising a SoC having a top side and a bottom side, the bottom side comprising a set of protrusions corresponding to a set of physical layer blocks, each protrusion having a specific position on the bottom side according to a protrusion definition, each protrusion corresponding to each physical layer block.
[0178] The first example apparatus also includes any of the previous examples, comprising a substrate having a set of reference channels for providing conductive paths between the SoC and the memory devices on the memory chip, the set of reference channels being embedded in the substrate, the SoC and the memory chip being mounted on the substrate to form a standard package.
[0179] The first example apparatus also includes any of the previous examples, comprising a substrate having an embedded silicon bridge having a set of reference channels to provide a conductive path between memory devices on the SoC and the memory chip, the SoC and memory chip being mounted on the substrate to form an advanced package.
[0180] The first example apparatus also includes any of the previous examples, comprising a substrate, an interposer mounted on the substrate, the interposer having a set of reference channels to provide conductive paths between memory devices on the SoC and the memory chip, the SoC and the memory chip being mounted on the interposer to form an advanced package.
[0181] The first example apparatus also includes any of the previous examples, comprising a substrate, an interposer mounted on the substrate, the interposer having a silicon bridge having a set of reference channels to provide a conductive path between the SoC and the memory devices on the memory chip, the SoC and the memory chip being mounted on the interposer to form an advanced package.
[0182] The first example apparatus also includes any of the previous examples, including where the UCIe interface for the UCIe interconnect is defined by the Universal Chip Interconnect Express (UCIe) specification.
[0183] A second example apparatus includes a system-on-chip (SoC) structure including a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect including interface logic to manage data transmission on the UCIe interconnect in a transmit mode, the interface logic to: decode memory signals from the SoC structure; map the memory signals to UCIe signals; and encode the UCIe signals for transmission on an asymmetric link of the UCIe interconnect.
[0184] The second example apparatus also includes any of the previous examples, including interface logic to manage data transmission on a UCIe interconnect in a receive mode, the interface logic being configured to: decode UCIe signals from an asymmetric link of the UCIe interconnect; map the UCIe signals to memory signals; and encode the memory signals for transmission on a SoC fabric.
[0185] The second example apparatus also includes any of the previous examples, including where the memory signal is a compute express link (CXL) memory signal.
[0186] The second example apparatus also includes any of the previous examples, including wherein the memory signals are Compute Express Link (CXL) signals, and the interface logic is to: map CXL command and data timing signals for the CXL memory interface to UCIe command and data timing signals for a UCIe adapter and physical layer (PHY) interface; and map UCIe command and data timing signals for the UCIe adapter and PHY interface to CXL command and data timing signals for the CXL memory interface.
[0187] The second example apparatus also includes any of the previous examples, including interface logic for: encoding UCIe signals for transmission over an asymmetric link of the UCIe interconnect to a memory chip in the same package as the SoC structure; and decoding UCIe signals from the asymmetric link of the UCIe interconnect from the memory chip in the same package as the SoC structure.
[0188] The second example apparatus also includes any of the previous examples, including interface logic for: encoding a cyclic redundancy check (CRC) or error check and correction (ECC) signal of a UCIe signal for transmission on an asymmetric link of a UCIe interconnect; and decoding a CRC signal or an ECC signal of a UCIe signal from the asymmetric link of the UCIe interconnect.
[0189] The second example apparatus also includes any of the previous examples, including interface logic to: encode UCIe signals for transmission on multiple transmit channels of an asymmetric link of the UCIe interconnect; and decode UCIe signals from multiple receive channels of the asymmetric link of the UCIe interconnect.
[0190] The second example apparatus also includes any of the previous examples, comprising interface logic to: encode Compute Express Link (CXL) input / output (I / O) signals of a UCIe signal for transmission on an asymmetric link of a UCIe interconnect; and decode the CXL I / O signals of the UCIe signal from the asymmetric link of the UCIe interconnect.
[0191] A first example method includes decoding memory signals from a memory controller of a system on chip (SoC), the system on chip including the memory controller and a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect including interface logic to manage data transfer on the UCIe interconnect in a transmit mode; mapping the memory signals to UCIe signals; and encoding the UCIe signals for transmission on a transmit channel of the UCIe interconnect.
[0192] The first example method also includes any of the previous examples, comprising: decoding, by interface logic in a receive mode, a UCIe signal from a receive channel of a UCIe interconnect; mapping the UCIe signal to a memory signal; and encoding the memory signal for transmission to a memory controller.
[0193] The first example method also includes any of the previous examples, including where the memory signal is a double data rate (DDR) memory signal or a high bandwidth memory (HBM) memory signal.
[0194] The first example method also includes any of the previous examples, including where the memory signals include double data rate (DDR) physical layer (PHY) interface (DFI) signals, and the interface logic is to: map the DFI command and data timing signals to UCIe command and data timing signals for a UCIe PHY interface; and map the UCIe command and data timing signals for the UCIe PHY interface to the DFI command and data timing signals.
[0195] The first example method also includes any of the previous examples, including interface logic for: encoding UCIe signals for transmission on a transmit channel of the UCIe interconnect to a memory chip located in the same package as the SoC; and decoding UCIe signals from a receive channel of the UCIe interconnect from the memory chip located in the same package as the SoC.
[0196] The first example method also includes any of the previous examples, including: encoding a cyclic redundancy check (CRC) or error check and correction (ECC) signal of a UCIe signal for transmission on a command channel of a UCIe interconnect; and decoding the CRC signal or ECC signal of the UCIe signal from the command channel of the UCIe interconnect.
[0197] The first example method also includes any of the previous examples, comprising: encoding a UCIe signal for transmission on a plurality of transmit channels of a UCIe interconnect; and decoding the UCIe signal from a plurality of receive channels of the UCIe interconnect.
[0198] The first example method also includes any of the previous examples, including: encoding a UCIe signal for transmission on a transmit channel of a second UCIe interconnect to a second memory chip located in the same package as the second SoC; and decoding a UCIe signal from a receive channel of the second UCIe interconnect from the second memory chip located in the same package as the second SoC.
[0199] The first example method also includes any of the previous examples, including: encoding a UCIe signal for transmission on a transmit channel of a UCIe interconnect through buffer logic of a first memory chip located in the same package as the SoC, for a second memory chip located in the same package as the SoC and the first memory chip; and decoding a UCIe signal from the second memory chip from a receive channel of the UCIe interconnect through the buffer logic of the first memory chip.
[0200] A second example method includes decoding memory signals from a system-on-chip (SoC) fabric including a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect including interface logic to manage data transfer over the UCIe interconnect in a transmit mode; mapping the memory signals to UCIe signals; and encoding the UCIe signals for transmission over an asymmetric link of the UCIe interconnect.
[0201] The second example method also includes any of the previous examples, including interface logic to manage data transmission on a UCIe interconnect in a receive mode, the interface logic to: decode UCIe signals from an asymmetric link of the UCIe interconnect; map the UCIe signals to memory signals; and encode the memory signals for transmission on the SoC structure.
[0202] The second example method also includes any of the previous examples, including where the memory signal is a compute express link (CXL) memory signal.
[0203] The second example method also includes any of the previous examples, including where the memory signals are Compute Express Link (CXL) signals, and mapping CXL command and data timing signals for the CXL memory interface to UCIe command and data timing signals for a UCIe adapter and physical layer (PHY) interface; and mapping the UCIe command and data timing signals for the UCIe adapter and PHY interface to CXL command and data timing signals for the CXL memory interface.
[0204] The second example method also includes any of the previous examples, including: encoding a UCIe signal for transmission over an asymmetric link of the UCIe interconnect to a memory chip located in the same package as the SoC structure; and decoding a UCIe signal from the asymmetric link of the UCIe interconnect from the memory chip located in the same package as the SoC structure.
[0205] The second example method also includes any of the previous examples, including: encoding a cyclic redundancy check (CRC) or error check and correction (ECC) signal of a UCIe signal for transmission on an asymmetric link of a UCIe interconnect; and decoding the CRC signal or ECC signal of the UCIe signal from the asymmetric link of the UCIe interconnect.
[0206] The second example method also includes any of the previous examples, comprising: encoding a UCIe signal for transmission on a plurality of transmit channels of an asymmetric link of the UCIe interconnect; and decoding the UCIe signal from the plurality of receive channels of the asymmetric link of the UCIe interconnect.
[0207] A second example method also includes any of the previous examples, comprising: encoding a compute express link (CXL) input / output (I / O) signal of a UCIe signal for transmission on an asymmetric link of a UCIe interconnect; and decoding the CXL I / O signal of the UCIe signal from the asymmetric link of the UCIe interconnect.
[0208] It is emphasized that the Abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the present technical disclosure. The Abstract of the present disclosure is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Furthermore, in the foregoing Detailed Description, it can be seen that, in order to streamline the present disclosure, various features are grouped together in a single embodiment. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than expressly recited in each claim. On the contrary, as reflected in the following claims, the inventive subject matter lies in fewer than all the features of a single disclosed embodiment. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment. In the appended claims, the terms "including" and "in which" are used as the plain-English equivalents of the terms "comprising" and "wherein," respectively. Furthermore, the terms "first," "second," "third," etc. are used merely as labels and are not intended to impose numerical requirements on their objects.
[0209] The foregoing description of example embodiments has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the present disclosure. The scope of the disclosure is not limited by this detailed description, but rather by the appended claims. Future applications claiming priority to the present application may claim the disclosed subject matter in different ways and may generally include any combination of one or more limitations as variously disclosed or otherwise illustrated herein.
Claims
1. A device comprising: A system on a chip (SoC) comprising a memory controller and a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect comprising interface logic to manage data transfers on the UCIe interconnect in a transmit mode, the interface logic being configured to: decoding memory signals from the memory controller; Mapping the memory signal to a UCIe signal; as well as The UCIe signal is encoded for transmission on a transmission channel of the UCIe interconnect.
2. The apparatus of claim 1 , wherein the interface logic is configured to manage data transmission over the UCIe interconnect in a receive mode, the interface logic being configured to: decoding a UCIe signal from a receive channel of the UCIe interconnect; Mapping the UCIe signal to a memory signal; and The memory signal is encoded for transmission to the memory controller.
3. The device according to claim 1 or 2, wherein: The memory signal is a double data rate (DDR) memory signal or a high bandwidth memory (HBM) memory signal.
4. The device according to claim 1 or 2, wherein: The memory signals include double data rate (DDR) physical layer (PHY) interface (DFI) signals, and the interface logic is to: Mapping the DFI command and data timing signals to UCIe command and data timing signals for the UCIe PHY interface; and The UCIe command and data timing signals for the UCIe PHY interface are mapped to DFI command and data timing signals.
5. The apparatus according to claim 1 or 2, wherein the interface logic is configured to: Encoding the UCIe signal for transmission over a transmit channel of the UCIe interconnect to a memory chip in the same package as the SoC; and A UCIe signal from a receive channel of the UCIe interconnect is decoded from the memory chip in the same package as the SoC.
6. The apparatus according to claim 1 or 2, wherein the interface logic is configured to: encoding a cyclic redundancy check (CRC) or an error checking and correction (ECC) signal of the UCIe signal for transmission over a command channel of the UCIe interconnect; and A CRC signal or an ECC signal of a UCIe signal from a command channel of the UCIe interconnect is decoded.
7. The apparatus according to claim 1 or 2, wherein the interface logic is configured to: encoding the UCIe signal for transmission over a plurality of transmit channels of the UCIe interconnect; and Decoding UCIe signals from a plurality of receive channels of the UCIe interconnect.
8. The device according to claim 1 or 2, comprising: A second SoC including a memory controller and a second UCIe interconnect, wherein the second UCIe interconnect includes second interface logic to manage data transfer on the second UCIe interconnect in a transmit mode and a receive mode, the second interface logic being configured to: Encoding a UCIe signal for transmission over a transmit channel of the second UCIe interconnect to a second memory chip in the same package as the second SoC; as well as A UCIe signal from a receive channel of the second UCIe interconnect is decoded from the second memory chip in the same package as the second SoC.
9. The apparatus according to claim 1 or 2, wherein the interface logic is configured to: encoding a UCIe signal for transmission over a transmit channel of the UCIe interconnect through buffer logic of a first memory chip in the same package as the SoC, the UCIe signal being intended for a second memory chip in the same package as the SoC and the first memory chip; and A UCIe signal from a receive channel of the UCIe interconnect from the second memory chip is decoded by the buffer logic of the first memory chip.
10. A device comprising: A system-on-chip (SoC) structure including a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect including interface logic to manage data transfer on the UCIe interconnect in a transmit mode, the interface logic configured to: decoding memory signals from the SoC structure; Mapping the memory signal to a UCIe signal; as well as The UCIe signal is encoded for transmission over an asymmetric link of the UCIe interconnect.
11. The apparatus of claim 10, wherein the interface logic is configured to manage data transmission over the UCIe interconnect in a receive mode, the interface logic being configured to: decoding a UCIe signal from an asymmetric link of the UCIe interconnect; Mapping the UCIe signal to a memory signal; and The memory signal is encoded for transmission over the SoC structure.
12. The device according to claim 10 or 11, wherein The memory signal is a compute express link (CXL) memory signal.
13. The device according to claim 10 or 11, wherein The memory signal is a compute express link (CXL) signal, and the interface logic is to: Mapping CXL command and data timing signals for the CXL memory interface to UCIe command and data timing signals for the UCIe adapter and physical layer (PHY) interface; as well as UCIe command and data timing signals for the UCIe adapter and PHY interface are mapped to CXL command and data timing signals for the CXL memory interface.
14. The apparatus according to claim 10 or 11, wherein the interface logic is configured to: encoding the UCIe signal for transmission over an asymmetric link of the UCIe interconnect to a memory chip in the same package as the SoC structure; and A UCIe signal from an asymmetric link of the UCIe interconnect is decoded from the memory chip in the same package as the SoC structure.
15. The apparatus according to claim 10 or 11, wherein the interface logic is configured to: encoding a cyclic redundancy check (CRC) or error checking and correction (ECC) signal of the UCIe signal for transmission over an asymmetric link of the UCIe interconnect; and A CRC signal or an ECC signal of a UCIe signal from an asymmetric link of the UCIe interconnect is decoded.
16. A method comprising: decoding memory signals from a memory controller of a system on a chip (SoC), the SoC including the memory controller and a Universal Chip Interconnect Express (UCIe) interconnect, the UCIe interconnect including interface logic to manage data transfer on the UCIe interconnect in a transmit mode; Mapping the memory signal to a UCIe signal; as well as The UCIe signal is encoded for transmission on a transmission channel of the UCIe interconnect.
17. The method according to claim 16, comprising: decoding, by the interface logic, a UCIe signal from a receive channel of the UCIe interconnect in a receive mode; Mapping the UCIe signal to a memory signal; as well as The memory signal is encoded for transmission to the memory controller.
18. The method according to claim 16 or 17, wherein The memory signal is a double data rate (DDR) memory signal or a high bandwidth memory (HBM) memory signal.
19. The method according to claim 16 or 17, wherein: The memory signals include double data rate (DDR) physical layer (PHY) interface (DFI) signals, and the interface logic is to: Mapping the DFI command and data timing signals to UCIe command and data timing signals for the UCIe PHY interface; and The UCIe command and data timing signals for the UCIe PHY interface are mapped to DFI command and data timing signals.
20. The method according to claim 16 or 17, comprising: Encoding the UCIe signal for transmission over a transmit channel of the UCIe interconnect to a memory chip in the same package as the SoC; as well as A UCIe signal from a receive channel of the UCIe interconnect is decoded from the memory chip in the same package as the SoC.
Citation Information
Patent Citations
wire knife for self-spinners
CH11002A
soot and spark arresting device on house chimneys
CH21004A