Firmware broadcast within multi-chip module
By adopting broadcast write and read operations in multi-chip modules, the problem of excessive firmware loading time is solved, efficient firmware loading is achieved, and the time requirements of the PCIe standard are met.
Patent Information
- Application Number
- CN202380085077.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-27
- Filing Date
- 2023-10-27
- Publication Date
- 2025-07-22
AI Technical Summary
In multi-chip modules, the firmware loading process occupies a large amount of startup time, especially when firmware needs to be loaded for multiple devices, resulting in startup time exceeding the limits of the PCIe standard.
By adopting broadcast writing and broadcast reading operations, the processor uses the broadcast address space of the multi-chip module to realize simultaneous writing and reading of multiple physical layer data channel circuits, shortening the firmware loading time.
It effectively reduces firmware loading time, improves the startup efficiency of multi-chip modules, and meets the time limit of the PCIe standard.
Smart Images

Figure CN120359506A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Application No. 63 / 381,264, filed on October 27, 2022, entitled "Securely Bootable PCIe Retimer", which is hereby incorporated by reference in its entirety for all purposes.
[0003] References
[0004] The following references are hereby incorporated by reference in their entirety for all purposes:
[0005] "PCIe Base Specification" (Version 1.0, Revision 6.0.1, September 13, 2022, available at: www.pcisig.com / specifications);
[0006] "PCIe Retimer Test Specification" (Version 1.0, Revision 4.0, June 10, 2022, available at: www.pcisig.com / specifications);
[0007] U.S. Patent No. 9,288,082, issued on March 15, 2016, entitled "Circuit for Vector Signaling Code for Differential and Efficient Detection of Inter-Chip Communication" with application number 13 / 895,206, filed on May 15, 2013, and inventors Roger Ulrich and Peter Hunt (hereinafter referred to as [Ulrich]). Background Art
[0008] Signals often attenuate when propagating within a wire, that is, the signal-to-noise ratio decreases. Such signal attenuation is typically measured in decibels (dB) and often increases with the length of the signal transmission wire.
[0009] Many standards for electronic devices specify the maximum loss of signal transmission between upstream and downstream components. For example, the Peripheral Component Interconnect Express (PCIe) 5.0 standard specifies that at 16 GHz, the transmission loss budget from an upstream component (usually a root complex or a switch) to a downstream component (usually an endpoint device or a switch) is -36 dB. Failure to meet this loss budget will result in non-compliance with the standard, which is undesirable. However, in practical situations, especially when the wire is long and the data rate is high, it may be difficult to meet the loss budget.
[0010] To solve this problem, a retimer can be adopted. The retimer is disposed in a signal path between an upstream component and a downstream component, and divides the link between the upstream component and the downstream component into two completely independent links. The retimer is used to adjust the signals received via an upstream pseudo-port, and then send the adjusted signals via a downstream pseudo-port. The retimer generally equalizes the input signals and restores the clock configuration of the input signals, so that the signals output by the retimer have high amplitude, low noise and low jitter. Therefore, the retimer can greatly reduce the total loss between the upstream and downstream components, enabling a link that originally did not meet the specifications to meet the specification requirements.
[0011] A retimer is generally an input / output device and can take the form of a multi-chip module (MCM) including multiple tiles communicatively connected in a certain way. Each tile may include multiple components, such as physical data channel circuits (PHY). Such components may need to be supported by certain configuration data (such as firmware) to operate properly, and each of the multiple components may have the same configuration data. When the components with configuration data are arranged serially, it may take too much time and cause corresponding performance problems for the MCM. SUMMARY OF THE INVENTION
[0012] This document describes a multi-chip module (MCM) that includes several devices communicatively connected by one or more buses and controlled by a processor such as a CPU or a microcontroller. Some of the devices require firmware to operate properly. The firmware is generally loaded during an initialization process called the "boot process" (or simply "boot"). The systems and methods described herein enable the firmware to be updated at any time. Various standards such as the Peripheral Component Interconnect Express (PCIe) standard generally have restrictions on the boot time. For example, the PCIe standard requires the MCM to start PCIe link training after a certain period of time (such as 100 ms or 120 ms) after boot. However, the firmware loading during the boot process may occupy a significant portion of the boot time, especially when there are multiple devices that need to load firmware. The techniques described herein can load the firmware into multiple MCM devices with high time efficiency. The MCM described herein takes a retimer as an example, but the present disclosure is not limited to retimers, and any functional MCM can adopt the techniques disclosed in this specification.
[0013] The disclosed technology is used to perform broadcast write operations within a single die package structure and a multi-die package structure including multiple dies. The package structure can provide any function, such as a retimer function. Among them, a broadcast address is simultaneously assigned to multiple devices of the MCM, and using this broadcast address, data packets are simultaneously sent to all devices. In the multi-die case, the first broadcast write operation is used to simultaneously write to the devices within the master die, and the second broadcast write operation is used to simultaneously write to the devices within one or more slave dies. These broadcast write operations are all performed using a broadcast address or address range shared by the master die device and the slave die device. In addition, a broadcast read operation performed within the MCM is also disclosed. Among them, by performing a bitwise "OR" operation on the read results, a total value can be generated for determining whether any device in any die issues an interrupt request.
[0014] According to one embodiment, a method includes: obtaining, by a processor of an MCM, a physical layer data channel circuit (PHY) configuration data packet, the MCM having multiple dies and multiple PHYs distributed on the multiple dies; and using, by the processor, the broadcast address space of the MCM, via a bus connected to the processor and the multiple PHYs, to simultaneously write the configuration data packet to the multiple PHYs, the broadcast address space of the MCM including at least one address assigned to all of the multiple PHYs, the MCM further having a non-broadcast address space that includes multiple unique addresses, each of the unique addresses being assigned to a corresponding one of the multiple PHYs.
[0015] According to one embodiment, an apparatus includes a multi-chip module (MCM) having multiple dies, multiple physical layer data channel circuits (PHYs) distributed on the multiple dies, a processor, and a bus connected to the processor and the multiple PHYs, the processor being configured to: obtain a PHY configuration data packet; and use the broadcast address space of the MCM, via the bus, to simultaneously write the configuration data packet to the multiple PHYs, the broadcast address space of the MCM including at least one address assigned to all of the multiple PHYs, the MCM further having a non-broadcast address space that includes multiple unique addresses, each of the unique addresses being assigned to a corresponding one of the multiple PHYs.
[0016] According to another embodiment, a method includes: a processor of a multi-chip module (MCM) simultaneously sending a broadcast read instruction to a master die register located within a master die and one or more slave die registers respectively located within one or more slave dies, the MCM having the master die and the one or more slave dies, the master die having a plurality of master physical data channel circuits (PHYs), and each slave die having a plurality of slave physical data channel circuits (slave PHYs); in response to the broadcast read instruction, receiving a master die read result from the master die register; in response to the broadcast read instruction, respectively receiving corresponding one or more slave die read results from the one or more slave die registers; and performing a bitwise "OR" operation on the master die read result and the one or more slave die read results to generate a total read value.
[0017] According to another embodiment, a device includes: a multi-chip module (MCM) having a plurality of master physical layer data channel circuits (PHYs) located within a master die of the MCM and a plurality of slave physical data channel circuits (slave PHYs) respectively located within each of one or more slave dies of the MCM; a master die register located within the master die; one or more slave die registers respectively located within the one or more slave dies; and a processor and a bus connected to the processor and the plurality of PHYs, wherein the processor is configured to: simultaneously send a broadcast read instruction to the master die register and the one or more slave die registers; in response to the broadcast read instruction, receive a master die read result from the master die register; in response to the broadcast read instruction, respectively receive corresponding one or more slave die read results from the one or more slave die registers; and perform a bitwise "OR" operation on the master die read result and the one or more slave die read results to generate a total read value. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A block diagram of a device including a multi-chip module suitable for implementing the embodiments described herein.
[0019] Figure 2 A block diagram of a retimer suitable for implementing the embodiments described herein.
[0020] Figure 3 A block diagram of a single-die retimer suitable for implementing the embodiments described herein.
[0021] Figure 4 A schematic diagram of the storage content of an external memory of a retimer, which can store data packets for writing into retimer components.
[0022] Figure 5A block diagram of a dual-die weight timer suitable for implementing the embodiments described herein.
[0023] Figure 6 is Figure 5 A block diagram of the slave die of the dual-die weight timer in
[0024] Figure 7 A block diagram of a four-die weight timer suitable for implementing the embodiments described herein.
[0025] Figure 8 A flowchart of performing a broadcast write operation within a single-die weight timer in one embodiment.
[0026] Figure 9 is Figure 7 Another block diagram of the four-die weight timer in
[0027] Figure 10 A flowchart of performing a broadcast write operation within a multi-die weight timer in one embodiment.
[0028] Figure 11 A flowchart of performing a broadcast read operation within a multi-die weight timer in one embodiment.
[0029] Figure 12 A block diagram of a multi-chip module capable of performing multicast operations in one embodiment.
[0030] Figure 13 is suitable for use in one embodiment for Figure 12 A schematic diagram of the address space of the embodiment.
[0031] Figure 14 is in one embodiment within Figure 12 A flowchart of performing a write operation within a multi-chip module such as a multi-chip module that supports broadcast and multicast operations. Detailed embodiments
[0032] The PCIe standard will be referred to from time to time in this specification to assist in understanding the present disclosure by describing certain features in the context of a specific standard. However, it should be understood that the content herein is equally applicable to situations outside the PCIe standard unless otherwise explicitly stated.
[0033] Figure 1Schematic diagram of a multi-chip module (MCM) 100 adapted to implement the embodiments described herein. The MCM 100 includes a main die 105 having a processor 110 disposed therein. In addition, slave dies 115a, 115b, 115c are provided. In the illustrated embodiment, there are a total of three slave dies, but it should be understood that this is for illustrative purposes only, and any number of slave dies, such as one, two, three, four, etc., can be provided. The main die 105 and the slave dies 115a to 115c are all part of the same package structure.
[0034] The difference between the main die 105 and the slave dies 115a to 115c is at least that the main die 105 has a processor 110 in an operating state. That is, the processor 110 performs computing tasks during at least part of the time when the MCM 100 is powered on. Although Figure 1 not shown in the figure, for the purpose of facilitating circuit design and manufacturing, each of the slave dies 115a to 115c may also have a processor disposed therein. However, when the slave die has a processor, these processors do not work (e.g., in a powered-off state) during the use of the MCM 100.
[0035] The main die 105 is communicatively connected to each of the slave dies 115a to 115c through respective connections 120a, 120b, 120c. The connections can be wires of a bus (such as a Serial Peripheral Interface (SPI) bus). More details in this regard will be given in the following description of this specification.
[0036] One or more physical layer entities, such as pseudo ports, Serializers / Deserializer (SerDes), etc., are provided on each die. These entities are referred to as physical layer data channel circuits (PHY) herein. In the illustrated embodiment, each die has four PHYs 125, but this number is not fixed, and any number of PHYs can be provided in addition (for the sake of clarity, Figure 1 only one PHY is marked in the figure, but it should be understood that Figure 1 other elements drawn in the same way in the figure are also PHYs). In addition, all dies may have the same number of PHYs, or the number of PHYs of some dies may be different from the number of PHYs of other dies. The number of PHYs of the main die 105 may be different from the number of PHYs of each of the slave dies 115a to 115c. In addition, the main die 105 may not have a PHY and mainly serves as the control die of the MCM. There are also other solutions in this regard.
[0037] In the illustrated embodiment, each PHY is capable of two-way communication with an entity external to the MCM 100 (such as a root complex, an endpoint device, or other devices). The present invention does not strictly require two-way communication, and the PHY can also support one-way communication.
[0038] In the following, taking the MCM 100 as a retimer as an example, each embodiment will be further described. However, it should be understood that unless the relevant description clearly relates to the retiming function, all "retimers" in this specification can be replaced with "MCM".
[0039] Figure 2 FIG. 5 is a schematic diagram of a system 200 including a retimer 210. The retimer 210 is a multi-channel PCIe retimer, that is, it is used to process (i.e., retime) PCIe traffic of multiple channels simultaneously. The retimer 210 is connected to an upstream component 205, and the upstream component 205 is usually a root complex or a switch. This connection is through the upstream pseudo-port 220a of the retimer 210. Similarly, the retimer 210 is connected to a downstream component 215 through a downstream pseudo-port 220b, and the downstream component 215 is usually a switch or an endpoint device. The upstream pseudo-port 220a and the downstream pseudo-port 220b are, for example, PHYs.
[0040] Therefore, from Figure 2 it can be seen that the retimer 210 is used to divide the link between the upstream component 205 and the downstream component 215 into two parts. The retimer 210 is used to adjust the signals received through the upstream pseudo-port 220a and provide a highly clean signal with low jitter and good signal-to-noise ratio for output through the downstream pseudo-port 220b. The retimer 210 is a bidirectional device, that is, it can also adjust the received signals input through the downstream pseudo-port 220b. In this case, the highly clean output signal is sent out through the upstream pseudo-port 220a.
[0041] Figure 3 FIG. 6 is a further detailed schematic diagram of the retimer 210. For ease of understanding, some components of the retimer 210 are not shown.
[0042] The retimer 210 includes a CPU core 300, which is also referred to as a processor in this article. This processor is equivalent to Figure 1 the processor 110 in FIG. 1. The CPU core 300 is used to execute various tasks to support the functions of the retimer 210. One of these tasks is to load firmware from an external non-volatile memory to start the ROM 305 during the startup process and load the firmware into the PHY of the retimer 210. Further details of this startup process are given below. The CPU core 300 runs according to the instructions stored in the instruction RAM 310 and performs operations on the data stored in the data RAM 315. The CPU core 300 is also connected to an interrupt request (IRQ) controller 320 so that the CPU core 300 can receive interrupt requests from other components of the retimer 210 and / or from external components.
[0043] The CPU core 300 is also connected to an Advanced Peripheral Bus (APB) interconnect 325. This APB interconnect enables the CPU core 300 to communicate with other components in the retimer 210 that are also connected to this bus. In this regard, see Figure 3 . It should be understood that the APB interconnect 325 can be replaced by other buses, such as the AHB or other AMBA buses, which are also within the scope of the present invention.
[0044] The APB interconnect 325 also enables other components of the retimer 210 to directly communicate with the instruction RAM 310 in a controlled manner (see the "access restrictions" in Figure 3 ). In this way, it is ensured that only the components that should have access to the instruction RAM 310 can access the instruction RAM 310, and that the instructions stored in the instruction RAM 310 by any such component are legitimate.
[0045] The retimer 210 also includes a non-volatile read-only memory, which can be the Figure 3 one-time programmable (OTP) memory 330 shown. In addition, other forms of non-volatile ROM can also be used. The OTP memory 330 stores public keys or public key hash values that can be used by the CPU core 300 to verify the authenticity of the firmware when the CPU core 300 loads the firmware.
[0046] The firmware is loaded from an external non-volatile memory. Here, "external" means that the memory is off-chip, that is, it is not on the same die 335 as the CPU core 300. The external non-volatile memory can be either part of the MCM package structure or outside the MCM package structure. In the illustrated embodiment, the external non-volatile memory is an SPI flash 340. The CPU core 300 communicates with the SPI flash 340 via the SPI bus, where the APB interconnect 325 is connected to the corresponding master SPI 345 to implement a complete communication channel between the CPU core 300 and the SPI flash 340. This structure is only an example and not the only feasible structure. For example, the external non-volatile memory can also use an EEPROM. In this case, the CPU core 300 can communicate with the EEPROM via the I 2 C bus connected to the APB interconnect 325 (see the master I Figure 3 C 350 in 2 ). In addition, there are other solutions, and it should be understood that any other solution that enables the CPU core 300 to communicate with the external non-volatile memory is within the scope of the present disclosure.
[0047] It should be noted that when applied to the retimer, the PCIe standard requires the use of I2 I²C bus. However, it has been recognized that, due to the relatively slow speed of the I²C interface, problems may occur when loading firmware from an external memory. Specifically, the combination of the I²C bus and the EEPROM may have difficulty meeting certain timing requirements of the PCIe specification. In contrast, the data transfer rate of the SPI interface is higher than that of the I²C interface. Therefore, the combination of the SPI bus and the SPI flash memory 340 can significantly shorten the firmware loading time. In view of this, in some contemplated embodiments of the present invention, the I²C bus can be completely deprecated. 2 I²C bus. However, it has been recognized that, due to the relatively slow speed of the I²C interface, problems may occur when loading firmware from an external memory. Specifically, the combination of the I²C bus and the EEPROM may have difficulty meeting certain timing requirements of the PCIe specification. In contrast, the data transfer rate of the SPI interface is higher than that of the I²C interface. Therefore, the combination of the SPI bus and the SPI flash memory 340 can significantly shorten the firmware loading time. In view of this, in some contemplated embodiments of the present invention, the I²C bus can be completely deprecated. 2 I²C bus. However, it has been recognized that, due to the relatively slow speed of the I²C interface, problems may occur when loading firmware from an external memory. Specifically, the combination of the I²C bus and the EEPROM may have difficulty meeting certain timing requirements of the PCIe specification. In contrast, the data transfer rate of the SPI interface is higher than that of the I²C interface. Therefore, the combination of the SPI bus and the SPI flash memory 340 can significantly shorten the firmware loading time. In view of this, in some contemplated embodiments of the present invention, the I²C bus can be completely deprecated. 2 I²C bus. However, it has been recognized that, due to the relatively slow speed of the I²C interface, problems may occur when loading firmware from an external memory. Specifically, the combination of the I²C bus and the EEPROM may have difficulty meeting certain timing requirements of the PCIe specification. In contrast, the data transfer rate of the SPI interface is higher than that of the I²C interface. Therefore, the combination of the SPI bus and the SPI flash memory 340 can significantly shorten the firmware loading time. In view of this, in some contemplated embodiments of the present invention, the I²C bus can be completely deprecated. 2 I²C bus.
[0048] The retimer 110 further includes a timer 355, general-purpose input / output pins (GPIO) 360, and a system management bus (SMBus) 365. All of these components are connected to the APB interconnect 325 to facilitate communication with other components of the retimer 210.
[0049] The timer 355 provides a programmable timing function to enable, for example, entering a low-power state during the intervals when performing periodic tasks. The GPIO 360 provides one or more general-purpose pins that are not used by default but can be used in a certain way under software control to, for example, extend the function of the retimer 210 in a certain way. The SMBus 365 provides the ability to send information (such as status, configuration, device name, type, etc.) related to the devices connected to the retimer 310 and to send commands to such devices. One or more of the timer 355, the GPIO 360, and the SMBus 365 may be omitted, or may be replaced by other components with similar functions, which also fall within the scope of the present disclosure.
[0050] The retimer 210 further includes a plurality of physical layer data channel circuits (PCIe PHY) 370, such as four or eight PHYs. Such PHYs are physical layer components, such as serial deserializers (SerDes). The PHYs 370 are connected to the APB interconnect 325 to provide a communication path to the CPU core 300 and any other components of the retimer 210 that are also connected to the APB interconnect 325. During the startup process, the PHYs 370 need to be initialized by the CPU core 300 in a manner that provides a PHY configuration data packet to the PHYs. The configuration data packet can be, for example, a PCIe PHY configuration data packet. The PHY configuration data packet can be loaded, for example, by the CPU core 300 from the SPI flash memory 340. Further information related to this process is given below.
[0051] The retimer 210 further includes a PCIe switch 375 connected to the APB interconnect 325. The PCIe switch 375 is used to implement the PCIe switching function specified by the relevant part of the PCIe standard, so that the retimer 210 can operate in the PCIe switching mode as required. It should be understood that when the retimer 210 does not need to provide the PCIe switching function, the PCIe switch 375 may not be provided.
[0052] Figure 3 Including "Peripheral Component N" 380 shown in a placeholder box, which is connected to the APB interconnect 325 to indicate that the retimer 210 is not limited to Figure 3 The series of specific peripheral components shown, but other peripheral components connected to the APB interconnect 325 can be further added as needed. Such other peripheral components include, for example, one or more PCIe Compute Express Link (CXL), Physical Coding Sublayer (PCS) components, packet inspection components, Joint Test Action Group (JTAG) interfaces, and / or the high-speed die-to-die interface described in
Ulrich
[0053] Figure 4 The above is a set of contents that can be saved in the SPI flash 340. In addition, there are other solutions. It should be understood that Figure 4 It is intended to assist in understanding the present disclosure and does not limit its scope.
[0054] The SPI flash 340 is divided into two regions (or partitions), an active region and a non-active region. Each region corresponds to a set of addresses of the SPI flash 340. This set of addresses does not necessarily have to be consecutive addresses. On the contrary, as Figure 4 shown, this set of addresses can be inserted at intervals. The active region is a set of memory addresses storing information for use by the CPU core 300 at the next startup, and the non-active region is a set of memory addresses storing information not used by the CPU core 300 at the next startup. The purpose of partitioning is to enable the updated firmware to be stored in the non-active region without interrupting the operation of the active region. That is, when the updated firmware image is not available (such as damaged or invalid), the retimer can still start according to the current firmware image stored in the active region.
[0055] The active state and the non-active state are set by one or more flags saved in the packet header 400. The header 400 can also save any other information considered useful, such as the size of each memory region in bits, the starting address of each region, the last update date of the SPI flash, version information, etc.
[0056] The activation area includes an activation firmware image 405, which is the firmware image to be used by the CPU core 300 when the retimer 210 is started next time. The activation firmware image 405 includes a configuration file 410, a PHY configuration data packet 415, and an application program 420. It should be understood that this is only an example, and the activation firmware image 405 may also include information different from that Figure 4 shown, or may further include Figure 4 information other than that shown.
[0057] The configuration file 410 stores information for the CPU core 300 to configure the retimer 210 during the startup process. For example, the configuration file 410 may include one or more values to be written into one or more registers of the retimer 210 respectively during the startup process. Protocol-specific information such as one or more PCIe vendor-defined message codes may be stored in the configuration file 410.
[0058] The PHY configuration data packet 415 is used to configure the PHY 370. The PHY configuration data packet 415 can be PHY firmware (i.e., a smaller firmware image contained in the activation firmware image 405) and / or PHY configuration data (such as the initial values of one or more registers of the PHY). The PHY configuration data packet 415 is used for the initialization and / or configuration of the PHY 370. For example, the CPU core 300 provides the PHY configuration data packet 415 to each PHY 370 during the startup process. The PHY configuration data packet 415 provides a secure and convenient channel for the configuration of the PHY 370 and the update of its firmware. After the PHY firmware is updated, it becomes a new firmware image and can be loaded into the SPI flash memory 340. Further information on the process of sending the PHY configuration data packet 415 to each PHY 370 will be given later in this specification.
[0059] The application program 420 is an executable file, and the CPU core 300 achieves correct startup by running this file. During the startup process, the CPU core 300 loads the application program 420. After loading, if all relevant security checks are passed successfully, the application program 420 is executed.
[0060] The activation firmware image 405 may further include a secondary boot loader (not shown). The secondary boot loader is an application program used to handle the loading of certain items such as a real-time operating system (RTOS) to assist the application program 420. If there is no need, the secondary boot loader can be omitted.
[0061] The inactive firmware image 425 is a copy of the active firmware image 405 and thus also includes the above configuration file, PHY configuration data packet, and application. As described above, the inactive firmware image 425 may differ from the active firmware image 405 in terms of firmware version, etc. For example, the PHY configuration data packets, configuration files, and / or applications in the inactive firmware image 425 and the active firmware image 405 may be different versions respectively.
[0062] The above description is limited to a single-die structure in which the components of the retimer 210 are provided within the same die 335 (except for the SPI flash 340 provided outside the die). Figure 5 and Figure 6 Shown is a multi-die structure introducing a second die. The devices of the second die are located on a second die 500 different from the above. As Figure 6 shown, the devices of the second die are substantially the same as those of the first die, so the second half of their reference numerals is the same as the second half of the Figure 3 devices. For relevant content, please refer to the above description.
[0063] In this article, the first die is referred to as the main die (or main control device), and the second die is referred to as the slave die (or slave device). In various embodiments, the difference between the main die and the slave die is that some devices of the slave die are in an inactive state, that is, completely turned off or in a low-power state. In one embodiment, in the slave die, at least the CPU core 600 is in an inactive state. In this way, the slave die is controlled and configured by the CPU core 300 of the main die. In another embodiment, in the slave die, at least the following devices are in an inactive state: CPU core 600, boot ROM 605, instruction RAM 610, data RAM 615, IRQ controller 620, OTP memory 630, main SPI 645, main I 2 C650, timer 655, GPIO 660, SMBus 665, and main T2T SPI 675. The reason for setting these devices on the slave die is that from a manufacturing perspective, it is easier to first manufacture identical dies and then designate one as the main device and the other as the slave device. However, as an alternative, the above devices may not be provided on the slave die. Similarly, the main die includes both a main T2T SPI 385 and a slave T2T SPI 390, and only the main T2T SPI 385 is in an active state. As described above, as an alternative, a manufacturing method may also be adopted in which only the main die has the main T2T and only the slave die has the slave T2T. In addition, during the chip testing process, when some die defects affecting the functions / circuits of the main die are detected, the die may still be regarded as a qualified slave die, thereby improving the production yield.
[0064] It should also be noted that since only the CPU core 300 of the main die is in the active state, there is no need to load any firmware into the CPU core 600 of the slave die in the non-active state, so there is no need to connect the slave die to any SPI flash memory (or other external memory).
[0065] The main die and the slave die communicate through a bus spanning between the two dies 335 and 500 (see Figure 5 ). In Figure 5 and Figure 6 shown cases, this bus is a die-to-die ("T2T") SPI bus. As an alternative, other types of buses can be used to replace the SPI bus according to requirements.
[0066] More specifically, the main die includes a main T2T SPI bus 670, and a corresponding slave T2T SPI bus 675 is provided on the slave die. These two devices are connected by wires extending between the main die and the slave die. Such wires can be, for example, circuit traces. In this article, the main T2T SPI 670 and the slave T2T SPI 675 are collectively referred to as the "T2T SPI bus". The main T2T SPI 670 is connected to the APB interconnect 625 to enable communication with other devices in the main die (such as the CPU core 300). Similarly, the slave T2T SPI 675 is connected to the APB interconnect 625 in the slave die to enable communication with other devices in the slave die (such as the PHY 685, the PCIe switch 675, and other peripheral components 680).
[0067] Following the above principle of manufacturing identical dies, in Figure 5 and Figure 6 , the slave die is shown as having both the main T2T SPI 670 and the slave T2T SPI 675. However, it should be understood that in Figure 6 the slave die, only the slave T2T SPI 675 is in the active state. Similarly, although the main die includes both the main T2T SPI 385 and the slave T2T SPI 390, only the main T2T SPI 385 is in the active state. As described above, as an alternative, a manufacturing method can also be adopted in which only the main die has the main T2T and only the slave die has the slave T2T.
[0068] The slave die has its own set of slave PHYs 685, PCIe switches 675, and other peripheral components 680. These devices are the same as the corresponding devices shown in Figure 3 , and for related content, please refer to the above description. The slave PHYs 685, PCIe switches 675, and other peripheral components 680 can be controlled by the CPU core 300 of the main die through the T2T SPI bus and the APB interconnect 625 of the slave die.
[0069] Two die can span more than one bus to form multiple communication channels between these two die. As an additional or alternative solution, an interface based on high-speed die-to-die SerDes as described in
Ulrich
Ulrich
[0070] The above dual-die structure can be extended to more die. Figure 7 Shown is a four-die structure. In this structure, there is a total of one master die and three slave die (die 1, die 2, and die 3). Each of the four die is disposed on its own die - the master die is disposed on die 335, the slave die 1 is disposed on die 500, the slave die 2 is disposed on die 700, and the slave die 3 is disposed on die 700'. Each slave die is the same as the Figure 5 and Figure 6 shown slave die and is as described above. The master die is the same as described above. The master T2T SPI 385 of the master die is connected to the corresponding slave T2T SPI of each slave die, that is, connected to the slave T2T SPI 675, the slave device 775, and the slave device 775'. In this way, the CPU core 300 can control any device of any slave die. Although Figure 7 not shown for clarity, the master die and each slave die have their own above-mentioned type of PHY, PCIe switch, and / or other peripheral components, and the CPU core 300 can control all of these devices.
[0071] Generally speaking, by connecting one master die and N - 1 slave die through an inter-die bus such as the above T2T SPI bus, it can be extended to N die. As needed, other types of buses (such as the I 2 C bus or the Universal Chiplet Interconnect (UCIe) bus) can be used instead of the SPI bus.
[0072] During the startup process of the retimer 210, the PHY 370 needs to be initialized. This initialization can include: loading configuration data, which is used by the PHY 370, for example, to set the initial values of registers to, for example, implement the selection of the operation mode. As an additional or alternative solution, this initialization can include: loading firmware for the processor of each PHY 370, for example, into the SRAM of each PHY. This operation can be achieved by sending a PHY configuration data packet 415 to each PHY 370 during the startup process.
[0073] The PHY 370 may include its own memory (such as SRAM), and the content of the PHY configuration data packet 415 sent to the PHY may be loaded into this memory. When performing data packet transmission, the CPU core 300 does not need to understand the working principle of the PHY, nor the content of the data packet, and only needs to obtain the data packet and the address of each destination PHY.
[0074] Since the number of PHYs can be multiple even in the case of a single die, it is necessary to send a PHY configuration data packet to each PHY. If this sending process is executed sequentially, it will be relatively time-consuming. Therefore, in order to avoid such unnecessary delays during the startup process, each of the multiple PHYs 370 is simultaneously assigned a broadcast address or an address range. In this way, the PHY configuration data packet can be sent to each of the multiple PHYs 370 simultaneously, thereby shortening the total time required to send the PHY configuration data packet to all PHYs. For further information on this "broadcast write" process, see below.
[0075] Figure 8 Shown is the process of sending a PHY configuration data packet (such as data packet 415) to the PHY 370 within a single-die system. This process can be part of the startup process. As an additional or alternative solution, this process can also be used in any other situation where it is necessary to provide the PHY configuration data packet to all PHYs 370 simultaneously.
[0076] In step 800, the processor of the multi-channel PCIe retimer 210 obtains the PHY configuration data packet. The processor can be, for example, the CPU core 300 or the processor 110. The PHY configuration data packet can be part of a firmware image (such as firmware image 405). The processor's acquisition of the configuration data packet can be, for example, part of a firmware update operation. The PHY configuration data packet can be obtained from a source outside the retimer 210 (such as the SPI flash 340), and the SPI flash 340 can receive data packets from another source via a network such as the Internet or a cellular network. Before loading, the data packet can be stored in the SPI flash 340 (or EEPROM) in the form of part of a firmware image 405. The loading process may involve one or more security checks to ensure the validity and authenticity of the firmware image before loading. In the case of not performing a firmware update (for example, not performing a firmware update during the startup process), the data packet can be obtained from the SPI flash 340 or some other memory.
[0077] When the PHY configuration data packet is part of a firmware image, the processor identifies the PHY configuration data packet from the firmware image. The PHY configuration data packet can be identified according to the information contained in the packet header (such as packet header 400). As needed, the data packet can also contain other information such as a version number and the target PHY suitable for use as the object of the PHY configuration data packet.
[0078] In step 805, the processor uses the broadcast address space of the multi-channel PCIe retimer to write the configuration data packet to the multiple PHYs simultaneously via the bus connected to the processor and the multiple PHYs. In Figure 3 the case of the structure shown, the bus is the APB interconnect 325. The structure shown in this figure is only one possible structure, which is intended to help understand the present invention and should not be regarded as a limitation. Moreover, there are variant schemes for this shown structure. The broadcast address space of the multi-channel PCIe retimer includes at least one broadcast address assigned to all the PHYs among the multiple PHYs. That is to say, since the at least one broadcast address is assigned to all the PHYs 370 among the multiple PHYs, all the PHYs 370 can obtain any data associated with the broadcast address space on the bus. Thus, the PHY configuration data packet can be sent to all the PHYs 370 simultaneously, thereby reducing the total time required to provide the PHY configuration data packet to the PHYs 370 compared with a series of write operations of writing to each PHY 370 separately in sequence.
[0079] In contrast, the multi-channel PCIe retimer also has a non-broadcast address space, which includes a plurality of unique addresses, and each address is respectively assigned to a corresponding one of the multiple PHYs. Thus, for a specific PHY among the multiple PHYs, it can be addressed by using the corresponding unique address of the specific PHY. The non-broadcast address space and the broadcast address space can be part of the same memory global mapping table that provides addresses to other devices of the main die of the multi-channel PCIe retimer (in a multi-die implementation, also provides addresses to devices of the slave die).
[0080] The above at least one broadcast address may include a 24-bit address assigned to all the PHYs 370. It should be understood that the present disclosure is not limited to a 24-bit address, and as an alternative, addresses with more or fewer bits than 24 can also be used. The broadcast address can be offset relative to the base address within the address space, and the address space can be a 32-bit address space. The base address can be equal to the product of a constant value and an identifier associated with the die (such as the die number).
[0081] In addition, in a multi-die structure, sometimes it is also necessary to send data to multiple PHYs. Figure 9 Shown are devices related to sending data to multiple PHYs within a multi-die structure, and for the sake of clarity, devices not related to this aspect are not shown. The PHYs within the slave die are called "slave PHYs".
[0082] As shown in FIG. 9, the main T2T SPI 385 has multiple slave die select lines, denoted as Each die select line is connected to a corresponding one of the slave T2T SPIs 675, 775, 775'. The boundaries between the dies are shown as dashed lines in the figure. For clarity, the SPI lines (such as the clock line, Leader Out Follower In (LOFI) line, Leader In Follower Out (LIFO) line, etc.) are not shown in the figure, but it should be understood that such lines can be further provided.
[0083] Each slave SPI uses the die select line to determine whether it is the intended recipient of the current data on the SPI bus. Specifically, when the die select line is in the active state or "set", it notifies the corresponding slave SPI that the current data on the bus is available for that slave SPI to read.
[0084] In the case of broadcast write, the master T2T SPI 385 simultaneously sets all the die select lines to the active state, so that all the slave SPIs 675, 775, 775' receive the data in the T2T SPI bus to implement the broadcast write operation. In the illustrated embodiment, the die select line is in the active state at a low level, but as an alternative, a scheme where it is in the active state at a high level can also be adopted.
[0085] As Figure 9 shown, the slave SPIs 675, 775, 775' are respectively connected to the corresponding PHYs 685, 785, 785' via the corresponding APB interconnects 625, 725, 725'. The SPI protocol does not support addressing, but the APB protocol supports addressing. A part of the data placed in the T2T SPI bus by the CPU core 300 is the APB address information, specifically the broadcast address corresponding to all the slave PHYs 685, 785, 785' (for example, depending on the number of PHYs in each die, there are 12 or 24 slave PHYs in total). Each PHY in each die is specifically assigned the same APB broadcast address or address range (i.e., the broadcast address or address range assigned to the PHY 370 in the master die). Since the corresponding broadcast address space involves multiple APB interconnects (such as 325, 625), the broadcast address space is described herein as being associated with a "multi-channel PCIe retimer" to aim to indicate that multiple physical instances of the APB bus are associated with the broadcast address space in the multi-die implementation.
[0086] Thus, by simultaneously setting all the die select lines to the active state and providing data including the APB broadcast address corresponding to all the PHYs of all the dies simultaneously in the T2T SPI bus, simultaneous writing to all the PHYs in all the dies can be achieved, thereby conforming to time-sensitive processes such as the startup process and shortening the time taken to configure and / or update all the PHYs.
[0087] In addition, each PHY (including the slave PHY) is also separately assigned a unique address / address range to enable writing to a specific PHY as needed. From the perspective of the master die processor, the entire multi-die module not only has a single global address space including corresponding regions for each PHY respectively, but also has a broadcast region for simultaneous writing to all the PHYs and slave PHYs.
[0088] The APB address space is a global address space across all the dies. That is, through this global address space, any register in any die can be addressed. In a specific structure, a base address in the form of the die identifier multiplied by a constant is provided to each die. The die identifier can be the die number, and the constant can be the base address of the master die. In addition, other storage space structures can also be adopted. Each register in each die is assigned a unique address or address range within the global address space. Thus, each PHY in PHYs 370, 685, 785, 785' is assigned a unique address or address range.
[0089] The broadcast write can be either a single-step operation or a two-step operation, which are described as follows.
[0090] In the single-step broadcast write operation, the SPI bus control information is included in the data provided to the T2T SPI bus. For ease of explanation, assuming the APB address is 24 bits and the data word size is 32 bits, the data provided to the T2T SPI bus can take the following format, which is referred to herein as the "control data packet".
[0091] r r r r r s s s a a a a a a a a a a a a a a a a a a a a a a a a
[0092] Among them, bits 0 to 23 are address bits ("a"), bits 24, 25, and 26 are die select bits ("s"), and bits 27 to 31 are reserved bits ("r"). In the specific case of this example, since the number of dies is three (i.e., the number of T2T SPIs is three), the number of die select bits is three. The reserved bits leave room for other die select bits - in this case, the number of reserved bits is five, so up to eight die select bits can be provided, thus supporting up to eight dies at most. This principle can be further extended to any number of dies by increasing the codeword size. In addition, other coding schemes can also be adopted. For example, in addition to representing the corresponding values of each specific die, the die select bits also represent a value corresponding to all dies. An embodiment of this scheme using two die select bits is as follows:
[0093] Selecting Bits from Dielets Corresponding Dielets 00 1 01 2 10 3 11 All
[0094] The address bits "a" form the APB address. In the case of a broadcast write, the same APB address is simultaneously assigned to all PHYs of all dies. Each T2T-SPI is respectively configured as the master bus of the corresponding APB interconnect, so that each T2T-Slave SPI can instruct its corresponding APB interconnect to perform a write operation on the corresponding PHY connected to the APB interconnect. In some cases, instead of using address data, a T2T-SPI bus that can obtain the address of the data to be written by automatically incrementing the address can be adopted.
[0095] Each die select bit respectively corresponds to a corresponding one of the T2T SPIs 675, 775, 775'. The value of each die select bit indicates whether the corresponding slave SPI receives data. Specifically, the master T2T SPI385 controls the die select line according to the value of the die select bit Each die select bit corresponds to one of the die select lines. In the case of a broadcast write, all three die select bits are set to a "valid" value (such as "1"), so that all three die select lines are simultaneously set to the active state. In the case of a non-broadcast write, only the die select bit corresponding to the die to be written is set to the valid value, while the other die select bits are set to the invalid value (such as "0").
[0096] In a two-step broadcast write operation, chiplet select control information is sent from the address data to the main T2T SPI 385 respectively. The chiplet select information can be sent in-band in the same manner as the above single-step operation, or other channels such as the System Management bus (SMBus) can be used. The address data can be sent separately before the PHY configuration data packet is sent. In some cases, instead of using address data, the T2T-SPI bus that can obtain the address of the data to be written by automatically incrementing the address can be adopted.
[0097] In each of the above single-step operation and two-step operation, data can be sent after the chiplet select information and address information are provided (the address information can be provided or not provided according to specific needs). The main T2T SPI 385 can keep the chiplet select line set to the active state until a new instruction related to the chiplet select line configuration is received. Similarly, the APB bus can continuously write to the specified address (or the auto-incremented address) before new addressing information is provided. In this way, the PHY configuration data packet can be broadcast to all PHYs of all chiplets simultaneously.
[0098] Next, in combination with Figure 10 , the process of sending the PHY configuration data packet to the main chiplet PHY and the slave PHYs of one or more chiplets in the multi-chiplet re-timer is described. The multi-chiplet re-timer can include two or more chiplets, one of which is the main chiplet and the rest are slave chiplets.
[0099] The structure of the multi-chiplet re-timer is as shown in Figure 7 and Figure 9 . It should be understood that only one slave chiplet can be set. Specifically, the re-timer includes a main chiplet, which has a main chiplet bus (such as the APB interconnect 325) and a processor (such as the CPU core 300) provided in the main chiplet. The re-timer also includes a slave chiplet, which has a plurality of slave PHYs (such as PHY 685) provided in the slave chiplet and a slave chiplet bus (such as the ABP interconnect structure 625) connected to these slave PHYs. The main chiplet and the slave chiplet are communicatively connected through a T2T bus (such as the above T2T SPI bus). The T2T bus has a T2T main bus (such as the main T2T-SPI 385) provided in the main chiplet and a slave T2T bus (such as the slave T2T-SPI 675) provided in the slave chiplet and connected to the slave chiplet bus.
[0100] In step 1000, the processor writes the configuration data packet to the multiple PHYs simultaneously via a bus connected to the processor and the multiple PHYs, using the broadcast address space of the multi-channel PCIe retimer. This step is the same as step 800, so please refer to the description of step 800 above. Step 1000 is part of the process of writing the configuration data packet to the main die PHY.
[0101] In step 1005, the processor writes the PHY configuration data packet to the multiple slave PHYs provided in one or more slave dies simultaneously via the T2T bus, using the broadcast address space of the multi-channel PCIe retimer. The PHY configuration data packet can be a PHY configuration data packet of the type described above (such as the configuration information and / or firmware of PHY 685, 785, and / or 785'). As described above, at least one address of the broadcast address space is assigned to all the slave PHYs among the multiple slave PHYs provided in the one or more slave dies.
[0102] In the case of setting multiple slave dies, step 1005 is performed on each slave die simultaneously. Regarding this, please refer to the above description again in combination with Figure 9 All the slave PHYs in all the slave dies have the same broadcast address / address range, that is, all the slave PHYs in all the slave dies are written simultaneously.
[0103] It should be noted that the write operation of the slave die involves the T2T bus, which enables the PHY configuration data packet to be sent from the master die to the slave die. Subsequently, the slave die places the PHY configuration data packet in the local APB interconnect of the slave die (such as 625, 725, 725') for further transmission to the slave PHY. The technology described above in combination with Figure 9 can be used to send data sub-packets in ABP format via the T2T bus.
[0104] Figure 10 The process shown enables the PHY configuration data packet to be sent to both the PHY 370 of the master die and the slave PHYs 685, 785, and / or 785' of the slave die. Among them, the sending to the master die and the slave die is performed separately. The reason is that since the master die PHY is directly connected to the CPU core 200 via the APB interconnect 325, the T2T bus does not need to communicate with the master die PHY 370. However, for the consideration of simplifying the broadcast write process from the perspective of the CPU core 300, the APB bus broadcast address of the master die PHY 370 can still be set to be the same as the APB bus broadcast address of the slave PHYs 685, 785, and / or 785'.
[0105] In addition, broadcast read operations can also be performed using a broadcast address space in a multi-die environment. One useful scenario for broadcast reads is to detect whether any interrupt requests have been issued.
[0106] Figure 11 The principle of broadcast read is shown. In a multi-die environment, the object of a broadcast read operation is a specific register shared by all dies. This register can be located within the PHY of each die, but this is not necessary since this operation can be used to read any register shared by all dies.
[0107] In step 1100, a processor (such as CPU core 200) of a multi-channel PCIe multi-die re-timer simultaneously sends a broadcast read operation to a master die register located within the master die and one or more slave die registers located within the one or more slave dies, respectively. This broadcast read operation can specify a broadcast read address that is simultaneously assigned to the master die register and the one or more slave die registers. The address assignment technique for broadcast reads is the same as the address assignment technique for broadcast writes. For related content, please refer to the above description of broadcast writes. Consistent with the above broadcast write scenario, in some cases, the relevant bus (such as APB interconnects 325, 625, 725, 725') supports an auto-increment function, so that it is not necessary to always provide a broadcast read address to the bus.
[0108] The broadcast read address can be, for example, a 24-bit address of the type described above. However, the present disclosure is not limited thereto, and other sizes of addresses can also be used instead.
[0109] In step 1105, the master die read result is received from the master die register. The master die read result is received along with the sending of the broadcast read instruction in step 1100. The master die read result can be received by other entities such as a processor or a storage buffer.
[0110] In step 1110, the corresponding one or more slave die read results are received from the one or more slave die registers, respectively. The one or more slave die read results are received along with the sending of the broadcast read instruction in step 1100. The result can be received by other entities such as a processor or a storage buffer. One result is received from each slave die. For example, in Figure 7 the case of the multi-die structure shown, there are a total of three slave die read results.
[0111] In step 1115, a bitwise OR operation is performed on the master die read result and the one or more slave die read results to generate a total read value. The bitwise OR operation is performed on all read results simultaneously. For example, in Figure 7 the case of three dies shown, the bitwise OR operation is as follows:
[0112] Total read value = bitwise OR (main die read result, dielet 1 read result, dielet 2 read result, dielet 3 read result)
[0113] If any one of the main die read result or the dielet read results is 1, the total read value generated by the bitwise OR operation is 1. The total read value is 0 only when all the main die read results and dielet read results are 0. This operation is useful for determining situations such as whether an interrupt request is issued. In this case, the main die register and the dielet registers can be used as interrupt request registers, and the processor can use Figure 11 the broadcast read operation shown to poll such interrupt request registers and determine whether any die has issued an interrupt request based on the total read value. When the total read value indicates that an interrupt request has been issued, the processor can take further measures to identify the specific details of the interrupt request, such as the device that generated the interrupt request. Subsequently, the processor can handle the interrupt request when the situation is appropriate.
[0114] The broadcast read operation can be used to read only a portion of the registers among a set of registers being read, for example, to read only the registers in one of the dielets. In this case, the main T2T SPI 385 can turn off the other registers that do not need to be read to prevent their signals from being included in the bitwise "OR" operation. Alternatively, the T2T SPI bus can be set to be active at a high level so that the registers that do not need to be read are at a low level during the broadcast read operation.
[0115] Since the above broadcast read operation is performed on all dielets simultaneously rather than sequentially, this read operation can shorten the total time spent in identifying specific situations such as pending interrupt requests, thereby, for example, improving the response speed when handling errors.
[0116] Figure 11 The operation shown can be performed by a multi-channel multi-die PCIe retimer that has multiple main physical layer data channel circuits (main PHY) located within the PCIe retimer main die and multiple secondary physical data channel circuits (secondary PHY) located in one or more dielets of the PCIe retimer, respectively. The PCIe retimer also has: a main die register provided within the main die; one or more secondary dielet registers provided in the one or more dielets; a processor; and a bus connected to the processor and the multiple PHYs.
[0117] The processor of the PCIe retimer is used to execute Figure 11The operations shown. Specifically, the processor is used to: simultaneously send a broadcast read instruction to the master die register and the one or more slave die registers; in response to the broadcast read instruction, receive a master die read result from the master die register; in response to the broadcast read instruction, respectively receive the corresponding one or more slave die read results from the one or more slave die registers; and perform a bitwise "OR" operation on the master die read result and the one or more slave die read results to generate a total read value. The processor may also be used to perform any other operation described above in combination with Figure 11 described. The processor may be, for example, CPU core 300.
[0118] In addition to the above broadcast write operation, or as an alternative to the broadcast write operation, the present disclosure also takes into account a so-called multicast operation. Multicast refers to sending data simultaneously to more than one PHY within the MCM, rather than the entire set of PHYs. That is, the object of the multicast transmission is a corresponding part of the PHYs within the MCM.
[0119] Figure 12 Shown is an MCM embodiment capable of performing a multicast operation. Figure 12 With Figure 1 There are common features, and in appropriate cases, a Figure 1 numbering method is used to represent such common features. Consistent with the above description of the present disclosure, Figure 12 the shown MCM may provide a retimer function. However, the MCM is not limited to this, and it may also provide other different functions.
[0120] Figure 12 In, the PHYs are grouped in pairs and labeled "PHY_Top". It should be understood that Figure 12 in, other components also labeled "PHY_Top" are exactly the same as PHY_Top 1200.
[0121] As shown in the figure, in this illustrated embodiment, each PHY_Top includes two PHYs. This number is for illustrative purposes only. Generally speaking, a PHY_Top may include any number of PHYs, such as three PHYs, four PHYs, or more PHYs. The significance of the number of PHYs in a PHY_Top is that since the multicast operation is targeted at a specific PHY_Top (i.e., all the PHYs within it), the number of PHYs determines the granularity of the multicast operation.
[0122] Although Figure 12Each die therein has two PHY_Tops and there are a total of four PHYs. However, it may also have other numbers of PHYs and other numbers of PHY_Tops. For example, as an alternative, each die may have four PHY_Tops and there are a total of eight PHYs, with each PHY_Top including two of the eight PHYs respectively. In addition, there are other variants. For example, each die may have two PHY_Tops and there are a total of eight PHYs, with each PHY_Top including four of the eight PHYs respectively.
[0123] As Figure 12 shown, each PHY has a configuration space "cfg" in the local memory of the PHY, and this configuration space contains data related to the configuration of the PHY. This data can be provided to the corresponding PHY in Figure 12 cases such as power-on and / or restart of the MCM shown. The power-on or startup process may need to meet specific timing constraints according to the regulations of protocols such as PCIe. The multicast operation described herein can shorten the total time required to load data into each PHY in the PHY_Top (e.g., load into the configuration space of the PHY), making it easier to meet the protocol requirements.
[0124] Communication between dies is achieved through a die-to-die interface. In Figure 12 an embodiment, this interface is a T2T SPI bus of the type described above in this specification. In addition to the T2T SPI bus, other types of buses such as UCIe and I 2 C can also be used. The functions of the T2T SPI bus are as described above in this specification, and for the sake of brevity, they will not be elaborated here.
[0125] Now refer to Figure 13 , which shows in a pictorial form an embodiment of the address space 1300 suitable for use in Figure 12 an embodiment. The address space 1300 is the address space of the entire MCM and is divided into four regions 1305a to 1305d, with each region corresponding to a respective one of the four dies 105, 115a to 115c. It should be understood that in connection with Figure 13 , "address" can refer to either a single address or an address range. In the case of an address range, the addresses that make up the range can be either consecutive addresses or non-consecutive addresses.
[0126] Each of the regions 1305a to 1305d contains a unique address for each PHY ("PHY_i address", i = 0 to 3) to enable the PHY to be the object of operation. In addition, each of the regions 1305a to 1305d also contains a unique address for each PHY_Top ("PHY_Top_j", j = 0 or 1) to enable the PHY_Top to be the object of operation. As can be seen from Figure 12 it, making a specific PHY_Top the object of operation means operating on all the PHYs that make up the PHY_Top. Thus, for example, when the PHY_Top_0 in Tile_0 is made the object of operation through the address range 1305, data is sent to the two PHYs, PHY_0 and PHY_1, in Tile_0.
[0127] The MCM further includes a broadcast address 1310. As described above, the broadcast address 1310 is not associated with any specific die, but can be used to operate on all the PHYs within all the dies in the same operation.
[0128] Although the above addresses are shown as consecutive addresses in Figure 13 they do not necessarily have to be so in reality. As an alternative, non-consecutive addresses can also be used. More generally, other address schemes that can implement Figure 13 the three-level mapping shown (i.e., single PHY operation, multicast operation, and broadcast operation) are also within the scope of the present disclosure.
[0129] Figure 14 Shown is the process of performing a write operation within an MCM that supports broadcast and multicast operations (such as the MCM shown in Figure 12 ). This process is described by taking the above T2T SPI bus as an example, but it should be understood that the process can also be changed to support other types of buses.
[0130] In step 1400, the die to be written is selected. In the case of the T2T SPI bus, this step may include setting the die select related to all the dies to be written to the "active state". When the die is the master die, the selection of the die can be achieved through the loopback path. If the write operation is a direct write to a single PHY, the die select is set to select the die where the single PHY is located. In the case of multicast write, the die select is set to operate on the die where the PHY_Top involved in the write operation is located. For broadcast write, the die select is set to operate on all the dies.
[0131] In step 1405, data is sent via the T2T SPI bus. Alternatively, in the case of not using the SPI bus, data is sent via other buses. It should be understood that the die select operation in step 1400 (such as the content of the die select setting) determines the die on which the data is to act.
[0132] In step 1410, each receiving die determines whether the relevant address of the data is a broadcast address. If it is not a broadcast address, it proceeds to step 1415, so that each receiving die writes the data to the PHY or PHY_Top corresponding to the address. As Figure 12 shown, the data can be written, for example, to the configuration space of each PHY.
[0133] If it is a broadcast address, it proceeds to step 1420, so that each receiving die writes the data to all PHYs within each die. As Figure 12 shown, the data can be written, for example, to the configuration space of each PHY.
[0134] It should be understood that for each die selected in step 1400, the operations described in steps 1410, 1415, and 1420 are performed separately and individually.
[0135] In addition to the above embodiments, the following give other embodiments of the present disclosure:
[0136] 1. A method, comprising: obtaining, by a processor of a multi-chip module (MCM), a physical layer data channel circuit (PHY) configuration data packet, the MCM having a plurality of dies and a plurality of PHYs distributed on the plurality of dies; and by the processor, using the broadcast address space of the MCM, via a bus connected to the processor and the plurality of PHYs, writing the configuration data packet to the plurality of PHYs simultaneously, the broadcast address space of the MCM including at least one address assigned to all PHYs among the plurality of PHYs, the MCM further having a non-broadcast address space including a plurality of unique addresses, each of the unique addresses being assigned to a corresponding one of the plurality of PHYs.
[0137] 2. The method according to item 1, wherein obtaining the PHY configuration data packet includes: receiving the PHY configuration data packet in a firmware image as part of a firmware update operation.
[0138] 3. The method according to item 1 or item 2, further comprising: by the processor, using the broadcast address space of the MCM, via a die-to-die (T2T) bus connected to a master die among the plurality of dies and a slave die among the plurality of dies, writing the PHY configuration data packet to at least one slave PHY located within the slave die among the plurality of PHYs, at least one PHY among the plurality of PHYs located within the master die of the MCM, and the non-broadcast address space including at least one unique address assigned to the die-to-die bus, and further including at least one unique address respectively assigned to the at least one slave PHY.
[0139] 4. The method according to item 1 or item 2 further includes: the processor uses the broadcast address space of the MCM to write the PHY configuration data packet simultaneously into multiple slave PHYs respectively located in the multiple slave die of the MCM, multiple PHYs in the master die of the MCM, and the non-broadcast address space includes at least one unique address allocated to the die-to-die (T2T) bus, and further includes multiple unique addresses, and each of the unique addresses is respectively allocated to a corresponding one of the multiple slave PHYs.
[0140] 5. The method according to item 4, wherein the T2T bus is a Serial Peripheral Interface (SPI) bus, the SPI bus has a master T2T bus located in the master die and multiple slave T2T buses respectively located in the multiple slave die, and wherein writing the configuration data packet simultaneously into the multiple slave PHYs includes: when all the slave die selection lines in the multiple slave die selection lines of the T2T bus are simultaneously set by the master T2T bus, the multiple slave die selection lines are respectively connected to the multiple slave T2T buses; and when all the slave die selection lines in the multiple slave die selection lines are simultaneously set, the master T2T bus sends the PHY configuration data packet.
[0141] 6. The method according to item 4, wherein the T2T bus is an SPI bus, the SPI bus has a master T2T bus located in the master die and multiple slave T2T buses respectively located in the multiple slave die, and wherein writing the configuration data packet simultaneously into the multiple slave PHYs includes: the processor sets multiple slave die selection bits for controlling data packet splitting to valid values, and each slave die selection bit corresponds to one of the multiple slave T2T buses respectively; the processor sends the control data packet splitting to the master T2T bus; the master T2T bus reads the slave die selection bits of the control data packet splitting; the master T2T bus simultaneously sets each of the multiple slave die selection lines of the T2T bus according to the corresponding value of the slave die selection bit; and when each of the multiple slave die selection lines is simultaneously set, the master T2T bus sends the PHY configuration data packet.
[0142] 7. The method according to any one of the above items, wherein obtaining the PHY configuration data packet further comprises: receiving a firmware image including the PHY configuration data packet; writing the firmware image into a non-volatile memory by the processor; obtaining firmware verification information from a read-only memory of the retimer by the processor; verifying the firmware image by the processor using the firmware verification information; and selectively preventing the firmware image from being executed and selectively preventing the PHY data packet from being distributed according to the verification result.
[0143] 8. An apparatus, comprising: a multi-chip module (MCM), the MCM having a plurality of dies, a plurality of physical layer data channel circuits (PHYs) distributed on the plurality of dies, a processor, and a bus connecting the processor and the plurality of PHYs, the processor being configured to: obtain a PHY configuration data packet; and use a broadcast address space of the MCM to write the configuration data packet to the plurality of PHYs simultaneously via the bus, the broadcast address space of the MCM including at least one address assigned to all of the plurality of PHYs, the MCM further having a non-broadcast address space including a plurality of unique addresses, each of the unique addresses being assigned to a corresponding one of the plurality of PHYs.
[0144] 9. The apparatus according to item 8, wherein the processor is further configured to: receive the PHY configuration data packet in a firmware image as part of a firmware update operation.
[0145] 10. The apparatus according to item 8 or item 9, further comprising a die-to-die (T2T) bus connecting a master die among the plurality of dies and a slave die among the plurality of dies, wherein: at least one master PHY among the plurality of PHYs is located in the master die, and at least one slave PHY among the plurality of PHYs is located in the slave die; the processor is further configured to use the broadcast address space of the MCM to write the configuration data packet to the at least one slave PHY simultaneously via the T2T bus, the non-broadcast address space further including at least one unique address assigned to the die-to-die bus and at least one address respectively assigned to the at least one slave PHY.
[0146] 11. The device as described in item 8 or item 9 further includes a die-to-die (T2T) bus connected to the master die among the multiple dies and the multiple slave dies among the multiple dies, where: at least one master PHY among the multiple PHYs is located within the master die, and at least one slave PHY among the multiple PHYs is respectively located within the corresponding slave die among the multiple slave dies, such that the MCM includes multiple slave PHYs; the processor is further configured to use the broadcast address space of the MCM to simultaneously write the configuration data packet to the multiple slave PHYs via the T2T bus, and the non-broadcast address space further includes at least one unique address allocated to the die-to-die bus and multiple unique addresses respectively allocated to the corresponding slave PHYs among the multiple slave PHYs one by one.
[0147] 12. The device as described in item 11, wherein the T2T bus is a Serial Peripheral Interface (SPI) bus, and the SPI bus includes: a T2T master bus located within the master die and connected to multiple slave die selection lines; and multiple slave T2T buses respectively located within the multiple slave dies and respectively connected to the multiple slave die selection lines, where the T2T master bus is configured to: simultaneously set all the slave die selection lines among the multiple slave die selection lines; and send the PHY configuration data packet when all the slave die selection lines among the multiple slave die selection lines are simultaneously set.
[0148] 13. The device as described in item 11, wherein the T2T bus is an SPI bus, and the SPI bus includes: a master T2T bus located within the master die and connected to multiple slave die selection lines; and multiple slave T2T buses respectively located within the multiple slave dies and respectively connected to the multiple slave die selection lines, where the processor is further configured to: set multiple slave die selection bits of a control sub-packet to valid values, each slave die selection bit corresponding to one of the multiple slave T2T buses respectively; and send the control sub-packet to the master T2T bus, where the master T2T bus is further configured to: respectively set each of the multiple slave die selection lines of the T2T bus simultaneously according to the corresponding value of the slave die selection bit; and send the PHY configuration data packet when each of the multiple slave die selection lines is simultaneously set.
[0149] 14. The apparatus according to any one of items 8 to 13 further includes: a non-volatile memory; and a read-only memory of the MCM, wherein the processor is configured to: receive a firmware image including the PCIe PHY configuration data packet; write the firmware image into the non-volatile memory; receive firmware verification information from the read-only memory; verify the firmware image using the firmware verification information; and selectively prevent the firmware image from being executed and selectively prevent the PHY data packet from being distributed according to the verification result.
[0150] 15. A method includes: simultaneously sending, by a processor of a multi-chip module (MCM), a broadcast read instruction to a master die register located within a master die and one or more slave die registers respectively located within one or more slave dies, the MCM having the master die and the one or more slave dies, the master die having a plurality of master physical data channel circuits (PHYs), and each slave die having a plurality of slave physical data channel circuits (slave PHYs); in response to the broadcast read instruction, receiving a master die read result from the master die register; in response to the broadcast read instruction, respectively receiving corresponding one or more slave die read results from the one or more slave die registers; and performing a bitwise "OR" operation on the master die read result and the one or more slave die read results to generate a total read value.
[0151] 16. The method according to item 15, wherein the broadcast read operation specifies a broadcast read address that is simultaneously assigned to the master die register and the one or more slave die registers.
[0152] 17. The method according to item 14 or 15, wherein the master die register and the one or more slave die registers are interrupt request registers, and the method further includes: determining, by the processor, whether an interrupt request is issued according to the total read value.
[0153] 18. A device, comprising: a multi-chip module (MCM), the MCM having a plurality of main physical layer data channel circuits (PHYs) located within a main die of the MCM and a plurality of slave physical data channel circuits (slave PHYs) respectively located within each of one or more slave dies of the MCM; a main die register located within the main die; one or more slave die registers respectively located within the one or more slave dies; and a processor and a bus connecting the processor and the plurality of PCIe PHYs, wherein the processor is configured to: simultaneously send a broadcast read instruction to the main die register and the one or more slave die registers; in response to the broadcast read instruction, receive a main die read result from the main die register; in response to the broadcast read instruction, respectively receive corresponding one or more slave die read results from the one or more slave die registers; and perform a bitwise "OR" operation on the main die read result and the one or more slave die read results to generate a total read value.
[0154] 19. The device according to item 18, wherein the broadcast read operation specifies a broadcast read address, and the broadcast read address is simultaneously allocated to the main die register and the one or more slave die registers.
[0155] 20. The device according to item 14 or 15, wherein the main die register and the one or more slave die registers are interrupt request registers, and the processor is further configured to: determine whether an interrupt request is issued according to the total read value.
[0156] 21. A method, comprising: obtaining, by a processor of a chip, a PHY configuration data packet, the chip including a die and a plurality of physical layer data channel circuits (PHYs) located within the die; and writing, by the processor, the configuration data packet to a suitable part of the plurality of PHYs via a bus connecting the processor and the plurality of PHYs, the multicast address space of the die including at least one address allocated to a suitable part of the plurality of PHYs, the die further having a non-broadcast address space including a plurality of unique addresses, each of the unique addresses being allocated to a corresponding one of the plurality of PHYs, and the suitable part of the plurality of PHYs including at least two PHYs.
[0157] 22. The method according to item 21, wherein the die further has a broadcast address space, and the broadcast address space includes addresses allocated to all of the plurality of PHYs.
[0158] 23. A method includes: obtaining, by a processor of a multi-chip module (MCM), a PHY configuration data packet, the MCM including a plurality of die and a plurality of physical layer data channel circuits (PHY) distributed among the plurality of die; and writing, by the processor, the configuration data packet into a suitable portion of the plurality of PHY via a bus connecting the processor and the plurality of die, the multicast address space of the MCM including at least one address assigned to a suitable portion of the plurality of PHY, the MCM further having a non-broadcast address space including a plurality of unique addresses, each of the unique addresses being respectively assigned to a corresponding one of the plurality of PHY, and the suitable portion of the plurality of PHY including at least two PHY.
[0159] 24. The method according to item 23, wherein the MCM further has a broadcast address space, and the broadcast address space includes addresses assigned to all of the plurality of PHY.
[0160] 25. An apparatus includes a die, the die including a processor, a plurality of physical layer data channel circuits (PHY), and a bus connecting the processor to each of the plurality of PHY, the processor being configured to: obtain a PHY configuration data packet; and write the PHY configuration data packet into a suitable portion of the plurality of PHY by using the multicast address space of the die, the multicast address space of the die including at least one address assigned to a suitable portion of the plurality of PHY, the die further having a non-broadcast address space including a plurality of unique addresses, each of the unique addresses being respectively assigned to a corresponding one of the plurality of PHY, and the suitable portion of the plurality of PHY including at least two PHY.
[0161] 26. The apparatus according to item 25, wherein the die further has a broadcast address space, and the broadcast address space includes addresses assigned to all of the plurality of PHY.
[0162] 27. A device includes a multi-chip module (MCM). The MCM includes a plurality of die, a processor within a master die among the plurality of die, a plurality of physical layer data channel circuits (PHY) distributed within the plurality of die, and a bus connecting the processor to each of the plurality of die. The processor is configured to: obtain a PHY configuration data packet; and write the PHY configuration data packet into a suitable portion of the plurality of PHY by using a multicast address space of the MCM. The multicast address space of the MCM includes at least one address assigned to a suitable portion of the plurality of PHY. The MCM further has a non-broadcast address space including a plurality of unique addresses, each of the unique addresses being respectively assigned to a corresponding one of the plurality of PHY. The suitable portion of the plurality of PHY includes at least two PHY.
[0163] 28. The device according to item 27, wherein the MCM further has a broadcast address space, and the broadcast address space includes an address assigned to all of the plurality of PHY.
[0164] It will be readily understood by those skilled in the art who benefit from the present disclosure that various modification schemes, extension schemes, replacement schemes, etc. may also exist for the technical solutions described herein, and such modified schemes are also within the scope of the present disclosure. It should also be noted that, unless otherwise clearly stated, such steps can be executed in any order when describing method steps.
Claims
1. A method, characterized in that, Comprising: Obtaining, by a processor of a multi-chip module, a physical layer data channel circuit configuration data packet, wherein the multi-chip module has a plurality of dies and a plurality of physical layer data channel circuits distributed on the plurality of dies; and Writing, by the processor, the configuration data packet into the plurality of physical layer data channel circuits simultaneously by using a broadcast address space of the multi-chip module via a bus connected to the processor and the plurality of physical layer data channel circuits, wherein the broadcast address space of the multi-chip module includes at least one address assigned to all the physical layer data channel circuits among the plurality of physical layer data channel circuits, wherein the multi-chip module further has a non-broadcast address space, and the non-broadcast address space includes a plurality of unique addresses, and each of the unique addresses is assigned to a corresponding one of the plurality of physical layer data channel circuits among the plurality of physical layer data channel circuits.
2. The method according to claim 1, characterized in that Obtaining the physical layer data channel circuit configuration data packet includes: receiving, as part of a firmware update operation, the physical layer data channel circuit configuration data packet in a firmware image.
3. The method according to claim 1, characterized in that Further comprising: Writing, by the processor, the physical layer data channel circuit configuration data packet into at least one slave physical layer data channel circuit located within a slave die among the plurality of physical layer data channel circuits simultaneously by using the broadcast address space of the multi-chip module via a die-to-die bus connected to a master die among the plurality of dies and the slave dies among the plurality of dies, wherein at least one physical layer data channel circuit among the plurality of physical layer data channel circuits located within the master die of the multi-chip module and the non-broadcast address space include at least one unique address assigned to the die-to-die bus and further include at least one unique address assigned to the at least one slave physical layer data channel circuit.
4. The method according to claim 1, characterized in that, Further comprising: Writing, by the processor, the configuration data packet into a plurality of slave physical layer data channel circuits located within the plurality of slave dies simultaneously by using the broadcast address space of the multi-chip module via a die-to-die bus connected to the master die of the multi-chip module and the plurality of slave dies among the plurality of dies, wherein a plurality of physical layer data channel circuits among the plurality of physical layer data channel circuits located within the master die of the multi-chip module and the non-broadcast address space include at least one unique address assigned to the die-to-die bus and further include a plurality of unique addresses, and each of the unique addresses is assigned to a corresponding one of the plurality of slave physical layer data channel circuits among the plurality of slave physical layer data channel circuits.
5. The method according to claim 4, characterized in that, The die-to-die bus is a serial peripheral interface bus, wherein the serial peripheral interface bus has a master die-to-die bus located within the master die and a plurality of slave die-to-die buses located within the plurality of slave dies, and wherein writing the configuration data packet into the plurality of slave physical layer data channel circuits simultaneously includes: Simultaneously set all of the multiple slave die select lines of the die-to-die bus from the master die to the die-to-die bus, wherein the multiple slave die select lines are connected to the multiple slave die-to-die buses; and When all of the multiple slave die select lines are simultaneously set, send the physical layer data channel circuit configuration data packet from the master die to the die-to-die bus.
6. The method according to claim 4, wherein The die-to-die bus is a Serial Peripheral Interface (SPI) bus, and the SPI bus has a master die-to-die bus within the master die and multiple slave die-to-die buses within the multiple slave dies. Wherein, simultaneously writing the configuration data packet to the multiple slave physical layer data channel circuits includes: Set multiple slave die select bits of a control sub-packet by the processor to valid values, where each slave die select bit corresponds to one of the multiple slave die-to-die buses; Send the control sub-packet from the processor to the master die-to-die bus; Read the slave die select bits of the control sub-packet by the master die-to-die bus; Set each of the multiple slave die select lines of the die-to-die bus by the master die-to-die bus according to respective values of the slave die select bits; and When each of the multiple slave die select lines is simultaneously set, send the physical layer data channel circuit configuration data packet from the master die-to-die bus.
7. The method according to claim 1, characterized in that, Obtaining the physical layer data channel circuit configuration data packet further includes: Receiving a firmware image including the physical layer data channel circuit configuration data packet; Writing the firmware image to a non-volatile memory by the processor; Obtaining firmware verification information by the processor from a read-only memory of a re-timer; Verifying the firmware image by the processor using the firmware verification information; and Selectively prevent the firmware image from being executed and selectively prevent the physical layer data channel circuit data packet from being distributed according to the verification result.
8. A device, characterized in that, Includes: A multi-chip module, where the multi-chip module has multiple dies, multiple physical layer data channel circuits distributed on the multiple dies, a processor, and a bus connected to the processor and the multiple physical layer data channel circuits. Wherein, the processor is configured to: Obtain a physical layer data channel circuit configuration data packet; and Simultaneously write the configuration data packet to the multiple physical layer data channel circuits via the bus using the broadcast address space of the multi-chip module, where the broadcast address space of the multi-chip module includes at least one address assigned to all of the multiple physical layer data channel circuits, and the multi-chip module also has a non-broadcast address space that includes multiple unique addresses, and each unique address is assigned to a corresponding one of the multiple physical layer data channel circuits.
9. The device according to claim 8, characterized in that, The processor is further configured to: receive the physical layer data channel circuit configuration data packet in the firmware image as part of a firmware update operation.
10. The device according to claim 8, characterized in that Further included are: a die-to-die bus connecting the master die among the multiple dies and the slave dies among the multiple dies, wherein: at least one master physical layer data channel circuit among the multiple physical layer data channel circuits is located within the master die, and at least one slave physical layer data channel circuit among the multiple physical layer data channel circuits is located within the slave die; the processor is further configured to use the broadcast address space of the multi-chip module to simultaneously write the configuration data packet to the at least one slave physical layer data channel circuit via the die-to-die bus, wherein the non-broadcast address space further includes: at least one unique address assigned to the die-to-die bus; and at least one address respectively assigned to the at least one slave physical layer data channel circuit.
11. The device according to claim 8, characterized in that, Further included are: a die-to-die bus connecting the master die among the multiple dies and multiple slave dies among the multiple dies, wherein: at least one master physical layer data channel circuit among the multiple physical layer data channel circuits is located within the master die, and at least one slave physical layer data channel circuit among the multiple physical layer data channel circuits is located within each of the multiple slave dies, such that the multi-chip module includes multiple slave physical layer data channel circuits; the processor is further configured to use the broadcast address space of the multi-chip module to simultaneously write the configuration data packet to the multiple slave physical layer data channel circuits via the die-to-die bus, and the non-broadcast address space further includes: at least one unique address assigned to the die-to-die bus; and multiple unique addresses, each unique address being assigned to a corresponding one of the multiple slave physical layer data channel circuits among the multiple slave physical layer data channel circuits.
12. The device according to claim 11, characterized in that, The die-to-die bus is a Serial Peripheral Interface (SPI) bus, wherein the Serial Peripheral Interface bus includes: a master die-to-die bus located within the master die and connected to multiple die select lines; and multiple slave die-to-die buses located within the multiple slave dies and connected to the multiple die select lines, wherein the master die-to-die bus is configured to: simultaneously set all the die select lines among the multiple die select lines; and send the physical layer data channel circuit configuration data packet when all the die select lines among the multiple die select lines are simultaneously set.
13. The device according to claim 11, wherein The die-to-die bus is a Serial Peripheral Interface (SPI) bus, wherein the Serial Peripheral Interface bus includes: a master die-to-die bus located within the master die and connected to multiple die select lines; and multiple slave die-to-die buses located within the multiple slave dies and connected to the multiple die select lines, wherein the processor is further configured to: set multiple die select bits of a control sub-packet to valid values, wherein each die select bit corresponds to one of the multiple slave die-to-die buses; and send the control sub-packet to the master die-to-die bus, wherein the master die-to-die bus is further configured to: Based on each value of the bits selected from the dies, simultaneously set each of the multiple die-select lines from the die to the die-to-die bus; and When each of the multiple die-select lines is simultaneously set, transmit the physical layer data channel circuit configuration data packet.
14. The device according to claim 8, characterized in that, Further includes: Non-volatile memory; And The read-only memory of the multi-chip module, Wherein, the processor is configured to: Receive a firmware image including the physical layer data channel circuit configuration data packet; Write the firmware image into the non-volatile memory; Receive firmware verification information from the read-only memory; Verify the firmware image by using the firmware verification information; and According to the verification result, selectively prevent the firmware image from being executed and selectively prevent the physical layer data channel circuit data packet from being distributed.
15. A method, characterized in that, Includes: The processor of the multi-chip module simultaneously sends a broadcast read instruction to the main die register located in the main die and one or more slave die registers located in one or more slave dies, wherein the multi-chip module has the main die and the one or more slave dies, the main die has multiple main physical data channel circuits, and each slave die has multiple slave physical data channel circuits; In response to the broadcast read instruction, receive the main die read result from the main die register; In response to the broadcast read instruction, receive one or more slave die read results from the one or more slave die registers; and Perform a bitwise "OR" operation on the main die read result and the one or more slave die read results to generate a total read value.
16. The method according to claim 15, wherein The broadcast read operation specifies a broadcast read address, wherein the broadcast read address is simultaneously allocated to the main die register and the one or more slave die registers.
17. The method according to claim 15, wherein The main die register and the one or more slave die registers are interrupt request registers, wherein the method further includes: The processor determines whether an interrupt request is issued according to the total read value.
18. A device, characterized in that, Includes: A multi-chip module, wherein the multi-chip module has multiple main physical layer data channel circuits located in the main die of the multi-chip module and multiple slave physical data channel circuits located in each of the one or more slave dies of the multi-chip module; A main die register located in the main die; One or more slave die registers located in the one or more slave dies; and A processor and a bus connected to the processor and the multiple physical layer data channel circuits, Wherein, the processor is configured to: Simultaneously send a broadcast read instruction to the main die register and the one or more slave die registers; In response to the broadcast read instruction, receive the main die read result from the main die register; In response to the broadcast read instruction, receive one or more slave die read results from the one or more slave die registers; and Perform a bitwise "OR" operation on the main die read result and the one or more slave die read results to generate a total read value.
19. The device according to claim 18, characterized in that, The broadcast read operation specifies a broadcast read address, where the broadcast read address is simultaneously allocated to the master die register and the one or more slave die registers.
20. The device according to claim 18, characterized in that The master die register and the one or more slave die registers are interrupt request registers, and the processor is further configured to: Determine that an interrupt request is issued according to the total read value.
Citation Information
Patent Citations
Circuits for efficient detection of vector signaling codes for chip-to-chip communication using sums of differences
US9288082B1