Speculative numbering of address spaces for bus equipment functions
The MPB addresses inefficiencies in PCIe bus function address space management by enabling flexible and rapid reassignment, optimizing resource use and scalability in dynamic systems.
Patent Information
- Application Number
- DE112016006065
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-03-24
- Filing Date
- 2016-11-26
- Publication Date
- 2025-11-27
- Estimated Expiration
- 2036-11-26
AI Technical Summary
Existing PCI Express (PCIe) systems face inefficiencies in the allocation and management of bus function address spaces, leading to wasted resources and slow reassignment processes, especially in deep hierarchies and systems with hot-add/remove components, due to the limitations of traditional segmentation mechanisms.
The implementation of a Mapping Portal Bridge (MPB) that uses hardware and/or software logic to provide flexible and efficient mapping between primary and secondary bus function address spaces, allowing for virtual segments and reducing the need for additional hardware resources, enabling rapid reassignment without system freezes.
This approach optimizes the use of bus function address spaces, enhances scalability, and minimizes resource waste, facilitating fast and efficient reassignment of bus numbers in dynamic systems, maintaining compatibility with existing PCIe system software stacks.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
REFERENCE TO RELATED REGISTRATIONS
[0001] This application claims priority over the preliminary US patent application with serial number 62 / 387,492, filed on December 26, 2015, entitled “SPECULATIVE ENUMERATION OF BUS-DEVICE-FUNCTION ADDRESS SPACE”, and over the US patent application with serial number 15 / 079,922, filed on March 24, 2016, entitled “SPECULATIVE ENUMERATION OF BUS-DEVICE-FUNCTION ADDRESS SPACE”. AREA
[0002] This disclosure relates to a computing system and in particular (but not exclusively) an address space mapping. STATE OF THE ART
[0003] A Peripheral Component Interconnect (PCI) configuration space is used by systems that utilize PCI, PCI-X, and PCI Express (PCIe) to perform configuration tasks for PCI-based devices. PCI-based devices have an address space for device configuration registers, referred to as the configuration space, and PCI Express introduces an extended configuration space for devices. Configuration space registers are typically mapped by the host processor to input / output locations in memory. Device drivers, operating systems, and diagnostic software access the configuration space and can read from and write information to configuration space registers.
[0004] One of the improvements the PCI Local Bus offered over other I / O architectures was its configuration mechanism. In addition to the normal memory-mapped I / O port spaces, each device function on the bus has a 256-byte configuration space that can be addressed by knowing the eight-bit PCI bus number, the five-bit device number, and the three-bit function number for the device (usually referred to as BDF or B / D / F, an abbreviation of Bus / Device / Function). This allows for up to 256 buses, each with up to 32 devices, each supporting eight functions. A single PCI expansion card can respond as one device and can implement at least function number zero. The first 64 bytes of the configuration space are standardized; the remainder is available for extensions defined in the specification and / or for vendor-defined purposes.
[0005] To allow more parts of the configuration space to be standardized without conflicting with existing usage, a list of capabilities can be defined in the first 192 bytes of the peripheral component interface configuration space. Each capability has one byte describing its nature and one byte pointing to the next capability. The number of additional bytes depends on the capability ID. If capabilities are used, a bit in the status register is set, and a pointer to the first capability in a linked list is provided. Previous versions of PCIe incorporated similar features, such as an extended PCIe capability architecture.
[0006] US 2011 / 0219164 A1 shows virtual functions (VFs) 602-1 to 602-N of an I / O device, which are separately assigned to a plurality of computers 1-1 to 1-N. Address exchange table 506 registers a root domain, which is an address space of computer 1, and mapping information of an I / O domain, the mapping information being an address space unique to I / O device 6.
[0007] The problem stated is solved according to the invention by the features of claims 1, 19 and 38. Further embodiments of the invention are described in the respective dependent claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates an embodiment of a computing system that includes a connection architecture. Fig. Figure 2 illustrates an embodiment of a connection architecture that includes a layered stack. Fig. Figure 3 illustrates an embodiment of a request or packet to be created or received in a connection architecture. Fig. Figure 4 illustrates an embodiment of a transmitter and receiver pair for a connection architecture. Fig. Figure 5 illustrates a representation of system buses. Fig. Figure 6 illustrates an example of the numbering of bus identifiers in a system. Fig. Figure 7 illustrates an embodiment of an imaging portal bridge (MPB). Fig. Figure 8 illustrates a representation of a numbering of bus identifiers in a system and corresponding address illustrations. Fig. Figure 9 illustrates a representation of at least part of an exemplary register of skills. Fig. 10A-10C are simplified block diagrams that illustrate an exemplary technique for numbering facilities within a system. Fig. Figure 11 is a simplified flowchart illustrating an exemplary technique for numbering facilities within a system. Fig. Figure 12 illustrates an embodiment of a block diagram for a computing system that includes a multi-core processor. Fig. Figure 13 illustrates another embodiment of a block diagram for a computing system. DETAILED DESCRIPTION
[0008] The following description presents numerous specific details, such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific dimensions / heights, specific processor pipeline stages and operations, etc., to enable a thorough understanding of the present invention. However, it is obvious to those skilled in the art that these specific details need not be used to carry out the present invention.In other cases, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / specific code for described algorithms, specific firmware code, a specific connection operation, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, a specific expression of algorithms in code, specific shutdown and cut-off techniques / logic, and other specific operating details of the computer system, were not described in detail to avoid unnecessary obfuscation of the present invention.
[0009] Although the following embodiments can be described with reference to energy saving and energy efficiency in certain integrated circuits, such as computing platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings from embodiments described herein can be applied to other types of circuits or semiconductor devices that can also benefit from improved energy efficiency and energy saving. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks™ and can also be used in other devices such as portable devices, tablets, other thin notebooks, system-on-a-chip (SoC) devices, and embedded applications.Some examples of portable devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, a network computer (NetPC), set-top boxes, network hubs, long-distance network (WAN) switches, or any other systems capable of performing the functions and operations taught below. Furthermore, the devices, procedures, and systems described herein are not limited to physical computing equipment but may also involve software optimizations for energy conservation and efficiency.
[0010] As computing systems evolve, their components become more complex. Consequently, the interconnect architecture for coupling and communication between these components also increases in complexity to ensure bandwidth requirements are met for optimal component operation. Furthermore, different market segments require different aspects of interconnect architectures to meet market needs. For example, servers require higher performance, while the mobile ecosystem can sometimes sacrifice overall performance for energy savings. However, the sole purpose of most interconnects is to provide the highest possible performance with maximum energy efficiency. Below, we discuss a number of interconnects that might benefit from aspects of the invention described herein.
[0011] One interconnect architecture is the Peripheral Component Interconnect (PCI) Express architecture (PCIe architecture). A primary goal of PCIe is to enable components and devices from various vendors to work together in an open architecture spanning multiple market segments: clients (desktops and mobile devices), servers (standard and enterprise), and embedded and communications devices. PCI Express is a high-performance, general-purpose I / O connection designed for a wide variety of future computing and communications platforms. Some PCI attributes, such as its usage model, load-store architecture, and software interfaces, have been maintained throughout its revisions, while previous parallel bus implementations have been replaced by a highly scalable, fully serial interface.Newer versions of PCI Express leverage the advantages of point-to-point connections, switch-based technology, and a packaged protocol to deliver a new level of performance and new features. Power management, Quality of Service (QoS), hot-swap support, data integrity, and error handling are among the advanced features supported by PCI Express.
[0012] With reference to Fig. Figure 1 illustrates an embodiment of a fabric consisting of point-to-point links connecting a set of components. A system 100 includes a processor 105 and system memory 110 coupled to a control node 115. The processor 105 contains any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or another processor. The processor 105 is coupled to the control node 115 by a front-side bus (FSB) 106. In one embodiment, the FSB 106 is a serial point-to-point link as described below. In another embodiment, the link 106 incorporates a serial differential link architecture that conforms to various link standards.
[0013] System memory 110 contains any memory device, such as random access memory (RAM), non-volatile (NV) memory, or other memory that devices in system 100 can access. System memory 110 is coupled to control node 115 via a memory interface 116. Examples of memory interfaces include a double data rate (DDR) memory interface, a dual-channel DDR memory interface, and a dynamic RAM memory interface (DRAM memory interface).
[0014] In one embodiment, the control node 115 is a root node, root complex, or root controller in a Peripheral Component Interconnect Express (PCIe) connection hierarchy. Examples of control nodes 115 include a chipset, a memory control node (MCH), a northbridge, a connection control node (ICH), a southbridge, and a root controller / root node. Often, the term chipset refers to two physically separable control nodes, i.e., a memory control node (MCH) coupled to a connection control node (ICH). It should be noted that current systems often have the MCH integrated into the processor 105, while the controller 115 communicates with I / O facilities in a manner similar to that described above. In some embodiments, peer-to-peer routing is optionally supported by the root complex 115.
[0015] Here, the control node 115 is coupled to a switch / bridge 120 via a serial link 119. Input / output modules 117 and 121, which can also be referred to as interfaces / ports 117 and 121, contain / implement a layered protocol stack to provide communication between the control node 115 and the switch 120. In one embodiment, multiple devices can be coupled to the switch 120.
[0016] The switch / bridge 120 forwards packets / messages upwards from a device 125, i.e., up a hierarchy to a root complex, to the control node 115, and downwards, i.e., down a hierarchy from a root control, from the processor 105 or from the system memory 110 to the device 125. In one embodiment, the switch 120 is referred to as a logical assembly of several virtual PCI-to-PCI bridge devices.The Device 125 contains any internal or external device or component that can be coupled to an electronic system, such as an I / O device, a network interface controller (NIC), an add-in card, an audio processor, a network processor, a hard disk, a storage device, a CD / DVD-ROM drive, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a FireWire device, a Universal Serial Bus (USB) device, a scanner, and other input / output devices. Such a device is often referred to as an endpoint in PCIe terminology. Although not specifically shown, the Device 125 may contain a PCIe-to-PCI / PCI-X bridge to support legacy PCI devices or PCI devices of a different version. Endpoint devices in PCIe are often classified as legacy, PCIe, or root complex integrated endpoints.
[0017] Furthermore, a graphics accelerator 130 is coupled to the control node 115 via a serial link 132. In one embodiment, the graphics accelerator 130 is coupled to an MCH, which is coupled to an ICH. The switch 120, and consequently the I / O device 125, is then coupled to the ICH. I / O modules 131 and 118 are also intended to implement a layered protocol stack for communication between the graphics accelerator 130 and the control node 115. Similar to the discussion of the MCH above, a graphics controller or the graphics accelerator 130 itself can be integrated into the processor 105.
[0018] On Fig. 2 With reference to this, an embodiment of a layered protocol stack is illustrated. The layered protocol stack 200 contains any form of layered communications stack, such as a QuickPath Interconnect (QPI) stack, a PCIe stack, a next-generation high-performance computing interconnect stack, or another layered stack. Although the immediately following discussion relates to Fig. Since 1-4 is implemented in conjunction with a PCIe stack, the same concepts can be applied to other connection stacks. In one embodiment, the protocol stack 200 is a PCIe protocol stack containing a transaction layer 205, a connection layer 210, and a physical layer 220. An interface, such as interfaces 117, 118, 121, 122, 126, and 131 in Fig. 1, can be represented as a communication protocol stack 200. A representation as a communication protocol stack can also be referred to as a module or interface that implements / contains a protocol stack.
[0019] PCI Express uses packets to communicate information between components. Packets are formed in the transaction layer (205) and the data link layer (210) to carry information from the sending component to the receiving component. As the sent packets pass through the other layers, they are augmented with additional information necessary for handling packets at those layers. On the receiving side, the reverse process takes place, and packets are transformed from their physical layer (220) representation to the data link layer (210) representation and finally (for transaction layer packets) to a form that can be processed by the transaction layer (205) of the receiving device. Transaction layer
[0020] In one embodiment, the transaction layer 205 provides an interface between a processing core of an institution and the link architecture, such as the data link layer 210 and the physical layer 220. In this respect, a primary task of the transaction layer 205 is the assembly and disassembly of packets (i.e., transaction layer packets or TLPs). The translation layer 205 typically manages credit-based flow control for TLPs. PCIe implements split transactions, i.e., transactions with time-separated requests and responses, which allows one link to carry other traffic while the destination institution gathers data for the response.
[0021] Furthermore, PCIe employs a credit-based flow control mechanism. In this scheme, an institution announces an initial credit amount for each of the receive buffers in transaction layer 205. An external institution at the opposite end of the link, such as control node 115 in Fig. 1. This counts the number of credits consumed by each TLP. A transaction can be sent as long as it does not exceed a credit limit. Upon receiving a response, a credit amount is restored. One advantage of a credit scheme is that the latency of credit return does not affect performance, provided the credit limit is not exceeded.
[0022] In one embodiment, an address space for four transactions contains a configuration address space, a memory address space, an input / output address space, and a message address space. Memory transactions contain one or more read and write requests to transfer data to / from a location mapped in memory. In one embodiment, memory transactions can use two different address formats, such as a short address format like a 32-bit address or a long address format like a 64-bit address. Configuration space transactions are used to access the configuration space of the PCIe devices. Configuration space transactions contain read and write requests. Message transactions are defined to support on-tape communication between PCIe intermediaries.
[0023] Therefore, in one embodiment, the transaction layer 205 assembles a packet header / payload 156. The format for current packet headers / payloads can be found in the PCIe specification on the PCIe specification website.
[0024] Shortly Fig. With reference to Section 3, an embodiment of a PCIe transaction descriptor is illustrated. In one embodiment, the transaction descriptor 300 is a mechanism for carrying transaction information. In this respect, the transaction descriptor 300 supports the identification of transactions in a system. Other possible uses include tracking modifications to a standard transaction sequence and mapping transactions to channels.
[0025] The transaction descriptor 300 contains a global identifier field 302, an attribute field 304, and a channel identifier field 306. The illustrated example shows a global identifier field 302 that includes a local transaction identifier field 308 and a source identifier field 310. In one embodiment, the global transaction identifier 302 is unique for all pending requirements.
[0026] After implementation, the local transaction ID field 308 is a field generated by a requesting intermediary and is unique for all pending requests that require fulfillment by that requesting intermediary. Furthermore, in this example, the source ID field 310 uniquely identifies the requesting intermediary within a PCle hierarchy. Accordingly, the local transaction ID field 308, together with the source ID 310, provides a global identification of a transaction within a hierarchy domain.
[0027] Attribute field 304 specifies properties and relationships of the transaction. In this respect, attribute field 304 may be used to provide additional information that allows modification of the default transaction handling. In one embodiment, attribute field 304 contains a priority field 312, a reserved field 314, an ordering field 316, and a no-snoop field 318. Here, the priority subfield 312 can be modified by an initiator to assign a priority to the transaction. The reserved attribute field 314 is reserved for future use or for a provider-defined use. Possible usage models that utilize priority or security attributes can be implemented using the reserved attribute field.
[0028] In this example, the order attribute field 316 is used to provide optional information about the type of order, which can modify the default ordering rules. According to an implementation example, an order attribute of "0" indicates that default ordering rules should be applied, while an order attribute of "1" indicates a more relaxed order, where writes can override other writes in the same direction and read completions can override writes in the same direction. The snoop attribute field 318 is used to determine whether transactions are snooped. As shown, the channel ID field 306 identifies a channel with which a transaction is associated. Compound layer
[0029] The data link layer 210, also known as the data link layer 210, acts as an intermediate layer between the transaction layer 205 and the physical layer 220. In one embodiment, one of the tasks of the data link layer 210 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components of a connection. One side of the data link layer 210 accepts TLPs composed by the transaction layer 205, applies a packet sequence identifier 211 (i.e., an identifier number or packet number), calculates and applies an error detection code (i.e., CRC 212), and passes the modified TLPs to the physical layer 220 for transmission over a physical layer to an external device. Physical layer
[0030] In one embodiment, the physical layer 220 includes a logical sub-block 221 and an electrical sub-block 222 for physically transmitting a packet to an external device. The logical sub-block 221 is responsible for the "digital" functions of the physical layer 221. In this respect, the logical sub-block comprises a transmit area for preparing outgoing information for transmission by the physical sub-block 222 and a receive area for identifying and preparing received information before it is transmitted to the link layer 210.
[0031] The physical block 222 comprises a sender and a receiver. The sender is supplied with symbols by the logical subblock 221, which the sender serializes and sends to an external device. The receiver is supplied with serialized symbols by an external device and converts the received signals into a bitstream. The bitstream is deserialized and supplied to the logical subblock 221. In one embodiment, an 8b / 10b transmit code is used, in which ten-bit symbols are sent / received. Special symbols are used to enclose a packet with frame 223. In addition, in one example, the receiver also provides a symbol clock derived from the incoming serial stream.
[0032] As stated above, although the transaction layer (205), the link layer (210), and the physical layer (220) were discussed with reference to a specific implementation of a PCIe protocol stack, a layered protocol stack is not so restricted. Rather, any layered protocol can be included / implemented. As an example, a port / interface represented as a layered protocol includes: (1) a first layer to assemble packets, i.e., a transaction layer; (2) a second layer to assemble packets in a sequence, i.e., a link layer; and (3) a third layer to send the packets, i.e., a physical layer. A layered Common Standard Interface (CSI) protocol is used as a specific example.
[0033] Next, with reference to Fig. Figure 4 illustrates an embodiment of a serial PCIe point-to-point interconnect. Although an embodiment of a serial PCIe point-to-point interconnect is shown, a serial point-to-point interconnect is not limited in this way, since it includes any transmit path for sending serial data. In the illustrated embodiment, a basic PCIe interconnect comprises two differentially operated low-voltage signal pairs: a transmit pair 406 / 411 and a receive pair 412 / 407. Accordingly, a device 405 comprises transmit logic 406 for sending data to a device 410, and receive logic 407 for receiving data from the device 410. In other words, two transmit paths, i.e., paths 416 and 417, and two receive paths, i.e., paths 418 and 419, are included in a PCIe interconnect.
[0034] A transmit path refers to any path used to send data, such as a transmit line, copper wire, optical line, wireless communication channel, infrared communication link, or other communication path. A connection between two facilities, such as facility 405 and facility 410, is called a link, such as link 415. A link can support a lane—each lane represents a set of differential signal pairs (one pair for transmitting, one pair for receiving). For bandwidth scaling, a link can aggregate multiple lanes, designated xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.
[0035] A differential pair refers to two transmission paths, such as lines 416 and 417, used to transmit differential signals. For example, when line 416 switches from a low voltage level to a high voltage level (i.e., a rising edge), line 417 transitions from a high logic level to a low logic level (i.e., a falling edge). Differential signals may exhibit better electrical characteristics, such as improved signal integrity (i.e., reduced cross-coupling, voltage boost / undershoot, oscillations, etc.). This allows for a better timing window, enabling faster transmission frequencies.
[0036] New and growing deployment models, such as PCIe-based storage arrays and Thunderbolt, are driving significant growth in the depth and breadth of the PCIe hierarchy. The PCI Express architecture (PCle architecture) built upon PCI, which defines a "configuration space" in which system firmware and / or software discover and enable / disable / control features. Addressing in this space is based on a 16-bit address (often referred to as the "BDF," or Bus Setup Function Number), consisting of an 8-bit bus number, a 5-bit setup number, and a 3-bit function number.
[0037] PCI allows systems to have multiple independent BDF spaces called "segments." Each segment can have specific resource requirements, such as a mechanism for generating PCI / PCIe configuration requests, including the Extended Configuration Access Mechanism (ECAM) defined in the PCIe specification. Additionally, input / output (I / O) memory management units (IOMMUs) (such as Intel VTd) can use the BDF space as an index but may not be configured to directly understand segments. Consequently, in some cases, a separate ECAM and IOMMU must be duplicated for each segment defined in a system. Fig. Figure 5 illustrates an example of a system containing multiple segments (e.g., 505a-c). For instance, in this example, a segment is defined for each of three switches 510, 515, and 520 connected to a root complex 525. A separate IOMMU and a separate ECAM can be implemented on the root complex 525 to enable each of the segments (e.g., 505a-c). Furthermore, in this example, a variety of switches (e.g., 530a-r), a variety of endpoints (EPs), and other equipment in each segment are connected to different buses. In some cases, a segment's configuration space can reserve multiple bus addresses for possible installation events during operation, which limits the total number of bus addresses available in each segment.Furthermore, the assignment of bus numbers in one or more segments can be performed according to an algorithm that barely considers dense address padding and compact use of the available bus address space. This can, in some cases, result in wasted configuration address space (i.e., BDF space).
[0038] Traditional PCIe systems are configured to allocate address space in a way that, when applied to modern and emerging use cases, results in inefficient use of BDF space, and especially bus numbers. While relatively few implementations might actually involve a single system consuming all unique BDF values (e.g., the 64K defined under PCIe), deep hierarchies, such as those found in deep hierarchies of PCIe switches, can very quickly consume the available bus numbers within the BDF space. Additionally, in applications that support hot-install, large portions of the BDF space are typically reserved for potential future use (i.e., when a future device is hot-installed into the system), further reducing the pool of bus numbers immediately available to a system.While segmentation mechanisms can be used to address this problem, the segments themselves have poor scalability because, as noted above, additional hardware resources (e.g., IOMMUs) must be incorporated into the CPU, platform control node (PCH), system-on-a-chip (SoC), root complex, etc., to support each segment. Therefore, using segments to address deep hierarchies results in scaling the system to meet a system requirement in the worst-case scenario, which is typically far more than most systems would require, leading to a significant waste of platform resources. Furthermore, creating segments outside of the system's root complex can be difficult (and in some cases virtually impossible), among other exemplary problems.
[0039] In some implementations, a system may be provided to enable more efficient use of BDF space and to solve at least some of the exemplary problems mentioned above. Among other exemplary advantages, this may allow the extension of PCIe, Thunderbolt, on-chip system fabrics (e.g., Intel On-Chip System Fabric (IOSF) and others), and other interconnects to very large topologies without requiring dedicated resources in the root complex, as would be the case with solutions that rely exclusively on segments or other alternatives. Fig. Figure 6 shows a simplified block diagram 600 illustrating an example system containing a root complex 615 to which endpoints (e.g., 605, 610) and switches (e.g., 620, 625, 630) are connected via several buses forming hierarchies of switch fabrics. The example of Fig. Figure 6 further illustrates an exemplary assignment of bus numbers to buses within the system following an exemplary PCIe BDF allocation. In this example, a system with two directly connected facilities 605 and 610, which are directly connected to a root complex 615, and two switch-based hierarchies (corresponding to switches 620 and 625) is numbered (or assigned) with approximately the densest possible bus number assignments (as indicated by circular labels (e.g., 650a-d, etc.)). In deep hierarchies, the available bus numbers in a single BDF space can be quickly consumed. Furthermore, real-world systems allocate bus numbers much less efficiently, resulting in a sparse (or "wasted") allocation of BDF space.
[0040] Another problem with use cases that support hot-add / remove components, such as Thunderbolt and, in some cases, PCIe-based storage, is that the bus number assignments in the BDF space are "reassigned" to accommodate changes in the hardware topology that occur in a running system. However, this reassignment can be very difficult for system software to perform, as in typical cases all PCI functions are then forced into a quiesced state to allow the system to renumber the BDF space, followed by re-enabling the PCI functions. This process can be quite slow and usually results in a system freeze, in the worst case for very long periods (e.g., long enough to be disruptive to running applications and easily noticeable to the end user).An improved system can also be provided to reduce the time required to apply a revised BDF space, so that the reassignment process can be performed in a span of hundredths of milliseconds or faster, without explicitly placing PCI functions into quiesced states.
[0041] Finally, very large systems or systems with (proprietary) mechanisms for supporting multiple root complexes that use segments can be defined. An improved system can also be applied in such use cases to provide facility management with a minimum of changes relative to what would be implemented using a system with a single root. More specifically, an improved system can provide a mapping portal bridge (MPB), implemented using hardware logic (and / or software logic) of one or more facilities in a system, to provide multiple views of a BDF space and remapping tables to allow a "bridge" (which is the logical view of a root or switch port) to translate one view of a BDF space to another, in both directions across the bridge, to effectively create a virtual segment.
[0042] An Mapping Portal Bridge (MPB) can be implemented as logical blocks (implemented in hardware, firmware, and / or software) deployed on one or more ports of a root node or switch to enable translation between two or more defined BDF spaces, for example, a primary and a secondary BDF space in some implementations. An MPB can be implemented on root ports and / or switch ports, using a consistent software model (such as that used by system software), and can be implemented recursively within a given topology, allowing for a high degree of scalability. Furthermore, an MPB need not be tied to a specific usage model (e.g., it can alternatively be used in either or both Thunderbolt (TBT) and standard PCIe use cases).Furthermore, MPB implementations can support implementation flexibility and compromise on design cost / performance. Additionally, consistency within an existing PCIe system software stack can be maintained, among other exemplary advantages.
[0043] The MPB employs a mapping mechanism that allows it to map all PCIe packets flowing between a primary BDF space and a secondary BDF space across the MPB. The primary BDF space refers to the view of the configuration address space seen on the primary side of the bridge (i.e., the side closer to the host CPU, such as the root complex). The secondary BDF space can refer to the view of the configuration address space created and maintained for facilities on the secondary side of the bridge (i.e., the side closer to the facilities and downstream of the root complex or CPU). In PCIe, the same BDF mappings used for facility configuration can also be used to identify the source (and sometimes the destination) of packets, report errors, and perform other functions.
[0044] Fig. Figure 7 illustrates a block diagram (700) that represents an implementation example of an MPB (705). The MPB (705) can contain a pointer to and / or a (complete or partial) local copy of two mapping tables: one for mapping the secondary address space to the primary address space (BDFsec→BDFpri) and another for mapping the primary address space to the secondary address space (BDFpri→BDFsec). In one example, the mapping tables can be stored in system memory (710). In this case, the MPB (705) (and the system software) can read the tables from system memory (at 715) and / or maintain local copies (720, 725) of at least one section of each mapping table (i.e., BDFsec→BDFpri and BDFpri→BDFsec) (for example, to improve performance).In other examples, the BDFsec→BDFpri and BDFpri→BDFsec mapping tables can be stored directly in the MPB 705 without maintaining copies in system memory 710. The MPB 705 can also perform and manage translations between the primary BDF space and one or more BDFsec spaces. In the case of multiple BDFsec spaces, a single mapping table can be used, containing one column to map not only the BDFsec address but also the specific BDFsec space to a BDFpri address. In other cases, separate mapping tables can be maintained for each BDFsec space. The MPB can be implemented in the hardware of a switch, node, or port, such as a port of a root complex. The MPB 705 can also include control logic 730 to enable / disable the mapping functionality (e.g.,to selectively configure the MPB 705 as an option on different ports of a switch or root complex, among other examples).
[0045] An MPB 705 can, in some cases, be configured to operate flexibly as either an MPB 705 (using the primary / secondary BDF space mapping mechanism) or a conventional bridge (using, for example, conventional PCle BDF addressing) (e.g., using control logic 730). When the system software wants to activate an MPB 705, it can configure the MPB to provide a unique one-to-one mapping between primary (BDFpri) and secondary (BDFsec) BDF addresses. In other words, a single BDF on the primary side can correspond to a single BDF on the secondary side. Among other exemplary advantages, this constraint can ensure that the MPB does not track outstanding requests, as such tracking would add significant costs to the MPB. In some implementations, multiple MPBs can be used in a single system.For example, a separate MPB 705 may be provided for each BDFsec space. In such cases, the BDFsec assignments behind these multiple different MPBs may be allowed to reuse the same BDF values (in their respective second BDF spaces) (which they likely will), provided these are mapped to unique values in the BDFpri space.
[0046] Fig. 8 is a simplified block diagram illustrating the example of Fig. Figure 6 shows that it has been modified by the use of MPB(s), which enables a BDFpri space / BDFsec space dichotomy. Fig. Figure 8, for example, shows how BDFsec spaces can be allocated in a system with two MPBs (a first MPB used to implement a virtual segment (vSEG) A (805a), and a second MPB used to implement vSEG B (805b)). In this example, bus numbers 1 and 2 (at 806 and 808) of root complex 625 can be maintained in the BDFpri space (810) and assigned to the two directly connected facilities 605 and 610. Hierarchies of bus 3-n connections can be handled in this example by BDFsec spaces (corresponding to vSEG A (805a) and vSEG B (805b)) and one or more corresponding MPBs. Buses 3-6 can, for example, be numbered towards a BDFsec space vSEG A (805a), thus providing a virtual segment to the buses and facilities connected to the root complex via buses 3-6.A second virtual segment can be provided by defining a second BDFsec space vSEG B (805b) that includes buses and facilities connected to the root complex via buses 7-n. Each BDFsec space can be assigned BDF addresses (and bus numbering) within that second space according to any suitable scheme, including schemes that assign these addresses inefficiently. In fact, different BDFsec spaces can assign addresses differently based on the endpoint types or routing facilities connected to the corresponding buses. Unlike the BDFsec spaces, the BDFpri space (i.e., the root complex's view of the configuration space) can be optimized to enable and control a compact and efficient assignment of bus addresses within the space (as in the example of [reference missing]). Fig. 6 illustrated).
[0047] Each BDFsec address in vSEG A and vSEG B can map to exactly one BDF address in the BDFpri space (e.g., according to the Fig. As an example, a first facility in vSEG A can be assigned the BDF "1:0:0" in the vSEG A BDFsec space, which is mapped to the primary BDF "4:0:0" (as in the Fig. (shown), among other examples. Different facilities in other BDFsec spaces of the system (e.g., vSEG B) can be assigned the same BDFsec values as those assigned in other BDFsec spaces (e.g., vSEG A). For example, a second facility in vSEG B can also be assigned the BDF "1:0:0", but in the BDFsec of vSEG B. However, the second facility will be mapped to a different BDF in the BDFpri of the system (i.e., BDF "7:0:1", as shown in the Fig. shown) and so on.
[0048] As noted above, in some implementations, mapping between BDFpri and BDFsec spaces can be achieved using mapping tables residing in system memory. Different packet types can be mapped differently (i.e., to mediate between BDFsec and BDFpri spaces). For bidirectional requests, for example, the appropriate requester ID can be remapped (e.g., according to the appropriate mapping table). Message requests routed by configuration and ID can be routed using ID bus / facility / function number fields. Bidirectional terminations can be routed using the requester and termination IDs, among other examples.
[0049] The MPB's mapping hardware can access mapping tables located in system memory, with optional caching within the MPB. In some implementations, a single mapping table can be used for traffic in both directions, with the MPB possessing logic to determine the mapping in both the forward (e.g., downstream) and reverse (e.g., upstream) directions (e.g., through reverse lookups). In other implementations, it may be more efficient to provide two separate mapping tables per MPB, one for the forward direction and the other for the reverse direction. This may, in some cases, be less resource-intensive than providing MPB hardware to perform a reverse lookup in either direction.
[0050] An MPB on the primary side (e.g., the port closest to the root complex) can be responsible for mapping a subset of the bus numbers in the BDFpri space to a corresponding BDFsec space. Accordingly, the range of bus numbers in the BDFpri space allocated to the MPB may be restricted to the range indicated by [secondary bus number to child bus number], since only packets in this range are routed to that particular MPB in the BDFpri space. Therefore, the BDFsec:BDFpri mapping table may involve a translation table large enough to cover the secondary to child range of bus numbers, but not necessarily larger. In some embodiments, a default table with 64K entries may be provided for ease of implementation.In fact, a 64K entry translation table can be provided to map BDFpri to BDFsec, making the full BDF space available on the secondary side, but this may be limited in some alternatives to reduce hardware / memory requirements.
[0051] Mapping tables can be maintained by system software that manages data communication in PCIe (or other connections implementing these functions). For example, in implementations that use two separate upstream and downstream tables, the two tables can be maintained by the system software. The system software can ensure consistency between the tables, so that a mapping from BDFx → BDFy → BDFx' always yields the result x = x', among other exemplary considerations.
[0052] In some implementations, an MPB can implement a cache to store at least a portion of the mapping tables locally on the MPB. The system software can assume that MPB cache translations are based on the mapping tables and can support cache management by providing the necessary cache management information to the hardware. It may be desirable to provide mechanisms for system software to ensure that the MPB cache is updated only under the control of the system software. Furthermore, a mechanism can be defined such as a specific register or registers in the space mapped in MPB memory (MMIO space) to provide pointers to the mapping tables in system memory.Registers can also be used for cache management, allowing system software to enable / disable caching, invalidate the MPB cache, and so on. Alternatively, in some implementations, the tables can be implemented directly in the MPB without maintaining copies in system memory. In this case, the system software updates the tables directly on the MPB as needed.
[0053] An additional advantage of the MPB mapping table mechanism is that system software can be executed to update the mapping tables atomically, for example, by creating new mapping tables and then "immediately" invalidating the MPB cache and redirecting the MPB from the old to the new tables. This can be achieved by defining the control register mechanism so that the MPB hardware must sample the register settings when indicated by the system software and then continue working with the sampled settings until instructed to sample again. This re-sampling can be performed by MPB hardware in such a way that the transition from the old to the new sample is atomically effective. By also providing a mechanism for system software to temporarily block traffic through the MPB (e.g., by...), the mechanism can be further enhanced by...To "pause" traffic at the MPB (to allow changes to the mapping tables), it becomes much more attractive to allow the system software to modify ("reshuffle") the bus number assignments in a running system, as the time required to switch the mapping tables can be kept very short. Alternatively, if the mapping tables are maintained directly in the MPB, a mechanism such as double buffering can be used, which allows the MPB to operate using a copy of the tables while an alternative set is being updated by the system software, and then, under the guidance of the system software, the MPB switches to operating with the updated tables (e.g., while local tables are being replaced with the updated versions).
[0054] Now on Fig. 9. Referring to this, a representation of an exemplary implementation of the register fields and bits for use in implementing the hardware / software interface for an exemplary MPB is shown. More specifically, in this particular example, an extended PCle capability for discovery / management in an MPB system can be defined. The extended capability can include fields such as: an outstanding requests (OR) field, which reflects the MPB's number of unsent (NP) requests in either direction (although the unsent requests would not be tracked individually or specifically); an MPB enable bit (E) to indicate whether the MPB functionality is enabled on a bridge that supports MPB (where address registers are to be sampled by the MPB when the E bit is set (0 → 1)); and a mapping table (TR) field, among other examples.In some cases, additional fields can be used in conjunction with a windowing mechanism to generate "sample" configRequests behind the MPB, such as a status bit so that the system software can understand when configWrites are complete. In some cases, the windowing mechanism may only support one configRequest at a time (e.g., no requests in the pipeline). In still other examples, the capability structure may provide values to describe the type of local translation tables maintained at an MPB (e.g., to indicate the table size if it is not full size, etc.), among other examples.
[0055] The primary / secondary BDF mapping and MPBs can be used in conjunction with segments in some implementations. As noted above, MPBs can be used to implement virtual segments (vSEGs) to at least partially replace and reduce the number of segments in a system design. In cases where segments need to be included, or when a system with multiple roots is created (e.g., using a proprietary load / store framing), MPBs can be extended to support mapping between different segments and between parts of a hierarchy that do not support segments. For example, in a system where BDFpri is extended to support segments, it effectively becomes a segment's BDFpri space (SBDFpri space), since the segment acts as a "prefix" to increase the number of available bus numbers.In such a system, a mechanism like a TLP prefix could be used to identify specific segments. However, since many facilities do not yet support a segment tagging mechanism, the MPB mapping mechanism can be used with extensions to support mapping a BDFsec space, which does not support segment tagging, to an SBDFpri space that does, among other examples.
[0056] Sampling can be included by MPBs in some implementations. Here, "sampling" can refer to reading the first data word (DW) of a function's configuration space to see if a valid vendor / facility ID exists. However, using the tables in memory during sampling can involve repeated updates, which can be particularly problematic due to repeated updating of the mapping tables if sampling occurs at runtime. To avoid this, a mechanism can be provided for the MPB to generate configuration requests for sampling. This mechanism may not be intended for use as the normal configuration generation path but can instead be configured for sampling (and in some cases, as a failsafe). In some examples, the mechanism might include a windowing mechanism within the extended MPB capability (e.g.,(Based on the CFC / CF8 configuration access mechanism defined for PCI, with support for a 4K configuration space). The system software can use this mechanism to discover a function, for example, by sampling. After discovering a function, the system software can then update the mapping tables to provide a translation for the function and proceed with further numbering / configuration through the MPB's table-based translation mechanism, among other exemplary functions.
[0057] System software can number the BDFsec space individually using any suitable or conventional algorithm. In some cases, before a specific BDF can be initially checked on the secondary side, the numberer configures the MPB to assign a "key" to BDFpri. This key can be used to generate and remap the configuration requirements to that specific BDF. The numberer must "understand" the BDFpri:BDFsec mapping during this initial numbering and can provide services to all modules using the BDF to perform this mapping during system operation. For example, a system software algorithm can be used to number PCI facilities under a mapping portal bridge (MPB).
[0058] The limitations of the traditional PCI numbering algorithm in resource allocation scalability are clearly evident in use cases for PCI-based SSDs (for example, NVMe-based) and in Thunderbolt hierarchies, where these limitations prevent the configuration of large / deep hierarchies. This leads to errors, and users cannot utilize facilities in such configurations. For example, there might be situations where more than 256 PCI-based solid-state drives need to be connected to a single system, or there might be on-the-fly installations, such as those for Thunderbolt, where the resources reserved for a particular section of the tree are insufficient to configure the Thunderbolt facilities installed while the system is running.In these scenarios, the conventional PCI numbering algorithm rapidly depletes scarce BDF resources due to allocation mechanisms such as evenly dividing and assigning available bus numbers among all connectors suitable for hot-swapping. The conventional approach of using redistribution to reallocate resources, as discussed above, is unsuitable for many use cases. The new algorithm proposed in this disclosure addresses these limitations related to MPB to significantly improve resource configuration and scalability in such scenarios.
[0059] Conventional resource numbering algorithms may be unable to number addresses (e.g., BDFs) for facilities located under a mapping portal bridge (MPB) because these algorithms lack insight into how to properly map these facilities between the secondary and primary BDF spaces. Traditionally, PCI facilities are numbered by the system software that samples the PCI bus using bus facility function (BDF) numbers, starting with BDF[0,0,0] through BDF[255,31,7]. In conventional numbering techniques, for each BDF that might correspond to a PCI function, the system software creates a configuration read transaction to read the vendor and facility ID of that specific function. A valid vendor and facility ID from the configuration read operation can indicate the presence of a function at that bus facility function (BDF).It may be necessary for each PCI device to implement Function 0. Therefore, the system software may have the freedom to skip the device's function number if Function 0 is not implemented. This conventional numbering algorithm does not work for devices due to a potential remapping of bus, device, and function numbers that exist under a mapping portal bridge (MPB). Accordingly, an improved numbering algorithm can be provided that the system software can use to detect the presence of an MPB and number devices located under the MPB.
[0060] As noted above, each logical function in a PCI subsystem can be identified by a BDF triad consisting of Bus (0-255), Facility (0-31), and Function Number (0-7). The BDF is a form of address that identifies each logical function within a PCI system. A PCI system, or "subsystem" (e.g., a subsystem of a wider system that includes non-PCI subsystems), can have 256 bus numbers within a segment. Each bus in the PCI subsystem can have 32 facilities, and each facility can have 8 functions. Traditionally, system software samples these bus, facility, and function numbers to number these facilities. Numbering can be performed by selecting a specific BDF and reading the vendor / facility ID, as discussed above.Based on the results of the vendor / facility ID read operation, the system software can record the results and perform additional tasks in the configuration space based on the requirements of the specific facility / function and the policies established for that particular system. For each BDF combination, the system software creates a configuration read transaction to read the vendor and facility IDs of that function. A valid response to this configuration read transaction indicates the presence of a function at that BDF. If no valid response is received, the system software can record this, and that BDF cannot be used (unless a facility is added and uses that BDF at a later time, such as during a hot-supply installation).
[0061] The PCI protocol uses a windowing mechanism to forward transactions from the primary side of a bridge to the secondary side of a bridge. Because of this windowing mechanism, the conventional software algorithm used for facility numbering must allocate enough bus numbers for all bridges suitable for on-the-fly installation, assuming that further hierarchies will be attached to the bridge suitable for on-the-fly installation. This exposes the problem of scaling scarce BDF resources.A secondary BDF space can be used to allow pre-allocation of parts of a BDF space without consuming resources in the primary BDF space (the space used by and natively referenced by the root complex) until the time when the resources are actually needed (whereby the corresponding mapping to the primary BDF space can be set up by the system software at that time).
[0062] The Mapping Portal Bridge (MPB) solves this problem of static (during startup) and dynamic (during installation and removal during operation) resource allocation of bus, equipment, and function by creating new bus, equipment, and function hierarchies that begin with BDF [0, 0, 0] in each hierarchy. Creating and supporting new BDF spaces (or hierarchies or views of the configuration space) complicates the numbering of these equipment within these hierarchies. For example, the system software might number a PCI bridge that has MPB capability using the conventional numbering method (as any other conventional equipment would be numbered and used within the same BDF space by the root complex). However, if the MPB capability is subsequently enabled (e.g.,In some cases, the BDF numbers generated by the system software to scan facilities below the MPB-enabled bridge may be invalid for facilities below the Mapping Portal Bridge (MPB) and may need to be renumbered (which involves building a corresponding mapping table). Indeed, because the MPB function uses a mapping table to map the BDF view of the primary side (the hierarchy above the MPB) to the BDF view of the secondary side (the hierarchy below the MPB), in some cases the entries in this mapping table are invalid, for example, after a system restart, until they are initialized by software. Each entry in the mapping table can be either valid or invalid, and only the entries explicitly initialized by the system software can be marked as "valid" mappings.Therefore, the BDF numbers of the primary side may not map correctly to the BDF numbers of the secondary side. Accordingly, an improved numbering algorithm can be provided to populate this mapping table and, furthermore, to support the numbering of facilities under the MPB.
[0063] In one implementation, an improved algorithm for numbering PCI facilities under a mapping portal bridge (MPB) can support the speculative creation of a mapping between a primary and a secondary BDF space. The mapping can be speculative in this context because the facilities / functions actually present in the system under the MPB may not yet have been discovered, but the system nevertheless configures the MPB to provide mappings of facilities / functions that might be discovered later in the numbering process. If these mappings are indeed used, they would typically be retained, whereas the unused ones could be "reused" by the system at some point, e.g.,BDFs mapped to one port of a switch can be reassigned to make them available to another port on the switch if hardware is installed under that second port during operation and requires more BDFs than originally allocated. Mapping gantry bridge (MPB) logic can create a new (e.g., secondary) hierarchy of bus device function numbers starting with BDF [0, 0, 0]. This is achieved by creating a mapping table that maps the BDFs of the primary side (the hierarchy above the MPB) to the BDFs of the secondary side (the hierarchy below the MPB). The algorithm uses a speculative approach to create mappings in the mapping gantry bridge's BDF mapping table. After creating these mappings, the algorithm uses the conventional PCI numbering algorithm to number PCI devices below the mapping gantry bridge (MPB).
[0064] In one example, an improved algorithm can build upon the principles of a conventional numbering algorithm and allow facilities under the MPB to be numbered in the same way as is conventionally done at other ports that do not use a secondary BDF space (e.g., by creating configuration transactions to read provider and facility identifiers from the configuration space, performing a tree search of the BDF space under the MPB), thus making the algorithm backward compatible. As mentioned above, the improved algorithm can allow the mapping table to be efficiently populated using the minimum number of mapping table entries to support a given size of hierarchy under the MPB, while optimizing / maximizing BDF usage on the primary side.The algorithm uses simple data structures, which makes its implementation in system software direct and easy.
[0065] In an implementation example, such as the one in the simplified block diagrams 1000a-c of the Fig. As illustrated in Figures 10A-10C, an improved algorithm for use in numbering facilities in a secondary BDF space can employ data structures that include, for example: • A global bus number inventory 1005 (bit vector 0-255) - This is used to maintain the bus number assignment history on the primary side of the MPB. • A numbering queue 1010 - This is used to hold bus numbers from the secondary side (e.g., based on a breadth-search algorithm).
[0066] As an illustrative example, the system software can begin numbering PCI devices connected to ports of a 615 root complex using the conventional algorithm. If the system software detects a Type 1 device (e.g., a PCI bridge) with MPB functionality, and if the system software intends to enable the MPB functionality, it can switch to using the improved numbering algorithm to manage a hierarchy of devices under the MPB. The numbering queue 1010 is used to track which bus number on the secondary side is to be sampled next. This may involve the system software first enabling the MPB functionality on the detected Type 1 device. Furthermore, the system software can create an empty mapping table (stored, for example, in system memory or the memory of a port of the MPB). Since the mapping table (e.g.,If 1015) is initially empty, the conventional numbering algorithm does not work on facilities (e.g., 630) under the MPB bridge, because no configuration requests flow from the primary to the secondary side of the MPB if there is no valid mapping in the MPB that maps the target BDF in the BDFpri space to the BDFsec space.
[0067] In one example, the BDF space of a hierarchy on the secondary side under a mapping portal bridge (or the view of the configuration space within the hierarchy space of the secondary side) always starts with bus number 0. In such an implementation, the improved algorithm can be used to add bus number 0 to the numbering queue 1010 (as in Fig. (shown in Figure 10B) to reflect the use of bus number 0 within the hierarchy space of the secondary side. The improved algorithm can remove the bus number from the numbering queue to indicate that this bus number will be sampled next on the secondary side. To perform the sampling, a speculative mapping is created between the next available bus number on the primary side and the bus number removed from the queue on the secondary side. Accordingly, for each bus number removed from the numbering queue, the algorithm retrieves the next available bus number from the global bus number pool. It then creates a speculative mapping between the next available bus number from the global bus number pool and the (or, respectively, the bus number removed from the queue)the possibly multiple) bus number(s) removed from the queue on the secondary side, assuming that facilities are found under the MPB bridge (as shown in Figure Table 1015 of . Fig. 10C shown).
[0068] For each speculative map, configuration read transactions can be created for all facility and function numbers within that speculative map to convert primary BDFs to secondary BDFs and read the facility and vendor IDs using the conventional numbering algorithm. Similarly, completing the configuration read operation can involve using the map to translate the BDF back from the BDFsec to the BDFpri. If a Type 1 facility is discovered during this operation, its secondary and child bus numbers are assigned, and the secondary bus number is placed in the numbering queue 1010. If a Type 0 facility is found, the required resource for it is assigned.
[0069] To illustrate an example, the following pseudocode represents an implementation example of an improved algorithm.
[0070] In some implementations, the system software can use "don't care" bits to reduce the number of entries in the mapping table. For example, if a number of BDFs are mapped under an MPB, a mapping can be constructed that applies only to some of the bits, instead of mapping each BDF individually. This implicitly ensures that the unmapped bits pass through the MPB unchanged. In such implementations, the system software can create speculative mappings for each bus setup function number combination that falls outside the "don't care" bitmask. For each speculative mapping, the system software then uses the conventional numbering algorithm for all bus setup function number combinations that fall within the "don't care" bitmask.
[0071] In one implementation example, the "Don't Care" bits can be set to 4. In this specific case, the system software can number two facility numbers (each with 8 functions) for each mapping. If the "Don't Care" bits are exhausted, new mappings are created. More rigorous implementations can be introduced to reuse the mapping entry if no facilities are found under the previously established speculative mapping. Furthermore, multiple bus numbers on the secondary side can be mapped under the same bus number on the primary side. These techniques can be used to maximize BDF usage on the primary side while keeping the mapping table concise.
[0072] On Fig. With reference to 11, a simplified flowchart 1100 is shown, illustrating an exemplary technique for numbering facilities within a system. A first facility can be identified at a first port of a root complex 1105. The first facility can be part of a first hierarchy. The first facility can be numbered (or "assigned") addresses 1110 according to a first or primary view of a configuration space (which contains other facilities in the hierarchy connected to the first port). A second facility can be identified as being connected to a second port of the root complex via a mapping portal bridge 1115. The mapping portal bridge can support addressing a second hierarchy of facilities connected below the mapping portal bridge according to a second view of the configuration space.In response to the detection of the mapping portal bridge connection, a mapping table can be created (1120) to map addresses in the secondary space to addresses in the primary space. The addresses of the second facility hierarchy can then be assigned using the mapping table (1125) (where, for example, the root complex and / or the system software translate address numbering between addresses in the primary and secondary address spaces when performing configuration tasks).
[0073] It should be noted that the devices, methods, and systems described above can be implemented in any electronic device or system. As specific illustrations, the figures below provide exemplary systems for implementing the invention as described herein. As the systems are described in more detail below, a number of different connections are disclosed, described, and reconsidered from the above discussion. And, as is evident, the advances described above can be applied to any of these connections, structures, or architectures.
[0074] Now on Fig. 12 With reference to this, an embodiment of a block diagram for a computing system containing a multi-core processor is shown. A Processor 1200 contains any processor or processing device, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a portable processor, an application processor, a coprocessor, a system-on-a-chip (SoC), or any other device for executing code. In one embodiment, the Processor 1200 contains at least two cores—cores 1201 and 1202—which may be asymmetric cores or symmetric cores (the illustrated embodiment). However, the Processor 1200 may contain any number of processing elements, which may be symmetric or asymmetric.
[0075] In one embodiment, a processing element refers to hardware or logic used to support a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a thread, a process unit, a context, a context unit, a logic processor, a hardware thread, a core, and / or any other element that can hold state for a processor, such as an execution state or architectural state. In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code, such as a software thread, an operating system, an application, or other code. A physical processor (or processor socket) typically refers to an integrated circuit, which may contain any number of other processing elements, such as cores or hardware threads.
[0076] A kernel often refers to logic residing on an integrated circuit that is capable of maintaining independent architectural states, each independently held architectural state associated with at least some dedicated execution resources. In contrast to kernels, a hardware thread typically refers to any logic residing on an integrated circuit that is capable of maintaining independent architectural states, where the independently held architectural states share access to execution resources. As can be seen, the boundary between the nomenclature of a hardware thread and a kernel overlaps when certain resources are shared and others are allocated to an architectural state.Nevertheless, a core and a hardware thread are often regarded by an operating system as individual logical processors, with the operating system being able to schedule operations on each logical processor individually.
[0077] The physical processor 1200, as in Fig. Figure 12 illustrates a configuration with two cores—cores 1201 and 1202. Here, cores 1201 and 1202 are considered symmetric cores, meaning cores with the same configurations, functional units, and / or logic. In another embodiment, core 1201 contains an out-of-order processor core, while core 1202 contains an in-order processor core. However, cores 1201 and 1202 can be individually selected from any core type, such as a native core, a software-managed core, a core adapted to run a native instruction set architecture (ISA), a core adapted to run a translated instruction set architecture (ISA), a code-designed core, or any other known core. In a heterogeneous core environment (i.e., asymmetric cores), some form of translation, such as binary translation, can be used to schedule or execute code on one or both cores.To continue the discussion, the functional units illustrated in core 1201 are discussed in more detail below, since the units in the embodiment shown in core 1202 operate in a similar manner.
[0078] As shown, the 1201 core contains two hardware threads, 1201a and 1201b, which can also be referred to as hardware thread slots 1201a and 1201b. Therefore, in one embodiment, software entities such as an operating system may view the 1200 processor as four separate processors—that is, four logical processors or processing units capable of executing four software threads concurrently. As mentioned above, a first thread is associated with architecture state registers 1201a, a second thread is associated with architecture state registers 1201b, a third thread may be associated with architecture state registers 1202a, and a fourth thread may be associated with architecture state registers 1202b. Here, each of the architecture state registers (1201a, 1201b, 1202a, and 1202b) can be referred to as a processing unit, thread slot, or thread unit, as described above.As illustrated, the architecture state register 1201a is reproduced in the architecture state register 1201b; therefore, individual architecture states / contexts can be stored for logical processor 1201a and logical processor 1201b. In core 1201, other smaller resources, such as instruction pointers and renaming logic in the allocation and renaming block 1230, can also be reproduced for threads 1201a and 1201b. Some resources, such as reorder buffers in a reorder / standby unit 1235, ILTB 1220, load / store buffers, and queues, can be shared through partitioning. Other resources such as general-purpose internal registers, page table base register(s), low-level data cache and data TLB 1215, execution unit(s) 1240 and parts of the out-of-order unit 1235 may be fully shared.
[0079] The 1200 processor often contains other resources that are fully shared, divided through partitioning, or dedicated to / for processing elements. Fig. Figure 12 illustrates an embodiment of a purely exemplary processor with illustrative logic units / resources of a processor. It should be noted that a processor may include or omit any of these functional units, as well as any other functional units, logic, or firmware not shown. As illustrated, core 1201 contains a simplified, representative out-of-order (OOO) processor core. However, in other embodiments, an in-order processor may be used. The OOO core contains a branch target buffer 1220 to predict branches to be executed / taken and an instruction translation buffer (1-TLB) 1220 to store address translation entries for instructions.
[0080] The 1201 core also contains the 1225 decode module, which is coupled to the 1220 retrieval unit to decode retrieved elements. In one embodiment, the retrieval logic includes individual flow controllers associated with the 1201a and 1201b thread slots, respectively. Typically, the 1201 core is associated with a first ISA that defines instructions executable on the 1200 processor. Often, machine code instructions that are part of the first ISA contain a section of the instruction (called an opcode) that references an instruction or operation to be performed. The 1225 decode logic contains circuitry that recognizes these instructions from their opcodes and forwards the decoded instructions in the pipeline for processing, as defined by the first ISA.As discussed in more detail below, the 1225 decoders, in one embodiment, contain logic, constructed or adapted, to recognize specific instructions, such as transaction instructions. As a result of recognition by the 1225 decoders, the 1201 architecture or core takes certain predefined actions to perform tasks associated with the respective instruction. It should be noted that any of the tasks, blocks, operations, and procedures described herein can be performed in response to one or more instructions; some of which may be new or old instructions. Note that, in one embodiment, the 1226 decoders recognize the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, the 1226 decoders recognize a second ISA (either a subset of the first ISA or a different ISA).
[0081] In one example, the allocation and rename block 1230 contains an allocation unit to reserve resources, such as register files, to store the results of instruction processing. Threads 1201a and 1201b may be capable of out-of-order execution, in which case allocation and rename block 1230 also reserves other resources, such as reorder buffers, to track instruction results. Unit 1230 may also contain a register rename unit to rename program / instruction reference registers to other registers internal to the processor 1200. The reorder / standby unit 1235 contains components, such as the reorder buffers, load buffers, and memory buffers mentioned above, to support out-of-order execution and subsequent in-order standby of instructions that are not executed sequentially.
[0082] In one embodiment, the planning and execution unit(s) 1240 includes a planning unit for scheduling instructions / operations on execution units. For example, a floating-point instruction is scheduled on a port of an execution unit that has an available floating-point execution unit. Register files associated with these execution units are also included for storing the results of the information instruction processing. Exemplary execution units include a floating-point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a memory execution unit, and other known execution units.
[0083] A low-level data buffer and a data translation buffer (D-TLB) 1250 are coupled to the execution unit(s) 1240. The data buffer stores recently used / modified elements, such as data operands, which may be held in memory coherence states. The D-TLB stores new address translations from virtual / linear to physical. As a specific example, a processor may contain a page table structure to partition physical memory into a multitude of virtual pages.
[0084] Here, cores 1201 and 1202 share access to a higher-level or more distant cache, such as a second-level cache associated with the on-chip interface 1210. It is important to note that the terms "higher-level" or "more distant" refer to cache levels that are progressively higher or further away from the execution unit(s). In one embodiment, a higher-level cache is a last-level data cache—the last cache in the memory hierarchy on the 1200 processor—like a second- or third-level data cache. However, a higher-level cache is not so restricted, as it can be associated with or contain an instruction cache.A trace cache—a type of instruction cache—can instead be coupled after decoder 1225 to store recently decoded traces. Here, an instruction might refer to a macro instruction (i.e., a general instruction recognized by the decoders) that can be decoded into a number of micro instructions (micro-operations).
[0085] In the configuration shown, the 1200 processor also includes an on-chip interface module 1210. Historically, a memory controller, described in more detail below, was incorporated into a computing system external to the 1200 processor. In this scenario, the on-chip interface 1210 communicates with external devices to the 1200 processor, such as the system memory 1275, a chipset (which often includes a memory control node to connect to the 1275 memory and an I / O control node to connect to peripheral devices), a memory control node, a northbridge, or another integrated circuit. And in this scenario, the bus 1205 can be any known connection, such as a multipoint bus, a point-to-point connection, a serial connection, a parallel bus, a coherent (e.g., a 5G bus), or a 1G bus.It includes a (with intermediate storage coherent) bus, a layered protocol architecture, a differential bus and a GTL bus.
[0086] The memory 1275 can be reserved for the processor 1200 or shared with other devices in a system. Common examples of memory types 1275 include DRAM, SRAM, non-volatile memory (NV memory), and other known storage devices. It should be noted that the device 1280 can include a graphics accelerator, a processor or card coupled to a memory control node, a data storage device coupled to an I / O control node, a wireless transceiver, a flash device, an audio control device, a network control device, or other known device.
[0087] However, it is now possible to integrate any of these facilities onto the 1200 processor, as more logic and facilities are integrated onto a single chip, such as a SoC. In one embodiment, for example, a memory control node is located on the same package and / or chip as the 1200 processor. Here, a section of the core (an in-core section) 1210 contains one or more controllers for coupling to other facilities, such as the 1275 memory or a 1280 graphics facility. The configuration that includes a connection and controllers for coupling to such facilities is often referred to as an in-core (or coreless) configuration. As an example, the in-chip interface 1210 includes a ring connection for in-chip communication and a high-speed serial point-to-point link 1205 for off-chip communication.In the SOC environment, even more facilities, such as the network interface, coprocessors, the 1275 main memory, the 1280 graphics processor, and any other known computer facilities / interfaces, can be integrated onto a single chip or integrated circuit to provide a small form factor with high functionality and low power consumption.
[0088] In one embodiment, the processor 1200 is capable of executing compiler, optimization, and / or translation code 1277 to compile, translate, and / or optimize application code 1276 to support or couple to the device and methods described herein. A compiler often contains a program or set of programs to translate source code into target code. Typically, compiling program / application code with a compiler involves multiple stages and passes to convert code in a high-level programming language into code in a low-level machine or assembly language. However, single-pass compilers can still be used for simple compilation.A compiler can use any known compilation techniques and perform any known compiler operations, such as lexical analysis, preprocessing, parsing, semantic analysis, code generation, code transformation, and code optimization.
[0089] Larger compilers often contain multiple phases, but these phases are most commonly contained in two general stages: (1) a front end, which generally involves syntactic processing, semantic processing, and some transformation / optimization operations, and (2) a back end, which generally involves analysis, transformations, optimizations, and code generation. Some compilers refer to a middle ground, illustrating a blurring of the boundary between a compiler's front end and back end. As a result, a reference to an insertion, association, generation, or other compiler operation can occur in any of the aforementioned phases or passes, as well as in any other known phases or passes of a compiler. As an illustrative example, a compiler might insert operations, calls, functions, and so on.Dynamic optimization can occur in one or more compilation phases, such as inserting calls / operations in a front-end stage of compilation and subsequently transforming these calls / operations into lower-level code during a transformation phase. It's worth noting that during dynamic compilation, the compiler code or dynamic optimization code can insert such operations / calls and optimize the code for execution at runtime. As a specific illustrative example, binary code (already compiled code) can be dynamically optimized at runtime. Here, the program code can contain the dynamic optimization code, the binary code, or a combination thereof.
[0090] Similar to a compiler, a translation unit, such as a binary translation unit, translates code either statically or dynamically to optimize and / or translate code. Therefore, a reference to code execution, application code, program code, or any other software environment can refer to: (1) either a dynamic or static execution of a compiler program or a binary translation unit.(1) Compiler programs, optimization code optimizers, or translation units for compiling program code, maintaining software structures, performing other operations, optimizing code, or translating code; (2) execution of main program code that includes operations / calls such as application code that has been optimized / compiled; (3) execution of other program code, such as libraries, associated with the main program code for maintaining software structures, performing other software-related operations, or optimizing code; or (4) a combination thereof.
[0091] Now on Fig. 13 With reference to this, a block diagram of a second system 1300 according to an embodiment of the present invention is shown. As in Fig. As shown in Figure 13, the multiprocessor system 1300 is a point-to-point interconnect system and includes a first processor 1370 and a second processor 1380, which are coupled via a point-to-point interconnect 1350. Each of the processors 1370 and 1380 can be a version of a processor. In one embodiment, 1352 and 1354 are part of a serial coherent point-to-point interconnect fabric, such as Intel's Quick-Path Interconnect (QPI) architecture. As a result, the invention can be implemented in the QPI architecture.
[0092] While the embodiment shown uses only two processors 1370 and 1380, it should be clear that the scope of the present invention is not limited thereto. In other embodiments, one or more additional processors may be present in a given processor.
[0093] The 1370 and 1380 processors are shown containing integrated memory control units 1372 and 1382, respectively. The 1370 processor also includes point-to-point (PP) interfaces 1376 and 1378 as part of its bus controller units; similarly, the 1380 processor includes PP interfaces 1386 and 1388. The 1370 and 1380 processors can exchange information via a point-to-point (PP) interface 1350 using the PP interface circuits 1378 and 1388. As shown in Fig. As shown in Figure 13, the IMCs 1372 and 1382 couple the processors to their respective main memory, namely a main memory 1332 and a main memory 1334, which can be parts of a main memory that are locally connected to the respective processors.
[0094] The 1370 and 1380 processors each exchange information with a 1390 chipset via individual PP interfaces 1352 and 1354, using point-to-point interface circuits 1376, 1394, 1386, and 1398. The 1390 chipset also exchanges information with a high-performance graphics circuit 1338 via an interface circuit 1392 along a high-performance graphics link 1339.
[0095] A shared cache (not shown) may be located inside one of the two processors or outside of both processors, but connected to the processors via a PP connection, so that when one processor enters a low-power mode, the local cache information from one or both processors can be stored in the shared cache.
[0096] The chipset 1390 can be coupled to a first bus 1316 via an interface 1396. In one embodiment, the first bus 1316 is a Peripheral Component Interconnect (PCI) bus or a bus such as a PCI Express bus or another third-generation I / O interconnect bus, although this does not limit the scope of the present invention.
[0097] As in Fig. Figure 13 shows various I / O devices 1314 coupled to the first bus 1316, along with a bus bridge 1318 that couples the first bus 1316 to a second bus 1320. In one embodiment, the second bus 1320 contains a low-pin-count (LPC) bus. Various devices are coupled to the second bus 1320, which includes, for example, a keyboard and / or mouse 1322, communication devices 1327, and a storage unit 1328, such as a disk drive or other mass storage device, which in one embodiment often contains instructions / code and data 1330. Furthermore, an audio I / O 1324 is shown coupled to the second bus 1320. It should be noted that other architectures are possible, with varying components and interconnection architectures. For example, instead of the point-to-point architecture of Fig. 13. Implement a multipoint bus or other such architecture.
[0098] While the present invention has been described with reference to a limited number of embodiments, those skilled in the art will appreciate the numerous modifications and variations thereof. It is intended that the appended claims cover all such modifications and variations that fall within the true spirit and scope of protection of this invention.
[0099] A design can go through several stages, from creation to simulation to fabrication. Data representing a design can represent it in a number of ways. First, the hardware can be represented using a hardware description language or another functional description language, as is useful in simulations. Additionally, a circuit-level model with logic and / or transistor gates can be fabricated at some stages of the design process. Furthermore, most designs reach a level of data at some stage that represents the physical placement of various components within the hardware model.In the case where conventional semiconductor manufacturing techniques are used, the data representing the hardware model can be data specifying the presence or absence of various features on different mask layers for masks used to fabricate the integrated circuit. In any design representation, the data can be stored on any form of machine-readable medium. Working memory or magnetic or optical storage, such as a disk, can be the machine-readable medium for storing information transmitted via an optical or electrical wave that is modulated or otherwise generated to transmit such information. When an electrical carrier wave is transmitted that indicates or carries the code or design, a new copy is made to the extent that copying, buffering, or retransmission of the electrical signal is performed.Therefore, a communications provider or network provider can at least temporarily store an object, such as information encoded in a carrier wave and employing techniques of embodiments of the present invention, on a tangible, machine-readable medium.
[0100] A module, as used herein, refers to any combination of hardware, software, and / or firmware. For example, a module includes hardware, such as a microcontroller, associated with a non-transient medium for storing code adapted to be executed by the microcontroller. Therefore, in one embodiment, a reference to a module refers to the hardware specifically configured to recognize and / or execute the code to be held on a non-transient medium. Furthermore, in another embodiment, the use of a module refers to the non-transient medium containing the code specifically adapted to be executed by the microcontroller to perform predetermined operations.And as can be deduced, in yet another embodiment, the term module (in this example) can refer to the combination of the microcontroller and the non-transient medium. Module boundaries often vary, frequently being illustrated as separate, and may overlap. For example, a first and a second module may share hardware, software, firmware, or a combination thereof, while possibly retaining some independent hardware, software, or firmware. In one embodiment, the use of the term logic includes hardware such as transistors, registers, or other hardware such as programmable logic devices.
[0101] The use of the phrase 'to' or 'configured to', in one embodiment, denotes arranging, assembling, manufacturing, offering for sale, importing, and / or constructing a device, hardware, logic, or element to perform an intended or predetermined task. In this example, a device or element thereof that is not operating is nevertheless 'configured' to perform an intended task if it is constructed, coupled, and / or connected to perform that intended task. As a purely illustrative example, a logic gate can provide a 0 or a 1 during operation. However, a logic gate that is 'configured' to provide an activation signal to a clock does not include every possible logic gate that can provide a 1 or a 0.Instead, the logic gate is one that is coupled in some way, so that during operation, the 1 or 0 output should activate the clock. It should be noted again that the use of the term 'configured to' does not require operation, but instead focuses on the latent state of a device, hardware, and / or element, where the device, hardware, and / or element is designed to perform a specific task in its latent state when the device, hardware, and / or element is operational.
[0102] Furthermore, the use of the terms 'capable of' and / or 'effective of' in an embodiment refers to a device, logic, hardware, and / or element that is designed to enable the use of the device, logic, hardware, and / or element in a predetermined manner. As noted above, the use of 'to,' 'capable of,' or 'effective of' in an embodiment refers to the latent state of a device, logic, hardware, and / or element, wherein the device, logic, hardware, and / or element is not in operation but is designed to enable the use of a device in a predetermined manner.
[0103] A value, as used herein, contains any known representation of a digit, a state, a logical state, or a binary logical state. Often, the use of logic levels, logic values, or logical values is also referred to as len and oen, which simply represent binary logical states. A 1, for example, denotes a high logic level, and a 0 denotes a low logic level. In one embodiment, a memory cell, such as a transistor or flash cell, may be capable of holding a single logical value or multiple logical values. However, other representations of values have been used in computer systems. The decimal number ten, for example, can also be represented as a binary value of 1010 and a hexadecimal letter A. Therefore, a value contains any representation of information that can be held in a computer system.
[0104] Furthermore, states can be represented by values or parts of values. For example, a first value, such as a logical one, can represent a default or initial state, while a second value, such as a logical zero, can represent a non-default state. Additionally, in one embodiment, the terms reset and set denote a default and an updated value, respectively. A default value, for instance, might contain a high logical value, i.e., reset, while an updated value might contain a low logical value, i.e., set. It should be noted that any combination of values can be used to represent any number of states.
[0105] The embodiments of methods, hardware, software, firmware, or code described above can be implemented by means of instructions or code stored on a machine-accessible, machine-readable, computer-accessible, or computer-readable medium and capable of being executed by a processing element. A non-transient machine-accessible / machine-readable medium contains any mechanism that provides (i.e., stores and / or transmits) information in a form readable by a machine, such as a computer or electronic system.A non-transient machine-accessible medium includes, for example, random access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; a magnetic or optical storage medium; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; another form of storage device for holding information received from transitory (propagated) signals (e.g., carrier waves, infrared signals, digital signals); etc., which are to be distinguished from the non-transient media that can receive information from them.
[0106] Instructions used to perform embodiments of the invention can be stored in a working memory within the system, such as DRAM, buffer memory, flash memory, or other storage medium. Furthermore, the instructions can be distributed over a network or by means of other computer-readable media. Therefore, a machine-readable medium can represent any mechanism for storing or transmitting information in a machine-readable format (e.g.,Computer-readable medium includes any type of tangible machine-readable medium capable of storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer), but is not limited to floppy disks, optical discs, write-protected compact discs (CD-ROMs) and magneto-optical disks, write-protected random access memory (ROMs), random access memory (RAM), erasable programmable write-protected memory (EPROM), electrically erasable programmable write-protected memory (EEPROM), magnetic or optical cards, flash memory, or any tangible machine-readable memory used in the transmission of information over the Internet via electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, computer-readable medium includes any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).can be read on a computer).
[0107] Example 1 discloses a method, a system and / or a machine-readable storage medium with executable code for determining that at least one first device is connected to a first of a plurality of ports of a root complex of a system, for assigning addresses corresponding to a first hierarchy of devices containing the first device, for determining that a second device is connected via a mapping portal bridge to a second of the plurality of ports of the root complex and that the second device is contained in another, second hierarchy of devices, and for triggering the creation of a mapping table corresponding to the mapping portal bridge.The mapping table defines a translation between an addressing used in a first view of a system configuration address space and an addressing used in a second view of the configuration address space, where the first view contains a view of the root complex and the second view contains a view corresponding to the second setup hierarchy, and the addresses assigned to the first setup hierarchy are assigned according to the first view.
[0108] In Example 2, the procedure, system, and medium from Example 1 can optionally also assign addresses for the second setup hierarchy that correspond to the first view of the configuration address space.
[0109] In Example 3, in the procedure, system and medium of one of Examples 1-2, each facility of the second facility hierarchy can optionally also be assigned a respective address according to the second view of the configuration address space.
[0110] In Example 4, in the procedure, system and medium, one of the addresses from Examples 1-3 in the first and second views of the configuration addresses can optionally be Bus Setup Function (BDF) numbers.
[0111] In Example 5, the addresses in the procedure, system and medium of Example 4 can optionally be assigned after the first view of the configuration address space in order to optimize the assignment of bus numbers used in the first view.
[0112] In Example 6, the addresses in the procedure, system, and medium of Example 5 are optionally assigned according to the second view of the configuration address space in accordance with a different, second address assignment scheme.
[0113] In Example 7, the second scheme in the procedure, system and medium of Example 6 can be agnostic to an optimization of the bus number assignment within the addresses of the second view.
[0114] In Example 8, the configuration address space in the method, system, and medium of Example 4 can optionally contain a PCIe configuration address space.
[0115] In Example 9, a first number of the bus numbers in the procedure, system and medium of Example 4 may optionally be allowed in the first view of the configuration address space, a second number of the bus numbers may optionally be assigned in the second view of the configuration address space, a third number of the bus numbers may optionally be assigned in the first view of the configuration address space, and the sum of the second and third numbers of the bus numbers may be greater than the first number.
[0116] In Example 10, the mapping portal bridge in the method, system and medium of one of Examples 1-9 can optionally be implemented in a switch facility that connects the facility hierarchy to the root complex.
[0117] In Example 11, the imaging portal bridge can optionally be implemented in the second connection in the method, system and medium of one of Examples 1-10.
[0118] In Example 12, the mapping portal bridge in the procedure, system and medium of one of Examples 1-11 uses the mapping table to support communication between the second facility hierarchy and the root complex.
[0119] In Example 13, facilities in the procedure, system and medium of one of Examples 1-12 are optionally discovered in the first and second facility hierarchy according to a respective search algorithm.
[0120] In Example 14, the search algorithm in the method, system and medium of Example 13 optionally includes a depth search.
[0121] In Example 15, the search algorithm in the method, system and medium of Example 13 optionally includes a breadth-first search.
[0122] In Example 16, the search algorithm in the procedure, system and medium of Example 13, which is used to discover facilities in the first hierarchy, is optionally different from the search algorithm used to discover facilities in the second hierarchy.
[0123] In Example 17, the search algorithm in the procedure, system and medium of Example 13, which is used to discover facilities in the first hierarchy, is the same as the search algorithm used to discover facilities in the second hierarchy.
[0124] In Example 18, in the procedure, system and medium of one of Examples 1-17, at least some of the addresses in the first view of the configuration address space may be reserved for installation during operation.
[0125] Example 19 discloses a system comprising a root complex containing a plurality of ports for coupling to a plurality of facility hierarchies, and system software. The system software can be executed by a processor to: determine that at least one first facility is connected to a first of the plurality of ports; assign addresses corresponding to a first hierarchy of facilities containing the first facility; determine that a second facility is connected via a mapping portal bridge to a second of the plurality of ports of the root complex, and that the second facility is contained in another, second hierarchy of facilities; and create a mapping table corresponding to the mapping portal bridge.The mapping table defines a translation between an addressing used in a first view of a system configuration address space and an addressing used in a second view of the configuration address space, where the first view contains a view of the root complex and the second view contains a view corresponding to the second setup hierarchy, and the addresses assigned to the first setup hierarchy are assigned according to the first view.
[0126] References to "a single embodiment" or "an embodiment" throughout this entire description mean that a particular feature, structure, or property described in connection with the embodiments is included in at least one embodiment of the present invention. Therefore, the appearance of the phrases "in a single embodiment" or "in an embodiment" at various points in this entire description does not necessarily always refer to the same embodiment. Furthermore, the individual features, structures, or properties may be combined in any suitable manner in one or more embodiments.Furthermore, the individual features, structures or properties can be combined in any suitable way in one or more embodiments.
Claims
[1] Machine-accessible storage medium or storage media containing code stored thereon, wherein the code, when executed on a machine, causes the machine to: determined that at least one first device is connected to a first of a plurality of ports of a root complex (615) of a system, assigns addresses that correspond to a first hierarchy (620) of facilities that include the first facility, determined that a second facility is connected via a mapping portal bridge (705) to a second of the multiple connections of the root complex (615) and that the second facility is contained in another, second hierarchy (625) of facilities, and triggers the creation of a mapping table (1015) corresponding to the mapping portal bridge (705), wherein the mapping table (1015) corresponds to a translation between an addressing used in a first view of a configuration address space of the system and an addressing used in a second view of the configuration address space, wherein the first view comprises a view of the root complex (615) and the second view comprises a view corresponding to the second setup hierarchy and the addresses (1110) assigned to the first setup hierarchy correspond to the first view, wherein the creation of the mapping table (1015) includes a speculative mapping between the addressing used in the first view of the configuration address space and the addressing used in the second view of the configuration address space, wherein the speculative mapping includes the following: Inserting a first value from a numbering queue (1010), where the first value represents an address of the second hierarchy (625) of facilities, Identification of a second value, where the second value represents an address of the second hierarchy (625) of facilities, by removal from the numbering queue (1010); and Mapping the second value to an address in the first hierarchy of institutions. [2] Storage medium according to claim 1, wherein the code is further executable to assign addresses for the second setup hierarchy (625) corresponding to the first view of the configuration address space. [3] Storage medium according to claim 2, wherein each device of the second device hierarchy (625) is also assigned a respective address according to the second view of the configuration address space. [4] Storage medium according to one of claims 1-3, wherein addresses in the first and second views of the configuration addresses each comprise respective Bus Setup Function (BDF) numbers. [5] Storage medium according to claim 4, wherein the addresses assigned according to the first view of the configuration address space are assigned to optimize an allocation of bus numbers used in the first view. [6] Storage medium according to claim 5, wherein the addresses assigned according to the second view of the configuration address space are assigned in accordance with another, second address assignment scheme. [7] Storage medium according to claim 6, wherein the second scheme is agnostic to an optimization of the bus number assignment within the addresses of the second view. [8] Storage medium according to claim 4, wherein the configuration address space comprises a PCIe configuration address space. [9] Storage medium according to claim 4, wherein a first number of the bus numbers is allowed in the first view of the configuration address space, a second number of the bus numbers is assigned in the second view of the configuration address space, a third number of the bus numbers is assigned in the first view of the configuration address space and a sum of the second and the third number of the bus numbers is greater than the first number. [10] Storage medium according to one of claims 1-9, wherein the imaging portal bridge (705) is implemented in a switch arrangement which connects the arrangement hierarchy to the root complex (615). [11] Storage medium according to one of claims 1-10, wherein the imaging portal bridge (705) is implemented in the second terminal. [12] Storage medium according to one of claims 1-11, wherein the mapping portal bridge (705) has to use the mapping table to support communication between the second setup hierarchy (625) and the root complex (615). [13] Storage medium according to one of claims 1-12, wherein the code is further executable to discover facilities in each of the first (620) and second facility hierarchy (625) according to a respective search algorithm. [14] Storage medium according to claim 13, wherein the search algorithm comprises a depth search, [15] Storage medium according to claim 13, wherein the search algorithm comprises a breadth-first search, [16] Storage medium according to claim 13, wherein the search algorithm used to discover facilities in the first hierarchy (620) is different from the search algorithm used to discover facilities in the second hierarchy (625). [17] Storage medium according to claim 13, wherein the search algorithm used to discover facilities in the first hierarchy (620) is the same as the search algorithm used to discover facilities in the second hierarchy (625). [18] Storage medium according to one of claims 1-17, wherein at least a part of the addresses in the first view of the configuration address space is reserved for installation during operation. [19] Procedures, including: Determine that at least one first device is connected to a first of a plurality of ports of a root complex (615) of a system, Assigning addresses that correspond to a first hierarchy (620) of establishments that include the first establishment, Determine that a second facility is connected via a mapping portal bridge (705) to a second of the multiple connections of the root complex (615) and that the second facility is contained in another, second hierarchy (625) of facilities, and Triggering the creation of a mapping table (1015) corresponding to the mapping portal bridge (705), wherein the mapping table (1015) corresponds to a translation between an addressing used in a first view of a configuration address space of the system and an addressing used in a second view of the configuration address space, wherein the first view comprises a view of the root complex (615) and the second view comprises a view corresponding to the second setup hierarchy and the addresses assigned to the first setup hierarchy of the first view, wherein the creation of the mapping table (1015) comprises a speculative mapping between the addressing used in the first view of the configuration address space of the system and the addressing used in the second view of the configuration address space of the system, wherein the speculative mapping comprises the following: Inserting a first value representing an address of the second hierarchy of facilities from a numbering queue (1010); Identification of a second value representing a second hierarchy (625) address of facilities by removing it from the numbering queue (1010); and Mapping the second value to an address of the first hierarchy (620) of institutions. [20] Method according to claim 19, wherein the code is further executable to assign addresses for the second setup hierarchy (625) corresponding to the first view of the configuration address space. [21] Method according to claim 20, wherein each device of the second device hierarchy (625) is also assigned a respective address according to the second view of the configuration address space. [22] Method according to one of claims 19-21, wherein addresses in the first and second views of the configuration addresses each comprise respective Bus Setup Function (BDF) numbers. [23] Method according to claim 22, wherein the addresses assigned according to the first view of the configuration address space are assigned to optimize an allocation of bus numbers used in the first view. [24] Method according to claim 23, wherein the addresses assigned according to the second view of the configuration address space are assigned in accordance with another, second address assignment scheme. [25] Method according to claim 24, wherein the second scheme is agnostic to an optimization of the bus number assignment within the addresses of the second view. [26] Method according to claim 22, wherein the configuration address space comprises a PCIe configuration address space. [27] Method according to claim 22, wherein a first number of the bus numbers is allowed in the first view of the configuration address space, a second number of the bus numbers is assigned in the second view of the configuration address space, a third number of the bus numbers is assigned in the first view of the configuration address space and a sum of the second and the third number of the bus numbers is greater than the first number. [28] Method according to one of claims 19-27, wherein the imaging portal bridge (705) is implemented in a switch device that connects the device hierarchy to the root complex (615). [29] Method according to one of claims 19-28, wherein the imaging portal bridge (705) is implemented in the second connection. [30] Method according to one of claims 19-29, wherein the mapping portal bridge (705) has to use the mapping table to support communication between the second facility hierarchy (625) and the root complex (615). [31] Method according to claims 19-30, wherein the code is further executable to discover facilities in each of the first and second facility hierarchies according to a respective search algorithm. [32] Method according to claim 31, wherein the search algorithm comprises a depth search. [33] Method according to claim 31, wherein the search algorithm comprises a breadth-first search. [34] Method according to claim 31, wherein the search algorithm used to discover facilities in the first hierarchy (620) is different from the search algorithm used to discover facilities in the second hierarchy (625). [35] Method according to claim 31, wherein the search algorithm used to discover facilities in the first hierarchy (620) is the same as the search algorithm used to discover facilities in the second hierarchy (625). [36] Method according to claim 19, wherein at least some of the addresses in the first view of the configuration address space are reserved for installation during operation. [37] System comprising means for carrying out the method according to any one of claims 19-36. [38] System, encompassing: a root complex (615) comprising a multitude of ports to couple to a multitude of hierarchies (620, 625) of facilities; System software executable by a processor (105, 1200) to: to determine that at least one first device is connected to a first of the multitude of connections, to assign addresses that correspond to a first hierarchy (620) of institutions that include the first institution, to determine that a second facility is connected via a mapping portal bridge (705) to a second of the multiple connections of the root complex (615) and that the second facility is contained in another, second hierarchy (625) of facilities, and to create a mapping table (1015) corresponding to the mapping portal bridge (705), wherein the mapping table (1015) corresponds to a translation between an addressing used in a first view of a configuration address space of the system and an addressing used in a second view of the configuration address space, wherein the first view comprises a view of the root complex and the second view comprises a view corresponding to the second setup hierarchy and the addresses assigned to the first setup hierarchy correspond to the first view; wherein the creation of the mapping table (1015) comprises a speculative mapping between the addressing used in the first view of the configuration address space of the system and the addressing used in the second view of the configuration address space of the system, wherein the speculative mapping comprises the following: Inserting a first value representing an address of the second hierarchy of facilities from a numbering queue (1010); Identification of a second value representing a second hierarchy (625) address of facilities by removing it from the numbering queue (1010); and Mapping the second value to an address of the first hierarchy (620) of institutions.
Citation Information
Patent Citations
I / O system and I / O control method
US20110219164A1