Sloping portal bridge.

The Flattening Portal Bridge (FPB) addresses inefficiencies in PCIe bus address space allocation by enabling efficient and dynamic re-allocation of resources, improving scalability and reducing system freezes during address space adjustments.

DE112017001148B4Active Publication Date: 2025-11-06INTEL CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE112017001148
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-09-30
Filing Date
2017-02-02
Publication Date
2025-11-06
Estimated Expiration
2037-02-02

AI Technical Summary

Technical Problem

Conventional PCI Express (PCIe) systems inefficiently utilize bus address space, leading to waste and inefficiencies in addressing components, particularly in deep hierarchies and dynamic usage scenarios where hot plugging occurs, resulting in slow renumbering processes that can freeze systems.

Method used

The implementation of a Flattening Portal Bridge (FPB) mechanism that allows for more efficient allocation and reallocation of bus and memory-mapped I/O addresses without disrupting ongoing operations, enabling dynamic renumbering and reducing resource waste.

Benefits of technology

FPB enhances scalability and efficiency in PCIe systems by allowing runtime re-allocation of resources, reducing the need for global balancing and minimizing system freezes during address space adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

At least one machine-accessible storage medium on which instructions are stored, wherein the instructions, when executed on a machine, cause the machine to do the following: Identifying a plurality of components in a system, wherein each component of the plurality of components in the system is connected by at least one of a plurality of logical buses; Assigning a respective address to each of the multitude of components, wherein assigning the address to a component includes the following: Determine whether the address is to be assigned according to a first addressing system or a second bus addressing system, wherein the first addressing system assigns a single bus number within a bus / component / puncture (BDF) address space to each component addressed in the first addressing system, and the second bus addressing system assigns a single bus component number within the BDF address space.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority over preliminary US patent application No. 62 / 303 487, filed on March 4, 2016. TECHNICAL AREA

[0002] This disclosure relates to a computing system and in particular (but not exclusively) address space mapping. STATE OF THE ART

[0003] Peripheral Component Interconnect (PCI) configuration space is used by systems employing PCI, PCI-X, and PCI Express (PCIe) to perform configuration tasks for PCI-based components. PCI-based components have an address space for component configuration registers, called the configuration space, and PCI Express introduces an extended configuration space for components. Configuration space registers are typically mapped from the host processor to memory-mapped input / output points. Device drivers, operating systems, and diagnostic software access the configuration space and can read and write information to configuration space registers.

[0004] One of the improvements the PCI Local Bus had over other I / O architectures was its configuration mechanism. In addition to the standard memory-mapped and I / O port spaces, each component function on the bus has a 256-byte configuration space, addressable by knowing the eight-bit PCI bus number, the five-bit component number, and the three-bit function number for the component (usually called BDF or B / D / F, short for Bus / Device / Function). This allows for up to 256 buses, each with a maximum of 32 components, each supporting eight functions. A single PCI expansion card can act as a single component and implement at least function number zero. The first 64 bytes of the configuration space are standardized; the remainder are available specification-defined extensions and / or reserved for vendor-defined purposes.

[0005] To allow more parts of the configuration space to be standardized without conflicting with existing uses, a list of capabilities can be defined within the upper 192 bytes of the Peripheral Component Interface (PCI) configuration space. Each capability has one byte describing its nature and one byte pointing to the next capability. The number of additional bytes depends on the capability ID. If capabilities are used, a bit is set in the status register, and a pointer to the first capability in a linked list is provided. Versions of PCIe have been provided with similar features, including an extended configuration space that increases the total configuration space size to 4096 bytes, and an extended PCIe capability structure.

[0006] US 2015 O 127 868 A1 relates to an information processing device comprising a processor, a plurality of devices, a plurality of bridges connecting the processor and the plurality of devices, a creator, and an assignor. The creator generates a management table that manages bridge addresses assigned to the plurality of bridges and, when an unassigned bridge address due to be assigned to a bridge of the plurality of bridges expires, cancels the assignment of the first bridge address of the bridge addresses assigned to another bridge of the plurality of bridges, reassigns the canceled first bridge address to the one bridge that was canceled, and updates the management table with respect to the canceled and reassigned first bridge address.When an access to one of the multitude of devices is detected, the assignor refers to the management table, performs the removal of the assignment of a second bridge address of the bridge addresses and the reassignment of the second bridge address to one or more of the multitude of bridges based on the referenced management table in order to enable the execution of the one access, and updates the management table with respect to the removed and reassigned second bridge address. SUMMARY

[0007] The solution according to the invention relates, according to main claim 1, to a machine-accessible storage medium on which instructions are stored. When executed on a machine, the instructions cause the machine to: identify a plurality of components in a system, wherein each component of the plurality of components in the system is connected by at least one of a plurality of logical buses; assign a respective address to each of the plurality of components. Assigning the address to a component comprises: determining whether the address is to be assigned according to a first addressing system or a second bus addressing system, wherein the first addressing system assigns a single bus number within a bus / component / puncture (BDF) address space to each component addressed in the first addressing system, and the second bus addressing system assigns a single bus component number within the BDF address space. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates an embodiment of a computer system that has a wiring architecture. Fig. Figure 2 illustrates an embodiment of a computer system that has a layered stack. Fig. Figure 3 illustrates an embodiment of a request or packet that is to be generated or received within a interconnection architecture. Fig. Figure 4 illustrates an embodiment of a transmitter-receiver pair for a circuit architecture. Fig. Figure 5 illustrates a representation of system buses. Fig. Figure 6 illustrates a representation of an enumeration of bus identifiers in a system. Fig. Figure 7A illustrates a representation of a system that uses instances of a flattening portal bridge (FPB). Fig.Figure 7B illustrates an exemplary implementation of an FPB. Fig. Figure 8 illustrates a detailed presentation of an exemplary FPB. Fig. Figure 9 illustrates exemplary addresses in BDF space and supported granularity. Fig. Figure 10 illustrates the layout of addresses in the memory address space below 4 GB for which the FPB MEM Low mechanism applies, and the effect of granularity on these addresses. Fig. Figure 11 illustrates an embodiment of a block diagram for a computer system that has a multicore processor. Fig. Figure 12 illustrates another embodiment of a block diagram for a computing system that includes a processor. DETAILED DESCRIPTION

[0008] The following description presents numerous specific details, such as examples of specific processor types and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific measurements / altitudes, specific processor pipeline stages and operation, etc., to provide a thorough understanding of the present invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to carry out the present invention.In other cases, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / code for described algorithms, specific firmware code, specific interconnection operation, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, specific expressions of algorithms in code, specific shutdown and gate techniques / logic, and other specific operational details of computer systems, were not described in detail in order to avoid unnecessary obfuscation of the present invention.

[0009] Although the following embodiments may be described with reference to energy conservation and energy efficiency in specific integrated circuits, such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings from embodiments described herein can be applied to other types of circuits or semiconductor devices that can also benefit from improved energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks™ and can also be used in other devices, such as handheld devices, tablets, other thin notebooks, system-on-a-chip (SoC) devices, and other embedded applications.Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, a network computer (NetPC), set-top boxes, network hubs, wide-area network (WAN) switches, or any other system capable of performing the functions and operations described below. Furthermore, the devices, procedures, and systems described here are not limited to physical computing equipment but may also include software optimizations for energy conservation and efficiency.

[0010] As computer systems evolve, their components become more complex. Consequently, the interconnect architecture for coupling and communicating between components becomes increasingly complex to ensure bandwidth requirements are met for optimal component operation. Furthermore, different market segments demand different aspects of interconnect architectures to suit market needs. For example, servers require higher performance, although the mobile ecosystem is sometimes able to sacrifice overall performance for power savings. Nevertheless, the overarching goal of most structures is to provide the highest possible performance with maximum power savings. The following section discusses a number of interconnects that would potentially benefit from aspects of the invention described herein.

[0011] A key interconnect architecture is the Peripheral Component Interconnect (PCI) Express (PCIe) architecture. A primary goal of PCIe is to enable components and devices from various manufacturers to interact within an open architecture spanning multiple market segments: clients (desktops and mobile), servers (standard and enterprise), and embedded and communications devices. PCI Express is a universal, high-performance I / O connection designed for a wide range of future computing and communications platforms. While some PCI attributes, such as the usage model, load memory architecture, and software interfaces, have been retained through revisions, earlier parallel bus implementations have been replaced by a highly scalable, fully serial interface.Newer versions of PCI Express leverage advancements in point-to-point interconnection, switch-based technology, and the packaged protocol to deliver new levels of performance and features. Power management, Quality of Service (QoS), hot-plug / hot-swap support, data integrity, and error handling are among the advanced features supported by PCI Express.

[0012] With reference to Fig.Figure 1 illustrates an embodiment of a structure consisting of point-to-point connections that interconnect a set of components. A system 100 includes a processor 105 and a system memory 110 coupled to a controller hub 115. The processor 105 includes any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or another processor. The processor 105 is coupled to the controller hub 115 by a front-side bus (FSB) 106. In one embodiment, the FSB 106 is a serial point-to-point interconnect, as described below. In another embodiment, the connection 106 has a serial differential interconnect architecture that conforms to different interconnect standards.

[0013] System memory 110 includes any storage device, such as random-access memory (RAM), non-volatile memory (NV memory), or other memory that devices in system 100 can access. System memory 110 is coupled to a controller hub 115 via memory interface 116. Examples of memory interfaces include a double-data-rate (DDR) memory interface, a dual-channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.

[0014] In one embodiment, the controller hub 115 is a root hub, root complex, or root controller in a Peripheral Component Interconnect Express (PCIe or PCIE) interconnect hierarchy. Examples of the controller hub 115 include a chipset, a memory controller hub (MCH), a northbridge, an interconnect controller hub (ICH), a southbridge, and a root controller / hub. Often, the term chipset refers to two physically separate controller hubs, that is, a memory controller hub (MCH) coupled to an interconnect controller hub (ICH). It should be noted that current systems often have the MCH integrated into the processor 105, while the controller 115 is intended to communicate with I / O devices in a manner similar to that described below. In some embodiments, peer-to-peer routing is optionally supported by the root complex 115.

[0015] Here, the controller hub 115 is coupled to a switch / bridge 120 via serial connection 119. Input / output modules 117 and 121, which can also be called interfaces / ports 117 and 121, comprise / implement a layered protocol stack to provide communication between the controller hub 115 and the switch 120. In one embodiment, multiple devices can be coupled to the switch 120.

[0016] The switch / bridge 120 routes packets / messages from the device 125 upstream, that is, a hierarchy upwards towards a root complex to the controller hub 115, and downstream, that is, a hierarchy downwards away from a root controller, from the processor 105 or system memory 110 to a device 125. In one embodiment, the switch 120 is called a logical arrangement of several virtual PCI-to-PCI bridge devices.Device 125 includes any internal or external device or component intended to connect to an electronic system, such as an I / O device, network interface controller (NIC), expansion card, audio processor, network processor, hard disk, storage device, CD / DVD-ROM drive, monitor, printer, mouse, keyboard, router, portable storage device, FireWire device, Universal Serial Bus (USB) device, scanner, and other input / output devices. Often, in PCIe terminology, such as device, it is referred to as an endpoint. Although not specifically shown, Device 125 may include a PCIe-to-PCI / PCI-X bridge to support established PCI devices or PCI devices of other configurations. Endpoint devices in PCIe are commonly classified as established PCIe or Root Complex Integrated Endpoints.

[0017] A graphics accelerator 130 is also coupled to the controller hub 115 via serial connection 132. In one embodiment, the graphics accelerator 130 is coupled to an MCH, which is coupled to an ICH. The switch 120, and consequently the I / O device 125, is then coupled to the ICH. I / O modules 131 and 118 are also intended to implement a layered protocol stack for communication between the graphics accelerator 130 and the controller hub 115. Similar to the MCH discussion above, a graphics controller or the graphics accelerator 130 itself can be integrated into the processor 105.

[0018] With reference to Fig.Figure 2 illustrates an embodiment of a layered protocol stack. A layered protocol stack 200 contains any form of layered communications stack, such as a Quick-Path Interconnect (QPI) stack, a PCle stack, a next-generation high-performance computer interconnect stack, or any other layered stack. Although the discussion immediately below refers to the Fig. Since sections 1 to 4 relate to a PCle stack, the same concepts are applicable to other interconnection stacks. In one embodiment, the protocol stack 200 is a PCle protocol stack comprising a transaction layer 205, a link layer 210, and a physical layer 220. An interface, such as interfaces 117, 118, 121, 122, 126, and 131 in Fig.1, can be represented as a communication protocol stack 200. The representation as a communication protocol stack can also be called a module or an interface that implements / has a protocol stack.

[0019] PCI Express uses packets to communicate information between components. Packets are formed in the transaction layer (205) and the data link layer (210) to carry information from the transmitting component to the receiving component. As the transmitted packets pass through the other layers, they are augmented with additional information necessary for handling packets at those layers. On the receiving side, the reverse process occurs, and packets are transformed from their physical layer (220) representation to the data link layer (210) representation and finally (for transaction layer packets) to a form that can be processed by the transaction layer (205) of the receiving device. Transaction layer

[0020] In one embodiment, the transaction layer 205 provides an interface between a component's processing core and the interconnect architecture, such as the data link layer 210 and the physical layer 220. In this respect, a primary responsibility of the transaction layer 205 is the assembly and disassembly of packages (i.e., transaction layer packages or TLPs). The transaction layer 205 typically manages credit-based flow control for TLPs. A PCIe implements split transactions, meaning transactions with a request and response separated by time, allowing one link to carry other traffic while the target component gathers data for the response.

[0021] Furthermore, PCIe employs credit-based guidance control. In this system, a component announces an initial credit amount for each of the receive buffers in transaction layer 205. An external component at the opposite end of the link, such as the controller hub 115 in Fig. 1. This counts the number of credits consumed by each TLP. A transaction can be transferred if it does not exceed a credit limit. Upon receiving a response, a credit amount is restored. One advantage of the credit system is that the latency of credit return does not affect performance, provided the credit limit is not reached.

[0022] In one embodiment, four transaction address spaces comprise a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions involve one or more read and write requests to transfer data to / from a memory-mapped location. In one embodiment, the memory space transactions are capable of using two different address formats, for example, a short address format such as a 32-bit address, or a long address format such as a 64-bit address. Configuration space transactions are used to access the configuration space of the PCIe components. Transactions to the configuration space involve read and write requests. Message transactions are defined to support in-band communication between PCIe agents.

[0023] In one embodiment, the transaction layer therefore assembles packet headers 205 and payload 156. The format for current packet headers / payloads can be found in the PCle specification on the PCle specification website.

[0024] With quick reference to Fig. Figure 3 illustrates an embodiment of a PCle transaction descriptor. In one embodiment, the transaction descriptor 300 is a mechanism for carrying transaction information. In this respect, the transaction descriptor 300 supports the identification of transactions in a system. Other potential uses include monitoring changes to default transaction orders and associations of transactions with channels.

[0025] The transaction descriptor 300 has a global identifier field 302, an attribute field 304, and a channel identifier field 306. In the illustrated example, a global identifier field 302 is shown, which includes a local transaction identifier field 308 and a source identifier field 310. In one embodiment, the global transaction identifier 302 is the same for all pending requests.

[0026] According to one implementation, the local transaction identifier field 308 is a field generated by a requesting agent and represents all pending requests that require completion for that requesting agent. Furthermore, in this example, the source identifier 310 uniquely identifies the requesting agent within a PCle hierarchy. Together with the source ID 310, the local transaction identifier field 308 thus provides global identification of a transaction within a hierarchy domain.

[0027] Attribute field 304 specifies characteristics and relationships of the transaction. In this respect, attribute field 304 is potentially used to provide additional information that allows for modification of the standard transaction processing. In one implementation, attribute field 304 includes a priority field 312, a reserved field 314, an ordering field 316, and a no-snoop field 318. Here, the priority subfield 312 can be modified by an initiator to assign a priority to the transaction. A reserved attribute field 314 is reserved for future use or for vendor-defined use. Possible usage models that utilize priority or security attributes can be implemented using the reserved attribute field.

[0028] In this example, the ordering attribute field 316 is used to provide optional information that conveys the ordering type, which can modify default ordering rules. According to one example implementation, an ordering attribute of "0" means that default ordering rules are applied, while an ordering attribute of "1" denotes looser ordering, where writes can propagate writes in the same direction and read completions can propagate writes in the same direction. The snooping attribute field 318 is used to determine whether transactions are snooped. As shown, the channel ID field 306 identifies a channel to which a transaction is associated. transmission layer

[0029] A transmission layer 210, also called data layer 210, acts as an intermediate layer between the transaction layer 205 and the physical layer 220. In one embodiment, the responsibility of the data layer 210 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components of a link. One side of the data layer 210 accepts TLPs assembled by the transaction layer 205, applies the single packet sequence identifier 211 (i.e., an identification number or packet number), calculates and applies an error detection code (CRC 212), and submits the modified TLPs to the physical layer 220 for transmission across a physical component to an external component. Physical layer (bit transmission layer)

[0030] In one embodiment, the physical layer 220 comprises a logical subblock 221 and an electrical subblock 222 for physically transferring a packet to an external component. Here, the logical subblock 221 is responsible for the "digital" functions of the physical layer 221. In this respect, a logical subblock comprises a transmit section to prepare outgoing information for transmission through the physical subblock 222, and a receive section to identify and prepare received information before passing it on to the transmission layer 210.

[0031] The physical block 222 comprises a transmitter and a receiver. The transmitter is supplied with symbols by the logical subblock 221, which the transmitter serializes and transmits to an external component. The receiver is supplied with serialized symbols from an external component and converts the received signals into a bitstream. The bitstream is deserialized and supplied to the logical subblock 221. In one embodiment, an 8b / 10b transmission code is used, in which ten-bit symbols are transmitted / received. Here, special symbols are used for framing a packet with frames 223. In one example, the receiver also provides a symbol clock that is recovered from the incoming serial stream.

[0032] As mentioned above, although transaction layer 205, transmission layer 110, and physical layer 220 are discussed in relation to a specific implementation of a PCIe protocol stack, a layered protocol stack is not so restricted. Any layered protocol can be included / implemented. As an example, a port / interface represented as a layered protocol has: (1) a first layer for assembling packets, that is, a transaction layer; a second layer for sequencing packets, that is, a transmission layer; and a third layer for transmitting the packets, that is, a physical layer. A layered Common Standard Interface (CSI) protocol is used as a specific example.

[0033] Now, with reference to Fig.Figure 4 illustrates an embodiment of a serial point-to-point PCIe structure. Although one embodiment of a serial point-to-point PCIe connection is illustrated, a serial point-to-point connection is not limited to this, as it includes any transmission path for transmitting serial data. In one embodiment, a basic PCIe connection has two differentially driven low-voltage signal pairs: a transmit pair 406 / 411 and a receive pair 412 / 407. The component 405 therefore has transmit logic 406 to transmit data to a component 410 and receive logic 407 to receive data from the component 410. In other words, a PCIe connection includes two transmit paths, namely paths 416 and 417, and two receive paths, namely paths 418 and 419.

[0034] A transmission path refers to any path for transmitting data, such as a transmission line, a copper wire, an optical line, a wireless communication channel, an infrared communication link, or any other communication path. A connection between two components, such as component 405 and component 410, is called a link, like Link 415. A link can support one lane, with each lane representing a set of differential signal pairs (one pair for transmitting, one pair for receiving). To scale bandwidth, a link can aggregate multiple lanes, designated xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.

[0035] A differential pair uses two transmission paths, such as lines 416 and 417, to transmit differential signals. For example, when line 416 transitions from a low-voltage level to a high-voltage level (a rising edge), line 417 transitions from a high logic level to a low logic level (a falling edge). Differential signals potentially exhibit better electrical characteristics, such as improved signal integrity (cross-coupling, boost / undershoot, ringing, etc.). This allows for a better timing window, enabling faster transmission frequencies.

[0036] New and increasing usage models, such as PCIe-based storage arrays and Thunderbolt, result in a significant increase in PCIe hierarchy depth and width. The PCI Express (PCIe) architecture was based on PCI, which defines a "configuration space" in which system firmware and / or software discover and enable / disable / control functions. Addressing within this space is based on a 16-bit address (usually called the "BDF" or "Bus Device Function Number"), consisting of an 8-bit bus number, a 5-bit device number, and a 3-bit function number. In PCIe, the bus number can refer to a logical bus rather than a physical bus.In addition to being used to address PCI functions in the configuration space and to identify specific functions for purposes such as fault reporting and IO virtualization, the space itself can be viewed as a resource type that is subject to allocation and management systems similar to other resources.

[0037] PCI allows systems to provide multiple, independent BDF spaces called "segments." Each segment can have specific resource requirements, such as a mechanism for generating PCI / PCIe configuration requests, including the Enhanced Configuration Access Mechanism (ECAM) defined in the PCIe specification. Additionally, input / output (I / O) memory management units (IOMMUs) (such as Intel VT-d) can use BDF space as an index but may not be configured to directly access segments. Consequently, in some cases, a separate ECAM and IOMMU are duplicated for each segment defined in a system. Fig.Figure 5 illustrates an example of a system with multiple segments (for example, 505a to c). In this example, a segment is defined for each of the three switches 510, 515, and 520, which are connected to a root complex 525. In this example, a separate IOMMU and a separate ECAM (for example, 530a to c) can be implemented on the root complex 525 to facilitate each of the segments (for example, 505a to c). Furthermore, in this example, a variety of endpoints (EPs) are connected to various buses in each segment. In some cases, the configuration of a segment may reserve multiple bus addresses for potential hot-plug events, which limits the total number of bus addresses available within each segment.Furthermore, the assignment of bus numbers in one or more segments can be performed according to an algorithm that pays little attention to dense address population and efficient use of the available bus address space. This can result in wasted configuration address space (i.e., BDF) in some cases.

[0038] Traditional PCIe systems are configured to allocate address space in a way that, when applied to modern and emerging use cases, tends to use BDF space, and especially bus numbers, inefficiently. While relatively few implementations actually involve a single system that consumes all 64K of unique BDF values ​​(defined, for example, under PCIe), deep hierarchies, such as those found in deep hierarchies of PCIe switches, can very quickly deplete available bus numbers. Furthermore, in applications that support hot-plugging, large portions of BDF space can typically be reserved for future potential use (that is, when a future component is plugged into the system while it is running) by taking additional swathes of bus numbers from the pool immediately available to a system.Although segmentation mechanisms can be used to address this problem, the segments themselves have limited scalability because, as mentioned above, additional hardware resources (e.g., IOMMUs) must be incorporated into the CPU, platform controller hub (PCH), system-on-a-chip (SoC), root complex, etc., to support each segment. Using segments to address deep hierarchies therefore results in scaling the system to meet worst-case system requirements, which is typically more than what would be necessary for most systems, leading to significant wasted platform resources. Furthermore, segments can be difficult (and in some cases essentially impossible) to create outside of the system's root complex.

[0039] In some implementations, a system can be provided to enable more efficient use of BDF space and to address at least some of the exemplary problems mentioned above. This can allow the expansion of PCIe, Thunderbolt, on-chip system architectures (for example, Intel On-Chip System Fabric (IOSF) and others), and other interconnects with very large topologies, but not without requiring dedicated resources in the root complex, as would be the case with solutions that rely solely on segments or other alternatives. Fig.Figure 6 illustrates an exemplary assignment of bus numbers to buses within the system according to an exemplary PCle BDF assignment. In this example, a system with two components 605 and 610 is directly connected to a root complex 615, and two switch-based hierarchies (corresponding to switches 620 and 625) are numbered with approximately the densest possible bus number assignments using conventional BDF assignment (as designated by circle labels (for example, 650a to d, etc.)). For deep hierarchies, the available bus numbers in a single BDF space can be quickly exhausted. Typically, real-world systems assign bus numbers much less efficiently, resulting in sparse (or "wasted") allocation of BDF space.

[0040] Another problem with use cases that support add / remove components while the system is running, such as Thunderbolt and, in some cases, PCIe-based storage, is that bus number assignments in BDF space are "balanced" to accommodate hardware topology changes that occur in a running system. However, this balancing can be very difficult for system software to perform, as in typical cases all PC functions are forced into a sleep state in conjunction with the balancing to allow the system to renumber the BDF space, followed by the reactivation of PCI functions. This process can be quite slow, however, and typically results in the system freezing for relatively long periods (for example, long enough to disrupt running applications and be noticeable to the end user).An improved system can also be provided to reduce the time required to apply a revised BDF space, allowing the balancing process to be executed in hundredths of a millisecond or faster, without explicitly placing the PCI functions in idle states. Finally, very large systems or systems with (protected) mechanisms to support multiple root complexes may be defined to require the use of segments.

[0041] As introduced above, the Flattening Portal Bridge (FPB) can be an optional mechanism used to address at least some of the exemplary problems mentioned above, including improving the scalability and runtime reassignment of Bus / Component / Function (BDF) and Memory-Mapped I / O (MMIO) spaces. The concept of the "BDF space" is related to the configuration address space but is generalized to acknowledge that the BDF is the basis for requester and completer IDs, the routing of completions, and can serve as an essential element in several mechanisms in addition to routing configuration requests.For functions associated with an upstream port, the function number portion of the BDF space address (for example, the 3-bit function number) can be determined by the architecture of the upstream port hardware, while the bus and component number portions can be determined by the downstream port above the upstream port. An example FPB can retain the existing architecture, with the upstream port determining the mapping of functions within the 3-bit function number portion of the BDF and operating only within the 13-bit bus / component number portion. In such cases, "BD space" can refer to the 13-bit bus / component number portion of the BDF. MMIO can specifically refer to memory read and write requests passed through a root port, switch port, or logic bridge if the FPB capability provides additional mechanisms to determine the address decoding of such requests.A bridge that implements the FPB capability can itself be called an FPB.

[0042] Fig. Figure 7A is a simplified block diagram illustrating the provision of FPB logic on each port of one or more switches within a system. Some switches may not have FPB logic and may only support bus numbering according to conventional BDF or MMIO space allocations. Furthermore, ports of a Root Complex 705 may also have FPB logic, such as ports designed to support potentially dynamic hot-plugging scenarios or, among other examples, to flexibly support different architectures (for example, one not predefined at the time of Root Complex design). Additionally, some ports of the Root Complex 705 may omit FPB logic, such as ports designed to support static scenarios, like ports with endpoints 710 and 715 in the example of the... Fig. 7A are connected.

[0043] FPB logic can be enabled or disabled on any port that supports it. FPB logic can be implemented in hardware, firmware, and / or software to support the allocation of BDF or MMIO space using a simplified approach. For example, each bus that connects switches, endpoints, and the root complex can be assigned a single bus number (e.g., as described above in conjunction with...). Fig.(discussed in section 6), with a number of components capable of being addressed by any one of the 32 possible component numbers under each bus number, and a number of possible function numbers that can be addressed by any one of the 8 possible function numbers under each component number. Flattening by FPB allows at least some of the buses to be uniquely addressed by a bus device (BD) number combination (rather than a single bus number), thus extending the maximum number of unique potential bus addresses from 256 (with conventional PCle bus numbering) to 8192 (with BD bus numbering). With FPB, the bus number can therefore be reused for several different buses, even though each bus number has a single BD number (that is, the bus number portion of the BD is the same, but the component number under that bus number is different).Some branches of the system can use conventional bus addressing (for example, via BDF bus numbers), while other branches use FPB-based addressing. This is because a mix of bus numbering systems can be used, resulting in the remaining unique bus addresses being used within the system.

[0044] With reference to Fig. Figure 7B illustrates a simplified block diagram of a high-level architecture of an exemplary implementation of an FPB logic block (as in the ports shown in the example of the Fig. Figure 7A illustrates this. A type 1 bridge function can be provided by any FPB module. The bridge function can support both established packet decoding / routing mechanisms (for example, conventional PCIe BDF decoding and routing) and FPB packet decoding / routing mechanisms.

[0045] FPB modifies how switches consume BDF resources to reduce waste by "flattening" the way bus numbers are used within switches and by downstream ports. FPB defines mechanisms for system software to allocate BDF and MMIO resources in non-adjacent areas, allowing the system software to allocate pools of BDF / MMIO from which it can map "bins" to functions below the FPB. This allows the system software to allocate BDF / MMIO required by a component hot-add without rebalancing other already allocated resource areas, and to revert to pool resources freed up, for example, by a hot-remove event.FPB is defined to allow established and new mechanisms to operate simultaneously, so that, for example, system firmware / software can implement a strategy where established mechanisms continue to be used in parts of the system where FPB mechanisms may not be required. In the example of... Fig. 7B, it can be assumed that the decoding logic provides a '1' output when a given TLP is decoded as associated with the secondary side of the bridge. The established decoding mechanism can apply as before, so that, for example, among other examples, only the bus number portion (bits 15:8) of a BDF address is tested by the established BDF decoding logic.

[0046] As in the example of the Fig.As illustrated in Figure 8, an instance of FPB logic can have both established decoding / routing mechanisms and FPB packet decoding / routing mechanisms. For established packet decoding / routing mechanisms, a TLP can be identified on the FPB-provided port, and established packet decoding / routing mechanisms can determine whether the TLP is routed to the secondary side of the port or whether the TLP routing is retained on the primary side based on conventional BDF routing. Similarly, the PB packet decoding / routing mechanism can determine whether the TLP is routed to the secondary side of the port or whether the TLP routing is retained on the primary side based on flattened BDF routing.If either the established packet decoding / routing mechanisms or the FPB packet decoding / routing mechanisms indicate that the packet should be forwarded to the secondary side, routing the packet traverses the secondary side bridge to its destination. Established packet decoding / routing mechanisms may include secondary / subordinate bus number registers for BDF decoding and, for memory (e.g., MMIO) decoding, memory base / limit registers, prefetch-enabled base / limit registers, a VGA enable bit, enhanced allocation, among other mechanisms that can be used by the FPB logic. FPB packet decoding / routing mechanisms may include BD secondary start, vector start, granularity, and related registers for use in conjunction with the BD vector for BD space decoding.Memory decoding can also employ vectors in FPB, such as, among other register examples, mechanisms, functionality and features, a MEM Low Vector for use in conjunction with MEM Low Vector Start, Granularity and related MEM Low registers, and a MEM High Vector for use in conjunction with MEM High Vector Start, Granularity and related MEM High registers.

[0047] In some cases, and although FPB can add additional paths for a specific bridge to decode a given TLP, FPB cannot change the fundamental functionalities of bridges within the switch and root complex architectural structures. In one example, FPB uses the same architectural concepts to provide management mechanisms for three different resource types: the bus / component space, bits 15:3, of BDF ("BD"); memory below 4GB ("MEM Low"); and memory above 4GB ("MEM High"). A hardware implementation of FPB is permissible to support any combination of these three mechanisms. For each mechanism, FPB uses a bit vector to indicate, for a specific subset range of the selected resource type, whether resources within that range are associated with the primary or secondary side of the FPB.Hardware implementations may be permitted to implement a small range of sizes for these vectors, and system firmware / software is capable of making the most effective use of the available vector by selecting an initial offset at which the vector is applied in increasing BD / address order, and a granularity for the individual bits within the vector to specify the size of the BD / address resource set for which the bits in a given vector apply.

[0048] For each of the BD / Mem-Low / Mem-High mechanisms, it may be particularly desirable for a Root Complex to provide a mechanism, for example configuration registers, by which hardware- or system-specific firmware / software can restrict the permissible range of BD / MMIO that the system software can allocate to the secondary side of a Root Port bridge.This can simplify the construction of multi-component root complexes by, for example, ensuring that system software does not attempt to apply FPB mechanisms to overlapping areas of the BD / MMIO space between root ports implemented on different components of the multi-component root complex, such that, among other examples, the root ports on one component are allowed to operate within a given range of BD / MMIO resources, and the root ports on another component are allowed to operate within a different range of BD / MMIO resources that does not overlap with the first.

[0049] In static use cases (or simply "static cases"), there are limits to hierarchy size and the number of endpoints due to bus and component number "waste" caused by the PCle / PCle architectural definition for switches, and by the conventional requirement that downstream ports associate an entire bus number with their link. In some implementations, this class of problems can be addressed by "flattening" BDF space usage so that switches and downstream ports can make more efficient use of the available space. For dynamic use cases (or simply "dynamic cases"), balancing by reserving large ranges of bus numbers and memory-mapped I / O (MMIO) in the bridge above the relevant endpoint(s) has been avoided in an attempt to satisfy any requirements within the pre-allocated ranges.However, this approach leads to additional waste, which exacerbates the shortcomings of conventional BDF mapping. Furthermore, this approach can be difficult to implement in the general case, even in relatively simple scenarios, such as when you have a solid-state drive (SSD) implementing a single endpoint, which is replaced by a unit with a switch by creating an internal hierarchy within the unit. Thus, although an initial mapping of just one bus would have sufficed, the initial mapping immediately breaks with the new unit.Furthermore, the pre-allocation approach can be problematic for MMIO if, during operation, connected endpoints may require the allocation of MMIO space below 4GB (a naturally limited resource), which is quickly consumed by pre-allocating even relatively small amounts. Pre-allocation is therefore not attractive due to the numerous system elements that impose requirements on system address space allocation below 4GB. Depending on various factors influencing the physical memory addressing capability of a given system, resource constraints on MMIO space above 4GB may also exist in some cases. The constraints that apply to MMIO space below 4GB may differ from those that apply above 4GB (and consequently, separate mechanisms may be optimized for each).

[0050] In some implementations, at least some of the problems in both static and dynamic use cases can be addressed by defining mechanisms to allow interrupted resource allocation (re)mapping for both BDF and MMIO. System software can maintain "resource pools" that can be allocated (and released) at runtime without interrupting other running operations, as is necessary for balancing.A Flattening Portal Bridge (FPB) can be provided as an optional capability, implemented by Type 1 (bridge) functions in root and switch ports, to support more efficient and denser BDF allocation, enable the reallocation of BDF resources without requiring balancing resources allocated elsewhere in a system, and allow for non-contiguous MMIO spaces, thus avoiding the need to balance MMIO resources. In some implementations, I / O space allocation may be left as is and not modified by the FPB.Among the potential exemplary benefits provided by an exemplary FPB are: BDF space allocation can be more efficient and denser, enabling larger hierarchies; runtime resource reassignment can be enabled for add / remove operations without the need for global resource balancing; the requirement for BDFs and MMIOs to be allocated in adjacent spaces can be eliminated; mixed systems can be supported, including components that support FPB alongside those that do not, all while not allowing any changes to existing discrete endpoints. For example, established root complexes, switches, bridges, and endpoints can be used in mixed system environments alongside RCs and switches that implement FPB.

[0051] In some implementations, FPB may involve the provision of new hardware and software that support FPB. However, this additional hardware and / or software can be optionally enabled in such a way that it has no effect unless activated, and it is disabled by default. In some cases, hardware changes to implement FPB hardware involving Type 1 functions may be required, while allowing endpoints and hardware functions supporting Type 0 functions to remain unaffected. FPB-enabled hardware may pass existing compliance and interoperability tests, and new tests may be developed to explicitly evaluate the additional FPB functionality. Software intended to work with components implementing FPB functionality may be configured to include the new, enhanced capability.Established software will continue to work with FPB hardware, but will not be able to utilize the FPB features.

[0052] FPB can continue to support the use of established resource allocation mechanisms for BDF and MMIO. In some cases, it may be desirable for system firmware to continue performing the initial system resource allocation using only the established mechanisms and to use FPB only after the operating system has booted. FPB can support this and specifically enable the system to continue using resources as allocated by the established mechanisms. FPB is specifically designed to enable system software to modify resource allocation within the system at runtime, requiring only that the hardware and processes associated with the resources being modified be put into a sleep state, and allowing all other hardware and processes to continue normal operation.

[0053] To support runtime use of FPB by system software, FPB hardware implementations should avoid stopping or introducing other types of interruptions to in-flight transactions, including during times when system software is changing the state of the FPB hardware. However, hardware is not expected to attempt to identify instances where system software mistakenly changes the FPB configuration in a way that affects in-flight transactions. As with established mechanisms, system software is responsible for ensuring that system operations are not corrupted due to a reconfiguration process.It is not explicitly required that system firmware / software perform the activation and / or deactivation of FPB mechanisms in a specific sequence; however, rules can be defined to implement resource allocation operations in a hierarchy in such a way that the hardware and software elements of the system are not corrupted or their failure is not caused.

[0054] If, in some implementations, system software violates any rules related to FPB, the hardware behavior may be undefined. FPB can be implemented in any PCI bridge (Type 1) function, and any function that implements FPB implements the extended FPB capability. If a switch implements FPB, the upstream port and all downstream ports of the switch implement FPB. A root complex may be allowed to execute FPB on some root ports but not others. A root complex may be allowed to execute FPB on an internal logical bus of the root complex. A Type 1 function is allowed to implement the FPB mechanism applied to any one, two, or three of these elementary mechanisms (BD, MEM Low, MEM High).System software may be allowed to enable any combination (including all or none) of the elementary mechanisms supported by a specific FPB. The error handling and reporting mechanisms, except where explicitly modified in this section, may remain unaffected by the FPB. In the case of an FPB function reset, the FPB hardware clears all bits in all implemented vectors. Once enabled (for example, by FPB-BD-Vector-Enable, FPB-MEM-Low-Vector-Enable, and / or FPB-MEM-High-Vector-Enable bits), if the system software subsequently disables an FPB mechanism, the values ​​of the inputs in the associated vector will be undefined. Furthermore, if the system software subsequently re-enables the FPB mechanism, the FPB hardware clears all bits in the associated vector.

[0055] In some implementations, the system software is explicitly permitted to modify an FPB vector when the corresponding FPB mechanism is enabled. If an FPB is implemented with the No_Soft_Reset bit cleared when it traverses D0→D3hot→D0, then, as in any other function configuration context, all FPB mechanisms are disabled, and the FPB clears all bits in all implemented vectors. If an FPB is implemented with the No_Soft_Reset bit set when it traverses D0→D3hot→D0, then, as in any other function configuration context, no FPB configuration states are modified, and the inputs to the FPB vectors are retained by the hardware. Hardware can be implemented such that there is no requirement to perform any type of limit check on FPB calculations, and the system software can ensure that the FPB parameters are programmed correctly.For example, system software may be allowed to program vector start values ​​that cause the higher-order bits of the corresponding vector to exceed the resource area associated with a given FPB, with the system software ensuring that these higher-order bits of the vector are cleared. Examples of errors that the system software must avoid include resource allocation duplication and combinations of start offsets with set vector bits that could create a "wrap-around" of boundary errors.

[0056] In some implementations of the FPB BD mechanism, FPB hardware assumes that a specific BDF is associated with the secondary side of the FPB if that BDF falls within the bus number range specified by the values ​​programmed in the secondary and subordinate bus number registers using a logical OR operation with the value in the corresponding entry in the BD vector. If only the FPB BD mechanisms are used for BDF decoding, system software can be employed to ensure that both the secondary and subordinate bus number registers are 0. System software can further ensure that the FPB routing mechanisms are configured such that configuration requests affecting the secondary side of FPB functions are routed from the primary to the secondary side of the FPB.The FPB BD mechanism can be applied with different granularities, which the system software can program via the FPB BD Vector Granularity Register in the FPB BD Vector Control 1 register. Fig. Figure 9 illustrates example addresses in BDF space and supported granularities. The representation in Fig. Figure 9 illustrates relationships between the layout of addresses in the BDF space and the supported granularities.

[0057] In some implementations, system software programs the FPB BD Vector Granularity and FPB BD Vector Start fields in the FPB BD Vector Control 1 register according to the requirements described in the descriptions of these fields. FPBs (other than those associated with upstream ports of the switches) can be restricted such that, if PCIe Alternative Routing ID Interpretation (ARI) forwarding is not supported, or if the ARI Forwarding Enable bit in the Device Control 2 register is cleared, FPB hardware should convert a Type 1 configuration request received on the primary side of the FPB to a Type 0 configuration request on the secondary side of the FPB if the BD address (bits 15:3 of the BDF) of the Type 1 configuration request matches the value in the BD Secondary Start field in the FPB BD Vector Control 2 register, and the system software must configure the FPB accordingly.If the ARI Forwarding Enable bit is set in the Device Control 2 register, the FPB hardware converts a Type 1 configuration request received on the primary side of the FPB into a Type 0 configuration request on the secondary side of the FPB if the bus number address (bits 15:8 of the BDF) of the Type 1 configuration request matches the value of the bus number address (bits 15:8) of the secondary start field in the FPB BD Vector Control 2 register, and the system software must configure the FPB accordingly.

[0058] In some implementations, FPB hardware for FPBs associated only with the upstream ports of switches can use the FPB Num Sec Dev field of the FPB Capability Register to specify the set of component numbers associated with the secondary side of the upstream port bridge. This can be used by the FPB, in addition to the BD Secondary Start field in the FPB BD Vector Control 2 register, to determine when a configuration request received on the primary side of the FPB is directed to the downstream ports of the switch. In effect, it determines when to convert such a request from a Type 1 configuration request to a Type 0 configuration request, with the system software configuring the FPB accordingly.If ACS Source Validation is enabled on a downstream port, the FPB checks the requester ID of each upstream request received by that port to determine if it maps to the FPB's secondary side. If the requester ID does not, this may represent a reported error (for example, an ACS violation) associated with the receiving port. FPBs can also implement bridge mapping for the virtual INTx wires.

[0059] To determine, for example, which input to the FPB BD vector applies to a given BDF address, FPB-equipped hardware and software can apply an algorithm such as the following:

[0060] In other words, the logic for determining which input to use in the FPB BD vector for a given BDF address can determine whether the BD address is below the value of the FPB BD Vector Start. If the BD address is below this value, the BD is out of range and should not be associated with the secondary side of the bridge. Otherwise, the logic can calculate the offset within the vector by first subtracting the value of FPB BD Vector Start, then dividing this by the value of the FPB BD Vector Granularity to determine the bit index within the vector. If the bit index value is greater than the length specified by the FPB BD Vector Size Supported, the BD is out of range (above this value) and should not be associated with the secondary side of the bridge.However, if the bit value within the vector at the calculated bit index position is 1b, the BD address is associated with the secondary side of the bridge; otherwise, the BD address is associated with the primary side of the bridge.

[0061] The FPB MEM Low mechanism can be applied with different granularities, which the system software can program through the FPB MEM Low Vector Granularity register in the FPB MEM Low Vector Control 1 register. Fig. Figure 10 illustrates the layout of addresses in the memory address space below 4 GB for which the FPB MEM Low mechanism applies, and the effect of granularity on these addresses. Fig. Section 10 also concerns the definition of the extended Flattening Portal Bridge (FPB) capability. System software can program the fields FPB MEM Low Vector Granularity and FPB MEM Low Vector Start in the FPB MEM Low Vector Control 1 register according to the specifications described in the descriptions of these fields.

[0062] In some instances of the FPB MEM Low mechanism, the FPB hardware may assume that a specific memory address is associated with the secondary side of the FPB if that memory address falls within any of the ranges specified by the values ​​programmed in other bridge memory decode registers (enumerated below), logically ORed with the value programmed in the corresponding input in the MEM Low vector. Other bridge memory decode registers may include: Memory Base / Limit register in the Type 1(Bridge) header; Prefetchable Base / Limit register in the Type 1(Bridge) header; VGA Enable bit in the Bridge Control Register of the Type 1(Bridge) header; Enhanced Allocation (EA) capability; FPB MEM High mechanism (if supported and enabled).To determine, for example, which entry in the FPB MEM vector applies to a given memory address, hardware and software can use an algorithm such as the following: .

[0063] In other words, the hardware and software used to determine which entry in the FPB MEM Low vector applies to a given memory address can determine whether the memory address is below the value of the FPB MEM Low Vector Start. If so, the memory address is out of range (below) and is not associated with the secondary side of the bridge. The logic can calculate the offset within the vector by first subtracting the value of FPB MEM Low Vector Start, then dividing this by the value of the FPB MEM Low Vector Granularity to determine the bit index within the vector. If the bit index value is greater than the length specified by the FPB MEM Low Vector Size Supported, the memory address is out of range (above) and is not associated with the secondary side of the bridge.On the other hand, if the bit value within the vector is at the calculated bit index position 1b, the memory address can be associated with the secondary side of the bridge; otherwise, the memory address is associated with the primary side of the bridge.

[0064] System software can program the FPB MEM High Vector Granularity and FPB MEM High Vector Start fields in the FPB MEM High Vector Control 1 register according to the specifications described in the descriptions of these fields. In some instances of the FPB MEM High mechanism, the FPB hardware can assume that a specific memory address is associated with the secondary side of the FPB if that memory address falls within one of the ranges specified by the values ​​programmed in other bridge memory decode registers (enumerated below), logically ORed with the value in the corresponding input in the MEM Low Vector.Other bridge memory decode registers may include Memory Base / Limit registers in the Type 1 (Bridge) header; Prefetchable Base / Limit registers in the Type 1 (Bridge) header; VGA Enable bits in the Type 1 (Bridge) Bridge Control Register header; Enhanced Allocation (EA) capability; and an FPB MEM Low mechanism (if supported and enabled). To determine, for example, which input to the FPB MEM High Vector applies to a given memory address, hardware and software may use an algorithm such as the following:

[0065] In other words, the hardware and software used to determine which input to use in the FPB MEM High Vector for a given memory address can determine whether the memory address is below the value of the FPB MEM High Vector Start. If so, it can be determined that the memory address is out of range (below) and not associated with the secondary side of the bridge. Otherwise, the offset within the vector can be calculated by first subtracting the value of the FPB MEM High Vector Start, and then dividing this result by the value of the FPB MEM High Vector Granularity to determine the bit index within the vector. If the bit index value is greater than the length specified by the FPB MEM High Vector Size Supported, the memory address is out of range (above) and not associated with the secondary side of the bridge.Otherwise, if the bit value within the vector is at the calculated bit index position 1b, the memory address is associated with the secondary side of the bridge, or the memory address is associated with the primary side of the bridge.

[0066] In some implementations, FPB can use a bit-vector mechanism to describe address spaces (BD space, MEM Low & MEM Hi). A bridge that supports FPB can have the following for each address space where it supports the use of FPB: a bit vector; a start address register; and a granularity register. These values ​​can be used by the bridge to determine whether a given address is part of the range that FPB decodes as associated with the secondary side of the bridge. An address that is not determined to be associated with the secondary side of the bridge by either the established decoding mechanism or the FPB decoding mechanism is (by default) associated with the primary side of the bridge. Here, the term "associated" might mean, for example, that the bridge applies the following handling to TLPs: • A TLP that is associated with and received at the primary page can be handled as an Unsupported Request (UR); • A TLP associated with the primary side and received at the secondary side can be treated as an upstream forward; • A TLP associated with the secondary side and received at the primary side can be treated as a downstream forward; • A TLP that is associated with and received at the secondary side can be handled as an Unsupported Request (UR), etc.

[0067] In FPB, each bit in the vector can represent a range of addresses, the size of which is determined by the selected granularity. If a bit in the vector is set, it indicates that packets addressed to an address within the corresponding range should be associated with the secondary side of the bridge. The specific range of addresses represented by each bit depends on that bit's index and the values ​​in the Start Address and Granularity registers. The Start Address register specifies the lowest address described by the bit vector. The Granularity register specifies the size of the range represented by each bit. Each subsequent bit in the vector applies to the next range, with the granularity increasing with each subsequent bit.

[0068] In some cases, downstream ports that do not have ARI forwarding enabled are associated only with component 0, with the component connected to the logical bus that represents the link from the port. Configuration requests targeting the bus number associated with a link specifying component number 0 are delivered to the component connected to that link. Configuration requests specifying any other component numbers (1 through 31) may therefore be terminated by the switch's downstream port or root port with an Unsupported Request Completion status (equivalent to Master Abort in PCI). In some cases, non-ARI components cannot assume that component number 0 is associated with their upstream port; instead, they capture their assigned component number and respond to all configuration read requests of type 0, regardless of the component number specified in the request.In some examples, when targeting an ARI component and the immediately upstream port is enabled for ARI forwarding, the component number is assumed to be 0, and the conventional component number field is used instead as part of an 8-bit function number field. If the configuration request is type 1, FPB logic can determine whether the Bus Number and Device Number fields (in the case of an Express PCI bridge) are equal to the bus number assigned to the secondary PCI bus, or, in the case of a switch or root complex, equal to the bus number and decoded component numbers assigned to the root (root complex) or downstream ports (switch). If so, the request can be forwarded to that downstream port (or PCI bus in the case of a PCI Express PCI bridge).If it is not the same as the bus number of any downstream port or secondary PCI bus, but is within the range of bus numbers that are assigned to either a downstream port or a secondary PC bus, the request can be forwarded to that downstream port interface without modification.

[0069] The enhanced Flattening Portal Bridge (FPB) capability can be an optional enhanced capability to be provided for any bridge feature or port that implements FPB. If a switch implements FPB, the upstream port and all downstream ports of the switch implement the enhanced FPB capability structure. A root complex may be allowed to execute the enhanced FPB capability structure on some root ports but not others. A root complex may be allowed to implement the FPB capability for internal logic buses in some implementations. The following description accesses the FPB registers through an enhanced PCI capability, but in other example implementations, the FPB registers may be accessed by other means, including, but not limited to, a PCI capability structure or a vendor-defined enhanced capability.In some implementations, the registers can be hosted in memory elements of the corresponding switches, bridges, root complex, or other components within the system.

[0070] Table 1 illustrates an example implementation of an extended FPB capability header. In one example, the extended FPB capability header can have an offset of 00h. Table 1: Header with enhanced FPB capability Bit position Register description 15:0 PCI Express Extended Capability ID - This field identifies the following structure as a structure with extended capability for a Flattening Portal Bridge (FPB) 19:16 Capability Version - This field is a PCI-SIG-defined version number that indicates the version of the capability structure. Must be 1h for this version of the specification. 31:20 Next Capability Offset This field contains the offset to the nearest PCI Express capability structure, or 000h if no other element exists in the linked list of capabilities. For extended capabilities implemented in configuration space, this offset relates to the beginning of the PCI-compliant configuration space and must therefore always be either 000h (to terminate the list of capabilities) or greater than 0FFh.

[0071] Table 2 illustrates an example implementation of an extended FPB capability header. In one example, the extended FPB capability header can have an offset of 04h. Table 2: FPB Capability Register Bit position Register description 0 FPB BD Vector Supported - If set, this indicates that the BD Vector mechanism is supported. 1 FPB MEM Low Vector Supported - If set, this indicates that the MEM Low Vector mechanism is supported. 2 FPB MEM High Vector Supported - If set, this indicates that the MEM High Vector mechanism is supported. 7:3 FPB Num Sec Dev - For upstream ports of switches only; this field specifies the set of component numbers associated with the secondary side of the upstream port's bridge. The set is determined by adding one to the numerical value of this field. Although the goal is for switch implementations to efficiently utilize function numbers, it is explicitly permitted for downstream ports to be assigned function numbers that are not contiguous within the specified range of component numbers, and system software is required to search for downstream port bridges for each function number within the specified set of component numbers associated with the upstream port's secondary side. This field is reserved for downstream ports. 10:8 FPB BD Vector Size Supported - Specifies the size of the FPBBD vector implemented in hardware and limits the permissible values ​​that software can write to the FPB BD Vector Granularity field. Defined codes are: Value Size Permissible Granularities 000b 256 bits 8, 16, 32, 64, 128, 256 001b 512 bits 8, 16, 32, 64, 128 010b 1K Bits 8, 16, 32, 64 011b 2K Bits 8, 16, 32 100b 4K bits 8, 16 101b 8K Bits 8 All other encodings are reserved. If the FPB (Vector Supported Bit) is cleared, the value in this field is undefined and must be ignored by software. 15:11 Reserved 18:16 FPB MEM Low Vector Size Supported - Specifies the size of the MEM Low vector implemented in hardware and limits the permissible values ​​that software can write to the FPB MEM Low Vector start field. Defined encodings are: Value Size Permissible Granularities 000b 256 bits 1, 2, 4, 8, 16 001b 512 bits 1, 2, 4, 8 010b 1K Bits 1, 2, 4 011b 2K Bits 1, 2 100b 4K bits 1 All other encodings are reserved. If the FPB MEM Low Vector Supported bit is cleared, the value in this field is undefined and must be ignored by software. 23:19 Reserved 26:24 FPB MEM High Vector Size Supported - Specifies the size of the MEM Low vector implemented in hardware. Defined codes are: 000b 256 bits 001b 512 Bits 010b 1K Bits 011b 2K Bits 100b 4K Bits 101b 8K Bits All other encodings are reserved. If the FPB MEM High Vector Supported bit is cleared, the value in this field is undefined and must be ignored by software. 31:27 Reserved

[0072] Table 3 illustrates an example implementation of an FPB BD Vector Control 1 register. In one example, the FPB BD Vector Control 1 register can have an offset of 08h. Table 3: FPB BD Vector Control 1 registers Bit position Register description 0 FPB BD Vector Enable - If set, the FPB BDVector mechanism is enabled. If the FPB BD Vector Supported bit is cleared, hardware is allowed to implement this bit as read-only (RO), and in this case, the value in this field is undefined. The default value of this bit is 0. 3:1 Reserved 6:4 FPB BD Vector Granularity - The value written to this field by the software controls the granularity of the FPB BD vector and the required orientation of the FPB BD vector start field (below). Defined encodings are: Value Granularity Start alignment 000b 8 BDF <keine Auflage> 001b 16 BDF ...0b 010b 32 BDF ...00b 011b 64 BDF ...000b 100b 128 BDF ...0000b 101b 256 BDF ...00000b All other encodings are reserved. Based on the implemented FPB BD Vector size, hardware is permitted to implement as RW only those bits of this field that can be programmed to non-zero values, in which case the higher-ranking bits are permitted (but not required) to be hardwired to 0. If the FPB BD Vector Supported bit is cleared, hardware is permitted to implement this bit as RO, and the value in this field is undefined. The default value for this field can be 0000b. 18:7 Reserved 31:19 FPB BD Vector Start - The value that the software writes to this field controls the offset within the BD space where the FPB BD vector is applied. The value represents a bus / component number (bits [15:3] of an address in the BDF space) such that bit 0 of the FPB BD vector defines the range from the value in this register up to the value plus the granularity. minus 1, and bit 1 represents the range from this register value plus granularity up to the value plus granularity minus 1, and so on. The function number offset (bits[2:0]) is set by hardware to 000b and cannot be changed. Software must program this field to a value that is naturally aligned according to the value in the FPB BD Vector Granularity field, as specified here: FPB BD Vector Granularity Start alignment pad 0000b <keine Auflage> 0001b ...0b 0010b ...00b 0011b ...000b 0100b ...0000b 0101b ...00000b If this requirement is violated, the hardware behavior is undefined. If the FPB BD Vector Supported bit is cleared, hardware is allowed to implement this bit as RO, and the value in this field is undefined. The default value for this field is 000h.

[0073] Table 4 illustrates an example implementation of an FPB BD Vector Control 2 register. In one example, the FPB BD Vector Control 2 register can have an offset of 0Ch. Table 4: FPB BD Vector Control 2 registers Bit position Register description 2:0 Reserved 15:3 BD Secondary Start The value written by the software to this field controls the offset within the BDF space at which type 1 configuration requests, passed downstream through the bridge, must be converted to type 0. The value represents a bus / component number (bits [15:3] of an address in BDF space). The function number offset (bits [2:0]) is set by hardware to 000b and cannot be changed. If the ARI Forwarding Enable bit in the DeviceControl 2 register is set, the software must write bits 7:3 of this field to 00000b. If the FPB BD Vector Supported bit is cleared, hardware is allowed to implement this bit as RO, and the value in this field is undefined. The default value for this field is 000h. 31:16 Reserved

[0074] Table 5 illustrates an example implementation of an FPB BD Vector Access Control register. In one example, the FPB BD Vector Access Control register can have an offset of 10h. Table 5: FPB BD Vector Access Control Registers Bit position Register description 7:0 FPB BD Vector Access Offset - The value in this field specifies the offset of the 32b section of the FPB BD vector, which can be read or written using the FPB BD Vector Access Data Register. The bits of this field map to the offset according to the value in the FPB BD Vector Size Supported field, as shown here: Offset Bits This field 000b 2:0 2:0 (7:3 not used) 001b 3:0 3:0 (7:4 not used) 010b 4:0 4:0 (7:5 not used) 011b 5:0 5:0 (7:6 not used) 100b 6:0 6:0 (7 not used) 101b 7:0 7:0 All other encodings are reserved. Bits in this field that are not used according to the table above must be written by the software as 0b, and it is permitted (but not mandatory) for them to be implemented as RO. If the FPB BD Vector Supported bit is cleared, hardware is permitted to implement this field as RO, and the value in this field is undefined. The default value for this field is 00h. 31:8 Reserved

[0075] Table 6 illustrates an example implementation of an FPB BD Vector Access Data register. In one example, the FPB BD Vector Access Data register can have an offset of 14h. Table 6: FPB BD Vector Access Data Registers Bit position Register description 31:0 FPB BD Vector Data - Reads from this register indicate the data point (DW) of data from the FPB BD vector at the position determined by the value in the FPB BD Vector Access Offset register. Writes to this register replace the DW of data from the FPB BD vector at the position determined by the value in the FPB BD Vector Access Offset register. If the FPB BD Vector Supported bit is cleared, hardware is permitted to implement this bit as a read-only (RO), and the value in this field is undefined. The default value for this field is 0000h.

[0076] Table 7 illustrates an example implementation of an FPB MEM Low Vector Control Register. In one example, the FPB MEM Low Vector Control Register can have an offset of 18h. Table 7: FPB MEM Low Vector Control Registers Bit position Register description 0 FPB MEM Low Vector Enable - If set, the FPB MEM Low Vector mechanism is enabled. If the FPB MEM Low Vector Supported bit is cleared, hardware is allowed to implement this bit as RO, and in this case, the value in this field is undefined. The default value of this bit is 0b. 3:1 Reserved 7:4 FPB MEM Low Vector Granularity - The value written to this field by the software controls the granularity of the FPB MEM Low vector and the Required alignment of the FPB MEM Low VectorStart field (below). Defined encodings are: Value Granularity Start alignment pad 000b 1MB <keine Auflage> 001b 2MB ...0b 010b 4MB ...00b 011b 8MB ...000b 100b 16MB ...0000b All other encodings are reserved. Based on the implemented FPB MEM Low Vector size, hardware is permitted to implement as RW only those bits of this field that can be programmed to non-zero values, in which case the higher-ranking bits are permitted (but not required) to be hardwired to 0. If the FPB MEM Low Vector Supported bit is cleared, hardware is permitted to implement this bit as RO, and the value in this field is undefined. The default value for this field can be 0000b. 19:8 Reserved 31:20 FPB MEM Low Vector Start The value the software writes to this field sets the base address at which the FPB MEM Low Vector is applied. The software must program this field to a value that is naturally aligned with the value in the FPB MEM Low Vector Granularity field, as specified in the field description (above). If this requirement is violated, the hardware behavior is undefined. If the FPB MEM Low Vector Supported bit is cleared, hardware is allowed to implement this field as RO, and the value in this field is undefined. The default value for this field is 0000h.

[0077] Table 8 illustrates an example implementation of an FPB MEM Low Vector Access Control register. In one example, the FPB MEM Low Vector Access Control register can have an offset of 1 ch. Table 8: FPB MEM Low Vector Access Control Registers Bit position Register description 6:0 FPB MEM Low Vector Access Offset - The value in this field specifies the offset of the 32b section of the FPB MEM Low Vector, which can be read or written using the FPB MEM Low Vector Access Data Register. The bits of this field map to the offset according to the value in the FPB MEM Low Vector Granularity field, as shown here: Offset bits This field 000b 2:0 2:0 (6:3 not used) 001b 3:0 3:0 (6:4 not used) 010b 4:0 4:0 (6:5 not used) 011b 5:0 5:0 (6 not used) 100b 6:0 6:0 Bits in this field that are not used according to the table above must be removed. Software can be written as 0b, and it is permitted (but not mandatory) for it to be implemented as RO. If the FPB MEM Low Vector Supported bit is cleared, hardware is allowed to implement this bit as RO, and the value in this field is undefined. The default value for this field is 00h. 31:7 Reserved

[0078] Table 9 illustrates an example implementation of an FPB MEM Low Vector Access Data register. In one example, the FPB MEM Low Vector Access Data register can have an offset of 20h. Table 9: FPB MEM Low Vector Access Data Register Bit position Register description 31:0 FPB MEM Low Vector Data - Reads from this register return the data write (DW) of data from the FPB MEM Low Vector at the location determined by the value in the FPBMEM Low Vector Access Offset register. Writes to this register replace the DW of data from the FPB MEM Low Vector at the location determined by the value in the FPB MEM Low Vector Access Offset register. If the FPB MEM Low Vector Supported bit is cleared, hardware is allowed to implement this bit as a read-only (RO), and the value in this field is undefined. The default value for this field is 0000h.

[0079] Table 10 illustrates an example implementation of an FPB MEM High Vector Control 1 register. In one example, the FPB MEM High Vector Control 1 register can have an offset of 24h. Table 10: FPB MEM High Vector Control 1 register Bit position Register description 0 FPB MEM High Vector Enable - If set, the FPB MEM High Vector mechanism is enabled. If the FPB MEM High Vector Supported bit is cleared, hardware is allowed to implement this field as RO, and in this case, the value in this field is undefined. The default value of this bit is 0b. 3:1 Reserved 7:4 FPB MEM High Vector Granularity The value written by the software to this field controls the granularity of the FPB MEM High vector and the required orientation of the FPB MEM High VectorStart Lower field. Software is permitted to select any permissible granularity from the following table, regardless of the value in the FPB MEM High Vector Size Supported field. Defined encodings are: Value Granularity Start alignment pad 000b 256MB <keine Auflage> 001b 512MB ...0b 010b 1GB ...00b 011b 2GB ...000b 100b 4GB ...0000b 101b 8GB ...00000b 110b 16 GB ...000000b 111b 32GB ...0000000b Based on the implemented FPB MEM High Vector size, hardware is permitted to implement only those bits of this field that can be programmed to non-zero values ​​as RW (Returnable). In this case, the higher-ranking bits are permitted (but not required) to be hardwired to 0. If the FPB MEM High Vector Supported bit is cleared, hardware is permitted to implement this field as RO (Returnable), and the value in this field is undefined. The default value for this field is 0000b. 27:8 Reserved 31:28 FPB MEM High Vector Start Lower - The value that the software writes to this field sets the less significant bits of the base address at which the FPBMEM High Vector is applied. Software must program this field to a value that is naturally aligned according to the value in the FPB MEM High Vector Granularity field as specified here (that is, the less significant bits are zeros): FPB MEMHigh Vector Granularity Edition 0000b <keine Auflage> 0001b ...0b 0010b ...00b 0011b ...000b 0100b ...0000b 0101b ...00000b 0110b ...000000b 0111b ...0000000b If this requirement is violated, the hardware behavior is undefined. If the FPB MEM High Vector Supported bit is cleared, hardware is allowed to implement this field as RO, and the value in this field is undefined. The default value for this field is 00h.

[0080] Table 11 illustrates an example implementation of an FPB MEM High Vector Control 2 register. In one example, the FPB MEM High Vector Control 2 register can have an offset of 28h. Table 11: FPB MEM High Vector Control 2 registers Bit position Register description 31:0 FPB MEM High Vector Start Upper The value written by the software to this field specifies bits 63 and 32 of the base address where the FPB MEM High Vector Supported bit is applied. If the FPB MEM High Vector Supported bit is cleared, hardware is allowed to implement this bit as a Return Out (RO), and the value in this field is undefined. The default value for this field is 00000000h.

[0081] Table 12 illustrates an example implementation of an FPB MEM High Vector Access Control register. In one example, the FPB MEM High Vector Access Control register can have an offset of 2Ch. Table 12: FPB MEM High Vector Access Control Registers Bit position Register description 7:0 FPB MEM High Vector Access Offset - The value in this field specifies the offset of the 32b section of the FPBMEM BD, MEMLow, or MEM High vector, which can be read or written using the FPBMEM High Vector Access Data register. The bits of this field represent the value in the FPBMEM High Vector Granularity field, as shown here: Offset bits This field 000b 2:0 2:0 (7:3 not used) 001b 3:0 3:0 (7:4 not used) 010b 4:0 4:0 (7:5 not used) 011b 5:0 5:0 (7:6 not used) 100b 6:0 6:0 (7 not used) 101b 7:0 7:0 Bits in this field that are not used according to the table above must be written by the software as 0b, and it is permitted (but not mandatory) for them to be implemented as RO. If the FPB MEM High Vector Supported bit is cleared, hardware is permitted to implement this field as RO, and the value in this field is undefined. The default value for this field is 00h. 13:8 Reserved 15:14 FPB Vector Select - The value written to this field selects the vector accessed at the specified FPB Vector Access offset, encoded as: 00: BD 01: MEM Low 10: MEM High 11: Reserved The default value for this field can be 00b. 31:16 Reserved

[0082] Table 13 illustrates an example implementation of an FPB MEM High Vector Access Data register. In one example, the FPB MEM High Vector Access Data register can have an offset of 30h. Table 13: FPB MEM High Vector Access Data Registers Bit position Register description 31:0 FPB MEM High Vector Data Reads from this register return the data write (DW) of data from the FPB MEM High Vector at the position determined by the value in the FPBMEM High Vector Access Offset register. Writes to this register replace the DW of data from the FPB MEM High Vector at the position determined by the value in the FPB MEM High Vector Access Offset register. If the FPB MEM High Vector Supported bit is cleared, hardware is allowed to implement this field as a read-only (RO), and the value in this field is undefined. The default value for this field is 0000h.

[0083] In an alternative implementation, instead of providing separate Vector Access Offset and Vector Data registers for each vector, a single Vector Access Offset register can be used with the addition of a field to specify which vector to access, and a single Vector Data register can be used to perform the read or write operations on the specified vector. In such an implementation, the indicator field can be implemented as a two-bit field encoded such that a value of 00 (binary) indicates access to the BD vector, a value of 01 (binary) indicates access to the MEM Low vector, a value of 10 (binary) indicates access to the MEM High vector, and a value of 11 (binary) indicates a reserved value.

[0084] It should be noted that the devices, methods, and systems described above can be implemented in any electronic component or system as mentioned above. The following figures provide specific illustrations of exemplary systems for the use of the invention as described herein. Since the following systems are described in more detail, a number of different circuit configurations are disclosed, described, and revisited from the discussion above. And, as is readily apparent, the advances described above can be applied to any of these circuit configurations, structures, or architectures.

[0085] With reference to Fig.Figure 11 shows an embodiment of a block diagram for a computing system comprising a multicore processor. A processor 1100 comprises any processor or processing components, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a handheld processor, an application processor, a coprocessor, a system-on-a-chip (SoC), or other components for executing code. In one embodiment, the processor 1100 comprises at least two cores, namely cores 1101 and 1102, which may be asymmetric cores or symmetric cores (the illustrated embodiment). However, the processor 1100 may comprise any number of processing elements, which may be symmetric or asymmetric.

[0086] In one embodiment, a processing element refers to hardware or logic to support a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a processing unit, a context, a context unit, a logic processor, a hardware thread, a core, and / or any other element capable of maintaining state for a processor, such as an execution state or architectural state. In other words, in one embodiment, a processing element refers to any hardware capable of being independently associated with code, such as a software thread, an operating system, an application, or other code. A physical processor (or processor socket) typically refers to an integrated circuit that may contain any number of other processing elements, such as cores or hardware threads.

[0087] A kernel often refers to logic residing on an integrated circuit capable of maintaining independent architectural states, each associated with at least some dedicated execution resources. In contrast to kernels, a hardware thread typically refers to any logic residing on an integrated circuit capable of maintaining independent architectural states, where these independent architectural states share access to execution resources. As can be seen, certain resources are shared, and others are dedicated to a single architectural state, with the line between the nomenclature of a hardware thread and a kernel overlapping.However, a core and a hardware thread are often treated by an operating system as individual logical processors, with the operating system being able to schedule operations on each logical processor individually.

[0088] The physical processor 1100, as in Fig.Figure 11 illustrates the configuration with two cores, namely core 1101 and core 1102. Here, core 1101 and core 1102 are considered symmetric cores, meaning cores with the same configurations, functional units, and / or logic. In another embodiment, core 1101 has an out-of-order processor core, while core 1102 has an in-order processor core. However, cores 1101 and 1102 can be individually selected from any core type, such as a native core, a software-managed core, a core adapted to run a native Instruction Set Architecture (ISA), a core adapted to run a translated Instruction Set Architecture (ISA), a co-designed core, or any other known core.In a heterogeneous kernel environment (that is, asymmetric kernels), a form of translation, such as binary translation, can be used to schedule or execute code on one or both kernels. However, to further the explanation, the functional units illustrated in kernel 1101 are described in more detail below, since the units in kernel 1102 operate in a similar manner in the illustrated embodiment.

[0089] As shown, the 1101 core has two hardware threads, 1101a and 1101b, which can also be called hardware thread slots 1101a and 1101b. Software entities, such as an operating system, therefore potentially see the 1100 processor in one embodiment as four separate processors, that is, four logical processors or processing elements capable of executing four software threads simultaneously. As indicated above, a first thread is associated with architecture state registers 1101a, a second thread is associated with architecture state registers 1101b, a third thread can be associated with architecture state registers 1102a, and a fourth thread can be associated with architecture state registers 1102b. Here, each of the architecture state registers (1101a, 1101b, 1102a and 1102b) can be called processing elements, thread slots or thread units, as described above.As illustrated, the architecture state registers 1101a are replicated to architecture state registers 1101b, allowing individual architecture states / contexts to be stored for logical processor 1101a and logical processor 1101b. Within core 1101, other smaller resources, such as instruction pointers and rename logic in an assigner and rename block 1130, can also be replicated for threads 1101a and 1101b. Some resources, such as reorder buffers in a reorder / retirement unit 1134, ILTB 1120, load / store buffers, and queues, can be shared through partitioning. Other resources, such as internal general-purpose registers, page table base registers, low-level data cache and data TLB 1115, execution unit(s) 1140 and sections of an out-of-order unit 1135, are potentially shared entirely.

[0090] The 1100 processor often has other resources that can be fully shared, shared through partitioning, or dedicated to / after processing elements. In Fig.Figure 11 illustrates an embodiment of a purely exemplary processor with illustrative logical units / resources of a processor. It should be noted that a processor may include or omit any of these functional units, as well as any other functional units, logic, or firmware not shown. As illustrated, core 1101 has a simplified representative out-of-order (OOO) processor core. However, an in-order processor may be used in different embodiments. The OOO core has a branch target buffer 1120 to predict branches to be executed / taken and an instruction translation buffer (I-TLB) 1120 to store address translation entries for instructions.

[0091] The 1101 core also includes a 1125 decode module, which is coupled to the 1120 fetch unit to decode fetched elements. In one embodiment, the fetch logic has individual sequencers, each associated with the 1101a and 1101b thread slots. Typically, the 1101 core is associated with a first ISA that defines / specifies instructions executable on the 1100 processor. Often, machine code instructions belonging to the first ISA include a section of the instruction (called an opcode) that references / specifies an instruction or operation to be executed. The 1125 decode logic includes circuitry that recognizes these instructions from their opcodes and passes the decoded instructions into the pipeline for processing as defined by the first ISA.As discussed in more detail below, in one embodiment, Decoder 1125 incorporates logic designed or adapted to recognize specific instructions, such as a transaction instruction. As a result of recognition by Decoder 1125, the Architecture or Core 1101 executes specific, predefined actions to perform tasks associated with the corresponding instruction. It is important to note that any of the tasks, blocks, operations, and procedures described here can be executed in response to one or more instructions, some of which may be new or old. It should be noted that in one embodiment, Decoder 1126 recognizes the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, Decoder 1126 recognizes a second ISA (either a subset of the first ISA or a separate ISA).

[0092] In one example, the allocation and rename block 1130 has an allocation block to reserve resources, such as register files, to store instruction processing results. Threads 1101a and 1101b are potentially capable of out-of-order execution, with the allocation and rename block 1130 also reserving other resources, such as reorder buffers, to track instruction results. Unit 1130 may also have a register renamer to rename program / instruction reference registers to other registers internally to the processor 1100. The reorder / retirement unit 1135 has components such as the reorder buffers mentioned above, load buffers, and memory buffers to support out-of-order execution and subsequent in-order retirement of instructions that were executed out of order.

[0093] In one embodiment, a scheduler and execution unit block 1140 includes a scheduler unit for scheduling instructions / operations to execution units. For example, a floating-point instruction is scheduled on a port of an execution unit that has an available floating-point execution unit. Register files associated with the execution units are also present for storing information instruction processing results. Exemplary execution units include a floating-point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a memory execution unit, and other known execution units.

[0094] The lower-level data cache and data translation buffer (D-TLB) 1150 are coupled to execution unit(s) 1140. The data cache is intended to store elements that have been recently used or on which operations have recently been performed, such as data operands, which are potentially held in memory coherence states. The D-TLB is intended to store the most recent virtual / linear-to-physical address translations. As a specific example, a processor might have a page table structure to organize physical memory into a multitude of virtual pages.

[0095] Here, cores 1101 and 1102 share access to a higher-level or "further-out" cache, such as a second-level cache associated with an on-chip interface. It's important to note that "higher-level" or "further-out" refers to cache levels that increase in size or are located further away from the execution unit(s). In one embodiment, a higher-level cache is a last-level data cache, the last cache in the memory hierarchy on processor 1100, such as a second- or third-level data cache. However, a higher-level cache is not limited to this, as it can be associated with or include an instruction cache. A trace cache, a type of instruction cache, can instead be coupled after decoder 1125 to store recently decoded traces.Here, an instruction potentially refers to a macro instruction (that is, a general instruction that is recognized by the decoders), which can be broken down into a number of micro instructions (micro-operations).

[0096] In the configuration shown, the 1100 processor also includes an on-chip interface module 1110. Historically, a memory controller, described in more detail below, was located in a computing system outside the 1100 processor. In this scenario, an on-chip interface 1110 communicates with components outside the 1100 processor, such as system memory 1175, a chipset (which often includes a memory controller hub to connect to the memory 1175 and an I / O controller hub to connect peripherals), a memory controller hub, a northbridge, or other integrated circuits. In this scenario, the 1105 bus can have a known interconnect, such as a multi-drop bus, a point-to-point interconnect, a serial interconnect, a parallel bus, a coherent (for example, cache-coherent) bus, a layered protocol architecture, a differential bus, and a GTL bus.

[0097] The 1175 memory component can be used by the 1100 processor alone or in conjunction with other components in a system. Conventional examples of this type of 1175 memory include DRAM, SRAM, non-volatile memory (NV memory), and other known storage devices. It should be noted that the 1180 component can include a graphics accelerator, a processor or card coupled with a memory controller hub, data storage coupled with an I / O controller hub, a wireless transmitter / receiver, a flash component, an audio controller, a network controller, or another known component.

[0098] Since more logic and components can be integrated onto a single die, such as a SoC, each of these components can now be integrated onto a Processor 1100. In one embodiment, for example, a memory controller hub is located on the same package and / or die as the Processor 1100. Here, a section of the core (an "on-core" section) 1110 has one or more controllers for interfaces with other components, such as the memory 1175 or a graphics component 1180. The configuration that includes interconnects and controllers for interfaces with such components is often called an "on-core" (or an on-core configuration). As an example, an on-chip interface 1110 has a ring interconnect for on-chip communication and a high-speed serial point-to-point link 1105 for off-chip communication.In the SOC environment, even more components, such as a network interface, co-processors, 1175 memory, 1180 graphics processor and any other known computer components / interfaces, can be integrated onto a single die or integrated circuit to provide a small form factor with high functionality and low power consumption.

[0099] In one embodiment, the processor 1100 is capable of executing compiler, optimizer, and / or translator code 1177 to compile, translate, and / or optimize application code 1176 to support or interface with the device and procedures described herein. A compiler often includes a program or set of programs for translating source text / code into target text / code. Typically, program / application code is compiled in multiple stages and passes to convert high-level programming language code into low-level machine or assembly language code. However, single-pass compilers can still be used for simple compilation.A compiler can use any known compilation techniques and perform any known compiler operations, such as lexical analysis, preprocessing, parsing, semantic analysis, code generation, code transformation, and code optimization.

[0100] Larger compilers often have multiple phases, but these phases are usually contained within two main phases: (1) a frontend, which generally involves syntactic processing, semantic processing, and some transformation / optimization, and (2) a backend, which generally involves analysis, transformations, optimizations, and code generation. Some compilers refer to a middle ground, which illustrates the blurring of clear distinctions between a compiler's frontend and backend. As a result, references to insertion, association, generation, or any other compiler operation can occur in any of the phases or passes mentioned above, as well as in any other known phases or passes of a compiler. As an illustrative example, a compiler potentially inserts operations, calls, functions, etc.Dynamic compilation involves one or more compilation phases, such as inserting calls / operations in a front-end compilation phase and then converting those calls / operations into low-level code during a transformation phase. It's worth noting that during dynamic compilation, compiler code or dynamic optimization code can insert such operations / calls and optimize the code for execution at runtime. As a specific illustrative example, binary code (already compiled code) can be dynamically optimized at runtime. Here, the program code can include the dynamic optimization code, the binary code, or a combination of these.

[0101] Similar to a compiler, a translator, such as a binary translator, translates code either statically or dynamically to optimize and / or translate code. References to code execution, application code, program code, or other software environments can therefore refer to: (1) executing one or more compiler programs, optimization code optimizers, or translators, dynamically or statically, to compile program code, maintain software structures, perform other operations, optimize code, or translate code; (2) executing main program code that includes operations / calls, such as application code that has been optimized / compiled; (3) executing other program code, such as libraries, associated with the main program code to maintain software structures, execute other software related to operations, or optimize code; or (4) a combination of these.

[0102] With reference to Fig.Figure 12 shows a block diagram of a system 1200 according to an embodiment of the present invention. As shown in Fig. As shown in Figure 12, the multiprocessor system 1200 is a point-to-point interconnect system and comprises a first processor 1270 and a second processor 1280, which is coupled via a point-to-point interconnect 1250. Each of the processors 1270 and 1280 can be any version of a processor. In one embodiment, 1252 and 1254 are part of a coherent serial point-to-point interconnect fabric, such as Intel's Quick Path Interconnect (QPI) architecture. As a result, the invention can be implemented within the QPI architecture.

[0103] Although it is shown with only two processors 1270, 1280, it must be understood that the scope of protection of the present invention is not limited thereto. In other embodiments, one or more additional processors may be present in a given processor.

[0104] The 1270 and 1280 processors are shown each having integrated memory controller units 1272 and 1282, respectively. The 1270 processor has point-to-point (PP) interfaces 1276 and 1278 as part of its bus controller units; the second processor, 1280, similarly has PP interfaces 1286 and 1288. The 1270 and 1280 processors can exchange data via a point-to-point (PP) interface 1250 using PP interface circuits 1278 and 1288. As shown in Fig.As shown in Figure 12, IMCs 1272 and 1282 couple the processors with respective memories, namely a memory 1232 and a memory 1234, which can be parts of a main memory that are attached locally to the respective processors.

[0105] The 1270 and 1280 processors can each exchange information with a 1290 chipset via individual PP interfaces 1252 and 1254 using point-to-point interface circuits 1276, 1294, 1286, and 1298. The 1290 chipset also exchanges information with a high-performance graphics circuit 1238 via an interface circuit 1292 along a high-performance graphics circuit 1239.

[0106] A shared cache (not shown) may be contained within both processors or outside of both processors, but connected to the processors via a PP interconnection, so that local cache information from one or both processors can be stored in the shared cache if one processor is put into a low-power mode.

[0107] The chipset 1290 can be coupled to a first bus 1216 via an interface 1296. In one embodiment, the first bus 1216 can be a PCI bus (PCI: Peripheral Component Interconnect) or a bus such as a PCI Express bus or another third-generation I / O interconnect bus, although the scope of protection of the present invention is not limited thereto.

[0108] As in Fig.As shown in Figure 12, various I / O devices 1214 can be coupled to the first bus 1216 together with a bus bridge 1218, which couples the first bus 1216 to a second bus 1220. In one embodiment, the second bus 1220 has an LPC (Low Pin Count) bus. Various devices are coupled to a second bus 1220 in one embodiment, which, for example, includes a keyboard and / or mouse 1222, communication devices 1227, and a storage unit 1228, such as a disk drive or other mass storage device, which often contains instructions / code and data 1230. Furthermore, an audio I / O 1224 is shown coupled to the second bus 1220. It should be noted that other architectures are possible, with varying components and interconnection architectures. For example, a system can use a point-to-point architecture instead of the... Fig. 12. Implement a multi-drop bus or other such architecture.

[0109] Although the present invention has been described with reference to a limited number of embodiments, the person skilled in the art will be able to appreciate numerous modifications and variants thereof. It is intended that the attached claims cover all such modifications and variants as fall within the true spirit and scope of protection of this present invention.

[0110] Aspects of the embodiments may include one or a combination of the following examples:

[0111] Example 1 is a system, method, device, or storage medium with instructions stored thereon that are executable to cause a machine to identify a plurality of components in a system and to assign a respective address to each of the plurality of components. Each component in the plurality of components is connected in the system by at least one of a plurality of buses, and the assignment of the address to a component involves determining whether the address is to be assigned according to a first addressing system or a second bus addressing system, wherein the first addressing system assigns a single bus number within a bus / component / function (BDF) address space to each component addressed in the first addressing system, and the second bus addressing system assigns a single bus component number within the BDF address space.

[0112] Example 2 can feature the subject of Example 1, reusing a special bus number to address two or more components in the second addressing system.

[0113] Example 3 can have the subject of any of Examples 1-2, wherein the assignment of addresses involves specifying a range of bus numbers in the BDF address space to be used to address components according to the second addressing system.

[0114] Example 4 may include the subject of Example 3, wherein the range of bus numbers is associated with a special switch and the bus numbers in the range of bus numbers are used in the bus component numbers to be assigned to each component connected to the switch.

[0115] Example 5 can include the object of Example 4, wherein the components connected to the switch have a segment.

[0116] Example 6 can have the subject of any of Examples 1 to 5, where the addresses are configuration addresses.

[0117] Example 7 can feature the subject of any of Examples 1 to 6, wherein the BDF address space is a Peripheral Component Interconnect (PCI) based address space.

[0118] Example 8 can include the subject of Example 7, wherein each bus component number has an eight-bit bus number and a five-bit component number.

[0119] Example 9 is a device having a port for receiving a special packet, wherein the port has a flattening gantry bridge (FPB), the FPB having a primary side and a secondary side, the primary side connecting to a first set of components addressed according to a first addressing system, and the secondary side connecting to a second set of components addressed according to a second addressing system. The FPB further determines whether the special packet is to be routed to the primary side or the secondary side based on address information in the special packet, wherein the first addressing system uses a single bus number within a bus / component / function (BDF) address space for each component in the first set of components, and the second bus addressing system uses a single bus component number for each component in the second set of components.

[0120] Example 10 can include the subject of Example 9, wherein the respective bus component numbers assigned to a multitude of components in the second set of components each have a special bus number and a distinct component number.

[0121] Example 11 can feature the subject of any of Examples 9-10, wherein the primary addressing system features an established addressing system.

[0122] Example 12 can feature the subject of any of Examples 9-11, wherein the BDF addressing space has a configuration space based on Peripheral Component Interconnect Express (PCIe).

[0123] Example 13 can have the subject of any of Examples 9 to 12, further comprising a plurality of ports, wherein the port has a special one of the plurality of ports and at least one other port in the plurality of ports has an FPB.

[0124] Example 14 can feature the subject of Example 13, wherein the plurality of ports includes at least one port without an FPB.

[0125] Example 15 can include the subject of Example 13, which further includes a switch, wherein the switch has the plurality of ports.

[0126] Example 16 can include the subject of Example 13, which further includes a root complex, wherein the root complex includes the plurality of ports.

[0127] Example 17 can have the subject of any of Examples 9 to 16, which furthermore has a BD Control 1 register.

[0128] Example 18 can include the subject of any of Examples 9 to 17, which further includes a BD Vector Control 2 register.

[0129] Example 19 can include the subject of any of Examples 9 to 18, which furthermore includes a BD Vector Access Control register.

[0130] Example 20 can include the subject of any of Examples 9 to 19, which furthermore includes a BD Vector Access Data register.

[0131] Example 21 can include the subject of any of Examples 9 to 20, which further includes a MEM Low Vector Control Register.

[0132] Example 22 can include the subject of any of Examples 9 to 21, which further includes a MEM Low Vector Access Control register.

[0133] Example 23 can include the subject of any of Examples 9 to 22, which further includes a MEM Low Vector Access Data register.

[0134] Example 24 can include the subject of any of Examples 9 to 23, which further includes a MEM High Vector Control-1 register.

[0135] Example 25 can include the subject of any of Examples 9 to 24, which further includes a MEM High Vector Control-2 register.

[0136] Example 26 can have the subject of any of Examples 9 to 25, which furthermore has a MEM High Vector Access Control register.

[0137] Example 27 can include the subject of any of Examples 9 to 26, which furthermore includes a MEM High Vector Access Data register.

[0138] Example 28 is a storage medium on which instructions are stored, wherein the instructions, when executed on a machine, cause the machine to configure registers of a component to support a primary bus addressing system in a bus / component / function (BDF) space, and an alternative bus addressing system that uses the same bus number within a memory-mapped input / output (I / O) (MMIO) space when numbering a multitude of different buses of a system.

[0139] Example 29 can include the subject of Example 28, with the instructions further being executable to restrict a permissible range of BD to be allocated to a secondary side of a root port bridge.

[0140] Example 30 is a system comprising a switching component, a hierarchy of components associated with the switching component, a set of one or more other components associated with the switching component, wherein the set of one or more other components is addressed according to a first addressing system, wherein the hierarchy of components is addressed according to a second addressing system, wherein the first addressing system uses a single bus number within a bus / component / function (BDF) address space for each component in the first set of components, and the second addressing system uses a single bus component number for each component in the second set of components.

[0141] Example 31 can include the subject of Example 30, wherein the switching device has a first port for connecting to the hierarchy of components, and the first port has bridge logic to determine whether a particular packet is to be routed to a primary side of the bridge using the first addressing system or to a secondary side of the bridge using the second addressing system.

[0142] Example 32 may include the subject of any of Examples 30-31, and may further include a capability register which is to be coded to selectively enable support for the secondary addressing system on a particular port of the switch.

[0143] Example 33 can feature the subject of any of Examples 30-32, wherein the switching component features a root complex.

[0144] A design can go through various stages, from conception to simulation to manufacturing. Data representing a design can depict it in a number of ways. First, the hardware, which is useful for simulations, can be represented using a hardware description language or another functional description language. Additionally, at some stages of the design process, a circuit-level model with logic and / or transistor gates can be created. Furthermore, most designs eventually reach a data level that represents the physical placement of various components within the hardware model.In the case where conventional semiconductor manufacturing techniques are used, the data representing the hardware model can be the data specifying the presence or absence of various features on different mask layers for masks used to fabricate the integrated circuit. In any representation of the design, the data can be stored on any form of machine-readable medium. A memory or magnetic or optical storage medium, such as a disc, can be the machine-readable medium for storing information transmitted via optical or electrical waves that are modulated or otherwise generated to transmit such information. When an electrical carrier wave specifying or carrying code or design is transmitted, a new copy is created by copying, buffering, or retransmitting the electrical signal.A communications provider or network provider can therefore store, at least temporarily, an element, such as information encoded in a carrier wave, on a tangible, machine-readable medium, embodying the techniques of embodiments of the present invention.

[0145] A module, as used here, refers to any combination of hardware, software, and / or firmware. For example, a module includes hardware, such as a microcontroller, associated with a non-volatile medium for storing code designed to be executed by the microcontroller. Thus, in one embodiment, a reference to a module refers to the hardware specifically configured to recognize and / or execute the code stored on the non-volatile medium. Furthermore, in another embodiment, the use of a module refers to the non-volatile medium containing the code specifically adapted for execution by the microcontroller to perform predetermined operations. As can be inferred, in yet another embodiment, the term "module" (in this example) can refer to the combination of the microcontroller and the non-volatile medium.Module boundaries, often depicted as separate, frequently vary in general terms and potentially overlap. For example, a first and a second module might share hardware, software, firmware, or a combination thereof, while potentially retaining some independent hardware, software, or firmware. In one embodiment, the term "logic" refers to hardware such as transistors, registers, or other hardware, such as programmable logic devices.

[0146] The use of the phrase "for" or "designed for" in an embodiment relates to arranging, assembling, manufacturing, offering for sale, importing, and / or designing a device, hardware, logic, or element to perform a designated or specific task. In this example, a device or element thereof that is not operating is still "designed" to perform a designated task if it is designed, coupled, and / or interconnected to perform the designated task. As a purely illustrative example, a logic gate may provide a 0 or a 1 during operation. However, a logic gate that is "designed" to provide an enable signal to a clock source does not have every potential logic gate that might provide a 1 or a 0.Instead, the logic gate is one that is coupled in some way, where, during operation, the output 1 or 0 is intended to enable the clock. It should be noted again that the use of the term "designed for" does not require operation, but instead focuses on the latent state of a device, hardware, and / or element, where the latent state of the device, hardware, and / or element is designed to perform a specific task when the device, hardware, and / or element is operational.

[0147] Furthermore, the use of the phrases "capable of" and / or "operable of" in an embodiment refers to any device, logic, hardware, and / or element that is designed to enable the use of the device, logic, hardware, and / or element in a specified manner. It should be noted, as above, that in an embodiment, the use of "capable of" or "operable of" refers to the latent state of the device, logic, hardware, and / or element, wherein the device, logic, hardware, and / or element is not operational but is designed to enable the use of the device in a specified manner.

[0148] As used here, a value has any known representation of a number, a state, a logical state, or a binary logical state. Often, the use of logic levels, logic values, or logical values ​​is also referred to as 1s and 0s, which simply represent binary logic states. For example, 1 refers to a high logic level, and 0 refers to a low logic level. In one embodiment, a memory cell, such as a transistor or flash cell, may be able to hold a single logical value or multiple logical values. However, other representations of values ​​have been used in computer systems. The decimal number ten, for example, can also be represented as a binary value of 1010 and a hexadecimal letter A. A value, therefore, has any representation of information capable of being held in a computer system.

[0149] Furthermore, states can be represented by values ​​or parts of values. For example, a first value, such as a logical one, can represent a default or initial state, whereas a second value, such as a logical zero, can represent a non-default state. Additionally, in one embodiment, the terms reset and set each refer to a default and an updated value or state, respectively. For example, a default value potentially has a high logical value, i.e., reset, whereas an updated value potentially has a low logical value, i.e., set. It should be noted that any combination of values ​​can be used to represent any number of states.

[0150] The embodiments of methods, hardware, software, firmware, or code described above may be stored by means of instructions or code that are executable by a processing element on a machine-accessible, machine-readable, computer-accessible, or computer-readable medium. A non-volatile machine-accessible / readable medium includes any mechanism that provides (that is, stores and / or transmits) information in a form readable by a machine, such as a computer or electronic system.For example, a non-volatile, machine-accessible medium includes random access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; a magnetic or optical storage medium; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; other forms of storage devices for holding information received from transitory (propagated) signals (for example, carrier waves, infrared signals, digital signals); etc., which must be distinguished from the non-volatile media that can receive information from them.

[0151] Instructions used to program logic for executing embodiments of the invention can be stored within a memory in the system, such as DRAM, cache, flash memory, or other storage. Furthermore, the instructions can be distributed over a network or via other computer-readable media.Thus, a machine-readable medium can include, among other things, any mechanism for storing or transmitting information in a form that can be read by a machine (for example, a computer), such as floppy disks, optical disks, compact discs, read-only storage (CD-ROMs) and magneto-optical disks, read-only storage (ROMs), random access storage (RAM), erasable programmable read-only storage (EPROM), electrically erasable programmable read-only storage (EEPROM), magnetic or optical cards, flash memory, or any tangible, machine-readable storage used in the transmission of information over the Internet via electrical, optical, acoustic, or other forms of propagated signals (for example, carrier waves, infrared signals, digital signals, etc.).The computer-readable medium therefore includes any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions and information in a form that can be read by a machine (for example, a computer).

[0152] Throughout this specification, reference to "(exactly) one embodiment" or "an embodiment" means that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the present techniques. The appearance of the phrases "in (exactly) one embodiment" or "in an embodiment" at various points throughout the specification therefore does not always necessarily refer to the same embodiment. Furthermore, the particular features, structures, or characteristics in one or more embodiments may be combined in any suitable manner.

Claims

[1] At least one machine-accessible storage medium on which instructions are stored, wherein the instructions, when executed on a machine, cause the machine to do the following: Identifying a plurality of components in a system, wherein each component of the plurality of components in the system is connected by at least one of a plurality of logical buses; Assigning a respective address to each of the multitude of components, wherein assigning the address to a component includes the following: Determine whether the address is to be assigned according to a first addressing system or a second bus addressing system, wherein the first addressing system assigns a single bus number within a bus / component / puncture (BDF) address space to each component addressed in the first addressing system, and the second bus addressing system assigns a single bus component number within the BDF address space. [2] Storage medium according to claim 1, wherein a special bus number is reused to address two or more components in the second addressing system. [3] Storage medium according to claim 1, wherein the assignment of addresses comprises defining a range of bus numbers in the BDF address space to be used to address components according to the second addressing system. [4] Storage medium according to claim 3, wherein the range of bus numbers is associated with a special switch and the bus numbers in the range of bus numbers are used in the bus component numbers to be assigned to each component connected to the switch. [5] Storage medium according to claim 4, wherein the components connected to the switch comprise a segment. [6] Storage medium according to claim 1, wherein the addresses include configuration addresses. [7] Storage medium according to claim 1, wherein the BDF address space comprises a Peripheral Component Interconnect (PCI)-based address space. [8] Storage medium according to claim 7, wherein each bus component number comprises an eight-bit bus number and a five-bit component number. [9] Device comprising the following: a port to receive a special packet, wherein the port comprises a flattening gantry bridge (FPB), the FPB comprising a primary side and a secondary side, the primary side connecting to a first set of components addressed according to a first addressing system, and the second side connecting to a second set of components addressed according to a second addressing system; wherein the FPB further determines whether the special packet is to be routed on the primary side or the secondary side based on address information in the special packet, wherein the first addressing system uses a single bus number within a bus / component / function (BDF) address space for each component in the first set of components, and the second bus addressing system uses a single bus component number in the BDF space for each component in the second set of components. [10] Device according to claim 9, wherein the respective bus component numbers assigned to a plurality of components in the second set of components each comprise a special bus number and a distinct component number. [11] Device according to claim 9, wherein the primary addressing system comprises an established addressing system. [12] Device according to claim 9, wherein the BDF addressing space comprises a Peripheral Component Interconnect Express (PCIe) configuration space. [13] Device according to claim 9, further comprising a plurality of ports, wherein the port comprises a particular one of the plurality of ports, and at least one other port in the plurality of ports comprises an FDB. [14] Device according to claim 13, wherein the plurality of ports includes at least one port without an FPB. [15] Device according to claim 13, further comprising a switch, wherein the switch comprises the plurality of ports. [16] Device according to claim 13, further comprising a root complex, wherein the root complex comprises the plurality of ports. [17] Device according to claim 16, wherein the FPB capability register indicates whether the second addressing system is enabled for one or more of the plurality of different resource types. [18] Device according to any one of claims 9 to 17, further comprising a BD Control 1 register. [19] Device according to any one of claims 9 to 18, further comprising a BD Vector Control 2 register. [20] Device according to any one of claims 9 to 19, further comprising a BD Vector Access Control register. [21] Device according to any one of claims 9 to 20, further comprising a BD Vector Access Data register. [22] Device according to any one of claims 9 to 21, further comprising a MEM Low Vector Control Register. [23] Device according to any one of claims 9 to 22, further comprising a MEM Low Vector Access Control register. [24] Device according to any one of claims 9 to 23, further comprising a MEM Low Vector Access Data Register. [25] Device according to any one of claims 9 to 24, further comprising a MEM High Vector Control 1 register. [26] Device according to any one of claims 9 to 25, further comprising a MEM High Vector Control 2 register. [27] Device according to any one of claims 9 to 26, further comprising a MEM High Vector Access Control register. [28] Device according to any one of claims 9 to 27, further comprising a MEM High Vector Access Data register. [29] At least one machine-accessible storage medium on which instructions are stored, wherein the instructions, when executed on a machine, cause the machine to do the following: Configuring a component's registers to support a primary bus addressing system in a Bus / Component / Function (BDF) space and an alternative bus addressing system that uses the same bus number within a Memory-Mapped Input / Output (I / O) (MMIO) space when numbering a variety of logically distinct logical buses of a system. [30] Storage medium according to claim 29, wherein the instructions are further executable to limit a permissible range of BD to be allocated to a secondary side of a root port bridge. [31] System comprising the following: a switching component; a hierarchy of components that is connected to the switching component; a set of one or more components that is / are connected to the switching component; wherein the set of one or more other components is addressed according to a first addressing system, the hierarchy of components is addressed according to a second addressing system, the first addressing system uses a single bus number within a bus / component / function (BDF) address space for each component in the first set of components, and the second addressing system uses a single bus component number in the BDF space for each component in the second set of components. [32] System according to claim 31, wherein the switching device comprises a first port for connecting to the hierarchy of components and the first port comprises bridge logic to determine whether a particular packet is to be routed on a primary side of the bridge using the first addressing system or on a secondary side of the bridge using the second addressing system. [33] System according to one of claims 31 to 32, further comprising a capability register which is to be coded to selectively enable support for the secondary addressing system on a particular port of the switch. [34] System according to one of claims 31 to 33, wherein the switching component comprises a root complex.

Citation Information

Patent Citations

  • Information processing device, control method, and non-transitory computer-readable recording medium having control program recorded thereon

    US20150127868A1