Directed Interrupt Virtualization
By using the mapping table of the bus attachment device in a multiprocessor system to convert the interrupt target ID into a logical processor ID and directly address the target processor, the problem of low interrupt signal routing efficiency is solved, and the efficiency and accuracy of interrupt signal processing is improved.
Patent Information
- Application Number
- CN202080014373.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-14
- Filing Date
- 2020-01-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-01-10
AI Technical Summary
In multiprocessor computer systems, the routing efficiency of interrupt signals is low, especially in a virtual machine environment, which makes the processor unsuitable for processing interrupt signals, affecting performance.
The interrupt target ID is converted into a logical processor ID through a bus attachment device using a mapping table and directly addressing the target processor, forwarding the interrupt signal to the most suitable processor.
The efficiency of interrupt signal processing is improved, cached traffic is reduced, and the interrupt signal can be processed quickly and accurately by the processor.
Smart Images

Figure CN113454590B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] The present invention generally relates to interrupt handling within a computer system, and more particularly to handling interrupts generated by a bus connection module in a multi-processor computer system.
[0002] Interrupts are used to signal events that require the attention of a processor. For example, hardware devices (e.g., hardware devices connected to a processor via a bus) use interrupts to convey that they require attention from the operating system. In the case where the receiving processor is currently performing some activity, the receiving processor may suspend its current activity, save its state, and handle the interrupt, e.g., by executing an interrupt handler, in response to receiving the interrupt signal. The interruption of the current activity of the processor caused by receiving the interrupt signal is only temporary. After the interrupt has been processed, the processor may resume its suspended activity. Thus, interrupts can improve performance by eliminating non-productive waiting times for the processor to wait for external events in a polling loop.
[0003] In a multi-processor computer system, interrupt routing efficiency issues may arise. The challenge is to efficiently forward interrupt signals sent by hardware devices such as bus connection modules to a processor among the multiple processors allocated for the operating system. This can be particularly challenging in the case where interrupts are used to communicate with a guest operating system on a virtual machine. A hypervisor or virtual machine monitor (VMM) creates and runs one or more virtual machines, i.e., guest machines. The virtual machine provides a guest operating system that runs on the same platform as the virtual operating platform while hiding the physical characteristics of the underlying platform. Using multiple virtual machines allows multiple operating systems to run in parallel. Since it is executed on a virtual operating platform, the view of the processor by the guest operating system Figure 1 generally may be different from that of the underlying (e.g., the physical view of the processor). The guest operating system uses a virtual processor ID (identifier) to identify the processor, which is usually inconsistent with the underlying logical processor ID. The hypervisor that manages the execution of the guest operating system defines a mapping between the underlying logical processor ID and the virtual processor ID used by the guest operating system. However, this mapping and the selection of the processor scheduled for use by the guest operating system are not static, but can be changed by the hypervisor while the guest operating system is running, without the knowledge of the guest operating system.
[0004] This challenge is typically addressed by using a broadcast to forward the interrupt signal. When using a broadcast, the interrupt signal is continuously forwarded among multiple processors until a processor suitable for handling the interrupt signal is encountered. However, in the case of multiple processors, the probability that the processor that first receives the broadcast interrupt signal is indeed suitable for handling that interrupt signal can be quite low. Additionally, being suitable for handling the interrupt signal does not necessarily mean that the corresponding processor is the best choice for handling the interrupt. SUMMARY OF THE INVENTION
[0005] Various embodiments provide a method, a computer system, and a computer program product for providing an interrupt signal to a guest operating system that is executed by one or more of a plurality of processors of a computer system allocated for use by the guest operating system, as described by the subject matter of the independent claims. Advantageous embodiments are described in the dependent claims. Embodiments of the present invention may be freely combined with each other without mutual exclusion.
[0006] In one aspect, the present invention relates to a method for providing an interrupt signal to a guest operating system that is executed by one or more of a plurality of processors of a computer system allocated for use by the guest operating system, the computer system further including one or more bus connection modules operably connected to the plurality of processors via a bus and bus-attached devices, each of the plurality of processors being assigned a logical processor ID by the bus-attached device for addressing the corresponding processor, and each of the plurality of processors allocated for use by the guest operating system further being assigned an interrupt target ID by the guest operating system and the one or more bus connection modules for locating the corresponding processor, the method including: receiving, by the bus-attached device, an interrupt signal having an interrupt target ID from one of the bus connection modules, the interrupt target ID identifying one of the processors that is assigned by the guest operating system as a target processor for handling the interrupt signal, converting, by the bus-attached device, the received interrupt target ID to the logical processor ID of the target processor using a mapping table included in the bus-attached device, the mapping table mapping the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the plurality of processors, directly addressing, by the bus-attached device, the target processor using the logical processor ID of the target processor, and forwarding the interrupt signal to be processed to the target processor.
[0007] On the other hand, the present invention relates to a system for providing an interrupt signal to a guest operating system, the guest operating system being executed by one or more of a plurality of processors of a computer system allocated for use by the guest operating system. The computer system further includes one or more bus connection modules operably connected to the plurality of processors via a bus and bus-attached devices. Each of the plurality of processors is assigned a logical processor ID by the bus-attached device for addressing the corresponding processor. Each of the plurality of processors allocated for use by the guest operating system is also assigned an interrupt target ID by the guest operating system and the one or more bus connection modules for finding the corresponding processor. The computer system is configured to execute a method, the method including: receiving, by the bus-attached device, an interrupt signal having an interrupt target ID from one of the bus connection modules, the interrupt target ID identifying one of the processors assigned by the guest operating system as the target processor for processing the interrupt signal; converting, by the bus-attached device, the received interrupt target ID into the logical processor ID of the target processor using a mapping table included in the bus-attached device, the mapping table mapping the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the plurality of processors; directly addressing, by the bus-attached device, the target processor using the logical processor ID of the target processor; and forwarding the interrupt signal to be processed to the target processor.
[0008] On the other hand, the present invention relates to a computer program product for providing an interrupt signal to a guest operating system, the guest operating system being executed by one or more processors of a computer system allocated for use by the guest operating system. The computer system further includes one or more bus connection modules operably connected to the plurality of processors via a bus and bus-attached devices. Each of the plurality of processors is assigned a logical processor ID by the bus-attached device for addressing the corresponding processor. Each of the plurality of processors allocated for use by the guest operating system is also assigned an interrupt target ID by the guest operating system and the one or more bus connection modules for finding the corresponding processor. The computer program product includes a computer-readable non-transitory medium readable by a processing circuit and stores instructions executed by the processing circuit for performing a method, the method including: receiving, by the bus-attached device, an interrupt signal having an interrupt target ID from one of the bus connection modules, the interrupt target ID identifying one of the processors allocated by the guest operating system as the target processor for processing the interrupt signal; converting, by the bus-attached device, the received interrupt target ID into a logical processor ID of the target processor using a mapping table included in the bus-attached device, the mapping table mapping the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the plurality of processors; directly addressing, by the bus-attached device, the target processor using the logical processor ID of the target processor; and forwarding the interrupt signal to be processed to the target processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Hereinafter, embodiments of the present invention will be explained in more detail by way of example only with reference to the accompanying drawings, in which:
[0010] Figure 1 A schematic diagram of an exemplary computer system is depicted.
[0011] Figure 2 A schematic diagram of an exemplary virtualization scheme is depicted.
[0012] Figure 3 A schematic diagram of an exemplary virtualization scheme is depicted.
[0013] Figure 4 A schematic diagram of an exemplary virtualization scheme is depicted.
[0014] Figure 5 A schematic diagram of an exemplary computer system is depicted.
[0015] Figure 6 A schematic flow chart of an exemplary computer system is depicted.
[0016] Figure 7 A schematic flow chart of an exemplary method is depicted.
[0017] Figure 8 Depicts a schematic flowchart of an exemplary method,
[0018] Figure 9 Depicts a schematic flowchart of an exemplary method,
[0019] Figure 10 Depicts a schematic diagram of an exemplary computer system,
[0020] Figure 11 Depicts a schematic flowchart of an exemplary method,
[0021] Figure 12 Depicts a schematic flowchart of an exemplary method,
[0022] Figure 13 Depicts a schematic diagram of an exemplary data structure,
[0023] Figure 14 Depicts a schematic diagram of an exemplary vector structure,
[0024] Figure 15 Depicts a schematic diagram of an exemplary vector structure,
[0025] Figure 16 depicts a schematic diagram of an exemplary vector structure,
[0026] Figure 17 depicts a schematic diagram of an exemplary vector structure,
[0027] Figure 18 Depicts a schematic diagram of an exemplary computer system,
[0028] Figure 19 Depicts a schematic diagram of an exemplary computer system,
[0029] Figure 20 Depicts a schematic diagram of an exemplary computer system,
[0030] Figure 21 Depicts a schematic diagram of an exemplary computer system,
[0031] Figure 22 depicts a schematic diagram of an exemplary unit, and
[0032] Figure 23 Depicts a schematic diagram of an exemplary computer system. Detailed implementation
[0033] Descriptions of different embodiments of the present invention will be given for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, practical applications, or technical improvements in the technology that have emerged in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.
[0034] Embodiments can have the beneficial effect of enabling a bus-attached device to directly address a target processor. Thus, an interrupt signal can be directed by an issuing bus connection module that selects a target processor ID to a specific processor, i.e., the target processor of a multiprocessor computer system. For example, a processor that has previously engaged in activities related to the interrupt can be selected as the target processor for the interrupt signal. Processing the interrupt signal by the same processor that performed the corresponding activity may result in a performance advantage because, in the case where the same processor also processes the interrupt signal, all data in the context of that interrupt may already be available to that processor and / or stored in the local cache, enabling the corresponding processor to access it quickly without requiring a large amount of cache traffic.
[0035] Thus, a broadcast of an interrupt signal to processors that cannot guarantee optimal handling of the interrupt from a performance perspective (such as minimizing cache traffic) can be avoided. Instead of submitting the interrupt signal to all processors, each of which attempts to handle the interrupt signal and one processor wins, the interrupt signal can be provided directly to the target processor to improve the efficiency of interrupt signal processing.
[0036] The interrupt mechanism can be implemented with directed interrupts. A bus-attached device can be enabled to directly address the target processor with the logical processor ID of the target processor when forwarding an interrupt signal to be processed to the target processor of the interrupt signal defined by the issuing bus connection module. Converting the interrupt target ID to the logical processor ID by the bus-attached device can further ensure that the same processor is always addressed from the perspective of the guest operating system, even if the mapping between the interrupt target ID and the logical processor ID or the selection of the processor to be scheduled for use by the guest operating system may be by the hypervisor.
[0037] To convert an interrupt target ID provided with an interrupt signal to a logical processor ID, a static mapping provided by a mapping table of a bus-attached device can be used. Thus, enabling the bus-attached device to identify and use the logical processor ID to directly address the target processor identified by the interrupt request. According to an embodiment, the interrupt signal is received in the form of a message signaled interrupt that includes the interrupt target ID of the target processor. Using message signaled interrupt (MSI) is a method by which a bus connection module (such as a Peripheral Component Interconnect (PCI) or a Peripheral Component Interconnect Express (PCIe) function) generates a Central Processing Unit (CPU) interrupt to notify a client operating system using the corresponding central processing unit of the occurrence of an event or the existence of a certain state. MSI provides an in-band method of signaling interrupts using special in-band messages, thus avoiding the need for a dedicated path separate from the main data path, such as dedicated interrupt pins on each device, to send such control information. MSI instead relies on the exchange of special messages indicating interrupts over the main data path. When the bus connection module is configured to use MSI, the corresponding module requests an interrupt by performing an MSI write operation of a specified number of bytes of data to a specific address. The combination of this specific address (i.e., the MSI address) and a unique data value (i.e., the MSI data) is called the MSI vector.
[0038] Modern PCIe standard adapters have the ability to submit multiple interrupts. For example, MSI-X allows a bus connection module to allocate up to 2048 interrupts. Thus, enabling individual interrupts to be directed to different processors, for example, in high-speed networking applications that rely on multi-processor systems. MSI-X allows the allocation of multiple interrupts, each with a separate MSI address and MSI data value.
[0039] To transmit an interrupt signal, an MSI-X message can be used. An MSI-X data table can be used to determine the required content of the MSI-X message. The MSI-X data table local to the bus connection module (i.e., the PCIe adapter / function) can be indexed by a number assigned to each interrupt signal (also known as an Interrupt Request (IRQ)). The content of the MSI-X data table is controlled by the client operating system and can be set to the operating system under the guidance of hardware and / or firmware. A single PCIe adapter can include multiple PCIe functions, and each PCIe function can have an independent MSI-X data table. This can be the case, for example, for a single root input / output virtualization (SR-IOV) or a multi-function device.
[0040] An interrupt target ID (such as, for example, a virtual processor ID) can be directly encoded as part of a message (such as an MSI-X message) sent by a bus connected module that includes an interrupt signal. The message (e.g., an MSI-X message) can include a requester ID, i.e., the ID of the bus connected module, the aforementioned interrupt target ID, a DIBV or AIBV index, an MSI address, and MSI data. The MSI-X message can provide 64 bits for the MSI address and 32 bits for the data. The bus connected module can use an MSI request for an interrupt by performing an MSI write operation of a specific MSI data value to a specific MSI address.
[0041] The device table is a shared table that can be fully indexed by the requester ID (RID) of the interrupt requester (i.e., the bus connected module). The bus attached device remaps and posts the interrupt, i.e., the bus attached device translates the interrupt target ID and uses the interrupt target ID to directly address the target processor.
[0042] The guest operating system can use the virtual processor ID to identify processors in a multiprocessor computer system. Thus, the view of the processors by the guest operating system may be different from the view of the underlying system that uses logical processor IDs. A bus connected module that provides resources used by the guest operating system can use the virtual processor ID as a resource for communicating with the guest operating system. For example, the MSI-X data table can be under the control of the guest operating system. As an alternative to the virtual processor ID, any other ID can be defined for addressing processors by the bus connected module.
[0043] The interrupt is submitted to the guest operating system or other software executing thereon, such as other programs, etc. As used herein, the term operating system includes operating system device drivers.
[0044] As used herein, the term bus connected module can include any type of bus connected module. According to an embodiment, the module can be a hardware module, such as a storage function, a processing module, a network module, a cryptographic module, a PCI / PCIe adapter, other types of input / output modules, etc. According to other embodiments, the module can be a software module, i.e., a function such as a storage function, a processing function, a network function, a cryptographic function, a PCI / PCIe function, other types of input / output functions. Thus, in the examples given herein, unless otherwise indicated, the module can be used interchangeably with a function (e.g., a PCI / PCIe function) and an adapter (e.g., a PCI / PCIe function).
[0045] Embodiments can have the following advantages: providing an interrupt signal routing mechanism (e.g., MSI-X message routing mechanism) that allows it to keep the bus-connected modules (e.g., PCIe adapters and functions) and the device drivers for operating or controlling the bus-connected modules unchanged. Additionally, it can prevent the hypervisor from intercepting the underlying architecture for implementing communication between the bus-connected module and the guest operating system, such as the PCIe MSI-X architecture. In other words, changes to the interrupt signal routing mechanism can be implemented outside the hypervisor and the bus-connected module.
[0046] According to an embodiment, the computer system further includes a memory, and the bus-attached device is operatively connected to the memory. The method further includes: retrieving, by the bus-attached device, a copy of a device table entry from a device table stored in the memory, the device table entry including a direct signaling indicator indicating whether to directly address a target processor. If the direct signaling indicator indicates direct forwarding of the interrupt signal, directly addressing the target processor with the logical processor ID of the target processor and performing forwarding of the interrupt signal. Otherwise, the bus-attached device forwards the interrupt signal to be processed by broadcast to the plurality of processors.
[0047] Embodiments can have the beneficial effect of using the direct signaling indicator to control whether to forward the interrupt signal by direct addressing or by broadcast. Using the direct signaling indicator for each bus-connected module, a separate predefined selection can be provided for performing direct addressing or broadcast on the interrupt signal received from that bus-connected module.
[0048] According to one embodiment, the direct signaling indicator is implemented with a single bit. Embodiments can have the following beneficial effects: the direct signaling indicator is provided in the form of minimal storage space and can be processed quickly and efficiently.
[0049] According to one embodiment, during initialization of the guest operating system, the direct signaling indicator is set to a static indicator for the guest operating system. According to an embodiment, the mapping of the interrupt target ID of the processor allocated for use by the guest operating system to the logical processor IDs of the plurality of processors is a static mapping defined by a mapping table. Embodiments can have the beneficial effect that the bus-attached device is equipped with a mapping that enables the attached device to perform direct addressing without having to obtain a large amount of data from the memory.
[0050] According to one embodiment, a bus-attached device checks whether a copy of a device table entry is cached in a local cache operably connected to the bus-attached device. If a copy of the device table entry is cached, the retrieval of the copy of the device table entry is retrieved from the corresponding cache. Otherwise, the retrieval of the device table entry is retrieved from the memory. The embodiment can have the effect of ensuring a fast and efficient retrieval of the copy of the device table entry. According to an embodiment, the memory further includes an interrupt summary vector, and the device table entry further includes an interrupt summary vector address indicator indicating the memory address of the interrupt summary vector. The interrupt summary vector includes an interrupt summary indicator for each bus connection module, and each interrupt summary indicator is assigned to a bus connection module, indicating whether there is an interrupt signal issued by the corresponding bus connection module to be processed. The method further includes the bus-attached device using the indicated memory address of the interrupt summary vector to update the interrupt summary indicator assigned to the bus connection module from which the interrupt signal is received, such that the updated interrupt summary indicator indicates that there is an interrupt signal issued by the corresponding bus connection module to be processed.
[0051] The embodiment can have the beneficial effect of monitoring and recording in which bus connection module there is an interrupt signal to be processed. This information can be particularly useful, for example, in the case where direct addressing fails or is unavailable and a broadcast must be performed as a backup. According to an embodiment, the interrupt summary vector is implemented as a contiguous region. The embodiment can have the following beneficial effects: the interrupt summary vector is provided in the form of minimal storage space and can be processed quickly and efficiently. The contiguous region can be, for example, a single cache line. According to one embodiment, each interrupt summary indicator is implemented as a single bit. The embodiment can have the following beneficial effects: the interrupt summary indicator is provided in the form of minimal storage space and is quickly and efficiently processable.
[0052] According to an embodiment, the memory further includes a directed interrupt summary vector, and the device table entry further includes a directed interrupt summary vector address indicator indicating the memory address of the directed interrupt summary vector. The directed interrupt summary vector includes a directed interrupt summary indicator for each interrupt target ID, and each directed interrupt summary indicator is assigned to an interrupt target ID, indicating whether there is an interrupt signal addressed to the corresponding interrupt target ID to be processed. Wherein, the method further includes the bus-attached device using the indicated memory address of the directed interrupt summary vector to update the interrupt summary indicator assigned to the bus connection module of the target processor ID addressed by the received interrupt signal, such that the updated interrupt summary indicator indicates that there is an interrupt signal addressed to the corresponding interrupt target ID to be processed.
[0053] Embodiments can have the beneficial effect of monitoring and recording in which bus connection module there is an interrupt signal to be processed. This information can be particularly useful, for example, in the case where direct addressing fails or is unavailable and broadcasting must be performed as a backup. If a directed interrupt summary indicator is assigned to a single interrupt target ID, it is only necessary to check this single indicator to determine whether there is an interrupt to be processed by a specific processor.
[0054] When, for example, an interrupt cannot be directly sent because the hypervisor does not schedule the target processor, the guest operating system may benefit from sending the interrupt with the initial expected association relationship, i.e., information about which processor the interrupt is expected to be processed by, in a broadcast. In this case, the bus-attached device can set the position bit in the DISB that specifies the target processor after setting the DIBV and before sending a broadcast interrupt request to the guest operating system. If the guest operating system receives a broadcast interrupt request, it can identify which target processors have an interrupt signal to be processed as indicated in the DIBV by scanning and disabling the direct interrupt summary indicators in the DISB (e.g., scanning and resetting the direct interrupt bits). Thus, the guest operating system can be enabled to decide whether the interrupt signal is to be processed by the current processor that received the broadcast or further forwarded to the original target processor.
[0055] According to an embodiment, the directed interrupt summary vector is implemented as a contiguous region. Embodiments can have the beneficial effect that the directed interrupt summary vector is provided in the form of minimal storage space and is quickly and efficiently processable. The contiguous region can be, for example, a single cache line. According to an embodiment, each of the directed interrupt summary indicators is implemented as a single bit. Embodiments can have the beneficial effect that the directed interrupt summary indicators are provided in the form of minimal storage space and are quickly and efficiently processable.
[0056] According to an embodiment, the memory further includes one or more interrupt signal vectors, the device table entry further includes an interrupt signal vector address indicator indicating the memory address of the interrupt signal vector in the one or more interrupt signal vectors, each of the interrupt signal vectors includes one or more signal indicators, each interrupt signal indicator is assigned to one of the one or more bus connection modules and an interrupt target ID, indicating whether an interrupt signal has been received from the corresponding bus connection module addressed to the corresponding interrupt target ID, and the method further includes: using the indicated memory address of the interrupt signal vector by the bus-attached device to select the interrupt signal indicator assigned to the bus connection module that issues the received interrupt signal and the interrupt target ID addressed by the received interrupt signal, and updating the selected interrupt signal indicator such that the selected interrupt signal indicator indicates that there is an interrupt signal to be processed issued by the corresponding bus connection module and addressed to the corresponding interrupt target ID.
[0057] According to an embodiment, the interrupt signal vectors each include an interrupt signal indicator for each interrupt target ID assigned to a corresponding interrupt target ID, each of the interrupt signal vectors being assigned to a single bus connection module, and the interrupt signal indicators of the corresponding interrupt signal vectors being further assigned to the corresponding single bus connection module. The embodiment can have the beneficial effect of enabling a guest operating system to keep track of which target processors the bus connection module has issued pending interrupt signals for.
[0058] According to an embodiment, the interrupt signal vectors each include an interrupt signal indicator for each bus connection module assigned to a corresponding bus connection module, each of the interrupt signal vectors being assigned to a single target processor ID, and the interrupt signal indicators of the corresponding interrupt signal vectors being further assigned to the corresponding target processor ID. The embodiment can have the beneficial effect of enabling a guest operating system to keep track of which bus attachment modules have issued interrupt signals to be processed by a particular target processor.
[0059] The interrupt signal vectors can be implemented as directed interrupt signal vectors sorted according to the target processor ID, i.e., optimized for tracking directed interrupts. In other words, the main order criterion is the target processor ID rather than the requester ID identifying the bus connection module that issued the interrupt request. Depending on the number of bus connection modules, each directed interrupt signal vector can include one or more directed interrupt signal indicators.
[0060] Therefore, it is possible to avoid sorting the interrupt signal indicators (e.g., in the form of interrupt signaling bits) that indicate that a single interrupt signal (e.g., in the form of an MSI-X message) has been received sequentially within a contiguous region of memory (e.g., a cache line) of a single bus connection module (such as a PCIe function). Enabling and / or disabling an interrupt signal indicator (e.g., by setting or resetting the interrupt signaling bit) requires transferring the corresponding contiguous region of memory to one of the processors to change the corresponding interrupt signal indicator accordingly.
[0061] A processor can process all indicators that are its responsibility from the perspective of the guest operating system, i.e., in particular, all indicators assigned to the corresponding processor. This can achieve a performance advantage because, in the case where each processor is processing all data assigned to the processor, the likelihood that the data required in this context is provided to the processor and / or stored in the local cache may be high, such that the corresponding data of the processor can be accessed quickly without a large amount of cache traffic.
[0062] However, each processor attempting to process all the indicators it is responsible for can still result in a relatively high cache traffic among the processors because each processor needs to write all the cache lines for all functions. Since the indicators assigned to each individual processor can be distributed across all contiguous regions, such as cache lines.
[0063] The interrupt signaling indicators can be reordered in the form of a directed interrupt signaling vector such that all the interrupt signaling indicators assigned to the same interrupt target ID are grouped in the same contiguous region of memory (e.g., a cache line). Thus, a processor intending to process the indicators assigned to the corresponding processor (i.e., the interrupt target ID) may only need to load a single contiguous region of memory. Thus, the contiguous region of each interrupt target ID is used instead of the contiguous region of each bus connection module. Each processor may only need to scan and update a single contiguous region of memory, e.g., the cache line of all the interrupt signals targeted at a particular processor identified by the interrupt target ID received from all available bus connection modules. According to an embodiment, the hypervisor may apply an offset to the guest operating system to align the bits to a different offset.
[0064] According to an embodiment, the interrupt signal vectors are each implemented as contiguous regions in memory. The embodiment can have the beneficial effect of occupying the minimum storage space in the form of the provided interrupt signal vectors and being quickly and efficiently processable. The contiguous region can be, for example, a cache line. According to an embodiment, the interrupt signal indicators are each implemented as a single bit. The embodiment can have the beneficial effect of occupying the minimum storage space in the form of the provided interrupt signal indicators and being quickly and efficiently processable.
[0065] According to an embodiment, the device table entry further includes a logical partition ID identifying the logical partition to which the guest operating system is assigned, and the forwarding of the interrupt signal by the bus-attached device further includes forwarding the logical partition ID along with the interrupt signal. The embodiment can have the beneficial effect of enabling the receiving processor to check which guest operating system the interrupt signal is addressed to.
[0066] According to an embodiment, the bus connection module includes a mapping table for each logical partition ID. The embodiment can have the beneficial effect of providing separate mapping tables, for example, for the hypervisor and / or the guest operating system.
[0067] According to an embodiment, the method further includes the bus-attached device retrieving an interrupt subclass ID identifying the interrupt subclass to which the received interrupt signal is assigned, and the forwarding of the interrupt signal by the bus-attached device further includes forwarding the interrupt subclass ID along with the interrupt signal.
[0068] According to an embodiment, a processor of a computer system is adapted to execute a plurality of guest operating systems, and a bus-attached device includes a mapping table for each of the plurality of guest operating systems.
[0069] According to an embodiment, the method further includes receiving, by the bus-attached device, a direct memory access request from a bus connection module to update status information of the bus connection module in a memory - a status update of the bus connection module triggers an interrupt signal; after receiving the request, performing, by the bus-attached device, a direct memory access to the memory to update the status information of the bus connection module in the memory.
[0070] According to an embodiment, instructions provided on a computer-readable non-transitory medium for execution by a processing circuit are configured to perform any embodiment of the method of providing an interrupt signal to a guest operating system as described herein.
[0071] According to an embodiment, a computer system is further configured to perform any embodiment of the method of providing an interrupt signal to a guest operating system as described herein.
[0072] Figure 1 An exemplary computer system 100 for providing an interrupt signal to a guest operating system is depicted. The computer system 100 includes a plurality of processors 130 for executing guest operating systems. The computer system 100 also includes a memory 140, also referred to as an internal memory or main memory. The memory 140 may provide memory space that is allocated for use by hardware, firmware, and software components included in the computer system 100, i.e., memory segments. The memory 140 may be used by the hardware and firmware as well as software (e.g., a hypervisor, host / guest operating systems, applications, etc.) of the computer system 100. One or more bus connection modules 120 are operably connected to the plurality of processors 130 and the memory 140 via a bus 102 and a bus attachment device 110. The bus attachment device 110 manages the communication between the bus connection module 120 and the processors 130 on one hand and the communication between the bus connection module 120 and the memory 140 on the other hand. The bus connection module 120 may be directly connected to the bus 102 or via one or more intermediate components such as a converter 104.
[0073] The bus connection module 120 can be provided, for example, in the form of a Peripheral Component Interconnect Express (PCIe) module (also known as a PCIe adapter or PCIe functionality provided by a PCIe adapter). The PCIe functionality 120 can issue requests that are sent to a bus-attached device 110, such as a PCI Host Bridge (PHB), also known as a PCI Bridge Unit (PBU). The bus-attached device 110 receives the requests from the bus connection module 120. These requests can include, for example, input / output addresses for performing direct memory access (DMA) to the memory 140 by the bus-attached device 110 or input / output addresses indicating interrupt signals (e.g., Message Signaled Interrupt (MSI)).
[0074] Figure 2 An exemplary virtual machine support provided by the computer system 100 is depicted. The computer system 100 can include one or more virtual machines 202 and at least one hypervisor 200. The virtual machine support can provide the ability to operate a large number of virtual machines, each capable of executing a guest operating system 204, such as z / Linux. Each virtual machine 201 can be capable of acting as a separate system. Thus, each virtual machine can be reset independently, execute the guest operating system, and run different programs, such as applications. The operating system or applications running in the virtual machine appear to be able to access the entire and complete computer system. However, in reality, only a portion of the available resources of the computer system are available for the corresponding operating system or application to use.
[0075] The virtual machine can use the V=V model, where the memory allocated to the virtual machine is supported by virtual memory rather than real memory. Thus, each virtual machine has a virtual linear memory space. The physical resources are owned by the hypervisor 200 (e.g., a VM hypervisor), and the hypervisor dispatches the shared physical resources to the guest operating system as needed to meet the processing requirements of the guest operating system. The V=V virtual machine model assumes that the interaction between the guest operating system and the physical shared machine resources is controlled by the VM hypervisor, as a large number of guests can prevent the hypervisor from simply partitioning the hardware resources and allocating the hardware resources to the configured guests.
[0076] The processor 120 can be allocated by the hypervisor 200 to the virtual machine 202. The virtual machine 202 can be allocated, for example, one or more logical processors. Each logical processor can represent all or part of a physical processor 120 that can be dynamically allocated by the hypervisor 200 to the virtual machine 202. The virtual machine 202 is managed by the hypervisor 200. The hypervisor 200 can be implemented, for example, in firmware running on the processor 120 or can be part of an operating system executing on the computer system 100. The hypervisor 200 can be, for example, a VM hypervisor, such as that provided by International Business Machines Corporation of Armonk, New York, USA
[0077] Figure 3 depicts an exemplary multi - level virtual machine support provided by computer system 100. In addition to Figure 2 the first - level virtualization, second - level virtualization with a second hypervisor 210 is provided, where the first - level guest operating system acts as the host operating system for the second hypervisor 210. The second hypervisor 210 can manage one or more second - level virtual machines 212, each of which is capable of executing a second - level guest operating system 212.
[0078] Figure 4 depicts an exemplary pattern showing the use of different types of IDs to identify processors at different levels in computer system 100. The underlying firmware 220 can provide a logical processor ID, ICPU 222, to identify the processor 130 of computer system 100. The first - level hypervisor 200 uses the logical processor ID ICPU 222 to communicate with the processor 130. The first - level hypervisor can provide a first virtual processor ID, vCPU 224, for use by the guest operating system 204 or the second - level hypervisor 210 executing on a virtual machine managed by the first - level hypervisor 200. The hypervisor 200 can group the first virtual processor IDs vCPU 224 to provide a logical partition (also known as a zone) for the guest operating system 204 and / or the hypervisor 210. The first - level hypervisor 200 maps the first virtual processor ID vCPU 224 to the logical processor ID lCPU 222. One or more first virtual processor IDs vCPU 224 provided by the first - level hypervisor 200 can be assigned to each guest operating system 204 or hypervisor 210 that executes using the first - level hypervisor 200. The second - level hypervisor 210 executing on the first - level hypervisor 200 can provide one or more virtual machines for executing software such as an additional guest operating system 214. To this end, the second - level hypervisor manages a second virtual processor ID, vCPU 226, for use by the second - level guest operating system 214 executing on a virtual machine of the first - level hypervisor 200. The second virtual processor ID vCPU 226 is mapped by the second - level hypervisor 200 to the first virtual processor ID vCPU 224.
[0079] The bus connection module 120 addressing the processor 130 used by the first / second - level guest operating system 204 can use the first / second virtual processor IDs vCPU 224, 226 or a target processor ID in an alternative ID form derived from the first / second virtual processor IDs vCPU 224, 226.
[0080] Figure 5 FIG. 1 depicts a simplified schematic setup of a computer system 100, showing the main participants in a method for providing an interrupt signal to a client operating system executing on the computer system 100. For illustrative purposes, the simplified setup includes a bus connection module (BCM) 120 that sends an interrupt signal to a client operating system executing on one or more processors (CPUs) 130. The interrupt signal is sent to a bus-attached device 110 together with an interrupt target ID (IT_ID) that identifies one of the processors 130 as the target processor. The bus-attached device 110 is an intermediate device that manages the communication between the bus connection module 120 and the processors 130 and the memory 140 of the computer system 100. The bus-attached device 110 receives the interrupt signal and uses the interrupt target ID to identify the logical processor ID of the target processor for directly addressing the corresponding target processor. The directed forwarding to the target processor can improve the efficiency of data processing, for example, by reducing cache traffic.
[0081] Figure 6 depicts Figure 5 the computer system 100. The bus-attached device 110 is configured to perform a status update on the status of the bus connection module 120 in a module specific area (MSA) 148 of the memory 140. This status update can be performed in response to a direct memory access (DMA) write received from the bus connection module that specifies the status update to be written to the memory 140.
[0082] The memory further includes a device table (DT) 144 having device table entries (DTEs) 146 for each bus connection module 120.
[0083] Upon receiving an interrupt signal (e.g., an MSI-X write message having an interrupt target ID identifying the target processor of the interrupt request and an interrupt requester ID identifying the origin of the interrupt request in the form of the bus connection module 120), the bus-attached device 110 extracts the DTE 146 assigned to the bus connection module 120 that issued the request. The DTE 146 can indicate, for example, using the dIRQ bit whether directed addressing of the target processor is enabled for the bus connection module 120 that issued the request. The bus-attached device updates the Directed Interrupt Signal Vector (DIBV) 162 and the Directed Interrupt Summary Vector (DISB) 160 to keep track of which of the processors in the processor 130 the interrupt signal has been received for. The DISB 160 can include one entry for each interrupt target ID, indicating whether there is an interrupt signal from any of the bus connection modules 120 to be processed by that processor 130. Each DIBV 162 is assigned to one of the interrupt target IDs (i.e., the processors 130) and can include one or more entries. Each entry is assigned to one of the bus connection modules 120. Thus, the DIBV indicates which bus connection modules the pending interrupt signals for a particular processor 130 are from. This can have the advantage that in order to check whether there are any interrupt signals or which bus connection modules 120 the pending interrupt signals for a particular processor 130 are from, only the signal entry (e.g., bit) or signal vector (e.g., bit vector) needs to be read from the memory 140. According to an alternative embodiment, an Interrupt Signal Vector (AIBV) and an Interrupt Summary Vector (AISB) can be used. The entries of the AIBV and the AISB are each assigned to a particular bus connection module 120.
[0084] The bus-attached device 110 uses a mapping table 112 provided on the bus connection module 110 to convert the interrupt target ID (IT_ID) to a logical processor ID (ICPU), and directly addresses the target processor using the logical processor ID, forwarding the received interrupt signal to the target processor. Each processor includes firmware (e.g., microcode 132) for receiving and processing direct interrupt signals. The firmware can further include, for example, the microcode and / or macrocode of the processor 130. It can include hardware-level instructions and / or data structures used in the implementation of higher-level machine code. According to various embodiments, it can include proprietary code that can be delivered as microcode including trusted software or microcode specific to the underlying hardware and controls access to the system hardware by the operating system.
[0085] Figure 7FIG. is a flow chart of an exemplary method for performing a status update of a bus connection module 120 via a bus attached device 110 using a DMA write request. In step 300, the bus connection module may decide to update its status and trigger an interrupt, for example, to indicate signal completion. In step 310, the bus connection module initiates a direct memory access (DMA) write to a memory segment (host memory) allocated to a host running on a computer system via the bus attached device to update the status of the bus connection module. DMA is a hardware mechanism that allows a peripheral component of a computer system to directly transfer its I / O data to and from the main memory without involving the system processor. To perform the DMA, the bus connection module sends a DMA write request to the bus attached device, for example, in the form of an MSI-X message. In the case of PCIe, the bus connection module may, for example, refer to a PCIe function provided on a PCIe adapter. In step 320, the bus connection module receives the DMA write request with the status update of the bus connection module and uses the received update to update the memory. The update may be performed in an area reserved in the host memory for the corresponding bus connection module.
[0086] Figure 8 is for using Figure 6 FIG. is a flow chart of an exemplary method for a computer system 100 to provide an interrupt signal to a guest operating system using. In step 330, the bus attached device receives an interrupt signal (e.g., in the form of an MSI-X write message) sent by the bus connection module. The transmission of this interrupt signal may be performed according to the specifications of the PCI architecture. The MSI-X write message includes an interrupt target ID that identifies the target processor of the interrupt. The interrupt target ID may be, for example, a virtual processor ID used by the guest operating system to identify a processor in a multi-processor computer system. According to an embodiment, the interrupt target ID may be any other ID agreed upon by the guest operating system and the bus connection module in order to be able to identify the processor. Such other ID may be, for example, the result of a mapping of the virtual processor ID. In addition, the MSI-X write message may also include an interrupt requester ID (RID) (i.e., the ID of the PCIe function that issued the interrupt request), a vector index that defines the offset of the vector entry within the vector, an MSI address (e.g., a 64-bit address), and an MSI data (e.g., a 32-bit data). The MSI address and MSI data may indicate that the corresponding write message is actually an interrupt request in the form of an MSI message.
[0087] In step 340, the bus-attached device obtains a copy of an entry of the device table stored in the memory. The device table entry (DTE) provides an address indicator of one or more vectors or vector entries to be updated to indicate that an interrupt signal has been received for the target processor. The address indicator of the vector entry may include, for example, the address of the vector in the memory and the offset within the vector. In addition, the DTE may provide a direct signaling indicator that indicates whether the target processor is to be directly addressed by the bus-attached device using the interrupt target ID provided with the interrupt signal. In addition, the DTE may provide a logical partition ID (also known as a zone ID) and an interrupt subclass ID. A corresponding copy of the device table entry may be obtained from the cache or from the memory. In step 350, the bus-attached device updates the vector specified in the DTE.
[0088] In step 360, the bus-attached device checks the direct signaling indicator provided with the interrupt signal. In the case where the direct signaling indicator indicates non-direct signaling, the bus-attached device forwards the interrupt signal by broadcasting using the zone identifier and the interrupt subclass identifier in order to provide the interrupt signal to the processor used by the guest operating system. In the case where the direct signaling indicator indicates non-direct signaling, in step 370, the interrupt signal is forwarded by the broadcast item processor. The broadcast message includes the zone ID and the interrupt subclass ID. When the processor receives the interrupt request, the interrupt request is enabled for the zone, for example, the status bit is atomically set according to the nested communication protocol. In addition, the firmware (e.g., microcode) on the processor interrupts its activity (e.g., program execution) and switches to execute the interrupt handler of the guest operating system. In the case where the direct signaling indicator indicates direct signaling, in step 380, the bus-attached device converts the interrupt target ID provided with the interrupt signal into the logical processor ID of the processor allocated for use by the guest operating system. To perform the conversion, the bus-attached device may use a mapping table included by the bus-attached device. The bus-attached device may include a mapping table or sub-table for each zone (i.e., logical partition).
[0089] In step 390, the bus-attached device directly addresses the target processor with the logical processor ID of the target processor and forwards the interrupt signal to be processed to the target processor, that is, sends a direct message. The direct message may further include the zone ID and / or the interrupt subclass ID. In step 3100, the firmware (e.g., microcode) of the target processor receives the interrupt. In response, the firmware may interrupt its activity, such as program execution, and switch to execute the interrupt handler of the guest operating system. The interrupt may be submitted to the guest operating system with the direct signaling indicator.
[0090] Figure 9 is further explained Figure 8Additional flowchart of the method. In step 400, an interrupt message is sent to the bus-attached device. In step 402, the interrupt message is received. In step 404, it is checked whether the DTE assigned to the interrupt requester (i.e., the bus connection module) is cached in the local cache operatively connected to the bus-attached device. In the case where the DTE is not cached, in step 406, the corresponding DTE is fetched from the memory by the bus-attached device. In step 408, the vector address indicator provided by the DTE is used to set the vector bit in the memory. In step 410, it is checked with the direct signaling indicator provided by the DTE whether the bus-attached device is to directly address the target processor using the interrupt target ID provided together with the interrupt signal. If the target processor is not to be directly targeted, the method continues, and in step 412, an interrupt request is broadcast to the processor. If the target processor is to be directly targeted, the method continues, and in step 414, the interrupt target ID is converted to a logical processor ID, and in step 416, a message for forwarding the interrupt signal is sent to the target processor. The target processor is directly addressed using the logical processor ID. In step 418, the processor receives the message. In step 420, the processor submits an interrupt request to the client operating system. Then, the processor continues its activity until the next interrupt message is received.
[0091] Figure 10 depicts Figure 5 Another embodiment of the computer system 100. Figure 10 The computer system 100 corresponds to Figure 6 The computer system 100, where the processor 130 further includes check logic 134 for checking whether the receiving processor is the same as the target processor according to the interrupt target ID forwarded by the bus-attached device 110 to the receiving processor 130. If the receiving processor 130 is not the target processor, i.e., if the received interrupt target ID does not match the reference interrupt target ID of the receiving processor 130, the interrupt signal is broadcast to the logical partition to find a processor for handling the interrupt signal.
[0092] Figure 11 is a flowchart of an exemplary method of providing an interrupt signal to a client operating system using Figure 10 the computer system 100. In addition to Figure 8 the steps shown, according to Figure 11 the method further includes the following steps: In step 390, the bus-attached device directly addresses the corresponding processor using the logical processor ID and forwards the interrupt signal to the target processor, i.e., sends a direct message. The direct message includes the interrupt target ID. The direct message may further include the region ID and / or the interrupt subclass ID. The receiving processor includes interrupt target ID check logic. In the case where the interrupt target ID is unique per logical partition only, the check logic may further take into account the logical partition ID.
[0093] In step 392, the check logic checks whether the received interrupt target ID and / or logical partition ID match the interrupt target ID and / or logical partition currently assigned to the receiving processor and accessible by the check logic. In the case of a mismatch, in step 394, the receiving firmware initiates a broadcast to broadcast the received interrupt request to the remaining processors by using the logical partition ID and / or interrupt subclass ID for identifying the valid target processor for handling the interrupt. In the case of an affirmative match, in step 3100, the receiving firmware (e.g., microcode) of the target processor accepts the directly addressed interrupt for submission to the guest operating system. In response, the firmware may interrupt its activity (e.g., program execution) and switch to execute the interrupt handler of the guest operating system. The submission of the interrupt to the guest operating system can be indicated by direct signaling and can be submitted to the guest operating system together with the direct signaling indication.
[0094] Figure 12 is a further explanatory Figure 11 flowchart of the method. In addition to Figure 9 the steps shown, according to Figure 12 the method further includes the following: The message includes an interrupt target ID, a logical partition ID, and an interrupt subclass ID. In step 418, the processor receives the message. In step 419, the processor checks whether the interrupt target ID and / or logical partition ID match the current interrupt target ID and / or logical partition ID provided as a reference for the check. In the case of a match, the processor submits the interrupt request to the guest operating system in step 420. In the case of a mismatch, the processor broadcasts the interrupt request to other processors in step 422. Then, the processor continues its activity until it receives the next interrupt message.
[0095] Figure 13 Depicts an exemplary DTE 146, including the logical partition ID (zone) assigned to the interrupt target ID and the within-DIBV and offset address (DIBVO). The DIBVO identifies the section or vector entry assigned to a particular bus connection module. The interrupt signal (e.g., MSI-X message) may provide a DIBV-Idx added to the DIBVO for identifying the specific entry of the vector assigned to the bus connection module. In addition, the number of interrupts (NOI) defining the maximum number of bits reserved in the DIBV for the corresponding bus connection module is also provided. Further details of the DIBV are shown in Figure 16A In the case of the AIBV, the DTE may provide the corresponding AIBV-specific parameters as shown in Figure 16B
[0096] Figure 14 Depicts the schematic structure of DISB 160 and multiple DIBV 162. DISB 160 can be provided in the form of a continuous memory segment, for example, in the form of a cache line, which includes an entry 161 for each interrupt target ID, such as a bit. Each entry indicates whether there is an interrupt request (IRQ) to be processed by the corresponding processor identified by the interrupt target ID. A DIBV 162 is provided for each interrupt target ID (i.e., the entry of DISB 160). Each DIBV 162 is assigned to a specific interrupt target ID and includes one or more entries 163 for each bus connection module MN A, MN B. DIBV162 can each be provided in the form of a continuous memory segment, for example, in the form of a cache line, which includes the entries 163 assigned to the same interrupt target ID. The entries of different bus connection modules can be in the order of different offset DIBVO of each bus connection module.
[0097] Figure 15 Depicts the schematic structure of AISB 170 and multiple AIBV 172. AISB 170 can be provided in the form of a continuous memory segment, for example, in the form of a cache line, which includes an entry 171 for each bus connection module MN A to MN D, such as a bit. Each entry indicates whether there is a pending interrupt request (IRQ) from the corresponding bus connection module. An AIBV 172 is provided for each bus connection module (i.e., the entry of AISB 170). Each AIBV 172 is assigned to a specific bus connection module and one or more entries 173 for each interrupt target ID. AIBV 172 can each be provided in the form of a continuous memory segment, for example, in the form of a cache line, which includes the entries 173 assigned to the same bus connection module. The entries regarding different bus connection modules can be in the order of different offset client AIBVO of each bus connection module.
[0098] Figure 17A and 17B Respectively show exemplary DISB 160 and AISB 170. The entries 161, 171 can be addressed using the base addresses DISB@ and AISB@ and the offset addresses DISBO and AISBO respectively. In the case of DISB 160, for example, DISBO can be the same as the interrupt target ID to which the corresponding entry 161 is assigned. The interrupt target ID can be provided, for example, in the form of a virtual processor ID (vCPU).
[0099] The client operating system can be implemented, for example, using a paged storage mode client. For example, The paged client in it can be interpretively executed at the second interpretive layer via the Start Interpretive Execution (SIE) instruction. For example, a logical partition (LPAR) hypervisor executes the SIE instruction to start a logical partition in physical, fixed memory. The operating system in this logical partition (e.g., ) can issue the SIE instruction to execute the client (virtual) machine in its virtual storage. Thus, the LPAR hypervisor can use level 1 SIE, the hypervisor can use level 2 SIE.
[0100] According to various embodiments, the computer system is a System server provided by International Business Machines Corporation. System is based on Regarding details of which are described in the publication entitled "z / Architecture Principles of Operation" ( Publication No. SA22-7832-11, August 25, 2017), the entire text of which is hereby incorporated by reference herein. and are registered trademarks of International Business Machines Corporation of Armonk, New York, USA. Other names used herein may be registered trademarks, trademarks or product names of International Business Machines Corporation or other companies.
[0101] According to various embodiments, computer systems of other architectures can implement and use one or more aspects of the present invention. For example, in addition to System Servers other than the server (such as Power Systems servers or other servers provided by International Business Machines Corporation) or servers of other companies implement, use, and / or benefit from one or more aspects of the present invention. Further, although in the examples herein, the bus connection module and the bus-attached device are considered part of the server, in other embodiments, they need not necessarily be considered part of the server, but can simply be considered as being coupled to the system memory and / or other components of the computer system. The computer system need not be a server. Further, although the bus connection module can be PCIe, one or more aspects of the present invention can be used with other bus connection modules. The PCIe adapter and PCIe functions are merely examples. Further, one or more aspects of the present invention can be applied to interrupt schemes other than PCIMSI and PCIMSI-X. Still further, although examples in which bits are set are described, in other embodiments, bytes or other types of indicators can be set. In addition, the DTE and other structures can include more, less, or different information.
[0102] In addition, other types of computer systems can benefit from one or more aspects of the present invention. As an example, a data processing system suitable for storing and / or executing program code including at least two processors directly or indirectly coupled to a memory element via a system bus is available. The memory element includes, for example, local memory employed during actual execution of the program code, mass storage devices, and a cache that provides at least some temporary storage of the program code to reduce the number of times code must be retrieved from the mass storage device during execution.
[0103] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, DASD, tapes, CDs, DVDs, thumb drives, and other storage media, etc.) can be coupled to the system directly or via an intermediate I / O controller. A network adapter can also be coupled to the system so that the data processing system can become coupled to other data processing systems or remote printers or storage devices via an intervening private or public network. Modems, cable modems, and Ethernet cards are merely a few of the available types of network adapters.
[0104] Refer to Figure 18, depicts representative components of a host computer system 500 for implementing one or more aspects of the present invention. The representative host computer 500 includes one or more processors (e.g., CPUs) 501 communicating with a computer memory 502, and an I / O interface connected to a storage media device 411 and a network 410 for communicating with other computers or a SAN, etc. The CPU 501 conforms to an architecture having an architectural instruction set and architectural functions. The CPU 501 may have a dynamic address translation (DAT) 503 for converting a program address, a virtual address into a real address of the memory. The DAT may include a translation lookaside buffer (TLB) 507 for cache translation, such that a later access to a block of the computer memory 502 does not require the latency of address translation. A cache 509 may be employed between the computer memory 502 and the CPU 501. The cache 509 may be hierarchically structured, thereby providing a large, high-level cache available for more than one CPU and smaller, faster, lower-level caches between the high-level cache and each CPU. In some embodiments, the lower-level cache may be partitioned to provide separate lower-level caches for instruction fetching and data access. According to an embodiment, instructions may be fetched from the memory 502 by an instruction fetch unit 504 via the cache 509. The instructions may be decoded in an instruction decoding unit 506 and, in some embodiments, together with other instructions, be dispatched to one or more instruction execution units 508. A number of execution units 508 may be employed, such as an arithmetic execution unit, a floating-point execution unit, and a branch instruction execution unit. The instructions are executed by the execution units, accessing operands from registers or memory specified by the instructions as needed. If an operand is to be accessed (e.g., loaded or stored) from the memory 502, a load / store unit 505 may handle the access under the control of the instruction being executed. The instructions may be executed in hardware circuitry or in internal microcode (i.e., firmware) or by a combination of both.
[0105] A computer system may include information in local or main storage devices, as well as addressing, protection, and referencing and changing records. Some aspects of addressing include the format of addresses, the concept of address spaces, different types of addresses, and the ways of converting one type of address into another type of address. Some of the main storage devices include permanently allocated storage locations. The main storage device provides the system with fast access storage that can be directly addressed. Both data and programs are loaded into the main storage device, for example, from input devices before they can be processed.
[0106] The main storage device may include one or more smaller, faster-accessed buffer memories, sometimes called caches. The cache may be physically associated with the CPU or an I / O processor. The physical construction of different storage media and the use, except for the performance impact, may generally be invisible to the programs being executed.
[0107] Separate caches can be maintained for instructions and for data operands. The information within the cache can be maintained in contiguous bytes at integral boundaries called cache blocks or cache lines. The model can provide an EXTRACT CACHE ATTRIBUTE instruction that returns the size of the cache line in bytes. The model can also provide PREFETCH DATA and PREFETCH DATA RELATIVE LONG instructions that implement prefetching from a storage device into a data or instruction cache or the release of data from the cache.
[0108] A storage device can be regarded as a long horizontal string of bits. For most operations, access to the storage device can be made in a left-to-right order. The bit string is subdivided into units of eight bits. These eight-bit units are called bytes, which are the basic building blocks of all information formats. Each byte position in the storage device can be identified by a unique non-negative integer, which is the address of the byte position, also called the byte address. Adjacent byte positions can have consecutive addresses starting from 0 on the left and continuing in a left-to-right order. The address is an unsigned binary integer, for example, it can be 24, 31, or 64 bits.
[0109] Information is transferred between the memory and the CPU one byte or a group of bytes at a time. Unless otherwise specified, in, for example, the group of bytes in the memory is addressed by the leftmost byte of the group. The number of bytes in the group is implied or explicitly specified by the operation to be performed. When used in a CPU operation, a group of bytes is called a field. Within each group of bytes, for example, in the bits are numbered in a left-to-right order. In In it, the leftmost bit is sometimes referred to as the "high-order" bit and the rightmost bit as the "low-order" bit. However, the bit number is not the storage address. Only bytes are addressable. To operate on individual bits of a byte in a storage device, the entire byte can be accessed. The bits in a byte can be numbered from 0 to 7 from left to right, for example, in z / Architecture. The bits in an address can be numbered 8 - 31 or 40 - 63 for a 24-bit address, or 1 - 31 or 33 - 63 for a 31-bit address; for a 64-bit address, they are numbered 0 - 63. Within any other fixed-length format of multiple bytes, the bits making up the format can be numbered continuously starting from 0. For error detection purposes, and preferably for correction, one or more parity bits can be transmitted along with each byte or group of bytes. Such parity bits are automatically generated by the machine and cannot be directly controlled by a program. Storage capacity is expressed in terms of the number of bytes. When the length of a storage operand field is implied by the operation code of an instruction, the field is said to have a fixed length, which can be one byte, two bytes, four bytes, eight bytes, or sixteen bytes. Larger fields can be implied for some instructions. When the length of a storage operand field is not implied but is explicitly stated, the field is said to have a variable length. The length of a variable-length operand can vary in increments of one byte or in increments of some instructions, in multiples of two bytes or other multiples. When placing information in a storage device, only the contents of those byte positions included in the specified field are replaced, even if the width of the physical path to the storage device may be greater than the length of the field being stored.
[0110] Certain information units will be on full boundaries in the storage device. When the storage address of information is a multiple of the unit length in bytes, the boundary is called an integer of the information unit. Fields of 2, 4, 8, and 16 bytes on full boundaries all have special names. A halfword is a group of two consecutive bytes on a two-byte boundary and is a basic building block of an instruction. A word is a group of four consecutive bytes on a four-byte boundary. A doubleword is a group of eight consecutive bytes on an eight-byte boundary. A quadword is a group of sixteen consecutive bytes on a sixteen-byte boundary. When the storage address specifies a halfword, word, doubleword, and quadword, the binary representation of the address contains one, two, three, or four rightmost zero bits respectively. Instructions are to be on two-byte integral boundaries. Most instructions have no boundary alignment requirement for storage operands.
[0111] On a device that implements separate caches for instruction and data operands, if a program stores into a cache line from which it subsequently fetches an instruction, a significant delay will be experienced, regardless of whether the store changes the subsequently fetched instruction.
[0112] In one embodiment, the present invention can be implemented by software, which is sometimes referred to as licensed internal code, firmware, microcode, nano-code, pico-code, etc., any of which will conform to the present invention. See Figure 18 , the software program code embodying the present invention can be accessed from a long-term storage medium device 511 such as a CD-ROM drive, a tape drive, or a hard disk drive. The software program code can be embodied on any of a variety of known media for a data processing system, such as a disk, a hard disk drive, or a CD-ROM. The code can be distributed on such media, or can be distributed from the computer memory 502 or a storage device of a computer system to a user via a network 510 connected to other computer systems for use by users of such other systems.
[0113] The software program code can include an operating system that controls the functions and interactions of different computer components and one or more application programs. The program code can be paged from the storage medium device 511 to a relatively high-speed computer memory 502, where it can be used by the processor 501 for processing. Well-known techniques and methods can be used for embodying the software program code in memory, on a physical medium, and / or for distributing the software code via a network. When created and stored on a tangible medium (including but not limited to an electronic memory module (RAM), flash memory, compact disk (CD), DVD, magnetic tape), the program code can be referred to as a "computer program product". The computer program product medium can be readable by a processing circuit preferably located in a computer system for execution by the processing circuit.
[0114] Figure 19 A representative workstation or server hardware system in which embodiments of the present invention can be implemented is shown. Figure 19 The system 520 includes a representative basic computer system 521, such as a personal computer, a workstation, or a server, including optional peripheral devices. The basic computer system 521 includes one or more processors 526 and a bus for connecting and enabling communication between the processor 526 and other components of the system 521 according to known techniques. The bus connects the processor 526 to a memory 525 and a long-term memory 527, which can include a hard disk drive (including, for example, any of magnetic media, CD, DVD, and flash memory) or a tape drive. The system 521 may also include a user interface adapter that connects the microprocessor 526 via the bus to one or more interface devices, such as a keyboard 524, a mouse 523, a printer / scanner 530, and / or other interface devices, which can be any user interface device, such as a touch screen, a digitizing tablet, etc. The bus also connects a display device 522 (such as an LCD screen or a monitor) to the microprocessor 526 via a display adapter.
[0115] System 521 can communicate with other computers or computer networks through a network adapter capable of communicating 528 with network 529. Example network adapters are communication channels, token rings, Ethernet, or modems. Alternatively, system 521 can communicate using a wireless interface (such as a Cellular Digital Packet Data (CDPD) card). System 521 can be associated with such other computers in a local area network (LAN) or wide area network (WAN), or system 521 can be a client in a client / server arrangement along with another computer, etc.
[0116] Figure 20 A data processing network 540 in which embodiments of the present invention can be implemented is shown. Data processing network 540 can include a plurality of separate networks, such as wireless networks and wired networks, each of which can include a plurality of separate workstations 541, 542, 543, 544. Additionally, as will be appreciated by those skilled in the art, one or more LANs can be included, where a LAN can include a plurality of intelligent workstations coupled to a host processor.
[0117] Still referring to Figure 20 , the network can also include a mainframe or server that can access a data repository and can also be directly accessed from workstation 545, such as a gateway computer (e.g., client server 546) or an application server (e.g., remote server 548). Gateway computer 546 can serve as an entry point into each separate network. A gateway may be required when connecting one networking protocol to another. Gateway 546 can preferably be coupled to another network, such as the Internet 547, through a communication link. Gateway 546 can also be directly coupled to one or more workstations 541, 542, 543, 544 using a communication link. An IBM eServerTM System server available from International Business Machines Corporation can be used to implement the gateway computer.
[0118] Also referring to Figure 19 and Figure 20 , the software programming code embodying the present invention can be accessed by the processor 526 of system 520 from a long-term storage medium 527 (such as a CD-ROM drive or a hard disk drive). The software programming code can be embodied on any of a variety of known media for a data processing system, such as a disk, a hard disk drive, or a CD-ROM. The code can be distributed on such media, or it can be distributed from the computer memory 502 or storage device of one computer system to users via a network connected to other computer systems for use by users of such other systems.
[0119] Alternatively, the programming code may be embodied in the memory 525 and accessed by the processor 526 using the processor bus. Such programming code may include an operating system that controls the functions and interactions of different computer components and one or more application programs 532. The program code may be paged from the storage medium 527 to the high-speed memory 525, where the program code may be used for processing by the processor 526. Well-known techniques and methods for embodying software programming code in memory, on physical media, and / or distributing software code via a network may be used.
[0120] The most readily accessible cache of the processor, i.e., the one that can be faster and smaller than other caches of the processor, is the lowest-level cache, also known as L1 or first-level cache. The main memory is the highest-level cache, and if there are n levels, the main memory is also referred to as the Ln-level cache. For example, if n = 3, it is called the L3-level cache. The lowest-level cache may be divided into an instruction cache and a data cache. The instruction cache is also known as the I cache and stores machine-readable instructions to be executed. The data cache is also known as the D cache and stores data operands.
[0121] Referring Figure 21 , an exemplary processor embodiment of the processor 526 is depicted. One or more levels of caches 553 may be employed to cache memory blocks in order to improve processor performance. The cache 553 is a buffer of cache lines that hold memory data that may be used. A cache line may be, for example, 64, 128, or 256 bytes of memory data. Separate caches may be employed to cache instructions and cache data. Cache coherence (i.e., synchronization of copies of lines in memory and the cache) may be provided by different suitable algorithms (e.g., the "snooping" algorithm). The main memory 525 of the processor system may be referred to as a cache. In a processor system with 4 levels of caches 553, the main memory 525 is sometimes referred to as the 5th-level (L5) cache because it can be faster and stores only a portion of the non-volatile memory available to the computer system. The main memory 525 "caches" data pages that are paged in and out of the main memory 525 by the operating system.
[0122] The program counter (instruction counter) 561 keeps track of the address of the current instruction to be executed. The program counter in the processor is 64 bits and can be truncated to 31 or 24 bits to support previous addressing limitations. The program counter can be embodied in the program status word (PSW) of the computer such that it persists during context switching. Thus, an ongoing program with a program counter value may be interrupted, for example, by the operating system, resulting in a context switch from the program environment to the operating system environment. The PSW of the program holds the program counter value when the program is inactive and uses the program counter in the operating system's PSW when the operating system is executing. The program counter can be incremented by an amount equal to the number of bytes of the current instruction. Reduced Instruction Set Computing (RISC) instructions can be of fixed length, while Complex Instruction Set Computing (CISC) instructions can be of variable length. IBM 's instructions are CISC instructions that are 2, 4, or 6 bytes long. For example, the program counter 561 can be modified by a context switch operation or a branch taken operation of a branch instruction. In a context switch operation, the current program counter value is saved in the program status word along with other state information about the program being executed (e.g., condition codes), and then a new program counter value pointing to the instruction of the new program module to be executed is loaded. A branch taken operation can be performed to allow the program to make decisions or loop within the program by loading the result of the branch instruction into the program counter 561.
[0123] An instruction fetch unit 555 can be employed to fetch instructions on behalf of the processor 526. The fetch unit fetches the "next sequential instruction", the target instruction of a branch taken instruction, or the first instruction of a program after a context switch. Modern instruction fetch units can employ prefetch techniques to speculatively prefetch instructions based on the likelihood of the prefetch instructions being used. For example, the fetch unit can fetch 16 bytes of instructions that include the next sequential instruction and additional sequential instructions.
[0124] The fetched instructions can then be executed by the processor 526. According to an embodiment, the fetched instructions can be passed from the fetch unit to a dispatch unit 556. The dispatch unit decodes the instructions and forwards information about the decoded instructions to appropriate units 557, 558, 560. An execution unit 557 can receive information about a decoded arithmetic instruction from the instruction fetch unit 555 and can perform arithmetic operations on the operands according to the opcode of the instruction. The operands can preferably be provided to the execution unit 557 from the memory 525, the architected registers 559, or from the immediate field of the instruction being executed. When the result of the execution is to be stored, it can be stored in the memory 525, the registers 559, or other machine hardware such as control registers, PSW registers.
[0125] The processor 526 can include one or more units 557, 558, 560 for performing the functions of the instructions. SeeFigure 22A , the execution unit 557 can communicate with the architected general registers 559, the decode / dispatch unit 556, the load / store unit 560, and other 565 processor units through the interface logic 571. The execution unit 557 can employ a number of register circuits 567, 568, 569 to hold information operated on by the arithmetic logic unit (ALU) 566. The ALU performs arithmetic operations such as addition, subtraction, multiplication, and division, as well as logical functions such as And, Or, Exclusive Or (XOR), rotate, and shift. Preferably, the ALU can support design-dependent specialized operations. Other circuits can provide other architected facilities 572 including, for example, condition code and recovery support logic. The result of the ALU operation can be saved in the output register circuit 570, which is configured to forward the result to various other processing functions. There are many ways to arrange the processor units, and this specification is only intended to provide a representative understanding of one embodiment.
[0126] Addition (ADD) instructions, for example, can be executed in the execution unit 557 with arithmetic and logic capabilities, while floating-point instructions, for example, will be executed in a floating-point execution unit with specialized floating-point capabilities. Preferably, the execution unit operates on the operands by performing the function defined by the opcode on the operands identified by the instruction. For example, an addition instruction can be executed by the execution unit 557 on the operands found in two registers 559 identified by the register fields of the instruction.
[0127] The execution unit 557 performs arithmetic addition on two operands and stores the result in a third operand, where the third operand can be a third register or one of the two source registers. The execution unit preferably utilizes the arithmetic logic unit (ALU) 566, which is capable of performing various logical functions such as shift, rotate, and XOR, as well as various algebraic functions including any one of addition, subtraction, multiplication, and division. Some ALUs 566 are designed for scalar operations, and some ALUs 566 are designed for floating point. Depending on the architecture, the data can be big endian (where the least significant byte is at the highest byte address) or little endian (where the least significant byte is at the lowest byte address). IBM is big endian. The signed field can be sign and magnitude, one's complement, or two's complement, depending on the architecture. Two's complement can be advantageous because the ALU does not need to be designed with a subtraction function, since negative or positive values in two's complement only require addition within the ALU. Numbers can be described in shorthand, where a 12-bit field defines the address of a 4,096-byte block and is described as, for example, a 4K byte (kilobyte) block.
[0128] See Figure 22BBranch instruction information for executing branch instructions can be sent to the branch unit 558, which typically uses a branch prediction algorithm (such as the branch history table 582) to predict the outcome of a branch before other conditional operations are completed. Before the conditional operation is completed, the target of the current branch instruction is extracted and executed speculatively. When the conditional operation is completed, based on the condition of the conditional operation and the speculative result, the speculatively executed branch instruction either completes or is discarded. Branch instructions can test the condition code, and if the condition code meets the branch requirement of the branch instruction, then the branch instruction can branch to the target address, and the target address can be calculated based on several numbers (including, for example, the numbers found in the register field or immediate field of the instruction). The branch unit 558 can employ an ALU 574 having multiple input register circuits 575, 576, 577 and an output register circuit 580. For example, the branch unit 558 can communicate with the general-purpose register 559, the decode / dispatch unit 556, or other circuitry 573.
[0129] The execution of a set of instructions can be interrupted for various reasons, including (for example) a context switch initiated by the operating system, a program exception or error that causes a context switch, an I / O interrupt signal that causes a context switch, or the multithreaded activity of multiple programs in a multithreaded environment. Preferably, the context switch action saves the state information about the currently executing program and then loads the state information about another program being called. The state information can be saved, for example, in hardware registers or in memory. The state information preferably includes the program counter value pointing to the next instruction to be executed, the condition code, the memory translation information, and the architected register contents. The context switch activity can be implemented by hardware circuitry, an application program, an operating system program, or firmware code (for example). Microcode, micro-microcode, or licensed internal code (LIC), either individually or in combination.
[0130] The processor accesses operands according to the method defined by the instruction. The instruction can use a value of a part of the instruction to provide an immediate operand, and can provide one or more register fields that explicitly point to general-purpose registers or special-purpose registers (such as floating-point registers). The instruction can utilize an implicit register identified by the opcode field as an operand. The instruction can utilize the memory location of the operand. The memory location of the operand can be provided by a register, an immediate field, or a combination of a register and an immediate field, such as illustrated by the long displacement facility, where the instruction defines a base register, an index register, and an immediate number field, i.e., the displacement field, and adds them together to provide the memory address of the operand. Unless otherwise indicated, locations herein can be implied to be locations in the main memory.
[0131] Referring to Figure 22C , the processor uses the load / store unit 560 to access memory. The load / store unit 560 can perform a load operation by obtaining the address of a target operand in memory 553 and loading the operand into register 559 or another memory 553 location, or can perform a store operation by obtaining the address of the target operand in memory 553 and storing the data obtained from register 559 or another memory 553 location in the target operand location in memory 553. The load / store unit 560 can be speculative and can access memory in an order that is out of order with respect to the instruction sequence; however, the load / store unit 560 will maintain the appearance that the program is executing instructions in order. The load / store unit 560 can communicate with the general-purpose registers 559, the decode / dispatch unit 556, the cache / memory interface 553, or other elements 583, and includes various register circuits, an ALU 585, and control logic 590 to compute the memory addresses and provide pipeline ordering to maintain ordered operation. Some operations may be out of order, but the load / store unit provides the function of making out-of-order operations appear to the program to have been executed in order.
[0132] Preferably, the addresses "seen" by an application program are generally referred to as virtual addresses. Virtual addresses are sometimes also referred to as "logical addresses" and "effective addresses". These virtual addresses are virtual because they are redirected to physical memory locations through one of a variety of dynamic address translation (DAT) techniques. Dynamic address translation techniques include, but are not limited to, simply prefixing the virtual address with an offset value, translating the virtual address via one or more translation tables, which preferably include at least one of a separate segment table, a separate page table, or a combination of a segment table and a page table, and preferably, the segment table has entries pointing to the page table. In a translation hierarchy is provided that includes a region first table, a region second table, a region third table, a segment table, and an optional page table. The performance of address translation is often improved by utilizing a translation lookaside buffer (TLB) that contains entries that map virtual addresses to associated physical memory locations. Entries are created when the DAT uses the translation table to translate a virtual address. Subsequent uses of the virtual address can then utilize the entries of the fast TLB instead of accessing the slow sequential translation table. The TLB contents can be managed by various replacement algorithms including least recently used (LRU).
[0133] Each processor in a multiprocessor system is responsible for maintaining the interlock of shared resources (such as I / O, cache, TLB, and memory) to achieve coherency. The so-called "snooping" technique can be used to maintain cache coherency. In a snooping environment, for the convenience of sharing, each cache line can be marked as being in any one of the shared state, exclusive state, modified state, invalid state, etc.
[0134] The I / O unit 554 can provide the processor with means for attaching to peripheral devices (including, for example, magnetic tapes, optical disks, printers, displays, and networks). The I / O unit is typically submitted to a computer program by a software driver. In a host such as System the host, the channel adapter and the open system adapter are the I / O units of the host that provide communication between the operating system and the peripheral devices.
[0135] Furthermore, other types of computer systems can benefit from one or more aspects of the present invention. As an example, a computer system can include an emulator, such as software or other emulation mechanisms, where emulation includes, for example, instruction execution, architectural functions (such as address translation), and a specific architecture of architectural registers or a subset thereof, such as being emulated on a native computer system having a processor and memory. In such an environment, one or more emulation functions of the emulator can implement one or more aspects of the present invention, even if the computer on which the emulator is executed can have an architecture different from the capabilities being emulated. For example, in the emulation mode, a specific instruction or operation being emulated can be decoded, and appropriate emulation functions can be constructed to implement the individual instruction or operation.
[0136] In an emulation environment, the host computer can include, for example: a memory for storing instructions and data; an instruction fetch unit for fetching instructions from the memory and optionally providing local buffering for the fetched instructions; an instruction decoding unit for receiving the fetched instructions and for determining the type of the instructions that have been fetched; and an instruction execution unit for executing the instructions. Execution can include loading data from the memory into registers, storing data from the registers back into the memory, and / or performing some type of arithmetic or logical operation as determined by the decoding unit. For example, each unit can be implemented in software. The operations performed by these units can be implemented as one or more subroutines within the emulator software.
[0137] More specifically, in a mainframe, a programmer (such as a "C" programmer) uses architectural machine instructions, for example, through a compiler application. These instructions stored in a storage medium can be executed natively on the server, or executed on a machine with a different architecture. They can be used in existing and future Mainframe servers and other machines in (e.g., Power Systems servers and System servers) are emulated. They can be executed on various machines running Linux that use hardware manufactured by AMD TM etc. In addition to being executed on this hardware under , machines using Linux and machines emulated by Hercules, UMX, or FSI (Fundamental Software, Inc) can be used, and these machines typically execute in emulation mode. In emulation mode, the emulation software is executed by the native processor to emulate the architecture of the emulated processor.
[0138] The native processor can execute emulation software including firmware or the native operating system to perform emulation of the emulated processor. The emulation software is responsible for fetching and executing the instructions of the emulated processor architecture. The emulation software maintains an emulation program counter to track instruction boundaries. The emulation software can fetch one or more emulated machine instructions at a time and convert the one or more emulated machine instructions into a corresponding set of native machine instructions for execution by the native processor. These converted instructions can be cached so that faster conversion can be achieved. Nevertheless, the emulation software will maintain the architectural rules of the emulated processor architecture to ensure that the operating systems and applications written for the emulated processor operate correctly. In addition, the emulation software will provide the resources identified by the emulated processor architecture, including but not limited to control registers, general-purpose registers, floating-point registers, dynamic address translation functions including, for example, segment tables and page tables, interrupt mechanisms, context-switching mechanisms, time-of-day (TOD) clocks, and an architected interface to the I / O subsystem, so that an operating system or application designed to run on the emulated processor can run on the native processor with the emulation software.
[0139] The specific instruction being emulated is decoded, and a subroutine is called to execute the function of a single instruction. For example, in a "C" subroutine or driver, or in some other method that provides a driver for the specific hardware, the emulation software function for emulating the functions of the emulated processor is implemented.
[0140] In Figure 23In [the text], an example of an emulation host computer system 592 of a host computer system 500' that emulates a host architecture is provided. In the emulation host computer system 592, a host processor (i.e., a CPU) 591 is an emulation host processor or a virtual host processor, and includes an emulation processor 593 having a native instruction set architecture different from the native instruction set architecture of the processor 591 of the host computer 500'. The emulation host computer system 592 has a memory 594 accessible by the emulation processor 593. In an exemplary embodiment, the memory 594 is divided into a host computer memory 596 portion and an emulation routine memory 597 portion. The host computer memory 596 is available for programs of the emulation host computer 592 in accordance with the host computer architecture. The emulation processor 593 executes native instructions of an architecture instruction set of an architecture other than the emulation processor 591, native instructions obtained from the emulation routine memory 597, and can access host instructions for execution from a program in the host memory 596 by using one or more instructions obtained in sequence, and access / decodes a routine that can decode the accessed host instructions to determine a native instruction execution routine for simulating the function of the accessed host instructions. Other facilities defined for the architecture of the host system 500' can be emulated by architecture facility routines, including facilities such as general-purpose registers, control registers, dynamic address translation, and I / O subsystem support, as well as processor caches. The emulation routines can also utilize functions available in the emulation processor 593 (such as general-purpose registers and dynamic translation of virtual addresses) to improve the performance of the emulation routines. Dedicated hardware and offload engines can also be provided to assist the processor 593 in simulating the functions of the host 500'.
[0141] It will be understood that one or more of the above-described embodiments of the present invention can be combined as long as the combined embodiments are not mutually exclusive. Ordinal numbers such as "first" and "second" are used herein to indicate different elements assigned the same name, but do not necessarily establish any order among the respective elements.
[0142] The present invention has been described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0143] The present invention can be a system, a method, and / or a computer program product. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0144] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0145] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0146] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the 'C' programming language or similar programming languages. The computer-readable program instructions may be executed entirely on a computer of a user computer system, partially on a computer of a user computer system, executed as a stand-alone software package, partially on a computer of a user computer system and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the computer of the user computer system through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present invention.
[0147] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0148] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, which, when executed by the processor of the computer or other programmable data processing apparatus, creates a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein includes an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0149] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0150] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions.
[0151] The above features may have the following possible combinations:
[0152] 1. A method for providing an interrupt signal to a guest operating system, the guest operating system being executed by one or more of a plurality of processors of a computer system allocated for use by the guest operating system, the computer system further including one or more bus connection modules operably connected to the plurality of processors via a bus and bus-attached devices,
[0153] Each of the plurality of processors is assigned a logical processor ID by the bus-attached device for addressing the corresponding processor,
[0154] Each of the plurality of processors allocated for use by the guest operating system is further assigned an interrupt target ID by the guest operating system and the one or more bus connection modules for locating the corresponding processor,
[0155] The method includes:
[0156] Receiving, by the bus-attached device, an interrupt signal having an interrupt target ID from one of the bus connection modules, the interrupt target ID identifying one of the processors that is allocated by the guest operating system as the target processor for processing the interrupt signal,
[0157] The bus-attached device converts the received interrupt target ID into the logical processor ID of the target processor using a mapping table included in the bus-attached device. The mapping table maps the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the multiple processors.
[0158] The bus-attached device directly addresses the target processor using the logical processor ID of the target processor and forwards the interrupt signal to be processed to the target processor.
[0159] 2. The method of item 1, wherein the interrupt signal is received in the form of a message-signaled interrupt, and the message-signaled interrupt includes the interrupt target ID of the target processor.
[0160] 3. The method of any one of the foregoing items, wherein the computer system further includes a memory, and the bus-attached device is operatively connected to the memory. The method further includes:
[0161] The bus-attached device retrieves a copy of the device table entry from the device table stored in the memory. The device table entry includes a direct signaling indicator indicating whether to directly address the target processor.
[0162] If the direct signaling indicator indicates direct forwarding of the interrupt signal, the bus-attached device directly addresses the target processor using the logical processor ID of the target processor and performs the forwarding of the interrupt signal.
[0163] Otherwise, the bus-attached device forwards the interrupt signal to be processed to the multiple processors by broadcasting.
[0164] 4. The method of item 3, wherein the direct signaling indicator is implemented with a single bit.
[0165] 5. The method of items 3 to 4, wherein during the initialization of the guest operating system, the direct signaling indicator is set to a static indicator for the guest operating system.
[0166] 6. The method of any one of items 3 to 5, wherein the mapping of the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the multiple processors is a static mapping defined by the mapping table.
[0167] 7. The method of any one of items 3 to 6, wherein the bus-attached device checks whether a copy of the device table entry is cached in a local cache operatively connected to the bus-attached device, wherein
[0168] If a copy of the device table entry is cached, the retrieval of the copy of the device table entry is performed from the corresponding cache.
[0169] Otherwise, the retrieval of the device table entry is performed from the memory.
[0170] 8. The method of any one of items 3 to 6, the memory further includes an interrupt summary vector, the device table entry further includes an interrupt summary vector address indicator indicating the memory address of the interrupt summary vector, the interrupt summary vector includes an interrupt summary indicator for each bus-connected module, each interrupt summary indicator is assigned to a bus-connected module, and indicates whether there is an interrupt signal issued by the corresponding bus-connected module pending processing.
[0171] The method further includes the bus-attached device using the indicated memory address of the interrupt summary vector to update the interrupt summary indicator assigned to the bus-connected module from which the interrupt signal is received, such that the updated interrupt summary indicator indicates that there is an interrupt signal issued by the corresponding bus-connected module pending processing.
[0172] 9. The method of item 8, the interrupt summary vector is implemented as a continuous region.
[0173] 10. The method of any one of items 8 to 9, each interrupt summary indicator is implemented as a single bit.
[0174] 11. The method of any one of items 3 to 10, the memory further includes a directed interrupt summary vector, the device table entry further includes a directed interrupt summary vector address indicator indicating the memory address of the directed interrupt summary vector, the directed interrupt summary vector includes a directed interrupt summary indicator for each interrupt target ID, each directed interrupt summary indicator is assigned to an interrupt target ID, and indicates whether there is an interrupt signal addressed to the corresponding interrupt target ID pending processing.
[0175] The method further includes the bus-attached device using the indicated memory address of the directed interrupt summary vector to update the interrupt summary indicator assigned to the bus-connected module of the target processor ID addressed by the received interrupt signal, such that the updated interrupt summary indicator indicates that there is an interrupt signal addressed to the corresponding interrupt target ID pending processing.
[0176] 12. The method of item 11, the directed interrupt summary vector is implemented as a continuous region.
[0177] 13. The method of any one of items 11 to 12, each directed interrupt summary indicator is implemented as a single bit.
[0178] 14. The method of any one of items 3 to 13, the memory further includes one or more interrupt signal vectors, the device table entry further includes an interrupt signal vector address indicator indicating the memory address of the interrupt signal vector in the one or more interrupt signal vectors, each of the interrupt signal vectors includes one or more signal indicators, and each interrupt signal indicator is assigned to one of the one or more bus connection modules and an interrupt target ID, indicating whether an interrupt signal has been received from the corresponding bus connection module addressed to the corresponding interrupt target ID.
[0179] The method further includes:
[0180] Using the indicated memory address of the interrupt signal vector by the bus-attached device to select the interrupt signal indicator assigned to the bus connection module that issues the received interrupt signal and the interrupt target ID addressed by the received interrupt signal.
[0181] Updating the selected interrupt signal indicator such that the selected interrupt signal indicator indicates that there is an interrupt signal to be processed issued by the corresponding bus connection module and addressed to the corresponding interrupt target ID.
[0182] 15. The method of item 14, each of the interrupt signal vectors includes an interrupt signal indicator for each interrupt target ID assigned to the corresponding interrupt target ID, each of the interrupt signal vectors is assigned to a single bus connection module, and the interrupt signal indicator of the corresponding interrupt signal vector is further assigned to the corresponding single bus connection module.
[0183] 16. The method of item 14, each of the interrupt signal vectors includes an interrupt signal indicator for each bus connection module assigned to the corresponding bus connection module, each of the interrupt signal vectors is assigned to a single target processor ID, and the interrupt signal indicator of the corresponding interrupt signal vector is further assigned to the corresponding target processor ID.
[0184] 17. The method of any one of items 14 to 16, each of the interrupt signal vectors is implemented as a continuous region in the memory.
[0185] 18. The method of any one of items 14 to 17, each of the interrupt signal indicators is implemented as a single bit.
[0186] 19. The method of any one of the foregoing items, the device table entry further includes a logical partition ID identifying the logical partition to which the client operating system is assigned, and the forwarding of the interrupt signal by the bus-attached device further includes forwarding the logical partition ID along with the interrupt signal.
[0187] 20. The method of item 19, the bus connection module includes a mapping table for each logical partition ID.
[0188] 21. A method of any one of the foregoing, the method further comprising retrieving, by a bus-attached device, an interrupt subclass ID that identifies an interrupt subclass to which the received interrupt signal is assigned, and the forwarding of the interrupt signal by the bus-attached device further comprises forwarding the interrupt subclass ID along with the interrupt signal.
[0189] 22. A method of any one of the foregoing, a processor of a computer system being adapted to execute a plurality of guest operating systems, and the bus-attached device including a mapping table for each guest operating system of the plurality of guest operating systems.
[0190] 23. A method of any one of the foregoing, the method further comprising:
[0191] receiving, by the bus-attached device, a direct memory access request from a bus connection module to update status information of the bus connection module in a memory, a status update of the bus connection module triggering an interrupt signal,
[0192] after receiving the request, performing, by the bus-attached device, a direct memory access to the memory to update the status information of the bus connection module in the memory.
[0193] 24. A system for providing an interrupt signal to a guest operating system, the guest operating system being executed using one or more of a plurality of processors of a computer system allocated for use by the guest operating system, the computer system further including one or more bus connection modules operably connected to the plurality of processors via a bus and a bus-attached device.
[0194] Each of the plurality of processors is assigned a logical processor ID by the bus-attached device for addressing the corresponding processor,
[0195] Each of the plurality of processors allocated for use by the guest operating system is further assigned an interrupt target ID used by the guest operating system and the one or more bus connection modules to locate the corresponding processor.
[0196] The computer system is configured to execute a method that includes:
[0197] receiving, by the bus-attached device, an interrupt signal having an interrupt target ID from one of the bus connection modules, the interrupt target ID identifying one of the processors that is assigned by the guest operating system as a target processor for processing the interrupt signal,
[0198] The bus-attached device converts the received interrupt target ID into the logical processor ID of the target processor using a mapping table included in the bus-attached device. The mapping table maps the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the multiple processors.
[0199] The bus-attached device directly addresses the target processor using the logical processor ID of the target processor and forwards the interrupt signal to be processed to the target processor.
[0200] 25. A computer program product for providing an interrupt signal to a guest operating system, the guest operating system being executed using one or more of a plurality of processors of a computer system allocated for use by the guest operating system. The computer system further includes one or more bus connection modules operably connected to the plurality of processors via a bus and a bus-attached device.
[0201] Each of the plurality of processors is assigned a logical processor ID by the bus-attached device for addressing the corresponding processor.
[0202] Each of the plurality of processors allocated for use by the guest operating system is further assigned an interrupt target ID by the guest operating system and the one or more bus connection modules for locating the corresponding processor.
[0203] The computer program product includes a computer-readable non-transitory medium readable by a processing circuit and stores instructions executed by the processing circuit for performing a method, the method including:
[0204] The bus-attached device receives an interrupt signal having an interrupt target ID from one of the bus connection modules. The interrupt target ID identifies one of the processors allocated by the guest operating system as the target processor for processing the interrupt signal.
[0205] The bus-attached device converts the received interrupt target ID into the logical processor ID of the target processor using a mapping table included in the bus-attached device. The mapping table maps the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the multiple processors.
[0206] The bus-attached device directly addresses the target processor using the logical processor ID of the target processor and forwards the interrupt signal to be processed to the target processor.
Claims
1. A method for providing an interrupt signal to a guest operating system, the guest operating system being executed using one or more of a plurality of processors of a computer system allocated for use by the guest operating system, the computer system further including one or more bus connection modules operably connected to the plurality of processors via a bus and bus-attached devices, Each of the plurality of processors is assigned a logical processor ID by the bus-attached device for addressing the corresponding processor, Each of the plurality of processors allocated for use by the guest operating system is further assigned an interrupt target ID by the guest operating system and the one or more bus connection modules for finding the corresponding processor, The method comprises: Receiving, by the bus-attached device, an interrupt signal having an interrupt target ID from one of the bus connection modules, the interrupt target ID identifying one of the processors allocated by the guest operating system as the target processor for processing the interrupt signal, Converting, by the bus-attached device, the received interrupt target ID into the logical processor ID of the target processor using a mapping table included in the bus-attached device, the mapping table mapping the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the plurality of processors, Retrieving, by the bus connection device, a copy of a device table entry from a device table stored in a memory operably connected to the bus connection device, the device table entry including a direct signaling indicator indicating whether to directly address the target processor, If the direct signaling indicator indicates direct forwarding of the interrupt signal, directly addressing, by the bus-attached device, the target processor using the logical processor ID of the target processor and forwarding the interrupt signal to be processed to the target processor, otherwise forwarding, by the bus-attached device, the interrupt signal to be processed to the plurality of processors by broadcast.
2. The method according to claim 1, wherein the interrupt signal is received in the form of a message-signaled interrupt, the message-signaled interrupt including the interrupt target ID of the target processor.
3. The method according to claim 1, wherein the direct signaling indicator is implemented as a single bit.
4. The method according to claim 1, wherein during initialization of the guest operating system, the direct signaling indicator is set as a static indicator for the guest operating system.
5. The method according to claim 1, wherein the mapping of the interrupt target IDs of the processors allocated for use by the guest operating system to the logical processor IDs of the plurality of processors is a static mapping defined by the mapping table.
6. The method according to claim 1, wherein the bus-attached device checks whether a copy of the device table entry is cached in a local cache operably connected to the bus-attached device, wherein If a copy of the device table entry is cached, the retrieval of the copy of the device table entry is performed from the corresponding cache, Otherwise, the retrieval of the device table entry is performed from the memory.
7. According to the method of claim 1, the memory further includes an interrupt summary vector, the device table entry further includes an interrupt summary vector address indicator indicating the memory address of the interrupt summary vector, the interrupt summary vector includes an interrupt summary indicator for each bus connection module, and each interrupt summary indicator is assigned to a bus connection module, indicating whether there is an interrupt signal issued by the corresponding bus connection module pending processing. The method further includes using the indicated memory address of the interrupt summary vector by the bus-attached device to update the interrupt summary indicator of the bus connection module from which the interrupt signal is received, such that the updated interrupt summary indicator indicates that there is an interrupt signal issued by the corresponding bus connection module pending processing.
8. According to the method of claim 7, the interrupt summary vector is implemented as a continuous region.
9. According to the method of claim 7, each of the interrupt summary indicators is implemented as a single bit.
10. According to the method of claim 1, the memory further includes a directed interrupt summary vector, the device table entry further includes a directed interrupt summary vector address indicator indicating the memory address of the directed interrupt summary vector, the directed interrupt summary vector includes a directed interrupt summary indicator for each interrupt target ID, and each directed interrupt summary indicator is assigned to an interrupt target ID, indicating whether there is an interrupt signal addressed to the corresponding interrupt target ID pending processing. The method further includes using the indicated memory address of the directed interrupt summary vector by the bus-attached device to update the interrupt summary indicator of the bus connection module assigned to the target processor ID to which the received interrupt signal is addressed, such that the updated interrupt summary indicator indicates that there is an interrupt signal addressed to the corresponding interrupt target ID pending processing.
11. According to the method of claim 10, the directed interrupt summary vector is implemented as a continuous region.
12. According to the method of claim 10, each of the directed interrupt summary indicators is implemented as a single bit.
13. According to the method of claim 1, the memory further includes one or more interrupt signal vectors, the device table entry further includes an interrupt signal vector address indicator indicating the memory address of the interrupt signal vector in the one or more interrupt signal vectors, and each of the interrupt signal vectors includes one or more signal indicators, and each interrupt signal indicator is assigned to a bus connection module in the one or more bus connection modules and an interrupt target ID, indicating whether an interrupt signal has been received from the corresponding bus connection module addressed to the corresponding interrupt target ID. The method further includes: selecting, by the bus-attached device, using the indicated memory address of the interrupt signal vector, the interrupt signal indicator assigned to the bus connection module that issued the received interrupt signal and the interrupt target ID to which the received interrupt signal is addressed; updating the selected interrupt signal indicator such that the selected interrupt signal indicator indicates that there is an interrupt signal issued by the corresponding bus connection module and addressed to the corresponding interrupt target ID pending processing.
14. The method according to claim 13, wherein the interrupt signal vectors each include an interrupt signal indicator for each interrupt target ID assigned to the corresponding interrupt target ID, each of the interrupt signal vectors being assigned to a single bus connection module, and the interrupt signal indicators of the corresponding interrupt signal vectors being further assigned to the corresponding single bus connection module.
15. The method according to claim 13, wherein the interrupt signal vectors each include an interrupt signal indicator for each bus connection module assigned to the corresponding bus connection module, each of the interrupt signal vectors being assigned to a single target processor ID, and the interrupt signal indicators of the corresponding interrupt signal vectors being further assigned to the corresponding target processor ID.
16. The method according to claim 13, wherein the interrupt signal vectors are each implemented as a contiguous region in memory.
17. The method according to claim 13, wherein the interrupt signal indicators are each implemented as a single bit.
18. The method according to claim 1, wherein the device table entry further includes a logical partition ID identifying the logical partition to which the guest operating system is assigned, and the forwarding of the interrupt signal by the bus-attached device further includes forwarding the logical partition ID along with the interrupt signal.
19. The method according to claim 18, wherein the bus connection module includes a mapping table for each logical partition ID.
20. The method according to claim 1, the method further comprising retrieving, by the bus-attached device, an interrupt subclass ID identifying the interrupt subclass to which the received interrupt signal is assigned, and the forwarding of the interrupt signal by the bus-attached device further includes forwarding the interrupt subclass ID along with the interrupt signal.
21. The method according to claim 1, wherein a processor of the computer system is adapted to execute a plurality of guest operating systems, and the bus-attached device includes a mapping table for each guest operating system of the plurality of guest operating systems.
22. The method according to claim 1, the method further comprises: receiving, by the bus-attached device, a direct memory access request from the bus connection module to update status information of the bus connection module in memory, a status update of the bus connection module triggering an interrupt signal, performing, by the bus-attached device after receiving the request, a direct memory access to memory to update the status information of the bus connection module in memory.
23. A computer system for providing interrupt signals to a guest operating system, the guest operating system being executed using one or more of a plurality of processors of the computer system allocated for use by the guest operating system, the computer system further comprising one or more bus connection modules operably connected to the plurality of processors via a bus and a bus-attached device, each of the plurality of processors being assigned a logical processor ID used by the bus-attached device to address the corresponding processor, each of the plurality of processors allocated for use by the guest operating system being further assigned an interrupt target ID used by the guest operating system and the one or more bus connection modules to locate the corresponding processor, the computer system being configured to execute the method according to any one of claims 1-22.
24. A computer program product for providing an interrupt signal to a guest operating system, the guest operating system being executed by one or more of a plurality of processors of a computer system allocated for use by the guest operating system, the computer system further including one or more bus connection modules operably connected to the plurality of processors via a bus and bus-attached devices, each of the plurality of processors is allocated a logical processor ID by the bus-attached device for addressing the corresponding processor, each of the plurality of processors allocated for use by the guest operating system is further allocated an interrupt target ID by the guest operating system and the one or more bus connection modules for locating the corresponding processor, The computer program product includes a computer-readable non-transitory medium readable by a processing circuit and storing instructions executed by the processing circuit for performing the method according to any one of claims 1-22.
Citation Information
Patent Citations
Interrupt Virtualization
US20110197003A1