Computer system, failure processing method, and program therefor

The computer system addresses the challenge of failure detection in virtual machines by using extended configuration registers to transmit failure information from the virtual machine guest to the host, facilitating quicker diagnosis and improving system reliability.

JP7683966B1Active Publication Date: 2025-05-27NEC PLATFROMS LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024011000
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-05-27
Estimated Expiration
2044-01-29

AI Technical Summary

Technical Problem

In virtual machine environments, when a failure occurs in a virtual machine, the failure information cannot be transmitted outside the virtual machine environment, making it difficult for the host to identify and recognize the failure, leading to delayed diagnosis and recovery.

Method used

A computer system with a first extended configuration register and a first diagnostic means that notifies the virtual machine host of a failure in a dedicated processor within the virtual machine guest, using information stored in the extended configuration register, and a second extended configuration register that allows the host to read this information and detect failures.

Benefits of technology

This solution enables quicker identification and diagnosis of failures within the virtual machine guest from the host side, reducing the time to recognize and address issues, thereby improving system availability and operation rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683966000001_ABST
    Figure 0007683966000001_ABST
Patent Text Reader

Abstract

Provide a computer system that makes it easier to identify suspected failures in the guest of a virtual machine even in the host of the virtual machine. 【Solution means】The computer system includes a virtual machine guest in which a dedicated processor incorporated in the virtual machine host is incorporated, the dedicated processor having a first diagnostic unit that notifies the virtual machine host of the occurrence of a failure in the dedicated processor via information included in a first extended configuration register, and a virtual machine host having a second diagnostic unit that detects a failure related to the dedicated processor from information included in a second extended configuration register that has the same content as the first extended configuration register and from which information written by the virtual machine guest to the first extended configuration register can be read from the virtual machine host via the second extended configuration register.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a computer system, a failure processing method, and a program therefor.

Background Art

[0002] For reasons such as being able to reduce the number of physical servers and expecting a cost reduction effect, virtual machines may be provided on one physical server. When a failure occurs in such a virtual machine, it is important to be able to quickly recover and increase the operation rate.

[0003] For example, Patent Document 1 describes a virtual computer system that provides failure monitoring means for each virtual machine, identifies the device in which a failure has occurred based on the detected failure information, and notifies the system administrator of the device failure.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way, in the construction of a virtual machine, a host for the virtual machine constructed on the host OS installed on the base computer may be provided, and further, a guest that becomes a virtual machine operating on the guest OS constructed on the host OS may be provided. Conventionally, when a failure is detected in a virtual machine, although the failure information is stored in the virtual machine environment, the failure information cannot be transmitted to the outside. Therefore, on the host side of the virtual machine, not only is it impossible to identify the suspect, but it is also impossible to recognize that a failure has occurred. As a result, symptoms such as abnormal job execution on the virtual machine occur, and the user of the virtual machine contacts the machine administrator to that effect. However, the machine administrator cannot immediately grasp the situation of what has happened, and will perform network confirmation, restart the virtual machine, etc., and it may take time to identify the cause.

[0006] An object of the present disclosure is to provide a computer system, a failure processing method, and a program therefor that solve the above problems.

Means for Solving the Problems

[0007] A computer system according to an aspect of the present disclosure includes a first extended configuration register, and a first diagnostic means for notifying a virtual machine host of the occurrence of a failure in a dedicated processor incorporated in a virtual machine guest via information included in the first extended configuration register. A virtual machine guest that is a virtual machine in which the dedicated processor is incorporated, a second extended configuration register, which has the same content as the first extended configuration register and allows the virtual machine host to read information written by the virtual machine guest to the first extended configuration register via the second extended configuration register, the second extended configuration register, and a second diagnostic means for detecting a failure related to the dedicated processor incorporated in the virtual machine guest from the information included in the second extended configuration register, and a virtual machine host that is a host machine.

[0008] According to one aspect of the present disclosure, a failure processing method causes a virtual machine guest, which is a virtual machine incorporating a dedicated processor, to notify a virtual machine host of the occurrence of a failure in the dedicated processor incorporated in the virtual machine host via information included in a first extended configuration register. The virtual machine host, which is the host machine, reads, via a second extended configuration register having the same content as the first extended configuration register, information written by the virtual machine guest to the first extended configuration register. In the second extended configuration register, a failure related to the dedicated processor incorporated in the virtual machine guest is detected from the information included in the second extended configuration register.

[0009] According to one aspect of the present disclosure, a program for failure processing causes a virtual machine guest, which is a virtual machine incorporating a dedicated processor, to notify a virtual machine host of the occurrence of a failure in the dedicated processor incorporated in the virtual machine host via information included in a first extended configuration register. The virtual machine host, which is the host machine, reads, via a second extended configuration register having the same content as the first extended configuration register, information written by the virtual machine guest to the first extended configuration register. In the second extended configuration register, a failure related to the dedicated processor incorporated in the virtual machine guest is detected from the information included in the second extended configuration register, and causes a computer to execute this.

Advantages of the Invention

[0010] According to the above aspect, it becomes easier to suspect and identify a failure in the guest of the virtual machine even in the host of the virtual machine.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Modes for Carrying Out the Invention

[0012] Hereinafter, each embodiment will be described with reference to the drawings. In all the drawings, the same or corresponding components are denoted by the same reference numerals, and common descriptions will be omitted. As an embodiment of the present disclosure, in a computer system in which a GPU (Graphics Processing Unit), a vector engine, etc. are PCI-connected, for example, when starting a virtual machine by a host-type hypervisor and using the GPU or vector engine from a guest environment with the PCI passthrough function, the case where a failure occurs in the GPU or vector engine will be described as an example. Here, the "host-type hypervisor" refers to a method in which it operates as application software on an OS, installs virtualization software on the basis of that OS, and operates a virtual machine with the virtualization software. Further, the "PCI passthrough function" is a function that enables direct access to PCI devices such as a GPU or vector engine, which is a physical machine, from a guest OS installed on a virtual machine.

[0013] Note that the GPU is a unit including a dedicated processor that processes calculations necessary for rendering images. The vector engine is a unit equipped with a processor having "vector instructions" for processing a large amount of data collectively. Hereinafter, "VE" is an abbreviation for "vector engine". On the other hand, the CPU (Central Processing Unit) is a general-purpose processor for performing general-purpose processing. Since a unit including a dedicated processor such as a GPU or VE is specialized for a specific process, there is a significant difference in processing speed in that specific process compared to a general-purpose processor such as a CPU. Note that hereinafter, a unit including a dedicated processor such as a GPU or vector engine will also be simply referred to as a "dedicated processor".

[0014] FIG. 1 is a diagram for explaining the configuration of a computer system as an embodiment in the present disclosure. In FIG. 1, a computer system (1) includes four vector engines (hereinafter abbreviated as VEs) (30 to 33) as dedicated processors, a CPU (10) as a general-purpose processor, and PCIeSWs (PCIe switches) (20) and (21) for connecting the VEs (30) to (33) to the CPU (10). Here, PCIe (Peripheral Component Interconnect Express) is a high-bandwidth expansion bus generally used for connecting peripheral devices such as the VEs (30) to (33). Also, the PCIeSWs (20) and (21) are switching devices having the aforementioned expansion bus. Also, it is assumed that the management numbers of the VEs (30) to (33) in the computer system (1) are managed in order as 0, 1, 2, and 3. Hereinafter, each vector engine may be denoted as VE#0, VE#1, VE#2, and VE#3 in the order of the vector engine numbers.

[0015] FIG. 2 is a diagram showing the software structure when a guest virtual machine is started on a host machine. In FIG. 2, two of the four VEs, numbered 2 and 3, are incorporated into a guest virtual machine, which is an example when it can be used in the environment of the guest virtual machine. On the other hand, the remaining two VEs numbered 0 and 1 are not incorporated into the guest virtual machine and are examples where they can be used in the host environment as before starting the virtual machine. Hereinafter, when starting the virtual machine, the environment that has started as the virtual machine is expressed as the VM (Virtual Machine) guest, and the host environment of the machine for the virtual machine is expressed as the VM host.

[0016] As shown in FIG. 2, the VM host is built on the host OS (1100). Also, in the application space (1200) of the VM host, diagnostic software (diagnostic SW) (1210), the PCI config space of the CPU (1220), the PCI config space on the VEx opposite side (1230), and the PCI config spaces (1240) to (1243) for VE#0 to VE#3 are provided. Further, in the OS (1100), VEDRVs (VE Drivers) (1110), (1111) for VE#0 and VE#1 used in the host environment are provided. Note that "x" in the PCI config space (1230) on the VEx opposite side represents an abbreviation of the management number of each VE. In the example of the present disclosure, there are four PCI config spaces on the opposite side of the VE, specifically, in the example of the present disclosure, there are a PCI config space on the opposite side of VE#0, a PCI config space on the opposite side of VE#1, a PCI config space on the opposite side of VE#2, and a PCI config space on the opposite side of VE#3. Also, "VEDRV" is a driver for the VE which is a dedicated processor, and is control software that serves as an interface between the application space and the VE.

[0017] The diagnostic SW (1210) resides in the VM host, monitors the failures of VEs and the Link connection status, and performs failure handling. For this purpose, the diagnostic SW (1210) is configured to be able to detect failures via the VE driver and further collect VE failure information via the VEDRV. Also, the diagnostic SW (1210) can read the Link status registers in the PCI config spaces (1220 and 1230) of the port on the PCIeSW side to which the VE is connected and the port on the CPU side to which the PCIeSW is connected in order to check the Link connection status with the PCI-connected devices. Furthermore, when the diagnostic SW (1210) detects a failure in the VE, it outputs information about the failure as a summary file and / or a system log (syslog). Hereinafter, the register is sometimes abbreviated as "reg". Note that information such as the Link speed and Link width can be obtained from the "Link status register". Also, in PCIe, eight types of Link widths, namely 2, 4, 8, 12, 16, 32, and 64 lanes, are defined. The VEDRV is a software interface for soft-connecting to the target VE as described above. Each VE can be used by the VM guest through the PCI passthrough function, but since the PCI passthrough function is a general function, its detailed description is omitted.

[0018] The VM guest is built on the VM software (1300) for building a virtual machine on the OS (1100). In FIG. 2, the VM guest in which VE#2 and VE#3 are incorporated also has the diagnostic SW (1710) resident in its application space (1700), and PCI config spaces (1742), (1743) for VE#2 and VE#3 are provided. Also, the guest OS (1600) is equipped with VEDRVs (1612), (1613) for VE#2 and VE#3 incorporated in the guest environment.

[0019] When the diagnostic SW (1710) detects a failure, it sets information indicating the occurrence of the failure in the extended configuration registers in the PCI configuration spaces (1742) and (1743). Also, when the preparation for starting the operation of the VM guest is complete, the diagnostic SW (1710) starts accumulating the counters included in the extended configuration registers.

[0020] Here, the extended configuration registers of the PCI configuration space (1242) for VE#2 configured in the application space (1200) of the VM host and the extended configuration registers of the PCI configuration space (1742) for VE#2 configured in the application space (1700) of the VM guest have the same content, and are configured such that values and information written from one of the VM host and the VM guest can also be read from the other. Similarly, in the case of the extended configuration registers of the PCI configuration space (1243) for VE#3 configured in the application space (1200) of the VM host and the extended configuration registers of the PCI configuration space (1743) for VE#3 configured in the application space (1700) of the VM guest, they have the same content, and are configured such that values and information written from one of the VM host and the VM guest can also be read from the other. Note that the configuration of the extended configuration registers will be described in detail separately. Also, each PCI configuration space is provided with a Link status register indicating its link status.

[0021] Figure 3 is a diagram showing the software structure when the virtual machine is not started. That is, it is an example where all four VEs are available in the host environment. The diagnostic SW (2210) is the same as the diagnostic SW (1210) in Figure 2. In addition, VEDRV (2110) to (2113), the PCI config space on the CPU side (2220), the PCI config space on the VEx opposite side (PCIeSW side) (2230), and the PCI config spaces for VE#1 to VE#3 (2240) to (2243) are also the same as VEDRV (1110) to (1113), the PCI config space on the CPU side (1220), the PCI config space on the VEx opposite side (PCIeSW side) (1230), and the PCI config spaces for VE#1 to VE#3 (1240) to (1243) in Figure 1, respectively.

[0022] Figure 4 is a diagram showing an example of the configuration of the extended config register in the PCI config space for the VE. In the PCI config space for the VE, an extended config register is provided. The extended config register has a dedicated area as an area that can be used by the user, and the dedicated area stores a serial number, a flag, a counter, and failure information. Here, the "serial number" is the serial number set for the VE. The "flag" is a value indicating the status of the operating state of the VM guest to be separately processed. The "counter" indicates a value that is the integrated result after the VE is ready to start operating. The "failure state" stores information regarding the VE failure to be separately described. As described above, the extended config register should be pre-arranged between the diagnostic SW (2210) in the VM guest and the diagnostic SW (1710) in the VM host so that the stored content can be read by the other party in response to the write of one of the VM host and the VM guest. In the following, the case of reading from the VM host in response to the write of the VM guest is described as an example.

[0023] FIG. 5 is a diagram showing an example of a table of VE numbers and serial numbers. The table shown in FIG. 5 shows an example of the correspondence between the numbers of VEs and the serial numbers set for those VEs. The table shown in FIG. 5 is provided in the VM host and is used to identify the VE in which a failure has occurred when the VM host detects a VE failure.

[0024] FIG. 6 is a diagram for explaining the flags provided in the dedicated area in the extended config register. In the example shown in FIG. 6, numbers 0 to 2 are associated as numbers indicating the status of the operation of the VM guest.

[0025] FIG. 7 is a diagram showing an example of the breakdown of the failure information provided in the dedicated area in the extended config register.

[0026] FIG. 8 is a diagram showing a registration example of the summary file output by the diagnostic SW (1210) on the VM host side or the diagnostic SW (1710) on the VM guest side.

[0027] FIG. 9 is a diagram showing an output example of the syslog file output by the diagnostic SW (1210) on the VM host side or the diagnostic SW (1710) on the VM guest side. Note that the summary file in FIG. 8 is the failure information obtained by translating the content of the syslog file in FIG. 9 so that it can be easily understood by the user or administrator of the computer system 1.

[0028] Next, the operation of the failure processing of the diagnostic SW (1210) on the VM host side and the operation of the failure processing of the diagnostic SW (1710) on the VM guest side will be described.

[0029] FIG. 10 is a flowchart showing the operation of the failure process of the diagnostic SW (1210) on the VM host side. Before starting up the virtual machine, the diagnostic SW (2210) reads out the VE-specific serial number from each VE, and stores the correspondence table between the management number of each VE and its serial number as a table (step 1100). FIG. 5 mentioned above is an example of a table showing the correspondence between the VE number and the serial number. After that, the VM host starts up the virtual machine. As described above, the extended config reg in the VM guest has the same content as the extended config reg in the VM host, and the value written from one side can be read from the other side. FIG. 2 is an example of the configuration when the virtual machine is started up. Here, as shown in FIG. 1, four VEs are connected to the CPU via the PCIeSW, and it is assumed that two of the VEs, VE#2 and VE#3, are incorporated into the VM guest as shown in FIG. 2.

[0030] The diagnostic SW (1210) in the VM host determines whether the VE is incorporated into the virtualization from VEDRV or the like (step 2100). If the VE is incorporated into the virtualization (step 2100: Yes), then next, the diagnostic SW (1210) monitors whether the value of the flag in the dedicated area in the extended config reg in the PCI config spaces (1242), (1243) indicates that the preparation is complete (step 2200), and when it recognizes that the preparation is complete (step 2200: Yes), it proceeds to the next process.

[0031] The diagnostic SW (1210) checks whether the Link connection state is normal (step 2300) from the Link state registers included in the PCI config space (1220) of the CPU and the Link state registers included in the PCI config space (1230) on the opposite side of the VEx. For example, the Link connection state targets the port on the PCIeSW side to which the VE is connected and the port on the CPU side to which the PCIeSW is connected, and checks whether the Link speed and Link width can maintain the expected values from the Link state registers. For example, in FIG. 1, the Link connection state with VE#3 (33) is checked for the port (22) provided in PCIeSW#1 (21) to which VE#3 (33) is connected and the port (11) provided in CUP (10) to which PCIeSW#1 (21) is connected. Then, the diagnostic SW (1210) checks whether the Link connection state is normal by referring to the Link state registers included in the PCI config space (1220) of the CPU and the Link state registers regarding PCIeSW#1 (21) included in the PCI config space (1230) on the opposite side of VE#3. Incidentally, in PCIeSW#1 (21), there are a total of three PCI config spaces for two ports for connecting to VE#2 (32) and VE#3 (33) respectively and one port for connecting to the CPU (10).

[0032] If the Link connection state is normal (step 2300: Yes), the diagnostic SW (1210) checks whether the counter in the dedicated area within the extended config register is being integrated (step 2400). If it can be confirmed that the counter is being integrated (step 2400: Yes), the diagnostic SW (1210) checks whether the flags in the dedicated area within the extended config register in the PCI config spaces (1242), (1243) indicate the occurrence of a failure in the VE incorporated into the VM guest (step 2500). If no failure has occurred (step 2500: No), it returns to step 2300 and steps 2300 to 2500 are repeated.

[0033] On the other hand, when a failure occurs (step2500: Yes), the diagnostic SW (1210) reads the failure information from the dedicated area in the extended config register containing the flag indicating the occurrence of the failure (step2501). Also, the diagnostic SW (1210) reads the serial number from the dedicated area in the extended config register containing the flag indicating the occurrence of the failure, retrieves it from the table where the management number of the VE at the time of non-virtualization was stored in step1100, and identifies the VE in which the failure has occurred (step2502). Then, the diagnostic SW (1210) registers in the summary file a timestamp which is the time when the failure occurred, the VE number, the fact that it is incorporated into the VM guest, and a failure suspicion indicating the functional block in which the failure has occurred (step2503). Furthermore, the diagnostic SW (1210) outputs to the syslog file the occurrence of the failure, the failure suspicion, etc. (step2504).

[0034] Through the processing from step2500 to step2504 above, the diagnostic SW (1210) detects the failure of the VE incorporated into the VM guest.

[0035] In step 2400, if it cannot be confirmed that the counter in the dedicated area of the extended config register is being incremented, that is, if the counter has stopped or cannot be read (step 2400: No), the diagnostic SW (1210) determines that there is a failure in the PCI interface part within the VE guest (step 2411). Then, the diagnostic SW (1210) reads the serial number from the dedicated area of the extended config register where the counter is not being incremented, and retrieves the VE with the failure from the table where the management number of the VE at the time of non-virtualization was stored in step 1100 (step 2412). And the diagnostic SW (1210) registers in the summary file a time stamp which is the time when the failure occurred, the VE number, that it is incorporated in the VM guest, and a failure suspicion indicating the functional block where the failure occurred (step 2413), and outputs to the syslog file that a failure has occurred and the failure suspicion, etc. (step 2414). Through the processing from step 2400 and step 2411 to step 2414 above, the diagnostic SW (1210) detects a failure in the PCI interface part within the VM guest in a state where the VE is incorporated in the virtual machine guest.

[0036] In step 2300, if it cannot be confirmed that the Link connection state is normal, that is, if the Link speed or Link width is not the expected value (step 2300: No), the diagnostic SW (1210) determines that there is a failure of incorrect Link state (step 2311). Then, the diagnostic SW (1210) collects the failure information and registers it in the log file (step 2312). Further, the diagnostic SW (1210) registers in the summary file a time stamp which is the time when the failure occurred, the VE number, that it is incorporated in the VM guest, and a failure suspicion indicating the functional block (PCIeSW or CPU) where the failure occurred (step 2313), and outputs to the syslog file that a failure has occurred and the failure suspicion, etc. (step 2314). Through the processing from step 2300 and step 2311 to step 2314 above, the diagnostic SW (1210) detects a failure of incorrect Link state in the port on the PCIeSW side to which the VE is connected and the port on the CPU side to which the PCIeSW is connected in a state where the VE is incorporated in the virtual machine guest.

[0037] If the VE is not incorporated into the virtual machine (2100: No), failure processing in a state where the virtualization machine has not been started for the VE will be performed. The failure processing in a state where the virtualization machine has not been started for the VE will be described below.

[0038] The diagnostic SW (1210) checks whether the Link connection state is normal from the Link state register (step 2110). This is equivalent to performing the same operation as step 2300. If the Link connection state is normal (step 2110: Yes), then next, the diagnostic SW (1210) continues monitoring via the VEDRV until a failure is detected (step 2111). If a failure is detected (step 2111: Yes), the diagnostic SW (1210) collects the failure information and registers it in the log file (step 2112). Further, the diagnostic SW (1210) registers in the summary file a time stamp that is the time when the failure occurred, the VE number, and a failure suspect indicating the functional block where the failure occurred (step 2113), and outputs to the syslog file information such as the occurrence of the failure and the failure suspect identified from the failure information (step 2114). Through the processing from step 2111 to step 2114 above, the diagnostic SW (1210) detects a failure within the VE in a state where the VE is not incorporated into the virtual machine guest.

[0039] In step 2110, if it cannot be confirmed that the Link connection state is normal, that is, if the Link speed or Link width is not the expected value (step 2110: No), the diagnostic SW (1210) determines that there is a Link state incorrect failure (step 2121), collects the failure information, and registers it in the log file (step 2122). Further, the diagnostic SW (1210) registers in the summary file a time stamp that is the time when the failure occurred, the VE number, and a failure suspect indicating the functional block (PCIeSW or CPU) where the failure occurred (step 2123), and outputs to the syslog file the occurrence of the failure and the failure suspect, etc. (step 2124). Through the processing from step 2110, step 2121 to step 2114 above, the diagnostic SW (1210) detects a Link state incorrect failure in the port on the PCIeSW side to which the VE is connected and the port on the CPU side to which the PCIeSW is connected in a state where the VE is not incorporated into the virtual machine guest.

[0040] As described above, by the diagnostic SW (1210) in the VM host outputting information related to the failure to the summary file and the syslog file by detecting the failure, the user or administrator of the VM host can grasp the details of the failure.

[0041] Next, the operation of the diagnostic SW (1710) in the VM guest will be described. FIG. 11 is a diagram showing the operation of the diagnostic SW (1710) in the VM guest. First, the diagnostic SW (1710) in the VM guest clears all the dedicated areas in the extended config registers in the PCI config spaces (1742), (1743) to zero (step 3100). Next, the diagnostic SW (1710) in the VM guest writes the serial numbers read from VE#2 and VE#3 to the dedicated areas in its extended config registers respectively (step 3200). Next, when the preparation for startup is complete, the diagnostic SW (1710) in the VM guest sets the flag in the dedicated area of its extended config register to the value "1" indicating "preparation complete" (step 3300). Also, the diagnostic SW (1710) in the VM guest starts accumulating the counter in the dedicated area of its extended config register (step 3400).

[0042] The diagnostic SW (1710) in the VM guest continues monitoring via VEDRV (1612), (1613) until a failure is detected (step 3411). When a failure is detected (step 3411: Yes), the diagnostic SW (1710) in the VM guest collects the failure information and registers it in the log file (step 3412). Furthermore, the diagnostic SW (1710) in the VM guest registers in the summary file a timestamp which is the time when the failure occurred, the VE number, and the suspected failure indicating the functional block where the failure occurred (step 3413), and outputs to the syslog file the fact that a failure has occurred and the suspected failure identified from the failure information, etc. (step 3414).

[0043] The diagnostic SW (1710) in the VM guest writes the failure information to the dedicated area in the extended config register corresponding to the VE where the failure occurred (step 3415). Also, the diagnostic SW (1710) in the VM guest sets the flag in the dedicated area of the extended config register corresponding to the VE where the failure occurred to the value "2" indicating "failure occurred" (step 3416).

[0044] As described above, by the diagnostic SW (1710) in the VM guest outputting information regarding the failure to the summary file and the syslog file, the user of the VM guest can grasp the details of the failure. Also, by the diagnostic SW (1710) in the VM guest setting a value indicating "failure occurred" to the flag in the dedicated area in the extended config register corresponding to the VE where the failure occurred and writing the failure information, it becomes possible to detect the occurrence of a failure in the built-in VE both on the VM host side and in the VM guest.

[0045] Note that in FIGS. 8 and 9, examples of the content registered in the summary file and the syslog file by the processing of step 2123 and step 2124 by the diagnostic SW (1210) of the VM host are (8-1) in FIG. 8 and (9-1) in FIG. 9, and examples of the content registered in the summary file and the syslog file by the processing of step 2113 and step 2114 are (8-2) in FIG. 8 and (9-2) in FIG. 9. Also, examples of the content registered in the summary file and the syslog file by the processing of step 2413 and step 2414 by the diagnostic SW (1210) of the VM host are (8-3) in FIG. 8 and (9-3) in FIG. 9, and examples of the content registered in the summary file and the syslog file by the processing of step 2503 and step 2504 are (8-4) in FIG. 8 and (9-4) in FIG. 9.

[0046] As described above, in the VE incorporated in the VM guest, by transmitting the failure information of the VE detected on the VM guest side to the VM host side through the extended config register, the VM host side can recognize the occurrence of a failure in the VE that occurred on the VM guest side and can also identify the suspect.

[0047] Furthermore, in a failure mode where the extended config register becomes blocked, that is, for a failure in the VE side PCIe interface section or a failure in the PCIeSW, the VM host can also detect such a failure by performing heartbeat monitoring (counter integration and its monitoring).

[0048] In the present disclosure, the VE (Vector Engine) is described as an example of a dedicated processor, but the present disclosure is not limited thereto. The same means can be applied using a GPU (Graphics Processing Unit) instead of the VE. In a computer system in which a VE or a GPU is incorporated into a virtual machine, when a failure occurs in the GPU or the VE, the occurrence of the failure can be quickly recognized, and the suspected failure can be specified, so that recovery can be achieved at an early stage, and the availability and the operation rate can be improved. Further, the technology of the present disclosure is applicable to various PCI devices such as a network card connectable by PCI and a fiber channel such as InfiniBand, in addition to the dedicated processors described above.

[0049] FIG. 12 is a diagram showing a configuration example of a computer system according to an embodiment of the present disclosure. A computer system (1) includes a virtual machine host (40) which is a host machine, and a virtual machine guest (50) which is a virtual machine in which a dedicated processor is incorporated. The virtual machine guest (50) includes a first extended configuration register (51), and a first diagnostic unit (first diagnostic means) (52) that notifies the virtual machine host of the occurrence of a failure in the dedicated processor incorporated in the virtual machine host via information included in the first extended configuration register. The virtual machine host (40) includes a second extended configuration register (41) having the same content as the first extended configuration register (51), and the information written by the virtual machine guest (50) to the first extended configuration register (51) can be read from the virtual machine host (40) via the second extended configuration register (41), the second extended configuration register (41), and a second diagnostic unit (second diagnostic means) (42) that detects a failure related to the dedicated processor incorporated in the virtual machine guest (50) from the information included in the second extended configuration register (41).

[0050] FIG. 13 is a block diagram showing an example of the hardware configuration in the computer system (1) illustrated in FIG. 1. Here, the computer system (1) includes a CPU (10), a PCIeSW #0 (20), a PCIeSW #1 (30), VEs #0 to #3 (30 to 33), and in addition, a ROM (Read Only Memory) (11), a recording device (12), a RAM (Random Access Memory) (13), and the like. In the ROM (11) and the recording device (12), a computer program and the like for realizing the functions of the computer system (1) are recorded. The RAM (13) is used as a work area or the like for temporarily storing data and the like used when the CPU (10) and the like are operating. The computer system (1) may include an input / output port (14) for connecting input / output devices such as a keyboard, a mouse, and a display device. These are connected by a bus or the like. Note that the ROM (11) is constituted by an EEPROM (Electrically Erasable Programmable Read-Only Memory) or the like, and the recording device (12) is constituted by a hard disk, an SSD, or the like, and the computer program for realizing the functions of the computer system (1) may be updated by these devices.

[0051] As described above, the present disclosure has been described with reference to the embodiments, but the present disclosure is not limited to the above-described embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. And each embodiment can be combined with other embodiments as appropriate.

[0052] Some or all of the above embodiments can be described as follows in the appended claims, but are not limited thereto.

[0053] (Appended Claim 1) A first extended configuration register, first diagnostic means for notifying the virtual machine host of the occurrence of a failure in a dedicated processor incorporated in the virtual machine host via information included in the first extended configuration register; A virtual machine guest that is a virtual machine incorporating the dedicated processor, and A second extended configuration register having the same content as the first extended configuration register, and enabling the virtual machine host to read, via the second extended configuration register, information written by the virtual machine guest to the first extended configuration register; the second extended configuration register; Second diagnostic means for detecting a failure related to the dedicated processor incorporated in the virtual machine guest from information included in the second extended configuration register; A virtual machine host that is a host machine, and A computer system comprising the same.

[0054] (Appendix 2) When the first diagnostic means of the virtual machine guest detects a failure via a driver for the incorporated dedicated processor, the first diagnostic means sets information indicating the occurrence of a failure of the incorporated dedicated processor in a flag included in the first extended configuration register and indicating the state of the virtual machine. The second diagnostic means of the virtual machine host detects a failure of the dedicated processor incorporated in the virtual machine guest from a flag included in the second extended configuration register and indicating the state of the virtual machine guest. The computer system according to Appendix 1.

[0055] (Appendix 3) When preparations for starting operation are complete, the first diagnostic means of the virtual machine guest starts accumulating a counter included in the first extended configuration register. When the accumulation of the counter included in the second extended configuration register stops or cannot be read, the second diagnostic means of the virtual machine host determines that there is a failure in the interface with the dedicated processor on the virtual machine guest side. The computer system according to Appendix 1 or Appendix 2.

[0056] (Appendix 4) Comprising a connection switch for connecting the dedicated processor and the general-purpose processor, When the link speed and / or link width on the general-purpose processor side of the connection switch, or the link speed and / or link width on the dedicated processor side of the connection switch do not reach the expected value, the second diagnostic means of the virtual machine host determines it as a fault of incorrect link state. The computer system according to any one of Appendices 1 to 3.

[0057] (Appendix 5) The second diagnostic means of the virtual machine host reads the unique serial number of the dedicated processor connected to the computer system from the connected dedicated processor, creates a table associating the management number of the connected dedicated processor with the serial number, The first diagnostic means of the virtual machine guest reads the serial number from the embedded dedicated processor and writes it into the first extended configuration register. When the second diagnostic means of the virtual machine host detects a fault of the dedicated processor embedded in the virtual machine guest from the flag included in the second extended configuration register, it searches the table using the serial number of the embedded dedicated processor included in the second extended configuration register to identify the management number of the embedded dedicated processor and identify the dedicated processor for which the fault has been detected. The computer system according to any one of Appendices 1 to 4.

[0058] (Appendix 6) When the first diagnostic means of the virtual machine guest detects a fault of the embedded dedicated processor, it collects fault information and writes the information about the detected fault into the first extended configuration register. When the second diagnostic means of the virtual machine host detects a failure of the dedicated processor incorporated in the virtual machine guest from a flag indicating the state of the virtual machine guest, the second diagnostic means reads information regarding the failure from the second extended configuration register. The computer system according to any one of Appendices 1 to 5.

[0059] (Appendix 7) When the second diagnostic means of the virtual machine host detects the occurrence of a failure, the second diagnostic means registers and outputs information suspected of being a failure, which is information regarding the failure, to a predetermined file in the virtual machine host. The computer system according to any one of Appendices 1 to 6.

[0060] (Appendix 8) The first diagnostic means of the virtual machine guest registers and outputs information suspected of being a failure, which is information regarding the failure of the incorporated dedicated processor, to a predetermined file in the virtual machine guest. The computer system according to any one of Appendices 1 to 7.

[0061] (Appendix 11) A virtual machine guest, which is a virtual machine incorporating a dedicated processor, notifies the virtual machine host of the occurrence of a failure in the dedicated processor incorporated in the virtual machine host via the information included in the first extended configuration register. In a second extended configuration register that has the same content as the first extended configuration register and is in a virtual machine host that is a host machine, and from which the virtual machine host can read the information written by the virtual machine guest to the first extended configuration register via the second extended configuration register, a failure related to the dedicated processor incorporated in the virtual machine guest is detected from the information included in the second extended configuration register. Failure handling method.

[0062] (Appendix 12) When a failure is detected by the first diagnostic means of the virtual machine guest via the driver for the dedicated processor incorporated therein, information indicating the occurrence of a failure in the incorporated dedicated processor is set in a flag included in the first extended configuration register and indicating the state of the virtual machine. The second diagnostic means of the virtual machine host detects a failure in the dedicated processor incorporated in the virtual machine guest from a flag included in the second extended configuration register and indicating the state of the virtual machine guest. The failure processing method according to Appendix 11.

[0063] (Appendix 13) When the first diagnostic means of the virtual machine guest is ready to start operation, the accumulation of a counter included in the first extended configuration register is started. When the second diagnostic means of the virtual machine host stops the accumulation of a counter included in the second extended configuration register or cannot read it, it is determined that there is a failure in the interface with the dedicated processor on the virtual machine guest side. The failure processing method according to Appendix 11 or Appendix 12.

[0064] (Appendix 14) When the second diagnostic means of the virtual machine host determines that the link speed and / or link width on the general-purpose processor side of the connection switch connecting the dedicated processor and the general-purpose processor, or the link speed and / or link width on the dedicated processor side of the connection switch, do not reach the expected value, it is determined that there is a failure in the link state. The failure processing method according to any one of Appendices 11 to 13.

[0065] (Appendix 15) The second diagnostic means of the virtual machine host reads out the unique serial number of the dedicated processor connected to the computer system from the connected dedicated processor, and creates a table associating the management number of the connected dedicated processor with the serial number. The first diagnostic means of the virtual machine guest reads out the serial number from the embedded dedicated processor and writes it to the first extended configuration register. When the second diagnostic means of the virtual machine host detects a failure of the dedicated processor embedded in the virtual machine guest from the flag included in the second extended configuration register, the table is searched using the serial number of the dedicated processor embedded in the second extended configuration register to identify the management number of the embedded dedicated processor and identify the dedicated processor in which the failure was detected. The failure processing method according to any one of Appendices 11 to 14.

[0066] (Appendix 16) When the first diagnostic means of the virtual machine guest detects a failure of the embedded dedicated processor, it collects failure information and writes the information about the detected failure to the first extended configuration register. When the second diagnostic means of the virtual machine host detects a failure of the dedicated processor embedded in the virtual machine guest from the flag indicating the state of the virtual machine guest, it reads out the information about the failure from the second extended configuration register. The failure processing method according to any one of Appendices 11 to 15.

[0067] (Appendix 17) When the second diagnostic means of the virtual machine host detects the occurrence of a failure, it registers and outputs the suspected failure information, which is information about the failure, to a predetermined file in the virtual machine host. The failure processing method according to any one of Appendices 11 to 16.

[0068] (Appendix 18) The first diagnostic means of the virtual machine guest registers and outputs information suspected of being a failure, which is information regarding a failure of the incorporated dedicated processor, to a predetermined file within the virtual machine guest. The failure processing method according to any one of Appendices 11 to 17.

[0069] (Appendix 21) The virtual machine guest, which is a virtual machine incorporating a dedicated processor, notifies the virtual machine host of the occurrence of a failure in the dedicated processor incorporated in the virtual machine host via the information included in the first extended configuration register. In a second extended configuration register having the same content as the first extended configuration register, which is a virtual machine host serving as a host, and from which the virtual machine guest can read the information written by the virtual machine guest to the first extended configuration register via the second extended configuration register, a failure related to the dedicated processor incorporated in the virtual machine guest is detected from the information included in the second extended configuration register. A program for causing a computer to execute failure processing.

[0070] (Appendix 22) When the first diagnostic means of the virtual machine guest detects a failure via the driver for the incorporated dedicated processor, information indicating the occurrence of a failure of the incorporated dedicated processor is set in a flag included in the first extended configuration register and indicating the state of the virtual machine. The second diagnostic means of the virtual machine host detects a failure of the dedicated processor incorporated in the virtual machine guest from a flag included in the second extended configuration register and indicating the state of the virtual machine guest. The program according to Appendix 21.

[0071] (Appendix 23) When the first diagnostic means of the virtual machine guest determines that the preparation for starting operation is complete, start accumulating the counter included in the first extended configuration register. When the second diagnostic means of the virtual machine host determines that the accumulation of the counter included in the second extended configuration register has stopped or cannot be read, it is determined that there is a failure in the interface with the dedicated processor on the virtual machine guest side. The program according to Appendix 21 or Appendix 22.

[0072] (Appendix 24) When the second diagnostic means of the virtual machine host determines that the link speed and / or link width on the general-purpose processor side of the connection switch connecting the dedicated processor and the general-purpose processor, or the link speed and / or link width on the dedicated processor side of the connection switch do not reach the expected value, it is determined that there is a failure in the link state. The program according to any one of Appendices 21 to 23.

[0073] (Appendix 25) The second diagnostic means of the virtual machine host reads the unique serial number of the dedicated processor connected to the computer system from the connected dedicated processor, and creates a table associating the management number of the connected dedicated processor with the serial number. The first diagnostic means of the virtual machine guest reads the serial number from the embedded dedicated processor and writes it into the first extended configuration register. When the second diagnostic means of the virtual machine host detects a failure of the dedicated processor incorporated into the virtual machine guest from the flag included in the second extended configuration register, the table is searched using the serial number of the dedicated processor incorporated in the second extended configuration register to identify the management number of the dedicated processor incorporated, and the dedicated processor in which the failure is detected is identified. The program according to any one of Appendices 21 to 24.

[0074] (Appendix 26) When the first diagnostic means of the virtual machine guest detects a failure of the incorporated dedicated processor, failure information is collected and the information regarding the failure collected is written into the first extended configuration register. When the second diagnostic means of the virtual machine host detects a failure of the dedicated processor incorporated into the virtual machine guest from the flag indicating the state of the virtual machine guest, information regarding the failure is read out from the second extended configuration register. The program according to any one of Appendices 21 to 25.

[0075] (Appendix 27) When the second diagnostic means of the virtual machine host detects the occurrence of a failure, information suspected of being a failure, which is information regarding the failure, is registered and output to a predetermined file in the virtual machine host. The program according to any one of Appendices 21 to 26.

[0076] (Appendix 28) When the first diagnostic means of the virtual machine guest registers and outputs information suspected of being a failure, which is information regarding the failure of the incorporated dedicated processor, to a predetermined file in the virtual machine guest. The program according to any one of Appendices 21 to 27.

Explanation of Signs

[0077] 1 Computer system 10 CPU 20,21 PCIeSW 30~33 VE(Vector Engine)

Claims

1. a first extended configuration register; a first diagnostic means for notifying a virtual machine host of a failure in a dedicated processor incorporated in a virtual machine guest via information contained in the first extended configuration register; a virtual machine guest that is a virtual machine incorporating the dedicated processor; a second extended configuration register, the second extended configuration register having the same contents as the first extended configuration register, and information written by the virtual machine guest to the first extended configuration register can be read from the virtual machine host via the second extended configuration register; a second diagnostic means for detecting a fault associated with a dedicated processor incorporated in the virtual machine guest from information included in the second extended configuration register; A virtual machine host comprising: A computer system comprising:

2. when the first diagnostic means of the virtual machine guest detects a failure via a driver for the built-in dedicated processor, the first diagnostic means sets information indicating the occurrence of a failure in the built-in dedicated processor to a flag that is included in the first extended configuration register and indicates a state of the virtual machine; the second diagnostic means of the virtual machine host detects a failure of the dedicated processor incorporated in the virtual machine guest from a flag that is included in the second extended configuration register and indicates a state of the virtual machine guest; 2. The computer system of claim 1.

3. the first diagnostic means of the virtual machine guest starts accumulating a counter included in the first extended configuration register when preparation for starting operation is completed; the second diagnostic means of the virtual machine host determines, when the accumulation of the counter included in the second extended configuration register stops or cannot be read, that a failure has occurred in the interface with the dedicated processor on the virtual machine guest side; 2. The computer system of claim 1.

4. a connection switch for connecting the dedicated processor and a general-purpose processor; the second diagnostic means of the virtual machine host determines that a link state error has occurred when the link speed and / or link width of the general-purpose processor side of the connection switch, or the link speed and / or link width of the dedicated processor side of the connection switch, does not reach an expected value; 2. The computer system of claim 1.

5. the second diagnostic means of the virtual machine host reads a unique serial number of a dedicated processor connected to a computer system from the dedicated processor connected to the computer system, and creates a table in which the management number of the dedicated processor connected to the computer system corresponds to the serial number; the first diagnostic means of the virtual machine guest reads the serial number from the embedded dedicated processor and writes the serial number to the first extended configuration register; the second diagnostic means of the virtual machine host, when detecting a failure of the dedicated processor incorporated in the virtual machine guest from a flag included in the second extended configuration register, identifies a management number of the incorporated dedicated processor by searching the table using the serial number of the incorporated dedicated processor included in the second extended configuration register, and identifies the dedicated processor in which a failure has been detected; 3. The computer system of claim 2.

6. when the first diagnostic means of the virtual machine guest detects a fault in the embedded dedicated processor, the first diagnostic means collects fault information and writes the collected fault information to the first extended configuration register; the second diagnostic means of the virtual machine host, when detecting a failure of the dedicated processor incorporated in the virtual machine guest from a flag indicating a state of the virtual machine guest, reads information regarding the failure from the second extended configuration register; 3. The computer system of claim 2.

7. the second diagnostic means of the virtual machine host, when detecting the occurrence of a failure, registers and outputs information on a suspected failure, which is information on the failure, to a predetermined file in the virtual machine host; A computer system according to any one of claims 1 to 6.

8. the first diagnostic means of the virtual machine guest registers and outputs information on a suspected fault, which is information on a fault in the built-in dedicated processor, to a predetermined file in the virtual machine guest; A computer system according to any one of claims 1 to 6.

9. notifying a virtual machine host of a failure in the dedicated processor incorporated in the virtual machine guest via information included in a first extended configuration register by a virtual machine guest, the virtual machine guest being a virtual machine incorporating a dedicated processor; a second extended configuration register having the same contents as the first extended configuration register, the second extended configuration register being read by the virtual machine host, which is a host machine, from information contained in the second extended configuration register, the second extended configuration register allowing information written by the virtual machine guest to the first extended configuration register to be read from the virtual machine host via the second extended configuration register, detecting a failure associated with a dedicated processor incorporated in the virtual machine guest from information contained in the second extended configuration register; How to handle failures.

10. notifying a virtual machine host of a failure in the dedicated processor incorporated in the virtual machine guest via information included in a first extended configuration register by a virtual machine guest, the virtual machine guest being a virtual machine incorporating a dedicated processor; a second extended configuration register having the same contents as the first extended configuration register, the second extended configuration register being read by the virtual machine host, which is a host machine, from information contained in the second extended configuration register, the second extended configuration register allowing information written by the virtual machine guest to the first extended configuration register to be read from the virtual machine host via the second extended configuration register, detecting a failure associated with a dedicated processor incorporated in the virtual machine guest from information contained in the second extended configuration register; A fault handling program that causes a computer to execute the following:

Citation Information

Patent Citations

  • Virtual computer system

    JP2008269194A

  • Virtual machine system and virtual machine program

    JP2017045371A

  • Method and system for capturing a frame buffer of a virtual machine in a GPU pass-through environment

    US20170004808A1

  • Server maintenance control device, system, control method, and program

    WO2022009438A1