Shared memory control program, shared memory control method, and information processing apparatus

The shared memory control program addresses unnecessary notifications by delaying failure alerts until a predetermined time, enhancing system reliability and efficiency by isolating affected virtual machines.

JP2025099190APending Publication Date: 2025-07-03FUJITSU LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023215642
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing shared memory systems notify all virtual machines of a failure, even when only a specific section is affected, leading to unnecessary notifications and processing, which can disrupt unaffected virtual machines.

Method used

A shared memory control program that identifies virtual machines connected to specific sections and delays failure notifications until a predetermined time, ensuring only affected machines receive the notification.

Benefits of technology

Reduces unnecessary notifications and processing in unaffected virtual machines, improving system reliability and efficiency by localized failure management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099190000001_ABST
    Figure 2025099190000001_ABST
Patent Text Reader

Abstract

To prevent unnecessary notifications from being provided to a virtual machine that can continue accessing a shared memory.SOLUTION: Connection management information 11a indicates an identifier of a virtual machine that connects to each of sections in a shared memory device 20. Time management information 11b holds the time at which the next notification regarding the shared memory 20 can be issued to the virtual machine, for each of the sections. When receiving a notification about a failure of the shared memory device 20 in response to an access from a virtual machine 13 to the shared memory device 20, a processing unit 12 acquires an identifier of the virtual machine 13. The processing unit 12 specifies, based on the identifier of the virtual machine 13 and connection management information 11a, a virtual machine 14 connected to a section 21a that the virtual machine 13 accesses. The processing unit 12 notifies the virtual machines 13, 14 about the failure of the shared memory device 20 after the time corresponding to the section 21a is reached, the time being included in the time management information 11b.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a shared memory control program, a shared memory control method, and an information processing apparatus.

Background Art

[0002] In an information processing system, virtualization technology for operating a plurality of virtual computers (virtual machines) on a physical computer (physical machine) is used. For a virtual machine, the processing power of a CPU (Central Processing Unit) and the storage area of a RAM (Random Access Memory) provided in the physical machine are allocated as resources for calculation.

[0003] In addition, in an information processing system, a shared memory device that provides a storage area that can be shared by a plurality of machines, which are physical machines or virtual machines, may be used. The shared memory device may also be called a system storage device or an inter-system shared memory. The storage area of the shared memory device can be accessed relatively quickly from each machine and can be used, for example, as the main storage area of each machine. For example, by sharing a part of the storage area of the shared memory device between two machines, a cluster can be configured in which one machine is an operation system and the other machine is a standby system, and the availability of the operation of an application executed on the machine can be improved.

[0004] Here, there is a proposal for a virtual environment operation support system that records the history of the usage relationship between a physical device shared by a plurality of physical machines and each virtual machine, and identifies the usage status of the physical device of each virtual machine at a plurality of dates and times based on the history.

[0005] In addition, there is a proposal for a computer system including a physical machine that executes configuration control based on configuration control rules when a physical machine is included in the influence range of a failure, and transmits a notification of the configuration control to the corresponding virtual machine when a virtual machine is included in the influence range of the failure.

Prior Art Documents

Patent Document

[0006]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0007] The storage area of a shared memory device can be logically divided into a plurality of sub-areas. The sub-areas are called sections. For example, a virtual machine on a physical machine can use a certain section of the shared memory device, and another virtual machine on the physical machine can use another section of the shared memory device. Also, by multiplexing such as doubling or quadrupling the shared memory device, the availability against failures of the shared memory device may be improved.

[0008] Here, the notification of a failure by the shared memory device may be performed only at the shared memory device unit. For example, even when a local failure has occurred in a specific section of the shared memory device and access to another section is normal, the shared memory device may only notify the failure of the shared memory device unit. In this case, the physical machine can detect a failure only at the shared memory device unit from the notification.

[0009] For example, in response to the detection of the failure, the physical machine can notify all virtual machines accessing the shared memory device that is the source of the failure notification of the failure of the shared memory device for the purpose of switching control of the connection destination shared memory device or the like. However, in this case, unnecessary notifications are made to virtual machines that can continue to access the normal section in the shared memory device.

[0010] On one side, the present invention aims to suppress unnecessary notifications to virtual machines capable of continuing access to a shared memory device.

Means for Solving the Problems

[0011] In one aspect, a shared memory control program is provided. The shared memory control program causes a computer to execute the following processing. The computer is a first shared memory device having a storage area, and when a failure of the first shared memory device is notified of an access by a first virtual machine among a plurality of virtual machines to the first shared memory device that enables each of a plurality of sections obtained by logically dividing the storage area to be shared by the plurality of virtual machines, the computer acquires an identifier of the first virtual machine. The computer specifies a second virtual machine connected to a first section to which the first virtual machine is connected among the plurality of sections based on the identifier of the first virtual machine and connection management information indicating identifiers of virtual machines connected to each of the plurality of sections. The computer notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device after a time corresponding to the first section reaches a time management information that holds, for each of the plurality of sections, a time when a notification regarding the first shared memory device can be executed next to the virtual machine.

[0012] Also, in one aspect, a shared memory control method executed by a computer is provided. Also, in one aspect, an information processing apparatus having a storage unit and a processing unit is provided.

Advantages of the Invention

[0013] On one side, it is possible to suppress unnecessary notifications to virtual machines capable of continuing access to a shared memory device.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0015] Hereinafter, this embodiment will be described with reference to the drawings. [First Embodiment] The first embodiment will be described.

[0016] FIG. 1 is a diagram for explaining an information processing apparatus according to the first embodiment. The information processing apparatus 10 is connected to the shared memory apparatus 20. The information processing apparatus 10 executes a plurality of virtual machines. The shared memory apparatus 20 has a plurality of RAMs and provides storage areas 21 of the plurality of RAMs to the information processing apparatus 10. The RAM included in the shared memory apparatus 20 may be a non-volatile semiconductor memory, that is, NVRAM (Non-Volatile RAM). The storage area 21 is logically divided into a plurality of sub-areas. The sub-areas are called sections. The sections are assigned to each virtual machine operating in the information processing apparatus 10. One section is assigned to one virtual machine. One section may be assigned to two or more virtual machines. In this case, the one section will be shared by the two or more virtual machines. For example, each virtual machine operating in the information processing apparatus 10 can use the section assigned to itself as a main storage area.

[0017] The information processing apparatus 10 has a storage unit 11 and a processing unit 12. The storage unit 11 may be a volatile semiconductor memory such as a RAM, or a non-volatile storage such as an HDD (Hard Disk Drive) or a flash memory. The processing unit 12 is, for example, a processor such as a CPU, a GPU (Graphics Processing Unit), or a DSP (Digital Signal Processor). However, the processing unit 12 may include an application-specific electronic circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The processor executes a program stored in a memory such as a RAM (which may also be the storage unit 11). A collection of a plurality of processors may be referred to as a "multiprocessor" or simply a "processor". The storage unit 11 and the processing unit 12 belong to the hardware layer in the information processing apparatus 10.

[0018] In addition, the information processing apparatus 10 has virtual machines 13, 14, 15, and 16. The virtual machines 13, 14, 15, and 16 belong to the VM (Virtual Machine) layer in the information processing apparatus 10. To the virtual machines 13, 14, 15, and 16, the storage area of the storage unit 11 and the computing power of the processing unit 12 provided in the information processing apparatus 10 are allocated as computing resources. For example, the processing unit 12 may function as a hypervisor that allocates computing resources to each virtual machine. The information processing apparatus 10 may have computing resources such as a RAM and a CPU other than the storage unit 11 and the processing unit 12. The information processing apparatus 10 may allocate computing resources from a RAM and a CPU other than the storage unit 11 and the processing unit 12 to the virtual machines 13, 14, 15, and 16.

[0019] In addition, the processing unit 12 pre-allocates sections of the shared memory device 20 to the virtual machines 13, 14, 15, and 16. The processing unit 12 stores connection management information 11a in the storage unit 11 in advance. The connection management information 11a indicates the identifiers of the virtual machines connected to each of the plurality of sections. The connection between a section and a virtual machine is a logical connection established based on the physical connection between the information processing apparatus 10 and the shared memory device 20.

[0020] Here, the identifier of the virtual machine 13 is "VM1". The identifier of the virtual machine 14 is "VM2". The identifier of the virtual machine 15 is "VM3". The identifier of the virtual machine 16 is "VM4". For example, the storage area 21 is divided into sections 21a and 21b. The identifier of the section 21a is "#1". The identifier of the section 21b is "#2".

[0021] In one example, the connection management information 11a includes a record indicating the following correspondence relationship between the identifier of the virtual machine (VM-ID (Identifier)) and the section. The first record indicates the correspondence relationship between the VM-ID "VM1" and the section "#1". The second record indicates the correspondence relationship between the VM-ID "VM2" and the section "#1". The third record indicates the correspondence relationship between the VM-ID "VM3" and the section "#2". The fourth record indicates the correspondence relationship between the VM-ID "VM4" and the section "#2". The first record and the second record of the connection management information 11a indicate that the section 21a is assigned to the virtual machines 13 and 14. That is, the virtual machines 13 and 14 are logically connected to the section 21a and can access the section 21a. The third record and the fourth record of the connection management information 11a indicate that the section 21b is assigned to the virtual machines 15 and 16. That is, the virtual machines 15 and 16 are logically connected to the section 21b and can access the section 21b.

[0022] In addition, the storage unit 11 stores time management information 11b. The time management information 11b is information that holds the next executable time for the virtual machine for a notification regarding the shared memory device 20 for each of a plurality of sections of the shared memory device 20. Examples of the notification regarding the shared memory device 20 include a notification of the hardware inspection result of the shared memory device 20 transmitted from the shared memory device 20 at an arbitrary timing, and a notification of a failure of the shared memory device 20 described later.

[0023] For example, the time management information 11b includes records indicating the following correspondence between sections and times. The first record indicates the correspondence between section "#1" and time "T1". The second record indicates the correspondence between section "#2" and time "T2". The first record of the time management information 11b indicates that, for the virtual machine accessing section 21a, the next notification regarding the shared memory device 20 can be executed after time T1. The second record of the time management information 11b indicates that, for the virtual machine accessing section 21b, the next notification regarding the shared memory device 20 can be executed after time T2.

[0024] Note that the processing unit 12 updates the time held in the time management information 11b each time a notification regarding the shared memory device 20 is sent to the virtual machine accessing the corresponding section. For example, when the processing unit 12 sends a notification regarding the shared memory device 20 to the virtual machine accessing the corresponding section at time Tx, the time associated with the corresponding section in the time management information 11b is updated to the time (Tx + α) obtained by adding a predetermined time α to time Tx. α is predetermined.

[0025] Here, when an access from the virtual machine to the shared memory device 20 occurs, the shared memory device 20 may notify the information processing device 10 of a failure of the shared memory device 20. In this case, the information processing device 10 executes the following processing. In the following description, as an example, it is assumed that when the virtual machine 13 accesses the shared memory device 20, a failure is notified from the shared memory device 20.

[0026] First, the virtual machine 13 accesses section 21a (step S1). The shared memory device 20 detects a failure in section 21a for the access and notifies the information processing device 10 of the failure of the shared memory device 20 (step S2). Here, the failure notification sent from the shared memory device 20 to the information processing device 10 may include, for example, the identification information of the shared memory device 20, but does not include the identifier of the corresponding section (here, section 21a).

[0027] When a failure of the shared memory device 20 is notified from the shared memory device 20 to the access by the virtual machine 13 to the shared memory device 20, the processing unit 12 acquires the identifier of the virtual machine 13. For example, the notification of the failure is received via arithmetic resources such as a virtual CPU used by the virtual machine 13. Therefore, the processing unit 12 can identify the virtual machine 13 corresponding to the arithmetic resource that has received the notification of the failure, and acquire the identifier of the virtual machine 13 that has performed the access that has caused the notification of the failure. Based on the identifier of the virtual machine 13 and the connection management information 11a, the processing unit 12 identifies the virtual machine 14 connected to the section 21a to which the virtual machine 13 is connected among the plurality of sections (step S3). Further, the processing unit 12 acquires the time T1 corresponding to the section 21a based on the time management information 11b.

[0028] Then, after the time T1 is reached, the processing unit 12 notifies the virtual machines 13 and 14 of the failure of the shared memory device 20 (step S4). Note that when the processing unit 12 performs the notification at the time T1, for example, the time corresponding to the section 21a in the time management information 11b is updated to the time (T1 + α).

[0029] In response to the notification of the failure of the shared memory device 20, the virtual machines 13 and 14 can perform predetermined processing regarding the shared memory device 20. For example, the virtual machines 13 and 14 record the notification of the failure of the shared memory device 20 in a log. Further, the virtual machines 13 and 14 can execute a process of disconnecting a logical connection to the shared memory device 20. For example, another shared memory device (not shown) may be connected to the information processing apparatus 10, and the shared memory devices may be multiplexed. In this case, the data in the section 21a is also synchronized with the section of the other shared memory corresponding to the section 21a. Therefore, by disconnecting the shared memory device 20, the virtual machines 13 and 14 can access the section of the other shared memory device corresponding to the section 21a and continue processing such as an application executed using the section 21a. At this time, the virtual machines 15 and 16 can continue to access the section 21b.

[0030] According to the information processing apparatus 10 in this way, when a failure of the shared memory apparatus 20 is notified for an access by the virtual machine 13 among a plurality of virtual machines to the shared memory apparatus 20, an identifier of the virtual machine 13 is acquired. Based on the identifier of the virtual machine 13 and connection management information 11a indicating identifiers of virtual machines connected to each of a plurality of sections, among the plurality of sections, a virtual machine 14 connected to a section 21a to which the virtual machine 13 is connected is specified. After a time T1 corresponding to the section 21a included in the time management information 11b is reached, the failure of the shared memory apparatus 20 is notified to the virtual machine 13 and the virtual machine 14.

[0031] Thereby, the information processing apparatus 10 can appropriately notify the failure of the shared memory apparatus 20. For example, the processing unit 12 notifies the failure to the virtual machines 13 and 14 affected by a local failure of the shared memory apparatus 20, and does not notify the failure to the virtual machines 15 and 16 not affected by the failure. For this reason, unnecessary failure notifications to the virtual machines 15 and 16 not affected by the failure are suppressed. Also, unnecessary processes such as recording a log related to the failure in the virtual machines 15 and 16 not affected by the failure or disconnecting the shared memory apparatus 20 for switching the connection to another shared memory apparatus are suppressed. In this way, the information processing apparatus 10 can appropriately notify the failure of the shared memory apparatus 20 so as to suppress unnecessary processes by the virtual machines 15 and 16 not affected by a local failure of the shared memory apparatus 20.

[0032] Further, the processing unit 12 can realize controlled control as a whole system by notifying the failure to the virtual machines 13 and 14 affected by the failure after the time T1 corresponding to the section 21a included in the time management information 11b is reached. Specifically, it is as follows.

[0033] As a method of the comparative example, every time a failure notification is received from the shared memory device 20, it is conceivable to notify the virtual machines 13 and 14 of the failure. However, if this is done, every time the virtual machines 13 and 14 access the shared memory device 20, a failure notification will be sent to both of the virtual machines 13 and 14.

[0034] In the method of the comparative example, for example, first, an access to the shared memory device 20 by the virtual machine 13 can cause the first failure notification to be sent to the virtual machines 13 and 14. Next, after a relatively short time has elapsed since the notification, an access to the shared memory device 20 by the virtual machine 14 can cause the second failure notification to be sent to the virtual machines 13 and 14. At this time, assuming that the virtual machine 13 responds to the first failure notification by, for example, disconnecting the shared memory device 20 or recording the disconnection in the log of the disconnection. Then, the virtual machine 13 may perform unnecessary processing such as detecting the failure of the shared memory device 20 and recording the failure of the shared memory device 20 in the log, even though the shared memory device 20 has been disconnected after the second failure notification. Thus, in the method of the comparative example, there is a possibility that inappropriate processing may be performed as a system.

[0035] Therefore, the information processing apparatus 10 can control the time at which each virtual machine accessing the section is notified for each section, so that a failure event of the same cause related to the section can be collectively notified to each corresponding virtual machine at once. In this way, the information processing apparatus 10 can appropriately notify the failure of the shared memory device 20 so as to reduce the possibility of performing inappropriate processing as described above by the virtual machines 13 and 14, for example.

[0036] [Second Embodiment] Next, the second embodiment will be described. FIG. 2 is a diagram showing an example of the information processing system according to the second embodiment.

[0037] The information processing system according to the second embodiment includes servers 100, 200, 300 and SSU (System Storage Unit) 400, 500. The servers 100, 200, 300 are each connected to the SSU 400, 500. For the connection between the servers 100, 200, 300 and the SSU 400, 500, for example, XAUI (10 Gigabit Attachment Unit Interface) or DXAUI (Double XAUI or Double rate XAUI) is used.

[0038] The servers 100, 200, 300 are server computers that execute applications. The servers 100, 200, 300 may be referred to as physical machines. In the example of the second embodiment, the server 100 executes a plurality of virtual machines. An application is executed by each of the plurality of virtual machines. The server 100 is an example of the information processing apparatus 10 according to the first embodiment. In the following description, the virtual machine is abbreviated as "VM".

[0039] The SSU 400, 500 are storage devices having a shared memory shared by the servers 100, 200, 300. In the example of the second embodiment, the SSU is duplicated by two SSU 400, 500, but multiplexing by four or more SSU may be performed. Each of the SSU 400, 500 is an example of the shared memory device 20 according to the first embodiment. "SSU" is also called a system storage device or an inter-system shared memory. The SSU 400, 500 provides a plurality of sections obtained by logically dividing the storage area of the shared memory to the servers 100, 200, 300. The sections are used as main memory areas by each of the servers 100, 200, 300.

[0040] For example, by sharing a certain section in the SSU400 between the server 200 and the VM on the server 100, a cluster can be formed by the server 200 and the said VM. Also, by sharing another section in the SSU400 between the server 300 and another VM on the server 100, a cluster can be formed by the server 300 and the said other VM.

[0041] Furthermore, the SSU400 is made redundant using the SSU500. That is, a section corresponding to each section of the SSU400 is provided by the SSU500. For example, when an update to data for a section of the SSU400 is made, the update is also reflected in the corresponding section of the SSU500 by the server performing the update. For example, the server 200 and the VM on the server 100 can also access the section of the SSU500 corresponding to the section of the SSU400 instead of the section of the SSU400.

[0042] FIG. 3 is a diagram showing a hardware example of the information processing system. The server 100 has a processor 101, a RAM 102, an HDD 103, a connection interface 104, a GPU 105, an input interface 106, a media reader 107, and a communication interface 108. These units of the server 100 are connected to each other by a bus inside the server 100. The processor 101 corresponds to the processing unit 12 of the first embodiment. The RAM 102 or the HDD 103 corresponds to the storage unit 11 of the first embodiment.

[0043] Processor 101 is an arithmetic unit that executes program instructions. Processor 101 is, for example, a CPU. Processor 101 loads at least a part of the programs and data stored in HDD 103 into RAM 102 and executes the programs. Note that Processor 101 may include a plurality of processor cores. Also, server 100 may have a plurality of processors. The processes described below may be executed in parallel using a plurality of processors or processor cores. Also, a set of a plurality of processors may be referred to as a "multiprocessor" or simply a "processor". Also, a processor may be referred to as a "processor circuitry".

[0044] RAM 102 is a volatile semiconductor memory that temporarily stores programs executed by Processor 101 and data used by Processor 101 for calculations. Note that server 100 may include other types of memory than RAM, and may include a plurality of memories.

[0045] HDD 103 is a non-volatile storage device that stores software programs such as an OS (Operating System), middleware, and application software, and data. Note that server 100 may include other types of storage devices such as flash memory and SSD (Solid State Drive), and may include a plurality of non-volatile storage devices.

[0046] Connection interface 104 is an interface for connecting to SSU 400 and 500. As described above, for example, XAUI or DXAUI is used for connection interface 104.

[0047] The GPU 105 outputs an image to the display 111 connected to the server 100 according to the instructions from the processor 101. As the display 111, any type of display such as a CRT (Cathode Ray Tube) display, a liquid crystal display (LCD), a plasma display, or an organic EL (OEL: Organic Electro-Luminescence) display can be used.

[0048] The input interface 106 acquires an input signal from the input device 112 connected to the server 100 and outputs it to the processor 101. As the input device 112, a pointing device such as a mouse, a touch panel, a touch pad, or a trackball, a keyboard, a remote controller, a button switch, etc. can be used. Also, a plurality of types of input devices may be connected to the server 100.

[0049] The media reader 107 is a reading device that reads programs and data recorded on the recording medium 113. As the recording medium 113, for example, a magnetic disk, an optical disk, a magneto-optical disk (MO), a semiconductor memory, etc. can be used. The magnetic disk includes a flexible disk (FD) and an HDD. The optical disk includes a CD (Compact Disc) and a DVD (Digital Versatile Disc).

[0050] The media reader 107 copies, for example, programs and data read from the recording medium 113 to other recording media such as the RAM 102 and the HDD 103. The read program is executed, for example, by the processor 101. Note that the recording medium 113 may be a portable recording medium and may be used for distributing programs and data. Also, the recording medium 113 and the HDD 103 may be referred to as computer-readable recording media.

[0051] The communication interface 108 is connected to the network 114 and communicates with other information processing devices via the network 114. The communication interface 108 may be a wired communication interface connected to a wired communication device such as a switch or a router, or may be a wireless communication interface connected to a wireless communication device such as a base station or an access point.

[0052] The SSU 400 has a processor 401, SSU memories 402, 402a, …, and a connection interface 403. The processor 401 is an arithmetic unit that controls the SSU 400. The processor 401 can be realized by, for example, a CPU, a DSP, an ASIC, or an FPGA.

[0053] The SSU memories 402, 402a, … are non-volatile semiconductor memories (NVRAMs). The storage areas in the SSU memories 402, 402a, … are logically divided into a plurality of sub-areas. The sub-areas are called sections. The sections can be used as the main storage areas of the servers 100, 200, 300. The sections can be shared by the servers 100, 200, 300.

[0054] The connection interface 403 is an interface used for connection to the servers 100, 200, 300. However, in FIG. 3, the illustration of the servers 200, 300 is omitted. As described above, for example, XAUI or DXAUI is used for the connection interface 403.

[0055] Note that the servers 200, 300 are realized by the same hardware as the server 100. Also, the SSU 500 is realized by the same hardware as the SSU 400. FIG. 4 is a diagram showing a functional example of the information processing system.

[0056] Server 100 includes a storage unit 120, a VM management unit 130, standby VMs 140 and 160, and development VMs 150 and 170. For the storage unit 120, the storage areas of the RAM 102 and the HDD 103 are used. The VM management unit 130, the standby VMs 140 and 160, and the development VMs 150 and 170 are realized by the program stored in the RAM 102 being executed by the processor 101. For the standby VMs 140 and 160 and the development VMs 150 and 170, for example, the computing resources of the processor 101 and the RAM 102 are allocated by a hypervisor executed on the server 100. For example, the CPU resources allocated to a VM are called virtual CPUs. Note that the program corresponding to the VM management unit 130 may be executed on the OS of the server 100 or on a VM (not shown) on the server 100.

[0057] Server 200 includes an application 210. The program of the application 210 is stored in the RAM of the server 200 and is executed by the processor of the application 210.

[0058] Server 300 includes an application 310. The program of the application 310 is stored in the RAM of the server 300 and is executed by the processor of the application 310.

[0059] Here, the SSU 400 includes sections 410 and 420. The SSU 500 includes sections 510 and 520. The sections are identified by section numbers. The section number is an example of an identifier of a section. The section number of the section 410 is "#1". The section number of the section 420 is "#2".

[0060] The section 510 corresponds to the section 410. The data stored in each of the sections 410 and 510 is synchronized. The section number of the section 510 is "#1". Section 520 corresponds to Section 420. The data stored in each of Sections 420 and 520 is synchronized. The section number of Section 520 is "#2".

[0061] The storage unit 120 stores data used for the processing of the VM management unit 130. Specifically, the storage unit 120 stores connection section management information and reflectable time management information. The connection section management information is information indicating the sections assigned to each VM. The reflectable time management information is information indicating, for each section, the next executable time for notifications regarding the SSU 400 and 500 to each VM.

[0062] The VM management unit 130 manages each VM on the server 100 based on the connection section management information and the reflectable time management information stored in the storage unit 120. The VM management unit 130 controls the logical connection relationship between each of the VMs and the SSU 400 and 500. When the VM management unit 130 receives a notification such as a failure from the SSU 400 and 500, it selects the VM that is the transmission target of the notification based on the connection section management information. The VM management unit 130 transmits the notification to the selected VM based on the reflectable time management information.

[0063] The standby VM 140 is a standby VM for the operation server 200. That is, the server 200 and the standby VM 140 form a cluster. The standby VM 140 executes the same application as the application 210, although not shown in the figure. For example, the standby VM 140 is hot standby.

[0064] Application 210 and standby VM 140 share sections 410 and 510 and access sections 410 and 510. For example, data writing by application 210 and standby VM 140 is performed for both sections 410 and 510. On the other hand, data reading by application 210 and standby VM 140 is performed for section 410, which is the pre-specified main access destination, when connections to both SSU 400 and 500 are valid. Data reading by application 210 and standby VM 140 is performed for the section belonging to the SSU with a valid connection destination among sections 410 and 510 when a connection to only one of SSU 400 and 500 is valid.

[0065] Application 210 and standby VM 140 exchange data via section 410. Application 210 periodically updates predetermined data in section 410. When standby VM 140 detects that the periodic update of the predetermined data has stopped, it determines that application 210 or server 200 has gone down due to a failure, and takes over the processing of application 210.

[0066] Development VM 150 accesses sections 410 and 510. In this case, together with application 210 and standby VM 140, development VM 150 also shares sections 410 and 510. Development VM 150 may provide a software development environment independent of application 210. The method of writing and reading data to / from sections 410 and 510 by development VM 150 is the same as the method of writing and reading data by application 210 and standby VM 140 described above.

[0067] Standby VM 160 is a standby VM for operation server 300. That is, server 300 and standby VM 160 form a cluster. Although not shown in the figure, standby VM 160 runs the same application as application 310. For example, standby VM 160 is hot standby.

[0068] The application 310 and the standby VM 160 share the sections 420 and 520 and access the sections 420 and 520. The method of writing and reading data to and from the sections 420 and 520 by the application 310 and the standby VM 160 is the same as the method of writing and reading data by the above-mentioned application 210 and the standby VM 140. Also, the method of switching from the operation on the application 310 on the server 300 to the operation on the application on the standby VM 160 is the same as the method of switching from the operation on the application 210 to the operation on the application on the standby VM 140.

[0069] The development VM 170 accesses the section 520. In this case, together with the application 310 and the standby VM 160, the development VM 170 will also share the sections 420 and 520. The development VM 170 may provide a software development environment independent of the application 310. The method of writing and reading data to and from the sections 420 and 520 by the development VM 170 is the same as the method of writing and reading data by the above-mentioned application 210 and the standby VM 140.

[0070] Here, each VM on the server 100 is identified by a VM identifier. The VM identifier of the standby VM 140 is "VM1". The VM identifier of the development VM 150 is "VM2". The VM identifier of the standby VM 160 is "VM3". The VM identifier of the development VM 170 is "VM4".

[0071] FIG. 5 is a diagram showing an example of connection section management information. The connection section management information 121 is stored in the storage unit 120. For example, the records of the connection section management information 121 are generated when the VM management unit 130 assigns sections to each VM and are registered in the connection section management information 121. The connection section management information 121 includes items for VM identifiers and section numbers. In the item for VM identifiers, VM identifiers are registered. The VM identifier is identification information for identifying a VM. In the item for section numbers, the section numbers of the sections assigned to the VM identified by the VM identifier are registered.

[0072] For example, the connection section management information 121 has a record with a VM identifier of "VM1" and a section number of "#1". This record indicates that sections 410 and 510 are assigned to the standby VM 140.

[0073] Also, the connection section management information 121 has a record with a VM identifier of "VM2" and a section number of "#1". This record indicates that sections 410 and 510 are assigned to the development VM 150.

[0074] Also, the connection section management information 121 has a record with a VM identifier of "VM3" and a section number of "#2". This record indicates that sections 420 and 520 are assigned to the standby VM 160.

[0075] Furthermore, the connection section management information 121 has a record with a VM identifier of "VM4" and a section number of "#2". This record indicates that sections 420 and 520 are assigned to the development VM 170.

[0076] Here, the connection section management information 121 is an example of the connection management information 11a in the first embodiment. FIG. 6 is a diagram showing an example of reflectable time management information.

[0077] The reflectable time management information 122 is stored in the storage unit 120. The reflectable time management information 122 includes items of a section number, SSU0, and SSU1. A section number is registered in the item of the section number. In the item of SSU0, the reflectable time corresponding to the section of SSU400 corresponding to the relevant section number is registered. In the item of SSU1, the reflectable time corresponding to the section of SSU500 corresponding to the relevant section number is registered. Note that "SSU0" is the identification information of SSU400. "SSU1" is the identification information of SSU500. Also, the reflectable time is the time when a notification regarding SSU400 or SSU500 can be executed next for each VM accessing the relevant section.

[0078] For example, the reflectable time management information 122 has a record of a section number "#1", SSU0 "time T11", and SSU1 "time T12". This record indicates that the time when a notification regarding SSU400 can be executed next for each VM (waiting VM 140 and development VM 150) accessing section 410 is after "time T11". Also, this record indicates that the time when a notification regarding SSU500 can be executed next for each VM (waiting VM 140 and development VM 150) accessing section 510 is after "time T12".

[0079] Also, the reflectable time management information 122 has a record of a section number "#2", SSU0 "time T21", and SSU1 "time T22". This record indicates that the time when a notification regarding SSU400 can be executed next for each VM (waiting VM 160 and development VM 170) accessing section 420 is after "time T21". Also, this record indicates that the time when a notification regarding SSU500 can be executed next for each VM (waiting VM 160 and development VM 170) accessing section 520 is after "time T22".

[0080] The time included in the reflectable time management information 122 is updated by the VM management unit 130 each time a notification regarding the SSU400 or SSU500 is sent to the VM. The VM management unit 130 updates the time included in the reflectable time management information 122 for each section. The reflectable time management information 122 is an example of the time management information 11b in the first embodiment.

[0081] The server 100 executes the following processing procedure. FIG. 7 is a flowchart showing a control example of an SSU failure notification to a VM. (S10) The VM management unit 130 determines whether it has received a failure notification from the SSU for an access to the SSU of a certain VM. If a failure notification is received from the SSU, the process proceeds to step S11. If a failure notification is not received from the SSU, the control of the SSU failure notification ends.

[0082] (S11) The VM management unit 130 acquires the VM identifier of the VM that accessed the SSU. For example, a failure notification from the SSU is received via a computing resource such as a virtual CPU of the VM that accessed the SSU. The VM management unit 130 identifies the VM corresponding to the computing resource used for receiving the notification as the VM that accessed the SSU, and acquires the VM identifier of the VM.

[0083] (S12) The VM management unit 130 identifies the failure section from the acquired VM identifier. Here, the failure section is the section that became the access destination in the access from the VM to the SSU that caused the failure notification from the SSU. It can also be said that the failure section is the section where an access failure has occurred in the SSU.

[0084] (S13) The VM management unit 130 acquires the reflectable time for the VM connected to the identified failure section based on the reflectable time management information 122. (S14) The VM management unit 130 determines whether the reflectable time obtained in step S13 has elapsed. If the reflectable time has elapsed, the process proceeds to step S15. If the reflectable time has not elapsed, the VM management unit 130 waits until the current time elapses the reflectable time, and when the current time reaches the reflectable time, it proceeds to step S15. Note that in FIG. 7, the waiting in the case of step S14 No is represented for convenience as proceeding to "end" of the procedure in FIG. 7.

[0085] (S15) The VM management unit 130 identifies the VMs using the faulty section based on the connection section management information 121. In step S15, the VM management unit 130 may identify a plurality of VMs using the faulty section.

[0086] (S16) The VM management unit 130 reflects the failure of the corresponding SSU on the VMs identified in step S15. That is, the VM management unit 130 notifies the VMs identified in step S15 of the failure of the corresponding SSU. In step S16, the failure of the corresponding SSU may be notified to the plurality of VMs identified in step S15.

[0087] Then, the VM management unit 130 updates the reflectable time corresponding to the faulty section of the faulty SSU in the reflectable time management information 122. For example, the VM management unit 130 may set, as the new reflectable time, the time obtained by adding a predetermined time α to the time set in the reflectable time management information 122 in the reflectable time management information 122. Alternatively, the VM management unit 130 may set, as the new reflectable time, the time obtained by adding a predetermined time α to the current time in the reflectable time management information 122. The predetermined time α is determined in advance. The predetermined time α is, for example, 400 milliseconds.

[0088] (S17) The VMs that have received the notification in step S16 disconnect the faulty SSU, that is, disconnect the faulty SSU. Thereby, the VM disconnects the logical connection with the faulty SSU and continues the processing of the VM using the section of the non-faulty SSU. Then, the control of the SSU failure notification ends.

[0089] Note that step S15 may be executed before step S13 or before step S14. For example, the VM management unit 130 can hold the VM identifier specified this time as the VM using the failed section in the storage unit 120 as the notification target of the failure of the SSU.

[0090] Also, in step S16, the VM management unit 130 may check whether each VM specified in step S15 is operating and connected to the failed SSU. Then, the VM management unit 130 may notify the failure of the SSU for the VMs that are operating and connected to the failed SSU among the VMs specified in step S15.

[0091] Also, in step S16, the VM management unit 130 may receive, from the SSU, notifications of a plurality of failures corresponding to accesses to the corresponding failed section for the plurality of VMs specified in step S15 from the time of the previous notification regarding the corresponding SSU to the current time. Even when the VM management unit 130 has received notifications of a plurality of failures from the SSU in this way, in step S16, it notifies the failure of the SSU only once for each of the plurality of VMs specified in step S15. In this way, the VM management unit 130 collectively transmits the notifications of a plurality of failures received from the SSU within a predetermined time to each VM at once. As a result, the server 100 can suppress the occurrence of inappropriate processing such as the failure of the failed SSU being notified again after the disconnection of the failed SSU in a certain VM and the failure of the SSU that should have been disconnected being recorded in the log.

[0092] Here, if an inappropriate event is recorded in the log, it may become difficult to analyze the log during troubleshooting or the like. By reducing the possibility of an inappropriate event being recorded in the log in the VM, the server 100 supports easy analysis of the log during troubleshooting or the like, and as a result, can contribute to improving the reliability of the entire system.

[0093] By the way, some SSU devices do not have a function to notify a failure in response to access from a VM. When such an SSU is used, since the VM management unit 130 can only receive a failure notification from the SSU at the timing of a machine check or the like that is performed irregularly by the SSU, it is not possible to identify a failure section as in steps S11 and S12. Therefore, in this case, the VM management unit 130 can save the memory area of the RAM 102 used as the storage unit 120 by holding the reflectable time in units of SSU without holding the reflectable time for each section.

[0094] Therefore, the VM management unit 130 checks whether the SSU has a function to notify a failure in response to access from a VM. If the SSU does not have such a function, as normal control, the VM management unit 130 holds the reflectable time in the storage unit 120 in units of SSU. On the other hand, when the SSU has a function to notify a failure in response to access from a VM, the VM management unit 130 expands the memory area used as the storage unit 120 so as to hold the reflectable time in units of section. Specifically, the VM management unit 130 executes the following procedure.

[0095] FIG. 8 is a flowchart showing an example of management of the reflectable time. The procedure of FIG. 8 can be executed, for example, immediately after the server 100 is started up. Also, the procedure of FIG. 8 can be executed at step S16 when changing the reflectable time of the failure section.

[0096] (S20) The VM management unit 130 determines whether the failure notification function at the time of VM access exists in the connected SSU. If the failure notification function at the time of VM access exists, the process proceeds to step S21. If the failure notification function at the time of VM access does not exist, the process ends. When the failure notification function at the time of VM access does not exist, as described above, as normal control, the reflectable time in units of SSU may be held in the storage unit 120.

[0097] Here, the VM management unit 130 can determine whether the destination SSU has a failure notification function during VM access from, for example, the hardware configuration information preset in the HDD 103 when the system is started (when the IPL (Initial Program Loader) is executed). Specifically, the hardware configuration information has a hardware technology improvement flag. When the hardware technology improvement flag for the destination SSU is ON, the SSU has a failure notification function during VM access. When the hardware technology improvement flag for the destination SSU is OFF, the SSU does not have a failure notification function during VM access. The SSUs 400 and 500 illustrated in the second embodiment are SSUs having a failure notification during VM access.

[0098] (S21) The VM management unit 130 expands the memory area for the reflectable time for each section. Thereby, for example, the reflectable time management information 122 is generated. At the stage when the reflectable time management information 122 is created, for example, the reflectable time of each section may be unset, or the time after a predetermined time from the current time may be set. When the reflectable time of each section is unset, the notification from the SSU is immediately notified to the VM corresponding to the section, and the time obtained by adding the time α to the time when the first notification is made is set as the next reflectable time for the section. Note that when the VM management unit 130 executes the procedure of FIG. 8 once at the time of starting the server 100 or the like and then executes the procedure of FIG. 8 for updating the reflectable time in step S16, step S21 may be skipped and the process may proceed to step S22.

[0099] (S22) The VM management unit 130 determines whether to start recording the next reflectable time. If the recording of the next reflectable time is started, the process proceeds to step S23. If the recording of the next reflectable time is not started, the process ends. For example, when the VM management unit 130 updates the reflectable time in step S16, the process proceeds to step S23, and in other cases (at the time of initial setting), the process ends.

[0100] (S23) The VM management unit 130 changes the recording of the next reflectable time in section units. That is, the VM management unit 130 determines the recording of the next reflectable time in section units instead of in SSU units.

[0101] (S24) The VM management unit 130 stores the next reflectable time in the memory area extended in step S21. For example, for the section of the SSU to be the change target of the reflectable time, the VM management unit 130 may set, in the reflectable time management information 122, as the new reflectable time, the time obtained by adding a predetermined time α to the time set in the reflectable time management information 122. Alternatively, the VM management unit 130 may set, in the reflectable time management information 122, as the new reflectable time, the time obtained by adding a predetermined time α to the current time.

[0102] In this way, when the connected SSU of the server 100 has a failure notification function at the time of VM access, the VM management unit 130 expands the memory area so as to hold the reflectable time for each section. Thereby, the VM management unit 130 suppresses the securing of an extra memory area for holding the reflectable time when the connected SSU does not have a failure notification function at the time of VM access, and can save the memory usage amount.

[0103] Here, the VM management unit 130 can identify the source VM accessing the SSU that has notified a failure, for example, as follows. FIG. 9 is a diagram showing a specific example of the source VM for the failure notification of the SSU.

[0104] In the example of FIG. 9, the standby VM 140 is illustrated, but the VM management unit 130 can similarly identify other VMs. Also, in FIG. 9, the development VMs 150 and 170 are not shown. First, when the VM management unit 130 starts the standby VM 140, it sets the VM identifier "VM1" of the standby VM 140 in a predetermined register for managing the started VMs (step ST1). The section 410 of the SSU 400 is allocated to the standby VM 140.

[0105] The standby VM 140 starts accessing section 410 (step ST2). The SSU 400 notifies the standby VM 140 of an SSU failure in response to the standby VM 140's access to section 410. For example, the VM management unit 130 recognizes the SSU failure based on a notification from the hardware corresponding to the computing resources of the standby VM 140 (step ST3).

[0106] Then, the VM management unit 130 identifies the standby VM 140 that has recognized the SSU failure as the access source VM to the corresponding SSU (step ST4). Specifically, the VM management unit 130 determines whether the identified standby VM 140 is operating. For example, the VM management unit 130 can determine whether the standby VM 140 is operating by checking the PSW (Program Status Word) corresponding to the standby VM 140. When the standby VM 140 is operating, the VM management unit 130 stores the VM identifier "VM1" of the register in step ST1 in the storage unit 120 as the access source VM to the corresponding SSU.

[0107] Note that the VM management unit 130 can identify the access source VM in the same manner as above even when the SSU 500 fails. Next, an example of expanding the memory area for holding the reflectable time for each section illustrated in FIG. 8 will be described.

[0108] FIG. 10 is a diagram showing an example of expanding the area for holding the reflectable time. The memory areas 123 and 124 are part of the storage unit 120 and are areas for holding the reflectable time management information 122. The information held in the memory areas 123 and 124 corresponds to the reflectable time management information 122.

[0109] Memory area 123 is associated with the identification information "SSU0" of SSU400. Memory area 124 is associated with the identification information "SSU1" of SSU500. In the example of FIG. 10, the case where the storage area of each SSU is logically divided into N is illustrated. N is an integer of 1 or more. In response to a change in the number of N, the memory area for the reflectable time management information 122 is dynamically changed.

[0110] Memory area 123 has areas A1, B1, C1, …, N1. Area A1 holds the time (for example, time T11) corresponding to the basic section of SSU400. Here, the basic section is the section set in the initial settings and is the section with section number "#1". Section 410 corresponds to the basic section.

[0111] Area B1 holds the time (for example, time T12) corresponding to the first extended section of SSU400. Here, the extended section is a section added in addition to the basic section. It can also be said that the extended section has a larger section number than the basic section. Also, the Nth extended section indicates the Nth added extended section. The section number of the first extended section is "#2". Section 420 corresponds to the first extended section.

[0112] Areas C1, …, N1 each hold the time (for example, times T13, …, T1N) corresponding to the second, …, Nth extended sections of SSU400. The section numbers of the second, …, Nth extended sections are "#3", …, "#N".

[0113] Memory area 124 has areas A2, B2, C2, …, N2. Area A2 holds the time (for example, time T21) corresponding to the basic section of SSU500. Section 510 corresponds to the basic section.

[0114] Region B2 holds the time (e.g., time T22) corresponding to the first extended section of SSU500. Section 520 corresponds to the first extended section. Regions C2, …, N2 each hold the time (e.g., times T23, …, T2N) corresponding to the second, …, Nth extended sections of SSU500. Data is synchronized and duplicated in two sections with the same section number in SSU400 and 500.

[0115] In this way, the VM management unit 130 can expand the memory area such as RAM102 according to the number of sections provided in SSU400 and 500, and hold the reflectable time management information 122. When the number of sections increases, the VM management unit 130 expands the memory area by the size corresponding to the increased number of sections. When the number of sections decreases, the VM management unit 130 releases the memory area by the size corresponding to the decreased number of sections. Thereby, the VM management unit 130 can save the memory area such as RAM102 and use the memory area efficiently.

[0116] Next, an example of expanding the memory area for holding the reflectable time will be described. FIG. 11 is a diagram showing an example of an expansion pattern of the area for holding the reflectable time. The area expansion pattern 125 exemplifies the pattern of expansion control of the memory area by the VM management unit 130. For this expansion control, a hardware technology improvement flag, conditions related to the hardware configuration and the sections of VM utilization are used.

[0117] The condition regarding the hardware technology improvement flag is whether the flag is “ON” or “OFF”. As described above, the hardware technology improvement flag is a flag indicating whether SSU400 and 500 have a function of notifying the failure of the SSU for access to the section. The hardware technology improvement flag “ON” indicates that the failure notification function is available. The hardware technology improvement flag “OFF” indicates that the failure notification function is not available.

[0118] The condition regarding the hardware configuration is whether or not a plurality of VMs managed by the VM management unit 130 can be connected to different sections. Such a configuration is called an inter-section cluster configuration. When it corresponds to the inter-section cluster configuration, the hardware configuration is "valid". When it does not correspond to the inter-section cluster configuration, the hardware configuration is "invalid". Note that when the hardware technology improvement flag is "OFF", the inter-section cluster configuration cannot be used, so the condition regarding the hardware configuration is not used. In the case where the hardware technology improvement flag is "OFF", the places where the condition regarding the hardware configuration is not used are indicated by the hyphen symbol "-" in FIG. 11.

[0119] The condition regarding the section for VM utilization is which section is assigned to the corresponding VM. The condition regarding the section for VM utilization is used only when the hardware technology improvement flag is "ON" and the hardware configuration is "valid". In the places where the condition regarding the section for VM utilization is not used, it is indicated by the hyphen symbol "-" in FIG. 11.

[0120] For example, when the hardware technology improvement flag is "ON", the hardware configuration is "valid", and the section for VM utilization is "Expansion N", as the dedicated data area for the reflectable time, area N1 is used for SSU400 (SSU0), and area N2 is used for SSU500 (SSU1). Here, Expansion N indicates the Nth expansion section. "Localizable" in the remarks column indicates that a VM that switches the connection destination SSU can be localized. Similarly, for example, when the section for VM utilization is "Expansion 1", areas B1 and B2 are used for SSU400 and 500, respectively, to hold the reflectable time.

[0121] Also, when the hardware technology improvement flag is "ON", the hardware configuration is "valid", and the section for VM utilization is "Basic", as the dedicated data area for the reflectable time, area A1 is used for SSU400 (SSU0), and area A2 is used for SSU500 (SSU1). In this case, the remarks column is also "localizable".

[0122] When the hardware technology improvement flag is "ON" and the hardware configuration is "valid", if the update of the reflectable time is possible, it is performed in units of sections of the SSU based on the logic indicated by the area expansion pattern 125. For example, when the VM management unit 130 receives a failure notification from the SSU 400 for the access of the standby VM 160 to the section 420 and issues the failure notification to the standby VM 160 and the development VM 170, the current time + α is set in the area B1.

[0123] Also, when the hardware technology improvement flag is "ON" and the hardware configuration is "invalid", area A1 is used for the SSU 400 (SSU0) and area A2 is used for the SSU 500 (SSU1) as the data area dedicated to the reflectable time. In this case, the remarks column is "Localization is not relevant". That is, since the extended section is not used, there is no need to localize the VM that switches the destination SSU.

[0124] Also, when the hardware technology improvement flag is "OFF", area A1 is used for the SSU 400 (SSU0) and area A2 is used for the SSU 500 (SSU1) as the data area dedicated to the reflectable time. In this case, the remarks column is "Not applicable (localization not possible)". That is, since the SSU does not have a function to notify the VM management unit 130 of a failure for access to the SSU, localization of the influence of the SSU failure by the VM management unit 130 cannot be applied.

[0125] For example, the VM management unit 130 can control the memory area for holding the reflectable time according to the area expansion pattern 125, so as to correspond to both the SSU having a failure notification function for VM access and the SSU not having a failure notification function for VM access.

[0126] Next, an example of switching the destination SSU by the server 100 will be described. FIG. 12 is a diagram showing an example of switching the destination SSU. For example, the standby VM 140 accesses section 410. Assume that a failure has occurred in section 410 at this time. In this case, when the VM management unit 130 receives a failure notification of the SSU 400 via the computing resources of the standby VM 140, it identifies that the standby VM 140 has recognized the failure of the SSU 400.

[0127] Then, the VM management unit 130 identifies section 410 accessed by the standby VM 140 based on the connection section management information 121. Also, the VM management unit 130 identifies the development VM 150 other than the standby VM 140 that accesses section 410 based on the connection section management information 121.

[0128] When the VM management unit 130 reaches the reflectable time corresponding to section 410 based on the reflectable time management information 122, it notifies the standby VM 140 and the development VM 150 of the failure of the SSU 400. The notification of the failure includes, for example, the identification information "SSU0" of the SSU 400. When the standby VM 140 and the development VM 150 receive the failure notification of the SSU 400 from the VM management unit 130, they disconnect the SSU 400. Then, the standby VM 140 and the development VM 150 continue their respective processes by accessing section 510 of the SSU 500. Here, in the figure, an X mark is attached to the path to be disconnected between each server and the SSU 400.

[0129] At this time, the VM management unit 130 does not notify the standby VM 160 and the development VM 170 of the failure of the SSU 400. Therefore, the standby VM 160 and the development VM 170 continue to access section 420 without disconnecting the SSU 400.

[0130] In this way, the VM management unit 130 can localize the VMs that switch the destination SSU. Note that SSU400 also notifies the server 200 of a failure in SSU400 in response to access to section 410 by application 210 of server 200. Therefore, the server 200 also disconnects SSU400. As a result, the main access destination of application 210 is also switched from section 410 to section 510. On the other hand, SSU400 does not notify the server 300 of a failure in SSU400 in response to access to section 420 by application 310 of server 300. Therefore, the server 300 does not disconnect SSU400. Accordingly, application 310 continues to access section 420.

[0131] FIG. 13 is a diagram showing a comparative example of switching the connection destination SSU. The information processing system of the comparative example includes servers 200, 300, 600 and SSU400, 500. Servers 200, 300, 600 are connected to SSU400, 500. The hardware of server 600 is the same as that of server 100.

[0132] Server 600 includes a VM management unit 630, standby VMs 640, 660 and development VMs 650, 670. The VM management unit 630 manages the standby VMs 640, 660 and the development VMs 650, 670. The VM management unit 630 can detect a failure in the corresponding SSU from SSU400, 500 irregularly or at the time of VM access. However, the VM management unit 630 does not have a function of specifying the accessing VM or the VM to be notified of the failure notification in response to the SSU failure notification. The VM management unit 630 issues a failure notification to all VMs connected to the corresponding SSU in response to the SSU failure notification.

[0133] The data in the sections with the same section number of SSU400, 500 are synchronized and duplicated. In the comparative example, server 200 and standby VM 640 form a cluster and access sections 410, 510. Development VM 650 accesses sections 410, 510.

[0134] In addition, the server 300 and the standby VMs 660 form a cluster and access sections 420 and 520. The development VMs 670 access sections 420 and 520. For example, the SSU 400 detects the occurrence of a failure in section 410. Then, the SSU 400 notifies the servers 200, 300, and 600 of the failure of the SSU 400. When receiving the notification of the failure of the SSU 400, the VM management unit 630 notifies all the VMs connected to the SSU 400, that is, the standby VMs 640 and 660 and the development VMs 650 and 670, of the failure of the SSU 400. The standby VM 640 and the development VM 650 disconnect from the SSU 400 and switch the main access destination to section 510 of the SSU 500. The standby VM 660 and the development VM 670 disconnect from the SSU 400 and switch the main access destination to section 520 of the SSU 500. In the figure, an X mark is attached to the path to be disconnected between each server and the SSU 400.

[0135] Note that when the servers 200 and 300 also receive the notification of the failure of the SSU 400 in the same way, they disconnect from the SSU 400 and use the SSU 500 as the main access destination. In this way, in the comparative example, when a failure occurs in section 410 of the SSU 400, the connection between the SSU 400 and all the machines accessing the SSU 400 on the server 600 is disconnected. In the comparative example, the VMs (the standby VMs 660 and the development VMs 670) and the server 300 that use the non-failed section 420 are also disconnected from the SSU 400, and the localization of the impact has not been achieved.

[0136] Here, since the SSU is multiplexed, when an SSU failure occurs, the failed SSU is disconnected and the SSU is degraded, that is, by changing from SSU duplication to SSU singleplex, data preservation of the SSU, that is, continuous operation of the system with only normal SSUs, is attempted.

[0137] However, as in the comparative example, if there is no mechanism to isolate the SSU close to the section where a failure has occurred, when an SSU failure occurs, the SSU is disconnected in all machines including the VMs connected to the SSU, and the localization of the impact cannot be achieved. Also, for the machines from which the SSU has been disconnected, restoration work of the SSU is required, which affects normal operations. For example, the restoration work of the SSU includes work to restore from SSU single redundancy to SSU dual redundancy to improve the reliability of SSU data.

[0138] Therefore, as exemplified in the second embodiment, the server 100 provides a mechanism to identify the failed section of the SSU and degrade the SSU limited to the affected VMs. Thereby, when an SSU failure occurs, the server 100 can perform degraded operation and maintenance work of the SSU limited to the VMs using the section where the SSU failure has occurred, and can localize the impact on the failure in terms of sections. That is, the server 100 can perform a limited degraded operation of disconnecting the virtual machines using the section where the SSU failure has occurred.

[0139] Also, the reliability of system operation of the server 100 can be improved. In particular, in a shared center among enterprises (such as multiple banks, etc.) that utilize the configuration of the information processing system of the second embodiment, the impact on enterprises using sections where no SSU failure has occurred can be suppressed.

[0140] Furthermore, in the comparative example of FIG. 13, when double-point failures occur in separate sections of SSU400 and 500 (for example, failures in sections 410 and 520, or failures in sections 420 and 510), both SSU400 and 500 are disconnected by standby VMs 640 and 660. When standby VMs 640 and 660 go down, if an abnormality occurs in operation system servers 200 and 300, the operation that takes over the processing in the standby system becomes impossible, and the reliability of the entire system cannot be ensured. Alternatively, the cluster configurations of server 200 via SSU400 and standby VM640 and of server 300 via SSU400 and standby VM660 cannot be maintained. In contrast, server 100 can prevent the standby VMs 140 and 160 from going down (or can maintain the cluster configuration), even when double-point failures occur in separate sections of, for example, SSU400 and 500, and can ensure the reliability of the entire system.

[0141] As described above, server 100 executes, for example, the following processes. The processor 101 executes a predetermined program stored in the RAM 102 to perform the following processes corresponding to the VM management unit 130.

[0142] When the processor 101 is notified of a failure of the first shared memory device with respect to access to the first shared memory device by the first virtual machine among a plurality of virtual machines, the processor 101 acquires the identifier of the first virtual machine. Here, the first shared memory device has a storage area, and each of a plurality of sections obtained by logically dividing the storage area can be shared by a plurality of virtual machines. The processor 101 identifies a second virtual machine connected to a first section to which the first virtual machine is connected, among the plurality of sections, based on the identifier of the first virtual machine and connection management information indicating the identifiers of the virtual machines connected to each of the plurality of sections. After the time corresponding to the first section included in the time management information is reached, the processor 101 notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device. Here, the connection management information is information that holds, for each of the plurality of sections, the time when the next execution is possible for the virtual machine with respect to a notification regarding the first shared memory device.

[0143] Thereby, the server 100 can appropriately notify of the failure of the shared memory device. For example, the processor 101 notifies the first and second virtual machines affected by the local failure of the first shared memory device of the failure, and does not notify other virtual machines not affected by the failure. For this reason, unnecessary notification of the failure to other virtual machines not affected by the failure is suppressed. Also, unnecessary processing such as recording a log regarding the failure in other virtual machines not affected by the failure or disconnecting the first shared memory device for switching the connection to another shared memory device is suppressed. In this way, the server 100 can appropriately notify of the failure of the first shared memory device so as to suppress unnecessary processing by other virtual machines not affected by the local failure of the first shared memory device.

[0144] Also, the processor 101 can notify each virtual machine accessing the section of the same failure event with the same cause in the section at once by controlling the time to notify each virtual machine accessing the section for each section. Therefore, the server 100 can suppress the same failure of the first shared memory device from being notified to the first and second virtual machines multiple times, and can reduce the possibility that inappropriate processing is performed in the first and second virtual machines due to the multiple notifications. In this way, the server 100 can appropriately notify the failure of the first shared memory device so as to reduce the possibility that inappropriate processing is performed in the first and second virtual machines.

[0145] Note that the SSU 400 is an example of the first shared memory device. The connection section management information 121 is an example of the connection management information. The reflectable time management information 122 is an example of the time management information.

[0146] When the processor 101 notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device, the processor 101 does not notify the third virtual machine among the plurality of virtual machines, which is connected to the second section among the plurality of sections in the first shared memory device, of the failure of the first shared memory device. That is, the processor 101 notifies only the first virtual machine and the second virtual machine among the plurality of virtual machines of the failure of the first shared memory device. Thereby, the server 100 can avoid making unnecessary notifications to the virtual machines that are not affected by the local failure of the first shared memory device, and can control so that inappropriate processing is not performed in the virtual machines.

[0147] For example, the processor 101 notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device, so that the first virtual machine and the second virtual machine execute disconnection of the connection with the first shared memory device. Thereby, the server 100 can prevent the first virtual machine and the second virtual machine connected to the failed section from accessing the failed section.

[0148] Also, the third section in the second shared memory device may hold a copy of the data held in the first section. After the detachment, the processor 101 causes each of the first virtual machine and the second virtual machine to access the copy of the data held in the third section of the second shared memory device, and continues the processing of each of the first virtual machine and the second virtual machine. Thereby, the server 100 can suppress the influence of a local failure of the first shared memory device on virtual machines other than the first virtual machine and the second virtual machine, and improve the availability of system operation. The SSU 500 is an example of the second shared memory device.

[0149] Also, when the processor 101 notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device, it updates the time corresponding to the first section included in the time management information. Thereby, the server 100 can appropriately manage the time when the next notification regarding the first shared memory device can be executed for the first virtual machine and the second virtual machine.

[0150] Furthermore, the processor 101 expands a memory area in a storage device such as the RAM 102 that stores the time management information according to the number of a plurality of sections in the first shared memory device. Thereby, the server 100 can efficiently use the memory area of the storage device such as the RAM 102.

[0151] Note that the information processing of the first embodiment can be realized by causing the processing unit 12 to execute a program. Also, the information processing of the second embodiment can be realized by causing the processor 101 to execute a program. The program can be recorded on a computer-readable recording medium 113.

[0152] For example, by distributing the recording medium 113 on which the program is recorded, the program can be circulated. Also, the program may be stored in another computer and distributed via a network. The computer may store (install) the program recorded on the recording medium 113 or the program received from another computer in a storage device such as the RAM 102 or the HDD 103, and read and execute the program from the storage device.

Explanation of Signs

[0153] 10 Information processing apparatus 11 Storage unit 11a Connection management information 11b Time management information 12 Processing unit 13, 14, 15, 16 Virtual machine 20 Shared memory device 21 Storage area 21a, 21b Section S1, S2, S3, S4 Step

Claims

1. A computer, to a first shared memory device having a storage area, wherein when a failure of the first shared memory device is notified in response to access by a first virtual machine among the plurality of virtual machines to the first shared memory device that enables each of a plurality of sections obtained by logically dividing the storage area to be shared by the plurality of virtual machines, an identifier of the first virtual machine is acquired, based on the identifier of the first virtual machine and connection management information indicating identifiers of virtual machines connected to each of the plurality of sections, a second virtual machine connected to a first section to which the first virtual machine is connected is specified among the plurality of sections, and after a time corresponding to the first section included in time management information that holds a next executable time for each of the plurality of sections for notifying the virtual machine of a notification regarding the first shared memory device has been reached, the failure of the first shared memory device is notified to the first virtual machine and the second virtual machine, a shared memory control program for causing execution of a process.

2. When notifying the first virtual machine and the second virtual machine of the failure of the first shared memory device, a third virtual machine among the plurality of virtual machines that is connected to a second section in the first shared memory device is not notified of the failure of the first shared memory device, The shared memory control program according to claim 1.

3. By notifying the first virtual machine and the second virtual machine of the failure of the first shared memory device, causing the first virtual machine and the second virtual machine to disconnect the connection with the first shared memory device, The shared memory control program according to claim 1.

4. A third section in a second shared memory device holds a copy of data held in the first section, after the disconnection, causing each of the first virtual machine and the second virtual machine to access the copy of the data held in the third section of the second shared memory device and continue the processing of each of the first virtual machine and the second virtual machine, The shared memory control program according to claim 3.

5. When notifying the first virtual machine and the second virtual machine of the failure of the first shared memory device, updating the time corresponding to the first section included in the time management information, The shared memory control program according to claim 1.

6. Expanding a memory area for storing the time management information according to the number of the plurality of sections The shared memory control program according to claim 1.

7. When a computer receives a notification of a failure of the first shared memory device having a storage area, where each of a plurality of sections obtained by logically dividing the storage area can be shared by a plurality of virtual machines, for access by a first virtual machine among the plurality of virtual machines, the computer acquires an identifier of the first virtual machine, identifies a second virtual machine connected to a first section to which the first virtual machine is connected, among the plurality of sections, based on the identifier of the first virtual machine and connection management information indicating identifiers of virtual machines connected to each of the plurality of sections, and notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device after a time corresponding to the first section, which is included in time management information holding, for each of the plurality of sections, a time when the virtual machine can execute next, has been reached. A shared memory control method.

8. An information processing apparatus including: a storage unit that stores connection management information indicating identifiers of virtual machines connected to each of a plurality of sections in a first shared memory device having a storage area, where each of the plurality of sections obtained by logically dividing the storage area can be shared by a plurality of virtual machines, and time management information holding, for each of the plurality of sections, a time when the virtual machine can execute next, and a notification regarding the first shared memory device; and a processing unit that, when a failure of the first shared memory device is notified for access to the first shared memory device by a first virtual machine among the plurality of virtual machines, acquires an identifier of the first virtual machine, identifies a second virtual machine connected to a first section to which the first virtual machine is connected, among the plurality of sections, based on the identifier of the first virtual machine and the connection management information, and notifies the first virtual machine and the second virtual machine of the failure of the first shared memory device after a time corresponding to the first section, which is included in the time management information, has been reached. An information processing apparatus having the above.

Citation Information

Patent Citations

  • Virtual environment operation support system

    JP2013206368A

  • Computer, computer control method, and computer control program

    WO2015072004A1