Primary Machine and Fault-Tolerant System

The primary machine and fault-tolerant system differentiate between hardware and software failures to avoid unnecessary control switching, ensuring continued operation by handling software failures locally and switching only for hardware failures.

JP7700765B2Active Publication Date: 2025-07-01YOKOGAWA ELECTRIC CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2022158954
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-07-01
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

Existing fault-tolerant systems fail to differentiate between hardware and software failures, leading to unnecessary control switching when software failures cannot be resolved by switching, thus affecting overall system operation.

Method used

A primary machine and fault-tolerant system that determine whether to switch control based on the type of failure, using a failure selection unit to distinguish between hardware and software failures, allowing the system to handle software failures without switching control and only switch for hardware failures.

Benefits of technology

This approach reduces unnecessary control switching, ensuring continued operation by addressing software failures locally and switching to a secondary machine only for hardware failures, thereby maintaining system integrity and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700765000001
    Figure 0007700765000001
  • Figure 0007700765000002
    Figure 0007700765000002
  • Figure 0007700765000003
    Figure 0007700765000003
Patent Text Reader

Abstract

To provide a primary machine and a fault tolerant system in which control switching is determined by a fault type.SOLUTION: A primary machine 100 has: a primary VM 120 which includes: a synchronization information generation part 124 that generates and outputs synchronization information on the basis of an instruction and an execution result of the instruction; and a fault selection part 128 that determines whether the fault information occurred during the execution of the instruction is information related to a hardware fault or a software fault. The primary VM 120 changes its behavior according to the determination result of the fault information type.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a primary machine and a fault-tolerant system.

Background Art

[0002] Conventionally, a fault-tolerant system that can operate at low load is known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a fault-tolerant system, when a failure that occurs in the primary machine cannot be resolved even if control is switched to the secondary machine, it is necessary to resolve the failure with the primary machine without switching control. It is required to switch control only when necessary without automatically switching control when a failure occurs.

[0005] The present disclosure has been made in view of the above points, and an object thereof is to provide a primary machine and a fault-tolerant system that determine whether to switch control depending on the type of failure.

Means for Solving the Problems

[0006] The (1) primary machine according to some embodiments includes a synchronization information generation unit that generates and outputs synchronization information based on instructions and execution results of the instructions, and a failure selection unit that determines the type of failure information generated during the execution of the instructions. The primary virtual machine changes its operation based on the determination result of the type of the failure information. By doing so, for example, unnecessary switching is avoided when an error that cannot be eliminated even by switching control occurs. As a result, the switching of control is determined according to the type of failure.

[0007] (2) In the primary machine of the above (1), the failure selection unit may determine whether the failure information is related to a hardware failure or a software failure. When the failure information is information related to the software failure, the primary virtual machine may execute at least one of determination of execution of error handling by an application operating on the primary virtual machine, execution of the error handling of the application, or instruction of the error handling for the application. By doing so, the failure is eliminated without switching control. As a result, unnecessary switching of control is avoided.

[0008] (3) In the primary machine of the above (2), the failure selection unit may determine the accuracy of the determination as to whether the failure information is related to a hardware failure or a software failure. The primary virtual machine may further change its operation based on the accuracy of the determination. By doing so, the accuracy of the determination of the switching of control is improved. As a result, unnecessary switching of control is avoided.

[0009] (4) In the primary machine of the above (2) or (3), the failure selection unit may acquire the return value of the system call or the error content as the failure information. The failure selection unit may determine whether the failure information relates to a hardware failure or a software failure based on a list that specifies which of the hardware failure or the software failure the return value of the system call or the error content corresponds to. By doing so, the type of the failure can be easily determined. As a result, the control switching is determined according to the type of the failure.

[0010] (5) A fault-tolerant system according to some embodiments may include any one of the primary machines from the above (1) to (4) and a secondary virtual machine that executes the instruction based on the synchronization information. By doing so, even when a failure occurs in the primary machine, the operation can be continued by the secondary machine. As a result, the processing is continued as a whole for the fault-tolerant system.

[0011] (6) In the fault-tolerant system of the above (5), the failure selection unit may determine whether the failure information relates to a hardware failure or a software failure. When it is determined that the failure information relates to the hardware failure, the primary virtual machine may switch the control to the secondary virtual machine. When it is determined that the failure information relates to the software failure, the primary virtual machine may not need to switch the control to the secondary virtual machine. By doing so, for example, unnecessary switching is avoided when an error that cannot be resolved even by switching the control occurs. As a result, the control switching is determined according to the type of the failure.

[0012] (7) In the fault-tolerant system of (6) above, the fault selection unit may determine the accuracy of the determination as to whether the fault information relates to a hardware fault or a software fault. Based on the accuracy of the determination, the primary virtual machine may further determine whether to switch control to the secondary virtual machine. By doing so, the accuracy of the determination for control switching is improved. As a result, unnecessary control switching is avoided.

Advantages of the Invention

[0013] According to the present disclosure, there are provided a primary machine and a fault-tolerant system in which control switching is determined according to the type of fault.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0015] (Comparative Example) As shown in FIG. 1, a fault-tolerant system 9 according to a comparative example includes a primary machine 800 and a secondary machine 900. The primary machine 800 and the secondary machine 900 are communicably connected via a network 300. Both the primary machine 800 and the secondary machine 900 execute the same processing. When a fault occurs in the primary machine 800, the secondary machine 900 takes over the processing. By doing so, the processing is continued for the entire fault-tolerant system 9.

[0016] The primary machine 800 includes hardware 840 and operates a primary operating system (OS) 830 or a hypervisor on the hardware 840. The primary machine 800 operates a primary virtual machine (VM) 820 on the primary OS 830 or the hypervisor. The primary VM 820 includes a synchronization information generation unit 824 and a failure detection unit 826. The primary machine 800 operates an application 810 on the primary VM 820.

[0017] The secondary machine 900 includes hardware 940 and operates a secondary OS 930 or a hypervisor on the hardware 940. The secondary machine 900 operates a secondary VM 920 on the secondary OS 930 or the hypervisor. The secondary VM 920 includes a synchronization execution unit 924 and a failure detection unit 926. The secondary machine 900 operates an application 910 on the secondary VM 920.

[0018] The primary machine 800 and the secondary machine 900 operate the application 810 and the application 910 so that the processing of the application 810 and the application 910 is the same processing. By doing so, when a failure occurs in the primary machine 800, the secondary machine 900 can take over the processing. When the applications 810 and 910 are not distinguishable, they are simply referred to as applications.

[0019] The hardware 840 includes a central processing unit (CPU), a memory, and a network interface controller (NIC). The hardware 940 includes a CPU, a memory, and a NIC. When the hardware 840 and 940 are not distinguishable, they are simply referred to as hardware.

[0020] The CPU may be composed of one or more processors. By executing a predetermined program, the processor may implement various functions. The processor may obtain a program from the memory or from the network 300.

[0021] The memory may be composed of, for example, a semiconductor memory or the like, or may be composed of a storage medium such as a magnetic disk. The memory may function as the working memory of the CPU. The memory may be included in the CPU.

[0022] The NIC may be configured to include a communication interface such as a LAN (Local Area Network).

[0023] The primary VM 820 includes a synchronization information generation unit 824 and a failure detection unit 826. The synchronization information generation unit 824 generates synchronization information for transmitting the execution status of the application 810 in the primary VM 820 to the secondary VM 920, and transmits it to the secondary VM 920. The failure detection unit 826 detects a hardware failure (HW failure) occurring in the hardware 840 or a software failure (SW failure) occurring in the primary OS 830.

[0024] The secondary VM 920 includes a synchronization execution unit 924 and a failure detection unit 926. The synchronization execution unit 924 receives synchronization information from the primary VM 820 and executes the application 910 in synchronization with the execution status of the application 810 in the primary VM 820. The failure detection unit 926 detects a hardware failure (HW failure) occurring in the hardware 940 or a software failure (SW failure) occurring in the primary OS 930.

[0025] The primary VM 820 operates on the primary OS 830. The secondary VM 920 operates on the secondary OS 930. The functions of the primary OS 830 are realized by the hardware 840. The functions of the secondary OS 930 are realized by the hardware 940. The primary OS 830 causes the primary VM 820 to execute the processing of the application 810. The secondary OS 930 causes the secondary VM 920 to execute the processing of the application 910.

[0026] An example of the operation flow until the application executes an instruction and obtains an operation result is described. In the primary machine 800, the application 810 outputs an instruction to the primary VM 820. The primary OS 830 that operates the primary VM 820 receives the instruction from the primary VM 820. The primary OS 830 converts the received instruction into a form executable by the hardware 840 and causes the hardware 840 to execute the instruction, and obtains the operation result of the hardware 840. The primary VM 820 outputs the operation result obtained by the primary OS 830 to the application 810. Through the above operations, the application 810 can obtain an operation result corresponding to the instruction.

[0027] In the primary machine 800, the synchronization information generation unit 824 of the primary VM 820 generates synchronization information including instructions and operation results, and outputs it to the synchronization execution unit 924 of the secondary machine 900. In the secondary machine 900, based on the synchronization information acquired by the synchronization execution unit 924, the application 910 outputs instructions to the secondary VM 920. The secondary OS 930 operating the secondary VM 920 receives instructions from the secondary VM 920. The secondary OS 930 converts the received instructions into a form executable by the hardware 940, causes the hardware 940 to execute the instructions, and acquires the operation results of the hardware 940. The secondary VM 920 outputs the operation results acquired by the secondary OS 930 to the application 910. Through the above operations, in the secondary machine 900, the application 910 can execute the same instructions as the application 810 of the primary machine 800 based on the synchronization information and acquire the operation results.

[0028] In the primary machine 800, the primary VM 820 detects hardware failures occurring in the hardware 840 and software failures occurring in the primary OS 830 by the failure detection unit 826. Similarly, in the secondary machine 900, the secondary VM 920 detects hardware failures occurring in the hardware 940 and software failures occurring in the secondary OS 930 by the failure detection unit 926.

[0029] In the fault-tolerant system 9 according to the comparative example, when a failure occurs in the primary machine 800, regardless of whether the occurred failure is a hardware failure or a software failure, the control is switched to the secondary machine 900. When the occurred failure is a hardware failure, the probability that the same hardware failure occurs in the secondary machine 900 is low. Therefore, by switching the control to the secondary machine 900, the fault-tolerant system 9 can avoid the failure of the primary machine 800 and continue the overall operation.

[0030] On the other hand, when the occurred failure is a software failure, since the secondary machine 900 is executing the same processing as the primary machine 800, the same software failure is likely to occur in the secondary machine 900 as well. Therefore, even after the control is switched to the secondary machine 900, the fault-tolerant system 9 may not be able to avoid the failure. When a software failure occurs, the primary machine 800 of the fault-tolerant system 9 needs to resolve the failure in the application 810 without switching the control. That is, in the fault-tolerant system 9 and the primary machine 800, the control switching may not be determined depending on the type of failure.

[0031] Therefore, the present disclosure will describe a primary machine 100 (see FIG. 2) and a fault-tolerant system 1 (see FIG. 2) in which control switching is determined depending on the type of failure.

[0032] (Embodiment) As shown in FIG. 2, a fault-tolerant system 1 according to an embodiment of the present disclosure includes a primary machine 100 and a secondary machine 200. The primary machine 100 and the secondary machine 200 are communicably connected via a network 300. The primary machine 100 and the secondary machine 200 both execute the same processing. When a failure occurs in the primary machine 100, the secondary machine 200 takes over the processing. By doing so, the processing is continued as a whole for the fault-tolerant system 1.

[0033] <Configuration Example> The primary machine 100 includes hardware 140 and operates a primary OS 130 on the hardware 140. The primary machine 100 operates a primary VM 120 (primary virtual machine) on the primary OS 130. The primary machine 100 operates an application 110 on the primary VM 120. The application 110 has a SW failure detection unit 116.

[0034] The secondary machine 200 includes hardware 240 and operates the secondary OS 230 on the hardware 240. The secondary machine 200 operates the secondary VM 220 (secondary virtual machine) on the secondary OS 230. The secondary machine 200 operates the application 210 on the secondary VM 220. The application 210 has a SW failure detection unit 216.

[0035] The primary machine 100 and the secondary machine 200 operate the application 110 and the application 210 so that the processing of the application 110 and the application 210 is the same processing. By doing so, when a failure occurs in the primary machine 100, the secondary machine 200 can take over the processing. When the application 110 and 210 are not distinguished, they are simply referred to as an application.

[0036] The hardware 140 includes a CPU, a memory, and a NIC. The hardware 240 includes a CPU, a memory, and a NIC. When the hardware 140 and 240 are not distinguished, they are simply referred to as hardware.

[0037] The CPU may be composed of one or more processors. The processor may realize various functions by executing a predetermined program. The processor may obtain a program from the memory or may obtain a program from the network 300.

[0038] The memory may be composed of, for example, a semiconductor memory or the like, or may be composed of a storage medium such as a magnetic disk. The memory may function as a working memory of the CPU. The memory may be included in the CPU.

[0039] The NIC may be configured to include a communication interface such as a LAN.

[0040] The primary VM 120 includes a synchronization information generation unit 124, a HW failure detection unit 126, and a failure selection unit 128. The synchronization information generation unit 124 generates synchronization information for transmitting the execution status of the application 110 in the primary VM 120 to the secondary VM 220 and transmits it to the secondary VM 220. The HW failure detection unit 126 detects a hardware failure (HW failure) that has occurred in the primary machine 100. The failure selection unit 128 outputs information regarding the hardware failure that has occurred in the hardware 140 to the HW failure detection unit 126 and outputs information regarding the software failure (SW failure) that has occurred in the primary OS 130 to the application 110.

[0041] The secondary VM 220 includes a synchronization execution unit 224, a HW failure detection unit 226, and a failure selection unit 228. The synchronization execution unit 224 receives synchronization information from the primary VM 120 and executes the application 210 in synchronization with the execution status of the application 110 in the primary VM 120.

[0042] <Operation Example of Application> The primary VM 120 operates on the primary OS 130. The secondary VM 220 operates on the secondary OS 230. The primary VM 120 and the secondary VM 220 are collectively referred to as VMs. The primary OS 130 and the secondary OS 230 are collectively referred to as OSs. That is, the VM operates on the OS. The functions of the OS and the VM are realized by hardware including a CPU.

[0043] The primary OS 130 causes the primary VM 120 to execute the processing of the application 110. The secondary OS 230 causes the secondary VM 220 to execute the processing of the application 210. That is, the OS causes the VM to execute the processing of the applications 110 and 210.

[0044] An example of the operation flow until the application executes an instruction and obtains an operation result is described. In the primary machine 100, the application 110 outputs an instruction to the primary VM 120. The primary OS 130 that operates the primary VM 120 receives the instruction from the primary VM 120. The primary OS 130 converts the received instruction into a form executable by the hardware 140, causes the hardware 140 to execute the instruction, and obtains the operation result of the hardware 140. The primary VM 120 outputs the operation result obtained by the primary OS 130 to the application 110. Through the above operations, the application 110 can obtain an operation result corresponding to the instruction.

[0045] In the primary machine 100, the synchronization information generation unit 124 of the primary VM 120 generates synchronization information including the instruction and the operation result, and outputs it to the synchronization execution unit 224 of the secondary machine 200. In the secondary machine 200, based on the synchronization information obtained by the synchronization execution unit 224, the application 210 outputs an instruction to the secondary VM 220. The secondary OS 230 that operates the secondary VM 220 receives the instruction from the secondary VM 220. The secondary OS 230 converts the received instruction into a form executable by the hardware 240, causes the hardware 240 to execute the instruction, and obtains the operation result of the hardware 240. The secondary VM 220 outputs the operation result obtained by the secondary OS 230 to the application 210. Through the above operations, in the secondary machine 200, the application 210 can execute the same instruction as the application 110 of the primary machine 100 based on the synchronization information and obtain the operation result.

[0046] In programming languages such as Java (registered trademark) or.Net, the VM is also called the Runtime. The VM may be implemented as a general-purpose programming language processing system. The general-purpose programming language processing system may include, for example, mruby or Micro Python. mruby is a lightweight Ruby language processing system for embedded systems and can operate in a memory-saving environment. The Ruby processing system is mainly implemented as an interpreter. The source code is compiled into bytecode during or before program execution. The interpreter executes the bytecode one instruction at a time.

[0047] The primary VM 120 and the secondary VM 220 store each bytecode at the same instruction address. The primary VM 120 and the secondary VM 220 acquire the bytecode from the same instruction address and execute the operation corresponding to the bytecode. Executing the operation corresponding to the bytecode is also referred to as executing the bytecode. The primary VM 120 may cause the synchronization information generation unit 124 to execute the bytecode. The secondary VM 220 may cause the synchronous execution unit 224 to execute the bytecode. The primary VM 120 and the secondary VM 220 synchronize every time they execute the operation corresponding to one bytecode. After the primary VM 120 and the secondary VM 220 execute the operation corresponding to one bytecode and synchronize, they execute the operation corresponding to the next bytecode. By doing so, the primary VM 120 and the secondary VM 220 can proceed with the processing while synchronizing with each other.

[0048] When the bytecode corresponds to an operation of acquiring data input from the outside or outputting data to the outside, only the primary VM 120 actually performs data input / output with the outside. On the other hand, the secondary VM 220 does not actually perform data input / output with the outside.

[0049] When the bytecode corresponds to an operation of acquiring data input from the outside, instead of acquiring the data input from the outside, the secondary VM 220 acquires the data input from the outside from the primary VM 120 with respect to the primary VM 120. When the bytecode corresponds to an operation of outputting data to the outside, the secondary VM 220 skips the execution of the bytecode.

[0050] Each time the primary VM 120 executes one bytecode, it transmits synchronization information to the secondary VM 220. The synchronization information may include the instruction address where the bytecode is stored, or the data input from the outside with respect to the primary VM 120. The synchronization information may include information representing the number of executed instructions. The synchronization information may include information for identifying the bytecode executed by the primary VM 120.

[0051] The secondary VM 220 receives the synchronization information from the primary VM 120 and proceeds with the processing of the bytecode based on the synchronization information. The secondary VM 220 proceeds with the processing of the bytecode that matches the instruction address received from the primary VM 120 or the number of executed instructions. After the processing of one bytecode is completed, the secondary VM 220 interrupts the processing until it receives the next synchronization information from the primary VM 120.

[0052] The primary VM 120 and the secondary VM 220 can synchronously proceed with the processing of the bytecode by transmitting and receiving the synchronization information as described above.

[0053] mruby simplifies the processing to reduce the program size of the VM in order to achieve memory saving of the VM. One of the functions for simplifying the processing is to make the program processing single-threaded. Being single-threaded means that multiple instructions are not executed in parallel at the same time and the processing is not interrupted by an external interrupt. Since the processing is not interrupted by an external interrupt, a complex timing adjustment function becomes unnecessary. As a result, the processing is simplified.

[0054] Here, even if no external interrupt occurs at the VM level, an external interrupt may occur at the OS level. Therefore, an external interrupt may occur while bytecode is being executed in the VM. However, the execution result of the bytecode does not change due to an external interrupt.

[0055] The synchronization information generation unit 124 of the primary VM 120 acquires the instruction address or the number of executed instructions of the executed bytecode as the bytecode is executed in the primary VM 120. When the primary VM 120 acquires data input from the outside, the synchronization information generation unit 124 acquires the data. The synchronization information generation unit 124 generates synchronization information including the acquired instruction address or the number of executed instructions or the data input from the outside, and transmits it to the secondary VM 220. An instruction for acquiring data input from the outside is also referred to as an input instruction. The synchronization information generation unit 124 outputs the data input from the outside as synchronization information by executing the input instruction as bytecode.

[0056] In addition, after transmitting the synchronization information to the secondary VM 220, the synchronization information generation unit 124 prevents the primary VM 120 from executing the next bytecode until it receives a response notification from the secondary VM 220. In other words, when the synchronization information generation unit 124 receives a response notification from the secondary VM 220, it permits the primary VM 120 to execute the next bytecode.

[0057] The synchronization execution unit 224 of the secondary VM 220 receives the synchronization information from the synchronization information generation unit 124 of the primary VM 120. The synchronization execution unit 224 controls the execution of the bytecode in the secondary VM 220 based on the received synchronization information. For example, the synchronization execution unit 224 may cause the secondary VM 220 to execute the bytecode stored at the instruction address included in the synchronization information. The synchronization execution unit 224 may cause the secondary VM 220 to execute the bytecode so as to match the number of executed instructions included in the synchronization information.

[0058] When the secondary execution unit 224 corresponds to an operation of obtaining data input from the outside for the bytecode to be executed next in the secondary VM 220, the secondary execution unit 224 causes the secondary VM 220 to skip the execution of the bytecode. In this case, the synchronization information includes the data input from the outside. The secondary VM 220 regards the data input from the outside included in the synchronization information as the data obtained as the execution result of the skipped bytecode, and proceeds to the processing of the next bytecode. Since the synchronization information includes the data input from the outside, the secondary machine 200 does not need to communicate with the outside. By doing so, the load on the fault tolerant system 1 is reduced. As a result, a fault tolerant system 1 that can operate with a low load is realized.

[0059] <An example of a program> Here, as an mruby program, a program that obtains a character string from the outside, concatenates another character string to the obtained character string, and outputs the concatenated character string to the outside will be described as an example. Suppose the two character strings are represented as X and Y. This program may be compiled into the following four bytecodes. Codes A, B, C, and D each correspond to one instruction. Code A: The VM assigns the string constant "X" to the first register. Code B: The VM obtains a character string as data input from the outside and assigns it to the second register. In this program example, "Y" is obtained as the character string. Code C: The VM concatenates the string in the first register and the string in the second register, and assigns the concatenated string to the first register. Code D: The VM outputs the string in the first register to the outside.

[0060] A configuration in which the primary machine 100 and the secondary machine 200 execute while synchronizing the above-described bytecodes will be described.

[0061] The primary machine 100 and the secondary machine 200 are communicably connected to an external device via the network 300. The primary machine 100 acquires input data from the external device. The primary machine 100 outputs output data to the external device. The primary machine 100 transmits synchronization information to the secondary machine 200. When the primary machine 100 acquires input data from the external device, the primary machine 100 outputs synchronization information including the input data to the secondary machine 200.

[0062] If a failure occurs in the primary machine 100, the secondary machine 200 can continue to execute the bytecode. The secondary machine 200 does not communicate with the external device while the primary machine 100 is operating, but when the primary machine 100 stops due to a failure, the secondary machine 200 can communicate with the external device to input and output data.

[0063] The primary VM 120 and the secondary VM 220 execute the above-described bytecode according to the procedure shown in FIG. 3.

[0064] The primary VM 120 executes code A (step S11). As an operation corresponding to code A, the primary VM 120 substitutes the string constant "X" into the first register. After the primary VM 120 executes code A in the procedure of step S11, the primary VM 120 transmits synchronization information A to the secondary VM 220. Note that the string constant "X" is included in code A and not in synchronization information A.

[0065] When the secondary VM 220 receives the synchronization information A from the primary VM 120, the secondary VM 220 executes code A based on the synchronization information A (step S21). As an operation corresponding to code A, the secondary VM 220 substitutes the string constant "X" into the first register. After the secondary VM 220 executes code A in the procedure of step S21, the secondary VM 220 transmits a response indicating that the execution of code A has been completed to the primary VM 120.

[0066] When the primary VM 120 receives a response from the secondary VM 220, it executes code B, which is the following bytecode (step S12). As an operation corresponding to code B, the primary VM 120 obtains the character string "Y" as input data from an external device and assigns it to the second register. After executing code B in the procedure of step S12, the primary VM 120 transmits synchronization information B including the character string "Y" as input data from the external device to the secondary VM 220.

[0067] When the secondary VM 220 receives the synchronization information B from the primary VM 120, it executes code B based on the synchronization information B (step S22). As an operation corresponding to code B, instead of obtaining input data from an external device, the secondary VM 220 assigns the character string "Y" included in the synchronization information B to the second register. After executing code B in the procedure of step S22, the secondary VM 220 transmits a response indicating that the execution of code B has been completed to the primary VM 120.

[0068] When the primary VM 120 receives a response from the secondary VM 220, it executes code C, which is the following bytecode (step S13). As an operation corresponding to code C, the primary VM 120 concatenates the character string in the first register and the character string in the second register, and assigns the concatenated character string to the first register. In this case, the character string assigned to the first register is "XY". After executing code C in the procedure of step S13, the primary VM 120 transmits synchronization information C to the secondary VM 220.

[0069] When the secondary VM 220 receives the synchronization information C from the primary VM 120, it executes the code C based on the synchronization information C (step S23). As an operation corresponding to the code C, the secondary VM 220 concatenates the string in the first register and the string in the second register, and substitutes the concatenated string into the first register. In this case, also in the secondary VM 220, the string substituted into the first register is "XY". After executing the code C in the procedure of step S23, the secondary VM 220 sends a response indicating that the execution of the code C has been completed to the primary VM 120.

[0070] When the primary VM 120 receives a response from the secondary VM 220, it executes the code D which is the next byte code (step S14). As an operation corresponding to the code D, the primary VM 120 outputs the string in the first register to an external device. In this case, the string acquired by the external device is "XY". After executing the code D in the procedure of step S14, the primary VM 120 sends the synchronization information D to the secondary VM 220.

[0071] When the secondary VM 220 receives the synchronization information D from the primary VM 120, it executes the code D based on the synchronization information D (step S24). As an operation corresponding to the code D, the secondary VM 220 does not output the string in the first register to the external device and does not execute anything. That is, the secondary VM 220 skips the operation corresponding to the code D. After skipping the operation corresponding to the execution of the code D in the procedure of step S24, the secondary VM 220 sends a response indicating that the execution of the code D has been completed to the primary VM 120.

[0072] After sending a response indicating that the execution of the code D has been completed to the primary VM 120, the secondary VM 220 ends the execution of the series of byte codes. By receiving a response from the secondary VM 220, the primary VM 120 ends the execution of the series of byte codes.

[0073] As described above, the primary VM 120 and the secondary VM 220 can execute bytecode while synchronizing with each other. Even if the primary VM 120 stops due to a failure during the execution of a series of bytecodes, the secondary VM 220 can continue to execute the bytecode. The secondary machine 200 can also continue to execute the bytecode corresponding to the operation of inputting and outputting data by being communicably connected to an external device via the network 300.

[0074] In the fault-tolerant system 1, when both the primary machine 100 and the secondary machine 200 are operating normally, redundancy of processing is realized. Here, the operation of the fault-tolerant system 1 when the primary machine 100 or the secondary machine 200 stops due to a failure will be described.

[0075] <When a failure occurs in the primary machine 100> When a failure occurs in the primary machine 100, the primary VM 120 may not be able to execute control processes such as the execution of bytecode normally. Here, when the failure that occurred in the primary machine 100 is a hardware failure, the possibility that the same hardware failure occurs in the secondary machine 200 is low. Therefore, the fault-tolerant system 1 can avoid the failure of the primary machine 100 and continue the overall operation by switching the control to the secondary machine 200.

[0076] On the other hand, when the failure that occurred in the primary machine 100 is a software failure, since the secondary machine 200 is executing the same processing as the primary machine 100, the possibility that the same software failure occurs in the secondary machine 200 is high. Therefore, the fault-tolerant system 1 may not be able to avoid the failure even if the control is switched to the secondary machine 200. Therefore, the fault-tolerant system 1 executes different operations depending on whether a hardware failure occurs or a software failure occurs.

[0077] In the primary machine 100, the primary VM 120 receives information regarding a hardware failure that occurred in the hardware 140 and a software failure that occurred in the primary OS 130. The information regarding the hardware failure and the information regarding the software failure are also collectively referred to as failure information.

[0078] The failure selection unit 128 of the primary VM 120 determines the type of the received failure information. The failure selection unit 128 may determine whether the received failure information is information regarding a hardware failure or information regarding a software failure. When the failure selection unit 128 determines that the received failure information is information regarding a software failure, the failure selection unit 128 outputs the information regarding the software failure to the application 110. Accordingly, the primary VM 120 outputs to the application 110 both the operation result of executing the instruction and the information regarding the software failure. The application 110 detects the information regarding the software failure by the SW failure detection unit 116. The application 110 responds to the software failure by error handling. The primary VM 120 may determine the execution of the error handling by the application 110. The primary VM 120 may instruct the application 110 to execute the error handling. The primary VM 120 may perform at least one of determining the execution of the error handling by the application 110 or instructing the application 110 to execute the error handling.

[0079] When the failure selection unit 128 determines that the received failure information is information regarding a hardware failure, it outputs the failure information to the HW failure detection unit 126. The HW failure detection unit 126 detects the occurrence of a hardware failure. When the occurrence of a hardware failure is detected in the fault-tolerant system 1, control processing such as the execution of bytecode is switched from the primary VM 120 of the primary machine 100 to the secondary VM 220 of the secondary machine 200. The fault-tolerant system 1 enters a single-operation state in which only the secondary machine 200 operates. In the single-operation state, the secondary VM 220 that substitutes for the operation of the primary VM 120 stops the synchronization process with the primary VM 120.

[0080] When the primary VM 120 can detect a hardware failure in the primary machine 100, it attempts to send a failure notification to the secondary VM 220 while stopping or restarting the primary machine 100. The failure notification may be sent through the same communication path as the synchronization information. When the primary VM 120 can send a failure notification to the secondary VM 220, the secondary VM 220 grasps that a hardware failure has occurred in the primary machine 100 by receiving the failure notification from the primary VM 120. When the primary VM 120 cannot send a failure notification to the secondary VM 220, the secondary VM 220 may grasp that a failure has occurred in the primary machine 100 by means of monitoring the primary machine 100. When the secondary VM 220 grasps that a failure has occurred in the primary machine 100, it takes over the control processing such as the execution of bytecode from the primary VM 120 and stops the synchronization process with the primary machine 100. Further, the secondary VM 220 executes the input / output processing of data with the outside on behalf of the primary VM 120 via the secondary OS 230.

[0081] The primary machine 100 may stop without the primary VM 120 detecting a failure and without sending a failure notification to the secondary VM 220. The secondary VM 220 may recognize that the primary machine 100 has stopped by means of monitoring the primary machine 100 and may recognize that a failure has occurred in the primary machine 100. When the secondary VM 220 recognizes that a failure has occurred in the primary machine 100, it takes over control processing such as the execution of bytecode from the primary VM 120 and stops the synchronization process with the primary machine 100. As described above, the primary VM 120 changes its operation according to the determination result of the type of failure information.

[0082] <When a failure occurs in the secondary machine 200> In the secondary machine 200, the secondary VM 220 receives information regarding a hardware failure that has occurred in the hardware 240 and a software failure that has occurred in the secondary OS 230. The failure selection unit 228 of the secondary VM 220 determines whether the received failure information is information regarding a hardware failure or information regarding a software failure. When the failure selection unit 228 determines that the received failure information is information regarding a software failure, it outputs the information regarding the software failure to the application 210. Accordingly, the primary VM 220 outputs to the application 210 both the operation result of executing an instruction and the information regarding the software failure. The application 210 detects the information regarding the software failure by the SW failure detection unit 216. The application 210 responds to the software failure by error processing.

[0083] When the received failure information is determined to be information regarding a hardware failure, the failure selection unit 228 outputs the failure information to the HW failure detection unit 226. The HW failure detection unit 226 detects the occurrence of a hardware failure. When a hardware failure occurs in the secondary machine 200, the fault tolerant system 1 enters a single operation state in which only the primary machine 100 operates. In the single operation state, the primary VM 120 stops the synchronization process with the secondary VM 220.

[0084] When the secondary VM 220 can detect a hardware failure in the secondary machine 200, it attempts to send a failure notification to the primary VM 120 while stopping or restarting the secondary machine 200. The failure notification may be sent through the same communication path as the synchronization information. When the secondary VM 220 can send a failure notification to the primary VM 120, the primary VM 120 grasps that a failure has occurred in the secondary machine 200 by receiving the failure notification from the secondary VM 220. When the secondary VM 220 cannot send a failure notification to the primary VM 120, the primary VM 120 may grasp that a failure has occurred in the secondary machine 200 by means of monitoring the secondary machine 200. When the primary VM 120 grasps that a failure has occurred in the secondary machine 200, it stops the synchronization process with the secondary machine 200 in control processing such as bytecode execution. The primary VM 120 may stop the synchronization process by stopping the operation of the synchronization information generation unit 124.

[0085] The secondary machine 200 may stop without the secondary VM 220 detecting a failure and sending a failure notification to the primary VM 120. The primary VM 120 may grasp that the secondary machine 200 has stopped by means of monitoring the secondary machine 200 and grasp that a failure has occurred in the secondary machine 200. When the primary VM 120 grasps that a failure has occurred in the secondary machine 200, it stops the synchronization process with the secondary machine 200.

[0086] The means by which the primary VM 120 monitors the secondary machine 200, or the means by which the secondary VM 220 monitors the primary machine 100, is realized, for example, as follows.

[0087] For example, the primary VM 120 and the secondary VM 220 may perform periodic communication for liveness monitoring such as heartbeats with each other. When there is no response from the secondary VM 220, the primary VM 120 may determine that a failure has occurred in the secondary machine 200. When there is no response from the primary VM 120, the secondary VM 220 may determine that a failure has occurred in the primary machine 100. The secondary VM 220 may also determine that no failure has occurred in the primary machine 100 by receiving synchronization information from the primary VM 120 in the synchronization process. The primary VM 120 may also determine that no failure has occurred in the secondary machine 200 by receiving a response from the secondary VM 220 in the synchronization process.

[0088] For example, a third machine different from the primary machine 100 and the secondary machine 200 may monitor the operations of the primary machine 100 and the secondary machine 200. The third machine may notify the secondary VM 220 of a failure that occurred in the primary machine 100, or may notify the primary VM 120 of a failure that occurred in the secondary machine 200. The third machine may also perform periodic communication for liveness monitoring such as heartbeats with the primary VM 120 and the secondary VM 220.

[0089] The primary VM 120 and the secondary VM 220 may erroneously determine that a failure has occurred in the primary machine 100 and the secondary machine 200 when the communication for liveness monitoring is interrupted due to a network failure. In order to avoid false detection of failures in the primary machine 100 and the secondary machine 200 due to network failures, the communication path for liveness monitoring may be multiplexed.

[0090] <Operation example of the failure selection unit> When the failure selection unit acquires information related to a failure, it may determine whether the failure is a hardware failure or a software failure based on, for example, the return value of a system call.

[0091] When the return value of the system call is an error, the failure selection unit may determine whether the failure is a hardware failure or a software failure based on the content of the error. For example, when the content of the error represents a failure of hardware such as storage, memory, CPU, or NIC, the failure selection unit may determine that the failure is a hardware failure. For example, when the content of the error represents a communication connection failure or transmission / reception failure, a memory shortage, or the non-existence of a file to be operated on by an application, the failure selection unit may determine that the failure is a software failure. Also, the failure selection unit may determine that a synchronization communication error in which synchronization information cannot be received due to a communication timeout is a hardware failure.

[0092] Even when the failure selection unit does not detect a failure in the primary machine 100 or the secondary machine 200 alone, it may compare the execution result of an instruction in the primary machine 100 with the execution result of an instruction in the secondary machine 200 based on the synchronization information. When the execution result of the instruction in the primary machine 100 does not match the execution result of the instruction in the secondary machine 200, the failure selection unit may determine that a software failure or a hardware failure has occurred in either the primary machine 100 or the secondary machine 200.

[0093] The failure selection unit may determine the accuracy of the determination when determining whether it is a hardware failure or a software failure. For example, when the failure selection unit determines that a synchronization communication error in which synchronization information cannot be received due to a communication timeout is a hardware failure, it may determine that the accuracy is high. Also, for example, when the failure selection unit determines that a failure has occurred when the execution result of the instruction in the primary machine 100 does not match the execution result of the instruction in the secondary machine 200, it may determine that the accuracy is low.

[0094] Based on the certainty of the determination as to whether the failure is a hardware failure or a software failure, the primary VM 120 may determine whether to switch control to the secondary machine 200 to stop or restart the primary machine 100. For example, when the primary VM 120 determines with a high certainty that the failure is a hardware failure, it may determine to switch control to the secondary machine 200 to stop or restart the primary machine 100. When the primary VM 120 determines with a high certainty that the failure is a software failure, it may determine to execute error handling by the application 110. The primary VM 120 may instruct the application 110 to execute error handling. The primary VM 120 may execute at least one of the determination of the execution of error handling by the application 110 or the instruction to cause the application 110 to execute error handling. When the primary VM 120 determines with a low certainty that the failure is a hardware failure or a software failure, it may determine whether to stop or restart the primary machine 100 after attempting error handling in the application 110. By determining the operation based on the certainty of the determination, the determination accuracy of the control switching is improved.

[0095] The failure selection unit may determine whether the failure is a hardware failure or a software failure based on a list that specifies which of a hardware failure or a software failure the return value of the system call corresponds to. The failure selection unit may determine whether the failure is a hardware failure or a software failure based on a list that specifies which of a hardware failure or a software failure the error content of the system call corresponds to. In the list, the certainty of the association between the return value or the error content and the type of failure may be associated. The failure selection unit may determine the certainty of the determination based on the list. By determining the type of failure based on the list, the failure selection unit can easily determine the type of failure.

[0096] For example, when error X of system call A is associated with a software failure in the list, the failure selection unit may determine that the failure is a software failure when it obtains error X as the error content of system call A. When error Y of system call A is associated with a hardware failure in the list, the failure selection unit may determine that the failure is a hardware failure when it obtains error Y as the error content of system call A. When there is another system call B in addition to system call A, the failure selection unit may determine whether a hardware failure or a software failure corresponds to the return value or error content of system call B based on the list related to system call B.

[0097] <Example of flowchart> Here, an example of the operation of the fault-tolerant system 1 when a failure occurs in the primary machine 100 will be described. The primary VM 120 may execute a control method including, for example, the procedure example of the flowchart shown in FIG. 4. The control method may be implemented as a control program to be executed by a processor that realizes the functions of the primary VM 120. The control program may be stored in a non-transitory computer-readable medium.

[0098] The primary VM 120 determines whether it has detected the occurrence of a hardware failure (HW failure) in the primary machine 100 (step S31). If the primary VM 120 has not detected the occurrence of a hardware failure (step S31: NO), it proceeds to the procedure of step S34. If the primary VM 120 has detected the occurrence of a hardware failure (step S31: YES), the primary VM 120 sends a failure notification to the secondary VM 220 that a hardware failure has occurred in the primary machine 100 (step S32). The primary VM 120 stops or restarts the primary machine 100 (step S33). After executing the procedure of step S33, the primary VM 120 ends the execution of the procedure of the flowchart in FIG. 4.

[0099] If the primary VM 120 does not detect the occurrence of a hardware failure (step S31: NO), it determines whether it has detected the occurrence of a software failure (step S34). If the primary VM 120 does not detect the occurrence of a software failure (step S34: NO), it ends the execution of the procedure of the flowchart in FIG. 4. If the primary VM 120 detects the occurrence of a software failure (step S34: YES), it executes error handling of the application 110 according to the software failure (step S35). The primary VM 120 may determine the execution of error handling by the application 110. The primary VM 120 may instruct the application 110 to execute error handling. The primary VM 120 may execute at least one of the determination of the execution of error handling by the application 110 or the instruction to cause the application 110 to execute error handling. After executing the procedure of step S35, the primary VM 120 ends the execution of the procedure of the flowchart in FIG. 4.

[0100] In the fault-tolerant system 1, the primary VM 120 can change its operation based on the determination result of the type of failure by executing the procedure of the flowchart illustrated in FIG. 4. Also, the primary VM 120 can avoid switching control to the secondary machine 200 when an error that cannot be resolved even by switching control occurs when a failure occurs in the primary machine 100.

[0101] (Summary) As described above, the primary machine 100 according to an embodiment of the present disclosure changes its operation based on the determination result of the type of failure. Also, the fault-tolerant system 1 according to an embodiment of the present disclosure determines whether the failure is a hardware failure or a software failure, and determines whether to switch control from the primary machine 100 to the secondary machine 200. By doing so, for example, unnecessary control switching is avoided when an error that cannot be resolved even by switching control occurs.

[0102] Also, in the fault-tolerant system 1 according to an embodiment of the present disclosure, the primary VM 120 may execute error processing when the failure information is information regarding a software failure. By doing so, the failure is resolved without switching control. As a result, unnecessary control switching is avoided.

[0103] As described above, the embodiments according to the present disclosure have been described with reference to the drawings. However, the specific configuration is not limited to this embodiment, and various modifications within the scope not departing from the gist of the present disclosure are also included.

Description of Reference Numerals

[0104] 1 Fault-tolerant system 100 Primary machine (110: Application, 120: Primary VM, 124: Synchronization information generation unit, 126: HW failure detection unit, 128: Failure selection unit, 130: Primary OS, 140: Hardware) 200 Secondary machine (210: Application, 220: Secondary VM, 224: Synchronous execution unit, 226: HW failure detection unit, 228: Failure selection unit, 230: Secondary OS, 240: Hardware) 300 Network

Claims

1. A primary virtual machine comprising: a synchronization information generation unit that generates and outputs synchronization information based on a command and an execution result of the command; and a failure selection unit that determines whether failure information generated during execution of the command relates to hardware failure or software failure. The primary virtual machine changes its operation based on a determination result as to whether the failure information relates to hardware failure or software failure. When the failure information relates to software failure, the primary virtual machine executes at least one of determination of execution of error handling by an application operating on the primary virtual machine or an instruction to cause the application to execute the error handling.

2. The failure selection unit determines the accuracy of determination as to whether the failure information relates to hardware failure or software failure. The primary virtual machine according to claim 1, wherein the primary virtual machine further changes its operation based on the accuracy of the determination.

3. A primary virtual machine comprising: a synchronization information generation unit that generates and outputs synchronization information based on a command and an execution result of the command; and a failure selection unit that determines whether failure information generated during execution of the command relates to hardware failure or software failure. The failure selection unit determines the accuracy of determination as to whether the failure information relates to hardware failure or software failure. The primary virtual machine changes its operation based on a determination result as to whether the failure information relates to hardware failure or software failure and the accuracy of the determination.

4. The failure selection unit according to claim 1, wherein the failure selection unit acquires a return value of a system call or error content as the failure information, and determines whether the failure information relates to hardware failure or software failure based on a list specifying which of hardware failure or software failure corresponds to the return value of the system call or the error content.

5. A primary virtual machine according to any one of claims 1 to 4, and a secondary machine having a secondary virtual machine that executes the command based on the synchronization information. A fault-tolerant system comprising

6. The primary virtual machine When it is determined that the failure information is information regarding the hardware failure, switches control to the secondary virtual machine, When it is determined that the failure information is information regarding the software failure, does not switch control to the secondary virtual machine, The fault-tolerant system according to claim 5.

Citation Information

Patent Citations

  • System control lsi for high-reliability computer and computer system using the same

    JP1994266574A

  • Fault-tolerant system

    JP2014059747A

  • Fault-tolerant system

    JP2014059748A

  • Fault-tolerant system

    JP2014059749A

  • Fault-tolerant system

    JP2014059750A