Host device, method, and non-transitory computer-readable recording medium for failover

The host device and method address the challenge of managing failovers in cloud database systems by transmitting signals to external host devices, deactivating connections, and initiating failover operations, resulting in efficient and rapid service continuity.

WO2025110457A1PCT designated stage expired Publication Date: 2025-05-30SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/014543
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-09-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing cloud database systems face challenges in efficiently managing failovers due to infrastructure failures, which can lead to data inconsistency and service downtime.

Method used

A host device and method that include a communication circuit and a processor, with instructions to transmit signals to external host devices in different availability zones, deactivate database connections upon failure, and initiate failover operations by instructing standby virtual machines to take over as primary after a specified period.

Benefits of technology

This solution enables rapid and efficient failover operations, minimizing data inconsistency and service downtime by ensuring seamless transition of database services to standby virtual machines in different availability zones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024014543_30052025_PF_FP_ABST
    Figure KR2024014543_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A host device is disclosed. The host device may include: a communication circuit; a processor; a virtual machine (275) that operates as a primary virtual machine to provide a service related to a database; and a memory storing instructions. The instructions, when executed by the processor, may cause the host device to transmit, through the communication circuit, a first signal to at least one external host device arranged in at least one second available area other than a first available area in which the host device is located. The instructions, when executed by the processor, may cause the host device to deactivate a connection related to the database with an electronic device on the basis that a response to the first signal is not obtained from the at least one external host device.
Need to check novelty before this filing date? Find Prior Art

Description

Host device, method, and non-transitory computer-readable recording medium for failover

[0001] The following descriptions relate to a host device, a method, and a non-transitory computer-readable recording medium for failover.

[0002] A cloud environment refers to an IT environment that provides virtualized servers accessible via a network, along with programs and databases running on those servers. Through the cloud environment, clients can access the computing resources they need from servers providing cloud services.

[0003] A cloud database (DB) is an example of a cloud service, a DB built to run in a cloud environment. A cloud DB service is a DB engine service provided and managed by a cloud service provider (CSP). The CSP deploys the infrastructure (or computing resources) required to run the DB engine on behalf of the client, and can install and manage the DB engine. Clients can access the DB engine through a DB endpoint (e.g., a uniform resource locator (URL)) provided by the CSP.

[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.

[0005] A host device is disclosed. The host device may include a communication circuit, a processor, and a memory storing instructions and a virtual machine (275) that operates as a primary virtual machine to provide database-related services. The instructions, when executed by the processor, may cause the host device to transmit a first signal to at least one external host device located in at least one second availability zone other than a first availability zone in which the host device is located, via the communication circuit. The instructions, when executed by the processor, may cause the host device to deactivate a connection related to the database with an electronic device based on a failure to obtain a response to the first signal from the at least one external host device.

[0006] A host device is disclosed. The host device may include a communication circuit, a processor, and a memory storing instructions. The instructions, when executed by the processor, may cause the host device to identify a health state of a first virtual machine operating as a primary virtual machine for providing database-related services to a plurality of electronic devices. The first virtual machine may be executed on a first external host device. The instructions, when executed by the processor, may cause the host device to, in response to identifying, based on the information, that the health state of the first virtual machine is a designated first state, transmit a signal to the second virtual machine through the communication circuit, the signal instructing, after a designated time, a second virtual machine running on a second external host device distinct from the first external host device to operate as the primary virtual machine.

[0007] A method is disclosed. The method may be performed by a host device including a communication circuit and a memory storing a virtual machine that operates as a primary virtual machine to provide database-related services. The method may include transmitting a first signal to at least one external host device located in at least one second availability zone other than a first availability zone in which the host device is located, via the communication circuit. The method may include disabling a connection related to the database with the electronic device based on a failure to obtain a response to the first signal from the at least one external host device.

[0008] A method is disclosed. The method can be performed by a host device including a communication circuit. The method can include an operation of identifying a health state of a first virtual machine operating as a primary virtual machine for providing a database-related service to a plurality of electronic devices. The first virtual machine can be executed on a first external host device. The method can include an operation of transmitting a signal to the second virtual machine through the communication circuit, the signal instructing a second virtual machine running on a second external host device distinct from the first external host device to operate as the primary virtual machine after a specified period of time, in response to identifying, based on the information, that the health state of the first virtual machine is in a specified first state.

[0009] A non-transitory computer-readable storage medium is disclosed. The non-transitory computer-readable storage medium can store a program including instructions. The instructions, when executed by a processor of a host device including a communication circuit and a memory storing a virtual machine operating as a primary virtual machine to provide database-related services, can cause the host device to transmit a first signal to at least one external host device located in at least one second availability zone other than a first availability zone in which the host device is located, through the communication circuit. The instructions, when executed by the processor, can cause the host device to deactivate a connection related to the database with an electronic device based on a failure to obtain a response to the first signal from the at least one external host device.

[0010] A non-transitory computer-readable recording medium is disclosed. The non-transitory computer-readable recording medium can store a program including instructions. The instructions, when executed by a processor of a host device including a communication circuit, can cause the host device to identify a health state of a first virtual machine operating as a primary virtual machine for providing database-related services to a plurality of electronic devices. The first virtual machine can be executed on a first external host device. The instructions, when executed by the processor, can cause the host device to, based on the information, in response to identifying the health state of the first virtual machine as being in a designated first state, transmit a signal to the second virtual machine through the communication circuit, the signal instructing a second virtual machine running on a second external host device distinct from the first external host device to operate as the primary virtual machine after a designated period of time.

[0011] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0012] Figure 2a illustrates an example of a cloud database environment.

[0013] Figure 2b illustrates an example of virtual machines (VMs) deployed in availability zones.

[0014] Figure 3 is a flowchart showing the connection deactivation operation of the primary DB VM.

[0015] Figure 4a is a flowchart showing a failover operation performed by a failover controller.

[0016] Figure 4b is a flowchart showing a failover operation performed by a failover controller.

[0017] Figure 5a is a diagram showing connection deactivation and failover operations over time.

[0018] Figure 5b is a diagram showing connection deactivation and failover operations over time.

[0019] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0020] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.

[0021] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0022] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0023] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0024] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0025] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0026] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0027] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0028] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0029] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0030] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0031] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0032] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0033] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0034] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0035] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0036] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0037] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0038] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for realizing eMBB, a loss coverage (e.g., 664 dB or less) for realizing mMTC, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 6 ms or less for round trip) for realizing URLLC.

[0039] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0040] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0041] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0042] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0043] Figure 2a illustrates an example of a cloud database environment. Figure 2b illustrates an example of virtual machines (VMs) deployed across availability zones.

[0044] Referring to FIG. 2A, a cloud database environment may include an electronic device (101), a plurality of availability zones (201, 204), a topology storage server (261), a network time server (263), and a service server (265). In one embodiment, the availability zone (201) may include a plurality of hosts (211 to 217). In one embodiment, the availability zone (204) may include a plurality of hosts (221 to 227). In one embodiment, the electronic device (101), the plurality of hosts (211 to 217) included in the availability zone (201), the plurality of hosts (221 to 227) included in the availability zone (204), the topology storage server (261), the network time server (263), and the service server (265) may be connected via a second network (199) (e.g., a long-distance wireless communication network).

[0045] In one embodiment, the availability zones (201, 204) may be located in the same region or in different regions. In one embodiment, a region may be a physical location where availability zones are gathered (e.g., Korea, North America). In one embodiment, an availability zone may be a physical data center within a region.

[0046] In one embodiment, the availability zones (201, 204) may include (or place) a plurality of hosts (211 to 217, 221 to 227). In one embodiment, the hosts may be servers (e.g., server 108 of FIG. 1). In one embodiment, each of the plurality of hosts (211 to 217, 221 to 227) may include a processor, a communication circuit, and a memory. In one embodiment, the processor, the communication circuit, and the memory of each of the plurality of hosts (211 to 217, 221 to 227) may correspond to the processor (120), the memory (130), and the communication module (190) of FIG. 1, respectively. For example, the host (211) of the availability zone (201) may include a processor (231), a communication circuit (233), and a memory (235). In one embodiment, the availability zones (201, 204) may include a plurality of hosts (211 to 217, 221 to 227), some of which may be mounted on a rack. The hosts mounted on a single rack may be connected to each other via a network switch.

[0047] In one embodiment, a virtual machine (VM) may be deployed on some of the plurality of hosts (211 to 217, 221 to 227). In one embodiment, the VM may be a virtual environment that operates as a virtual computer system by utilizing the physical hardware resources of the host (e.g., processor, communication circuit, and memory). In one embodiment, the VM may be classified as a service VM or a database (DB) VM, but is not limited thereto.

[0048] In one embodiment, a service VM may be a VM that runs a process for providing a cloud service. In one embodiment, a service VM may be created when a cloud service is established. In one embodiment, a service VM may not be deleted or modified from a host except for physical failure reasons. In one embodiment, a service VM may be deployed by a cloud service provider (CSP). In one embodiment, referring to FIGS. 2A and 2B , an availability zone (201) may include a service VM (241). In one embodiment, referring to FIG. 2B , an availability zone (204) may include a service VM (271). In one embodiment, referring to FIG. 2B , an availability zone (207) may include a service VM (279).

[0049] In one embodiment, a service VM may include a failover controller to perform a failover procedure to overcome a failure that occurred in a DB VM. In one embodiment, one failover controller among the failover controllers of different hosts may operate as an active process, and the remaining failover controllers may operate as standby processes. For example, a failover controller selected as a leader through a designated distributed consensus algorithm (e.g., raft algorithm) among the failover controllers may operate as an active process, and the remaining failover controllers may operate as standby processes. In one embodiment, referring to FIGS. 2A and 2B , a service VM (241) in an availability zone (201) may include a failover controller (251). In one embodiment, referring to FIG. 2B , a service VM (271) in an availability zone (204) may include a failover controller (281). In one embodiment, referring to FIG. 2B, a service VM (279) of an availability zone (207) may include a failover controller (289). In one embodiment, among the failover controllers (251, 281, 289), the failover controller (281) may operate as an active process according to a specified distributed consensus algorithm. In one embodiment, among the failover controllers (251, 281, 289), the failover controllers (251, 289) may operate as standby processes according to a specified distributed consensus algorithm.

[0050] In one embodiment, the DB VM may be a VM that drives a process for providing a cloud database service to a client (e.g., a user of an electronic device (101)). In one embodiment, the DB VM may be deployed in one or more availability zones (201, 204) according to a setting (or request) of a cloud service user. For example, the DB VM may be deployed in one or more availability zones (201, 204) according to an availability level of the DB specified by the client (e.g., a user of an electronic device (101)). When the DB VM is deployed in two or more availability zones (201, 204), the DB VM deployed in one availability zone (e.g., availability zone (201)) may operate as a primary DB VM, and the DB VMs deployed in the remaining availability zones (e.g., availability zone (204)) may operate as standby DB VMs. In one embodiment, referring to FIGS. 2A and 2B, an availability zone (201) may include a DB VM (245). In one embodiment, referring to FIG. 2B, an availability zone (204) may include a DB VM (275).

[0051] In one embodiment, the standby DB VM may replicate the primary DB VM. For example, the standby DB VM may periodically obtain (or collect) (or receive) (or update) (or store) updated data from the primary DB VM.

[0052] In one embodiment, the standby DB VM can be promoted to the primary DB VM according to a failover procedure to overcome the failure in the case of a failure in the primary DB VM. In one embodiment, the standby DB VM can be deployed in a new availability zone according to the failover procedure.

[0053] In one embodiment, the DB VM may include an access breaker, a topology cache, and / or a health checker. In one embodiment, the access breaker may be a process running in the primary DB VM. In one embodiment, the access breaker may disconnect (or down) (or deactivate) the connection (or link) between the DB VM and the electronic device (101) depending on the network status of the primary DB VM. In one embodiment, the topology cache may obtain (or receive) (or store) information about VMs (e.g., service VMs and / or DB VMs) (e.g., information about deployed availability zones and / or information for accessing the VMs) from the topology storage server (261). In one embodiment, the health checker may check the health status of the DB VM (or the health status of the host). In one embodiment, the health checker may periodically transmit information indicating the health status to the failover controller (or, a failover controller operating as an active process). In one embodiment, referring to FIGS. 2A and 2B, a DB VM (245) in an availability zone (201) may include an access breaker (253), a topology cache (255), and / or a health checker (257). In one embodiment, referring to FIG. 2B, a DB VM (275) in an availability zone (204) may include an access breaker (283), a topology cache (285), and / or a health checker (287).

[0054] In one embodiment, the topology storage server (261) may store information about VMs (e.g., service VMs, and / or DB VMs) (e.g., information about deployed availability zones, and / or information for access to the VMs).

[0055] In one embodiment, the network time server (263) may be a server for synchronizing time information between VMs (e.g., service VMs and / or DB VMs). In one embodiment, each of the VMs (e.g., service VMs and / or DB VMs) may be synchronized with the time information of the network time server (263) at a specified period. Accordingly, the VMs (e.g., service VMs and / or DB VMs) may operate within a specified error range.

[0056] In one embodiment, the service server (265) may be a server operated by a cloud service provider (CSP). In one embodiment, the service server (265) may provide an application programming interface (API) for a specified function to VMs (e.g., a service VM and / or a DB VM).

[0057] In one embodiment, the electronic device (101) may correspond to the electronic device (101) of FIG. 1. In one embodiment, the electronic device (101) may access a DB VM (245) operating as a primary DB VM through a second network (199).

[0058] As described above, the DB VM (245, 275) may have difficulty providing cloud data services due to a failure (e.g., a failure of the DB VM itself or a failure caused by the cloud infrastructure). A failure of the DB VM itself may occur when the VM does not operate normally due to a configuration error or a bug in the VM itself. A failure of the DB VM itself can be resolved through the API (or cloud API) provided by the service server (265). For example, a failure of the DB VM itself can be resolved by reinstalling (or reconfiguring) the DB VM through the cloud API. A failure caused by the cloud infrastructure can be classified into a failure occurring at the host level, a failure occurring at the rack level, or a failure occurring at the data center (or availability zone) level. A failure occurring at the host level can be caused by a network cable failure, a power supply failure, a disk failure, or a motherboard failure. A failure occurring at the rack level can be caused by a power supply failure of the host, a network failure between racks, or between hosts within a rack. Failures at the data center (or availability zone) level can be caused by power supply failures for racks, network failures between data centers, or between racks within a data center. Such cloud infrastructure failures can be physically resolved by administrators.

[0059] However, if a failure occurs in the DB VM (245, 275) due to the cloud infrastructure, a failover operation may be required to detect the failure caused by the cloud infrastructure and maintain a consistent DB for the user when the failure is detected. Hereinafter, with reference to FIG. 3, an operation for handling a failure caused by the cloud infrastructure in a DB VM operating as a primary DB VM may be described. In addition, below, with reference to FIGS. 4a and 4b, an operation for handling a failure caused by the cloud infrastructure in a failover controller operating as an active process may be described.

[0060] Figure 3 is a flowchart showing the connection deactivation operation of the primary DB VM.

[0061] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0062] The description of FIG. 3 can be explained with reference to FIG. 1, FIG. 2a, and FIG. 2b.

[0063] Referring to FIG. 3, in operation 310, the access breaker (253) can identify (or obtain) information for connection of VMs (e.g., service VMs and / or DB VMs) stored in the topology cache (255) (e.g., information on deployed availability zones and / or information for access to VMs).

[0064] In operation 320, the access breaker (253) may identify at least one VM (303, 305, 307) based on the information. In one embodiment, the access breaker (253) may identify, based on the information, VMs (305, 307) located in other availability zones (301, 302) other than the availability zone (201) in which it is located. In one embodiment, the access breaker (253) may identify, based on the information, at least one VM in each of the other availability zones (301, 302). For example, the access breaker (253) may identify, based on the information, one VM (305) among the VMs in the availability zone (301) and one VM (307) among the VMs in the availability zone (302). However, the present invention is not limited thereto.

[0065] In operation 331, the access breaker (253) may transmit a signal. In one embodiment, the signal may be a signal to check connectivity (or reachability) between a DB VM (245) operating as a primary DB VM and a VM (305) located in another availability zone (301). For example, the signal may be a ping (packet internet groper) signal.

[0066] In operation 335, the access breaker (253) may transmit a signal. In one embodiment, the signal may be a signal to check connectivity (or accessibility) between a DB VM (245) operating as a primary DB VM and a VM (307) located in another availability zone (302).

[0067] In operation 340, the access breaker (253) may determine whether a response to the signal is received. In one embodiment, the access breaker (253) may determine whether a response to the signal is received during a specified connectivity timeout (or accessibility timeout).

[0068] In one embodiment, the access breaker (253) may determine that a response to a signal has not been received if no response is received from all of the VMs (305, 307) to which the signal was transmitted. In one embodiment, the access breaker (253) may determine that a response to a signal has been received if a response is received from at least one of the VMs (305, 307) to which the signal was transmitted.

[0069] In response to (or based on) receiving a response to the signal at operation 340, the access breaker (253) may perform operation 310 again. In response to (or based on) not receiving a response to the signal at operation 340, the access breaker (253) may perform operation 350.

[0070] In operation 350, the access breaker (253) can disconnect (or down) (or deactivate) the connection (or link) of the DB VM (245). For example, the access breaker (253) can disconnect (or down) (or deactivate) the network interface for connection between the DB VM (245) and the electronic device (101). For example, the access breaker (253) can disconnect (or down) (or deactivate) the network interface for connection between the DB VM (245) and the electronic device (101) using a specified command (e.g., ifconfig down, or ifdown).

[0071] As described above, the access breaker (253) may determine that a communication failure (e.g., a failure of the DB VM (245) itself or a failure caused by the cloud infrastructure) has occurred in the DB VM (245) based on the fact that no response to a signal is received during a specified connectivity time-out (or accessibility time-out). If a communication failure occurs in the DB VM (245), the access breaker (253) may disconnect (down) (or deactivate) the DB VM (245) without inactivating (or shutting down) the DB VM (245). Accordingly, the access breaker (253) may only affect its own DB VM (245) operating as the primary DB VM, and thus may not affect other VMs (e.g., service VMs (241)) included in the same host (211) as the DB VM (245). Therefore, cloud service providers do not need to deploy VMs on separate hosts according to their types in order to manage DB VMs, and thus the physical resources of the hosts can be efficiently managed and / or utilized.

[0072] Figure 4a is a flowchart showing a failover operation performed by a failover controller.

[0073] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0074] The description of Fig. 4a can be explained with reference to Fig. 1, Fig. 2a, and Fig. 2b.

[0075] The operations of Fig. 4a can be assumed to be performed after the network of the DB VM (245) operating as the primary DB VM is disconnected. The operations of Fig. 4a can be performed by the failover controller (401) operating as an active process at least after the network of the DB VM (245) is disconnected. The operations of Fig. 4a can be assumed to be performed on the same host by the DB VM (245) and the failover controller (401) of the service VM. However, the present invention is not limited thereto.

[0076] Referring to FIG. 4A, in operation 410, the health checker (257) may transmit information indicating a health status. In one embodiment, the health checker (257) may transmit information indicating a health status of a DB VM (245) operating as a primary DB VM to a failover controller (401) operating as an active process (or selected as a leader). In one embodiment, the health status may indicate a network connection status of the DB VM (245) (or a network connection status of a host (211) of the DB VM (245)). For example, the health status may indicate a status of a network interface of the DB VM (245) (or a status of a network interface of the host (211) of the DB VM (245). In one embodiment, the health status may indicate a speed of a network interface and / or a number of transmitted and received packets. In one embodiment, the health status may indicate the current state of the network interface (e.g., listen, established, time_wait, closed), but is not limited thereto. In one embodiment, the health status may be identified from the responses of the VMs (305, 307) to a signal (e.g., ping) transmitted by the access breaker (253), but is not limited thereto. In one embodiment, the health status may be identified based on other designated commands (e.g., tcping, netstat).

[0077] In operation 420, the failover controller (401) may determine whether the health state is a designated first state. In one embodiment, the failover controller (401) may identify (or confirm) (or determine) the health state of the DB VM (245) based on information indicating the health state from the health checker (257). For example, the failover controller (401) may identify (or confirm) (or determine) the health state of the DB VM (245) based on the speed of the network interface and / or the number of transmitted and received packets. For example, the failover controller (401) may identify the health state of the DB VM (245) as an unhealthy state (e.g., unhealthy) if the speed of the network interface is lower than the designated speed. For example, the failover controller (401) may identify the health status of the DB VM (245) as healthy (e.g., healthy) when the speed of the network interface exceeds a specified speed. In one embodiment, the first state may indicate an unhealthy state (e.g., unhealthy). In one embodiment, the failover controller (401) may indicate a healthy state (e.g., healthy) as the second state.

[0078] In response to (or based on) that the health state is not the designated first state in operation 420, the failover controller (401) may perform operation 420 again. In response to (or based on) that the health state is the designated first state in operation 420, the failover controller (401) may perform operation 430.

[0079] In operation 430, the failover controller (401) may send a termination request to the service server (265). The failover controller (401) may request the service server (265) to terminate the DB VM (245) operating as the primary DB VM through a designated cloud API (e.g., a cloud API for termination).

[0080] In operation 435, the service server (265) may transmit a termination response to the failover controller (401). In one embodiment, the service server (265) may instruct the DB VM (245) to terminate based on a designated cloud API (e.g., a cloud API for termination). In one embodiment, the service server (265) may transmit a response including information indicating the status of the DB VM (245) to the failover controller (401). In one embodiment, the service server (265) may periodically transmit information indicating the status of the DB VM (245) to the failover controller (401). In one embodiment, the status of the DB VM (245) may indicate termination or unknown.

[0081] In operation 440, the failover controller (401) can determine whether the DB VM (245) has been terminated. In one embodiment, the failover controller (401) can determine (or identify) whether the DB VM (245) has been terminated based on the status of the DB VM (245) indicated by the service server (265).

[0082] In one embodiment, the failover controller (401) can determine (or identify) that the DB VM (245) has been terminated when the status of the DB VM (245) indicated by the service server (265) indicates termination.

[0083] In one embodiment, the failover controller (401) may determine (or identify) that the DB VM (245) is not terminated when the status of the DB VM (245) indicated by the service server (265) indicates unknown. For example, when the network of the DB VM (245) is disconnected, the DB VM (245) may not receive a termination instruction from the service server (265), and the service server (265) may not be able to identify the status of the DB VM (245). Alternatively, when the network of the DB VM (245) is disconnected, even when the DB VM (245) is terminated in response to a termination instruction from the service server (265), the service server (265) may not be able to identify the status of the DB VM (245). In such cases, the status of the DB VM (245) indicated by the service server (265) may indicate unknown.

[0084] In response to (or based on) that the DB VM (245) is terminated at operation 440, the failover controller (401) may perform operation 460. In response to (or based on) that the DB VM (245) is not terminated at operation 440 (or is in an unknown state), the failover controller (401) may perform operation 450.

[0085] In operation 450, the failover controller (401) may wait for a specified period of time. In one embodiment, the specified period of time may include the time required for the access breaker (253) of the DB VM (245) operating as the primary DB VM to check connectivity (or accessibility). In one embodiment, the specified period of time may include the time required for the access breaker (253) of the DB VM (245) operating as the primary DB VM to disconnect (or down) (or deactivate) the connection (or link) of the DB VM (245). In one embodiment, the specified period of time may include the time required for a response from a specified cloud API. However, the present invention is not limited thereto.

[0086] In operation 460, the failover controller (401) may transmit a signal for operation as a primary DB VM to a DB VM (275) operating as a standby DB VM. In one embodiment, the failover controller (401) may transmit a signal for operation as a primary DB VM to a DB VM (275) selected from among a plurality of DB VMs operating as standby DB VMs.

[0087] In one embodiment, the failover controller (401) may update connection information (e.g., DB endpoint) pointing to the DB VM (245) to point to the DB VM (275). For example, the connection information (e.g., DB endpoint) may indicate a uniform resource locator (URL) of the DB VM (275).

[0088] In operation 470, the failover controller (401) may request the creation of a standby DB VM from the host (405). In one embodiment, the failover controller (401) may request the creation of a standby DB VM from a host (405) selected from among hosts included in a different availability zone from the host (211) of the DB VM (245). In one embodiment, the failover controller (401) may request the creation of a standby DB VM from a host (405) selected from among hosts included in a different availability zone from the host of the DB VM (275).

[0089] As described above, the failover controller (401) can automatically complete failover operations within a certain period of time without administrator intervention when the network status of the primary DB VM is poor, thereby reducing operating costs and enabling rapid failover processing. In addition, the failover controller (401) can reduce the possibility of stale reads by waiting for at least one required amount of time to disable the connection of the primary DB VM after the network of the primary DB VM is disconnected.

[0090] Figure 4b is a flowchart showing a failover operation performed by a failover controller.

[0091] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0092] The description of Fig. 4b can be explained with reference to Fig. 1, Fig. 2a, Fig. 2b, and Fig. 4a.

[0093] The operations of Fig. 4b can be assumed to be performed after the network of the DB VM (245) operating as the primary DB VM is disconnected. The operations of Fig. 4b can be performed by the failover controller (401) operating as an active process at least after the network of the DB VM (245) is disconnected. The operations of Fig. 4b can be assumed to be performed when the failover controller (401) of the DB VM (245) and the service VM are executed on different hosts (or hosts in different availability zones). However, the present invention is not limited thereto.

[0094] Referring to FIG. 4B, in operation 410, the health checker (257) may transmit information indicating a health status. In one embodiment, the health checker (257) may transmit information indicating a health status of a DB VM (245) operating as a primary DB VM to a failover controller (401) operating as an active process (or selected as a leader).

[0095] In one embodiment, information indicating the health status in operation 410 may not be transmitted to the failover controller (401). For example, if the network of the DB VM (245) including the health checker (257) is disconnected, information indicating the health status may not reach the failover controller (401).

[0096] In operation 425, the failover controller (401) may determine whether there is a timeout for health status information (e.g., a health event timeout). In one embodiment, the failover controller (401) may determine whether information indicating a health status from the health checker (257) was received during the health event timeout.

[0097] In one embodiment, at operation 425, if information indicating a health status from the health checker (257) is not received during the health event timeout, the failover controller (401) may perform operation 430 in response to (or based on) that information indicating a health status from the health checker (257) is not received during the health event timeout. In one embodiment, at operation 425, in response to (or based on) that information indicating a health status from the health checker (257) is received during the health event timeout, the failover controller (401) may perform operation 425 again. However, the present invention is not limited thereto. At operation 425, in response to (or based on) that information indicating a health status from the health checker (257) is received during the health event timeout, the failover controller (401) may perform operation 420 of FIG. 4A instead of the operations of FIG. 4B.

[0098] In operation 431, the failover controller (401) may transmit a termination request to the service server (265). In operation 436, the service server (265) may transmit a termination response to the failover controller (401). In operation 441, the failover controller (401) may determine whether the DB VM (245) has been terminated. In response to (or based on) that the DB VM (245) has been terminated in operation 441, the failover controller (401) may perform operation 461. In response to (or based on) that the DB VM (245) has not been terminated (or is in an unknown state), the failover controller (401) may perform operation 451.

[0099] In operation 451, the failover controller (401) may wait for a specified period of time. In operation 461, the failover controller (401) may transmit a signal to the DB VM (275) operating as the standby DB VM to operate as the primary DB VM. In operation 471, the failover controller (401) may request the host (405) to create a standby DB VM.

[0100] As described above, the failover controller (401) can automatically complete failover operations within a certain period of time without administrator intervention when the primary DB VM is in a network disconnected state, thereby reducing operating costs and enabling rapid failover processing. In addition, the failover controller (401) can reduce the possibility of stale reads by waiting for at least one required amount of time to disable the connection of the primary DB VM after the network disconnection of the primary DB VM.

[0101] Figure 5a is a diagram showing connection deactivation and failover operations over time.

[0102] The description of Fig. 5a can be explained with reference to Fig. 1, Fig. 2a, Fig. 2b, Fig. 3, Fig. 4a, and Fig. 4b.

[0103] Referring to FIG. 5a, at point in time (W0), the network of a DB VM (245) operating as a primary DB VM may be disconnected. At least after point in time (W0) when the network of the DB VM (245) is disconnected, a failover controller (511) may be selected as an active process by a designated consensus algorithm.

[0104] At time point (W1), the health checker (257) of the DB VM (245) can transmit information indicating the health status to the failover controller (511). At time point (W1) of transmitting information indicating the health status, the health checker (257) can transmit information indicating the health status to the failover controller (511). If the network of the DB VM (245) is disconnected, the information indicating the health status may not reach the failover controller (511).

[0105] At time point (W2), the access breaker (253) of the DB VM (245) can transmit a signal to the VMs (501, 503, 505) stored in the topology cache (255) to check connectivity (or accessibility). In one embodiment, the access breaker (253) can transmit a signal to check connectivity (or accessibility) between the DB VM (245) and VMs (501, 503, 505) located in different availability zones.

[0106] At time point (W3), the failover controller (511) may determine that a timeout (e.g., health event timeout (T0)) for health status information has elapsed. In one embodiment, the health event timeout (T0) may be a time corresponding to W3-W1.

[0107] At time point (W3), the failover controller (511) may determine that a communication failure (e.g., failure of the DB VM (245) itself or failure caused by the cloud infrastructure) has occurred in the DB VM (245) based on (or in response to) the health event timeout (T0) expiring.

[0108] At time point (W3), the failover controller (511) may request the service server (265) to terminate the DB VM (245) through a cloud API (e.g., a cloud API for termination) designated for the service server (265) based on (or in response to) the health event timeout (T0) elapsed. In one embodiment, the termination timeout (T2) of the DB VM (245) by the cloud API may be a time corresponding to W5-W3. In one embodiment, the termination timeout (T2) of the DB VM (245) may be a time required for the DB VM (245) to terminate.

[0109] In one embodiment, the failover controller (511) may determine (or identify) whether the DB VM (245) has been terminated based on the status of the DB VM (245) indicated by the service server (265) until the termination timeout (T2) after requesting the termination of the DB VM (245) to the service server (265) at time point (W3). In one embodiment, the failover controller (511) may determine (or identify) whether the DB VM (245) has been terminated based on the status of the DB VM (245) indicated by the service server (265) until the termination timeout (T2) of the DB VM (245). In one embodiment, the failover controller (511) may determine (or identify) that the DB VM (245) has been terminated if the status of the DB VM (245) indicated by the service server (265) indicates termination before the termination timeout (T2). In one embodiment, the failover controller (511) may continuously check the status of the DB VM (245) until a termination timeout (T2) if the status of the DB VM (245) indicates unknown after requesting the service server (265) to terminate the DB VM (245). In one embodiment, the failover controller (511) may determine (or identify) that the DB VM (245) is not terminated if the status of the DB VM (245) indicated by the service server (265) indicates unknown until a termination timeout (T2).

[0110] In one embodiment, the failover controller (511) may perform failover without waiting for the waiting time (T4) for the operation of the access breaker (253) if the status of the DB VM (245) indicated by the service server (265) indicates termination before the termination timeout (T2) of the DB VM (245). For example, if the status of the DB VM (245) indicated by the service server (265) indicates termination before the timeout (T2), the controller (511) may transmit a signal for operation as a primary DB VM to a DB VM (e.g., DB VM (275)) operating as a standby DB VM. For example, if the status of the DB VM (245) indicated by the service server (265) before the timeout (T2) indicates termination, the failover controller (511) may request the creation of a standby DB VM to a host (e.g., host (501)) selected from among the hosts included in a different availability zone from the host (211) of the DB VM (245).

[0111] At time point (W4), the access breaker (253) may determine that no response is received from any of the VMs (501, 503, 505) during a timeout for accessibility (e.g., accessibility timeout (T1)). In one embodiment, the access breaker (253) may determine that the accessibility timeout (T1) has elapsed without a response being received from any of the VMs (501, 503, 505). In one embodiment, the accessibility timeout (T1) may be a time corresponding to W4-W2.

[0112] At time point (W4), the access breaker (253) may determine that a communication failure (e.g., failure of the DB VM (245) itself or failure caused by the cloud infrastructure) has occurred in the DB VM (245) based on (or in response to) the elapsed accessibility timeout (T1).

[0113] At time point (W4), the access breaker (253) may perform a connection disable operation for the DB VM (245) based on (or in response to) the accessibility timeout (T1) expiring. The access breaker (253) may release (or down) (or disable) the connection (or link) of the DB VM (245) based on (or in response to) the accessibility timeout (T1). In one embodiment, the operation time (T3) of the access breaker (253) may be a time corresponding to W6-W4.

[0114] At time point (W5), the failover controller (511) can identify the termination timeout (T2) of the DB VM (245).

[0115] At point (W5), if the status of the DB VM (245) indicates unknown, the failover controller (511) can wait for a waiting time (T4) for the operation of the access breaker (253).

[0116] At a point in time (W7) after the waiting time (T4) for the operation of the access breaker (253), the failover controller (511) can perform a failover. For example, at a point in time (W7) after the waiting time (T4) for the operation of the access breaker (253), the failover controller (511) can transmit a signal for operation as a primary DB VM to a DB VM (e.g., DB VM (275)) operating as a standby DB VM. For example, at a point in time (W7) after the waiting time (T4) for the operation of the access breaker (253), the failover controller (511) can request the creation of a standby DB VM to a host (e.g., host (501)) selected from among hosts included in an availability zone different from the host (211) of the DB VM (245).

[0117] Figure 5b is a diagram showing connection deactivation and failover operations over time.

[0118] The description of Fig. 5b can be explained with reference to Fig. 1, Fig. 2a, Fig. 2b, Fig. 3, Fig. 4a, Fig. 4b, and Fig. 5a.

[0119] Referring to FIG. 5b, at point (W0), the network of the DB VM (245) operating as the primary DB VM may be disconnected.

[0120] At time point (W1), the health checker (257) of the DB VM (245) can transmit information indicating the health status to the failover controller (511). At time point (W1) of transmitting information indicating the health status, the health checker (257) can transmit information indicating the health status to the failover controller (511). If the network of the DB VM (245) is disconnected, the information indicating the health status may not reach the failover controller (511).

[0121] At time point (W3), the failover controller (511) can receive (or acquire) information indicating the health status from the health checker (257) of the DB VM (245). At time point (W3), the failover controller (511) can identify the health status of the DB VM (245) as unhealthy (e.g., unhealthy) based on the information indicating the health status.

[0122] At time point (W3), the failover controller (511) may determine that a communication failure (e.g., failure of the DB VM (245) itself or failure caused by the cloud infrastructure) has occurred in the DB VM (245) based on (or in response to) identifying the health status of the DB VM (245) as unhealthy (e.g., unhealthy) based on information indicating the health status.

[0123] At time point (W3), the failover controller (511) may, based on (or in response to) identifying the health status of the DB VM (245) as unhealthy (e.g., unhealthy), request the service server (265) to terminate the DB VM (245) through a cloud API (e.g., a cloud API for termination) designated for the service server (265). In one embodiment, the termination timeout (T2) of the DB VM (245) by the cloud API may be a time corresponding to W5-W3. In one embodiment, the termination timeout (T2) of the DB VM (245) may be a time required for the DB VM (245) to be terminated.

[0124] In one embodiment, the failover controller (511) may determine (or identify) whether the DB VM (245) has been terminated based on the status of the DB VM (245) indicated by the service server (265) until the termination timeout (T2) after requesting the termination of the DB VM (245) to the service server (265) at time point (W3). In one embodiment, the failover controller (511) may determine (or identify) whether the DB VM (245) has been terminated based on the status of the DB VM (245) indicated by the service server (265) until the termination timeout (T2) of the DB VM (245). In one embodiment, the failover controller (511) may determine (or identify) that the DB VM (245) has been terminated if the status of the DB VM (245) indicated by the service server (265) indicates termination before the termination timeout (T2). In one embodiment, the failover controller (511) may continuously check the status of the DB VM (245) until a termination timeout (T2) if the status of the DB VM (245) indicates unknown after requesting the service server (265) to terminate the DB VM (245). In one embodiment, the failover controller (511) may determine (or identify) that the DB VM (245) is not terminated if the status of the DB VM (245) indicated by the service server (265) indicates unknown until a termination timeout (T2).

[0125] In one embodiment, the failover controller (511) may perform failover without waiting for the waiting time (T4) for the operation of the access breaker (253) if the status of the DB VM (245) indicated by the service server (265) indicates termination before the termination timeout (T2) of the DB VM (245). For example, if the status of the DB VM (245) indicated by the service server (265) indicates termination before the timeout (T2), the controller (511) may transmit a signal for operation as a primary DB VM to a DB VM (e.g., DB VM (275)) operating as a standby DB VM. For example, if the status of the DB VM (245) indicated by the service server (265) before the timeout (T2) indicates termination, the failover controller (511) may request the creation of a standby DB VM to a host (e.g., host (501)) selected from among the hosts included in a different availability zone from the host (211) of the DB VM (245).

[0126] At time point (W2), the access breaker (253) of the DB VM (245) can transmit a signal to the VMs (501, 503, 505) stored in the topology cache (255) to check connectivity (or accessibility).

[0127] At time point (W4), the access breaker (253) can identify that no response is received from any of the VMs (501, 503, 505) during the accessibility timeout (e.g., accessibility timeout (T1)).

[0128] At time point (W4), the access breaker (253) may determine that a communication failure (e.g., failure of the DB VM (245) itself or failure caused by the cloud infrastructure) has occurred in the DB VM (245) based on (or in response to) the elapsed accessibility timeout (T1).

[0129] At time point (W4), the access breaker (253) may perform a connection disable action for the DB VM (245) based on (or in response to) the accessibility timeout (T1) expiring.

[0130] At time point (W5), the failover controller (511) can identify the termination timeout (T2) of the DB VM (245).

[0131] At point (W5), if the status of the DB VM (245) indicates unknown, the failover controller (511) can wait for a waiting time (T4) for the operation of the access breaker (253).

[0132] At a point in time (W7) after the waiting time (T4) for the operation of the access breaker (253), the failover controller (511) can perform failover.

[0133] In one embodiment, in order to prevent two primary DB VMs from existing simultaneously even when the failover controller (511) and the DB VM (245) are not network connected to each other, the failover operation of the failover controller (511) must be performed after the connection of the access breaker (253) is disabled. Accordingly, the failover operation of the failover controller (511) can be performed immediately when the status of the DB VM (245) indicates termination, but when the status of the DB VM (245) indicates unconfirmed, the failover operation of the failover controller (511) must be performed after the operation time (T3) required for the connection of the access breaker (253) to be disabled. Accordingly, in order to ensure that the failover operation of the failover controller (511) is performed after the connection of the access breaker (253) is disabled, it may be necessary to establish a relationship between time intervals (e.g., health event timeout (T0), accessibility timeout (T1), termination timeout (T2), operation time (T3), and standby time (T4) of FIG. 5a or 5b) according to events that may occur after the network of the DB VM (245) is disconnected.

[0134] For example, if the operating time (T3) and the standby time (T4) are the same, the start time (W4) of T3 must always be before the start time (W5) of T4 so that the failover operation of the failover controller (511) can be performed after the connection deactivation of the access breaker (253). For example, if the operating time (T3) is longer than the standby time (T4), the start time (W4) of T3 must always be before the start time (W5) of T4 and the sum of the time difference between T3 and T4 so that the failover operation of the failover controller (511) can be performed after the connection deactivation of the access breaker (253). For example, if the operation time (T3) is shorter than the standby time (T4), the start time (W5) of T4 must always be later than the sum of the start time (W4) of T3 and the time difference between T3 and T4 so that the failover operation of the failover controller (511) can be performed after the connection of the access breaker (253) is disabled.

[0135] For example, referring to FIGS. 5A and 5B , the transmission cycle of information indicating a health state related to a health event timeout may not be considered in the relationship between the start time of T3 (W4) and the start time of T4 (W5) when the health state is considered as an unhealthy state (e.g., unhealthy). For example, if the operation time (T3) and the standby time (T4) are the same, the sum of the check cycle of connectivity (or accessibility) and the time length of the accessibility timeout (T1) must be shorter than the end timeout (T2) so that the start time of T3 (W4) always exists before the start time of T4 (W5).

[0136] In summary, when the operation time (T3) and the waiting time (T4) are equal, the sum of the length of the connectivity (or accessibility) check period and the accessibility timeout (T1) must be shorter than the termination timeout (T2) to ensure that the start time of T3 (W4) always starts before the start time of T4 (W5).

[0137] As described above, the host device (211) may include a communication circuit (233), a processor (231), a virtual machine (275) that operates as a primary virtual machine to provide database-related services, and a memory (235) that stores instructions. The instructions, when executed by the processor (231), may cause the host device (211) to transmit a first signal to at least one external host device (221) located in at least one second availability zone (204) other than the first availability zone (201) in which the host device (211) is located, through the communication circuit (233). The instructions, when executed by the processor (231), may cause the host device (211) to deactivate a connection related to the database with the electronic device (101) based on a failure to obtain a response to the first signal from the at least one external host device (221).

[0138] The above instructions, when executed by the processor (231), may cause the host device (211) to obtain information for connection with the at least one external host device (221) disposed in the at least one second availability zone (204) through the communication circuit (233). The above instructions, when executed by the processor (231), may cause the host device (211) to transmit the first signal to the at least one external host device (221) based on the obtained information.

[0139] The above information may include information for connection with at least one virtual machine (271, 275, 279) running on the at least one external host device (221). The instructions, when executed by the processor (231), may cause the host device (211) to transmit the first signal to the at least one virtual machine (271, 275, 279) based on the information for connection with the at least one virtual machine (271, 275, 279).

[0140] The above instructions, when executed by the processor (231), may cause the host device (211) to transmit a signal instructing the electronic device (101) to operate a virtual machine (275) as a primary virtual machine for providing a service related to the database based on obtaining the response from the at least one external host device (221) to the first signal.

[0141] The above instructions, when executed by the processor (231), may cause the host device (211) to obtain information indicating a health status of the virtual machine (245) operating as the primary virtual machine in a first cycle. The above instructions, when executed by the processor (231), may cause the host device (211) to transmit the information obtained in the first cycle through the communication circuit (233) to a first external host device (221) on which a failover controller (281) is executing.

[0142] The above instructions, when executed by the processor (231), may cause the host device (211) to transmit information acquired in the first cycle to the first external host device (221) through the communication circuit (233), and then to acquire another signal from the server (265) through the communication circuit (233) instructing termination of the virtual machine (245). The above instructions, when executed by the processor (231), may cause the host device (211) to transmit a response indicating a termination result of the virtual machine (245) to the server (265) in response to termination of the virtual machine (245) based on the other signal.

[0143] As described above, the host device (221) may include a communication circuit, a processor, and a memory storing instructions. The instructions, when executed by the processor, may cause the host device (221) to identify a health status of a first virtual machine (245) that operates as a primary virtual machine for providing database-related services to a plurality of electronic devices (101). The first virtual machine (245) may be executed on a first external host device (211). The above instructions, when executed by the processor, may cause the host device (221) to, based on the information, in response to identifying that the health state of the first virtual machine (245) is a designated first state, transmit a signal to the second virtual machine (275) through the communication circuit (233) instructing a second virtual machine (275) running on a second external host device distinct from the first external host device (211) to operate as the primary virtual machine after a designated time.

[0144] The above instructions, when executed by the processor, may cause the host device (221) to obtain information indicating a health status of the first virtual machine (245) from the first external host device (211) through the communication circuit (233). The above instructions, when executed by the processor, may cause the host device (221) to identify the health status of the first virtual machine (245) based on the obtained information.

[0145] The above instructions, when executed by the processor, may cause the host device (221) to identify the health state of the first virtual machine (245) as being the specified first state based on the fact that information indicating the health state of the first virtual machine (245) is not acquired from the first external host device (211) through the communication circuit (233) for a specified period of time.

[0146] The above instructions, when executed by the processor, may cause the host device (221) to transmit another signal to the first external host device (211) requesting termination of the first virtual machine (245) in response to identifying that the health state of the first virtual machine (245) is a designated first state.

[0147] The instructions, when executed by the processor, may cause the host device (221) to transmit the signal to the second virtual machine (275) instructing the second virtual machine (275) to operate as the primary virtual machine, based on the termination of the first virtual machine (245) being identified within a specified other time period after transmitting the other signal. The instructions, when executed by the processor, may cause the host device (221) to transmit the signal to the second virtual machine (275) instructing the second virtual machine (275) to operate as the primary virtual machine, after a specified time period, based on the termination of the first virtual machine (245) not being identified within a specified other time period after transmitting the other signal.

[0148] The above instructions, when executed by the processor (231), may cause the host device (211) to identify a third external host device located in a second availability zone (204) other than the first availability zone (201) in which the first external host device (221) is located, based on the fact that no response to the other signal is obtained from the first external host device (211). The above instructions, when executed by the processor, may cause the host device (221) to transmit a request to the third external host device to execute a standby virtual machine for the primary virtual machine.

[0149] The first external host device (211) and the second external host device may be located in different availability areas.

[0150] The above-mentioned specified time may include a time allocated for the first virtual machine running on the first external host device (211) to deactivate the connection related to the database with the electronic device (101).

[0151] The above-mentioned specified time may include a time allocated for identifying connectivity of the first virtual machine running on the first external host device (211) with at least one external host device located in at least one second availability zone (204) other than the first availability zone (201) in which the first external host device (211) is located.

[0152] As described above, the method may be performed by a host device (211) including a communication circuit (233) and a memory (235) storing a database. The method may include an operation of transmitting a first signal to at least one external host device (221) located in at least one second availability zone (204) other than a first availability zone (201) in which the host device (211) is located, through the communication circuit (233). The method may include an operation of deactivating a connection related to the database with the electronic device (101) based on a failure to obtain a response from the at least one external host device (221) to the first signal.

[0153] The method may include an operation of acquiring information for connection with the at least one external host device (221) disposed in the at least one second availability area (204) through the communication circuit (233). The method may include an operation of transmitting the first signal to the at least one external host device (221) based on the acquired information.

[0154] The above information may include information for connection with at least one virtual machine (271, 275, 279) running on the at least one external host device (221). The method may include an operation of transmitting the first signal to the at least one virtual machine (271, 275, 279) based on the information for connection with the at least one virtual machine (271, 275, 279).

[0155] The method may include an operation of transmitting a signal to the electronic device (101) to instruct a virtual machine to operate as a primary virtual machine for providing a service related to the database, based on obtaining the response from the at least one external host device (221) to the first signal.

[0156] The method may include an operation of acquiring information indicating the health status of the virtual machine operating as the primary virtual machine in a first cycle. The method may include an operation of transmitting the information acquired in the first cycle through the communication circuit (233) to a first external host device (221) on which a failover controller is running.

[0157] As described above, a non-transitory computer readable storage medium can store a program including instructions. The instructions, when executed by a processor (231) of a host device (211) including a communication circuit (233) and a memory (235) storing a database, can cause the host device (211) to transmit a first signal to at least one external host device (221) located in at least one second availability zone (204) other than a first availability zone (201) in which the host device (211) is located, via the communication circuit (233). The instructions, when executed by the processor (231), can cause the host device (211) to deactivate a connection related to the database with the electronic device (101) based on a failure to obtain a response to the first signal from the at least one external host device (221).

[0158] As described above, a non-transitory computer-readable recording medium can store a program including instructions. The instructions, when executed by a processor of a host device (221) including a communication circuit, can cause the host device (221) to identify a health status of a first virtual machine (245) that operates as a primary virtual machine for providing database-related services to a plurality of electronic devices (101). The first virtual machine (245) can be executed on a first external host device (211). The above instructions, when executed by the processor, may cause the host device (221) to, based on the information, in response to identifying that the health state of the first virtual machine (245) is a designated first state, transmit a signal to the second virtual machine (275) through the communication circuit (233) instructing a second virtual machine (275) running on a second external host device distinct from the first external host device (211) to operate as the primary virtual machine after a designated time.

[0159] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0160] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0161] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0162] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0163] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., by download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0164] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In the host device (211), Communication circuit (233), At least one processor (231) comprising a processing circuit, and A virtual machine (275) that operates as a primary virtual machine to provide database-related services and a memory (235) that stores instructions, wherein the instructions, when individually or collectively executed by at least one processor (231), cause the host device (211) to Through the above communication circuit (233), a first signal is transmitted to at least one external host device (221) located in at least one second availability zone (204) other than the first availability zone (201) where the host device (211) is located, Causing to disable the connection related to the database with the electronic device (101) based on the fact that no response is obtained from the at least one external host device (221) to the first signal. Host device.

2. In claim 1, The above instructions, when individually or collectively executed by the at least one processor (231), cause the host device (211) to: Through the above communication circuit (233), information for connection with at least one external host device (221) disposed in the at least one second available area (204) is obtained, Based on the information obtained above, causing the at least one external host device (221) to transmit the first signal, Host device.

3. In claim 2, The above information includes information for connection with at least one virtual machine (271, 275, 279) running on at least one external host device (221), The above instructions, when individually or collectively executed by the at least one processor (231), cause the host device (211) to: Based on the information for connection with at least one virtual machine (271, 275, 279), causing the at least one virtual machine (271, 275, 279) to transmit the first signal, Host device.

4. In any one of claims 1 to 3, Based on the acquisition of the response from the at least one external host device (221) to the first signal, the virtual machine (275) is maintained as a primary virtual machine. Host device.

5. In claim 4, The above instructions, when individually or collectively executed by the at least one processor (231), cause the host device (211) to: Information indicating the health status of the virtual machine (245) operating as the primary virtual machine is acquired in the first cycle, Causing the information acquired in the first cycle through the above communication circuit (233) to be transmitted to the failover controller (281). Host device.

6. In claim 5, The above instructions, when individually or collectively executed by the at least one processor (231), cause the host device (211) to: After transmitting the information acquired in the first cycle through the communication circuit (233) to the failover controller, another signal instructing termination of the virtual machine (245) is acquired from the server (265) through the communication circuit (233), In response to terminating the virtual machine (245) based on the other signal, causing the server (265) to transmit a response indicating the termination result of the virtual machine (245). Host device.

7. In the host device (221), communication circuit, At least one processor comprising a processing circuit; and A memory storing instructions, comprising one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor, cause the host device (221) to: Identifying the health status of a first virtual machine (245) that operates as a primary virtual machine for providing database-related services to multiple electronic devices (101), wherein the first virtual machine (245) is executed on a first external host device (211), Based on the above information, in response to identifying that the health state of the first virtual machine (245) is a designated first state, a signal is transmitted to the second virtual machine (275) through the communication circuit (233) to instruct a second virtual machine (275) running on a second external host device distinct from the first external host device (211) to operate as the primary virtual machine after a designated time. Host device.

8. In claim 7, The above instructions, when individually or collectively executed by the at least one processor, cause the host device (221) to: Through the above communication circuit (233), information indicating the health status of the first virtual machine (245) is obtained from the first external host device (211), Based on the information obtained above, causing the health status of the first virtual machine (245) to be identified, Host device.

9. In claim 7 or claim 8, The above instructions, when individually or collectively executed by the at least one processor, cause the host device (221) to: Through the above communication circuit (233), based on the information indicating the health status of the first virtual machine (245) not being acquired from the first external host device (211) for a specified other time, causing the health status of the first virtual machine (245) to be identified as the specified first state. Host device.

10. In any one of claims 7 to 9, The above instructions, when individually or collectively executed by the at least one processor, cause the host device (221) to: In response to identifying that the health state of the first virtual machine (245) is a designated first state, causing the first external host device (211) to transmit another signal requesting termination of the first virtual machine (245). Host device.

11. In claim 10, The above instructions, when individually or collectively executed by the at least one processor, cause the host device (221) to: After transmitting the other signal, based on the termination of the first virtual machine (245) being identified within a specified other time, transmitting the signal to the second virtual machine (275) to instruct the second virtual machine (275) to operate as the primary virtual machine, After transmitting the other signal, based on the termination of the first virtual machine (245) not being identified within the other specified time, causing the second virtual machine (275) to transmit the signal instructing the second virtual machine (275) to operate as the primary virtual machine after the other specified time. Host device.

12. In any one of claims 7 to 11, The above instructions, when individually or collectively executed by the at least one processor, cause the host device (221) to: Based on the fact that no response is obtained from the first external host device (211) to the other signal, a third external host device is identified that is located in a second availability zone (204) other than the first availability zone (201) where the first external host device (221) is located. Causing the third external host device to transmit a request to run a standby virtual machine for the primary virtual machine to the third external host device; Host device.

13. In any one of claims 7 to 12, The first external host device (211) and the second external host device are located in different availability areas. Host device.

14. In any one of claims 7 to 13, The above specified time includes the time allocated for the first virtual machine running on the first external host device (211) to disable the connection related to the database with the electronic device (101). Host device.

15. In any one of claims 7 to 14, The above-mentioned specified time includes a time allocated to identify connectivity of the first virtual machine running on the first external host device (211) with at least one external host device located in at least one second availability zone (204) other than the first availability zone (201) in which the first external host device (211) is located. Host device.

Citation Information

Patent Citations

  • Cloud platform virtual machine recovery method and computer equipment

    CN112732406A

  • Application continuous high availability solution

    US11036530B2

  • Managing primary region availability for implementing a failover from another primary region

    US11397652B2

  • Heartbeat monitoring of virtual machines for initiating failover operations in a data storage management system, including virtual machine distribution logic

    US20180095845A1

  • Server clustering in a computing-on-demand system

    US20230367682A1