Cloud server control method and apparatus, storage medium, and electronic device

CN115904642BActive Publication Date: 2026-08-28DOUYIN VISION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110957018.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-19
Publication Date
2026-08-28
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

对于手动处理的方式,需要耗费较多的人力和时间,对于内存错误的处理效率较低

Benefits of technology

[0017] The above technical solution allows for the determination of error types and total number of detected memory errors when they are detected. Based on these factors, the cloud scheduler can then select a corresponding scheduling service from a database containing a first and a second scheduling service. The first service restarts the virtual machine on its original host, while the second service migrates it to another host. Compared to directly restarting the virtual machine in related technologies, this approach addresses memory errors more specifically, reducing their impact on cloud computing services. Furthermore, scheduling services based on error types and total number of errors enables automated memory error handling, reducing manpower and time spent on error resolution and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115904642B_ABST
    Figure CN115904642B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a cloud service control method and device, a storage medium and an electronic device, to specifically handle memory errors and improve the processing efficiency of memory errors. The method comprises: when a memory error is detected, determining a memory address at which the memory error occurs, and determining an error type of the memory error according to the memory address, the error type being used to identify whether the memory error will cause a virtual machine to crash; determining a total error number of the detected memory error; and according to the error type of the memory error and the total error number, controlling a cloud scheduler to schedule a corresponding scheduling service from a scheduling service library, the scheduling service library storing a first scheduling service and a second scheduling service, the first scheduling service being used to control restarting a virtual machine on an original host to which the virtual machine belongs, and the second scheduling service being used to control migrating the virtual machine from the original host to which the virtual machine belongs to another host.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of cloud computing technology, and more specifically, to a cloud service control method, apparatus, storage medium, and electronic device. Background Technology

[0002] With the rise of cloud computing services, a large number of enterprise and personal services are deployed on cloud servers, making effective management of cloud servers extremely important. Memory errors (MCE, Machine Check Error), as a common hardware failure, can affect the normal operation of virtual machines to varying degrees. Some memory errors can cause the virtual machine monitor (hypervisor) to stop running, while others can cause the virtual machine to restart.

[0003] In related technologies, after a memory error is detected, operations and maintenance personnel need to log in to the machine to check the specific cause and manually handle it, or directly restart the virtual machine. Manual handling requires significant manpower and time, and is inefficient for dealing with memory errors. Directly restarting the virtual machine cannot address the specific memory error, thus impacting the normal operation of cloud computing services. Summary of the Invention

[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, this disclosure provides a cloud service control method, the method comprising:

[0006] When a memory error is detected, the memory address where the memory error occurred is determined, and the error type of the memory error is determined based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash.

[0007] Determine the total number of memory errors detected;

[0008] Based on the error type of the memory error and the total number of errors, the cloud scheduler is controlled to schedule the corresponding scheduling service from the scheduling service library. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from its original host machine to another host machine.

[0009] Secondly, this disclosure provides a cloud service control device, the device comprising:

[0010] The first determining module is used to determine the memory address where the memory error occurred when a memory error is detected, and to determine the error type of the memory error based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash.

[0011] The second determination module is used to determine the total number of memory errors detected;

[0012] The third determining module is used to control the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library according to the error type of the memory error and the total number of errors. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from the original host machine to another host machine.

[0013] Thirdly, this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in the first aspect.

[0014] Fourthly, this disclosure provides an electronic device, comprising:

[0015] A storage device on which computer programs are stored;

[0016] A processing device for executing the computer program in the storage device to implement the steps of the method in the first aspect.

[0017] The above technical solution allows for the determination of error types and total number of detected memory errors when they are detected. Based on these factors, the cloud scheduler can then select a corresponding scheduling service from a database containing a first and a second scheduling service. The first service restarts the virtual machine on its original host, while the second service migrates it to another host. Compared to directly restarting the virtual machine in related technologies, this approach addresses memory errors more specifically, reducing their impact on cloud computing services. Furthermore, scheduling services based on error types and total number of errors enables automated memory error handling, reducing manpower and time spent on error resolution and improving efficiency.

[0018] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0020] Figure 1 This is a flowchart illustrating a cloud service control method according to an exemplary embodiment of the present disclosure;

[0021] Figure 2 This is a flowchart illustrating a cloud service control method according to another exemplary embodiment of the present disclosure;

[0022] Figure 3 This is a block diagram illustrating a cloud service control device according to an exemplary embodiment of the present disclosure;

[0023] Figure 4 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifiers "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless explicitly stated otherwise in the context, they should be understood as "one or more".

[0028] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0029] As mentioned in the background section, after detecting a memory error, the relevant technology requires operations and maintenance personnel to log in to the machine to check the specific cause and manually handle it, or directly restart the virtual machine. Manual handling is time-consuming and labor-intensive, and its efficiency in dealing with memory errors is low. Restarting the virtual machine directly cannot address the specific memory error, thus impacting the normal operation of cloud computing services.

[0030] In view of this, this disclosure provides a cloud service control method to automatically schedule services to handle memory errors based on the error type and total number of errors, thereby reducing the manpower and time spent in memory error handling and improving memory error handling efficiency.

[0031] Figure 1 This is a flowchart illustrating a cloud service control method according to an exemplary embodiment of this disclosure. (Refer to...) Figure 1 The method includes:

[0032] Step 101: When a memory error is detected, determine the memory address where the memory error occurred, and determine the error type based on that memory address. The error type indicates whether the memory error will cause the virtual machine to crash.

[0033] Step 102: Determine the total number of memory errors detected.

[0034] Step 103: Based on the error type and total number of memory errors, control the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on its original host machine, and the second scheduling service is used to control the migration of the virtual machine from its original host machine to another host machine.

[0035] Compared to directly restarting virtual machines, the above method allows for targeted handling of memory errors, thereby reducing their impact on cloud computing services. Furthermore, by scheduling services based on the error type and total number of errors, memory error handling can be automated, reducing manpower and time spent on error resolution and improving efficiency.

[0036] To enable those skilled in the art to better understand the cloud service control method provided in this disclosure, the above steps are illustrated in detail below.

[0037] It should be understood that the cloud service control method provided in this disclosure can be applied to cloud computing scenarios. For example, steps 101 to 103 can be executed through a server running cloud computing services. Alternatively, to be further subdivided, the server runs virtual machines, a virtual machine monitor, and a cloud scheduler, where the virtual machine monitor is an intermediate software layer between the server and the virtual machines, allowing multiple virtual machines to share the server. In this scenario, the server can execute step 101 through the virtual machine monitor to determine the error type of the memory error. Then, the virtual machine monitor reports the detailed information and error type of the memory error to the cloud scheduler, which then executes steps 102 and 103.

[0038] In one possible approach, prior to step 101, if a memory error occurs, a memory error notification can be sent to the virtual machine monitor via the cloud server's operating system kernel. Upon receiving the memory error notification, the virtual machine monitor confirms that a memory error has been detected. Afterward, steps 101 through 103 can be executed.

[0039] In other words, when a memory error occurs, the server's operating system kernel will notify the virtual machine monitor. When the virtual machine monitor receives the kernel's notification, it will determine that a memory error has been detected.

[0040] It should be understood that in a cloud computing scenario, server memory is divided into two parts: one part corresponds to the virtual machine, and the other part corresponds to the virtual machine monitor. Therefore, after detecting a memory error, it is possible to further determine the memory address where the memory error occurred, and thus determine the error type based on that memory address.

[0041] For example, in cloud computing services, some memory errors will cause the virtual machine monitor to stop running, but the virtual machine can continue to run, while other memory errors will cause the virtual machine process to crash, requiring a restart of the virtual machine. Therefore, error types can include crash errors or non-crash errors. Crash errors are memory errors that can cause the virtual machine to crash, while non-crash errors are memory errors that allow the virtual machine to continue running.

[0042] Among the possible methods, the error type of a memory error can be determined as follows: if the memory address falls within the memory address range corresponding to the virtual machine monitor, determine whether the memory error can be repaired by memory erasure coding. If the memory error can be repaired by memory erasure coding, then the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. If the memory error cannot be repaired by memory erasure coding, then the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

[0043] For example, memory erasure coding can be used to repair memory errors caused by an anomaly in a single bit of memory. If a memory error can be repaired by memory erasure coding, it indicates that the memory error was caused by an anomaly in a single bit of memory and will not cause the virtual machine process to crash, thus confirming that the memory error is a non-crash error. Conversely, if a memory error cannot be repaired by memory erasure coding, it indicates that the memory error may be caused by anomalies in multiple bits of memory, which may ultimately cause the virtual machine process to crash, thus confirming that the memory error is a crash error.

[0044] In another possible approach, if the memory address falls within the memory address range corresponding to the virtual machine, it can be determined whether a memory error can be injected into the virtual machine. If the memory error is successfully injected into the virtual machine, the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. If the memory error is not successfully injected into the virtual machine, the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

[0045] For example, if the memory address where the memory error occurred belongs to a virtual machine, the memory error might be caused by an application running on the virtual machine. To further determine whether the memory error would cause the virtual machine process to crash, the memory error can be attempted to be injected into the virtual machine. If the injection is successful, it means the virtual machine can automatically repair the memory error, thus confirming that the memory error is not a crash error. Conversely, if the injection fails, it means the virtual machine cannot automatically repair the memory error, thus confirming that the memory error is a crash error.

[0046] After determining the type of memory error, the total number of memory errors detected by the virtual machine monitor can be determined. Then, by combining the type of memory error with the total number of errors, the cloud scheduler can be controlled to schedule the corresponding scheduling service from the scheduling service library.

[0047] In one possible approach, if the memory error type indicates that the memory error will not cause the virtual machine to crash, then when the total number of errors reaches a preset threshold, the cloud scheduler can be controlled to schedule a second scheduling service from the scheduling service library. This second scheduling service is used to control the migration of the virtual machine from its original host to another host.

[0048] For example, the preset threshold can be set according to actual conditions, and this embodiment does not limit this. The second scheduling service is used to control the migration of virtual machines from their original host to another host, such as hot migration or cold migration of virtual machines from their original host to another host depending on whether the virtual machine has crashed. Hot migration completely saves the running state of the virtual machine during its operation and quickly restores it to the original hardware platform or a different hardware platform. After restoration, the virtual machine continues to run smoothly, and the user will not perceive any difference. Cold migration refers to migrating a virtual machine to another host while it is in a shutdown state.

[0049] In this embodiment, if the memory error type indicates that the memory error will not cause the virtual machine to crash (i.e., the memory error is a non-crash error), the virtual machine can continue to run. Further, it can be determined whether the total number of errors has reached a preset threshold. If the total number of errors has not reached the preset threshold, the total number of errors continues to be recorded. If the total number of errors reaches the preset threshold, it indicates that the host machine to which the virtual machine belongs is in an unhealthy state and is not suitable for the normal operation of the virtual machine. Therefore, a second scheduling service can be scheduled through the cloud scheduler, that is, the virtual machine can be migrated to another healthy host machine. In this case, since the virtual machine has not crashed, the running state of the virtual machine can be hot-migrated to another healthy host machine while the virtual machine is running. The host machine to which the virtual machine belongs can be understood as the server mentioned above that executes the method of this disclosure, and the other host machine can be understood as another server, distinct from this server, that can ensure the normal operation of the virtual machine. In this way, the normal operation of the virtual machine can be guaranteed, thereby ensuring the normal operation of the cloud computing service.

[0050] In other possible ways, if the error type of the memory error indicates that the memory error will cause the virtual machine to crash, then the cloud scheduler can be controlled to schedule the first scheduling service from the scheduling service library if the total number of errors has not reached the preset threshold, or the cloud scheduler can be controlled to schedule the second scheduling service from the scheduling service library if the total number of errors has reached the preset threshold.

[0051] For example, if a memory error is a crash error, it means that the memory error will cause the virtual machine process to crash. In this case, if the total number of errors does not reach the preset threshold, it means that the host machine to which the virtual machine belongs is still in a healthy state. The host machine can ensure the normal operation of the virtual machine, and thus the cloud scheduler can be controlled to schedule the first scheduling service from the scheduling service library, that is, to restart the virtual machine on the original host machine. If the total number of errors reaches the preset threshold, it means that the host machine to which the virtual machine belongs is in an unhealthy state. The host machine cannot ensure the normal operation of the virtual machine, so the cloud scheduler can be controlled to schedule the second scheduling service from the scheduling service library, that is, to migrate the virtual machine to another host machine. Furthermore, in this case, since the virtual machine process crashes, a cold migration method can be used.

[0052] In one possible way, controlling the cloud scheduler to schedule the first scheduling service from the scheduling service library could be: determining whether the memory address where the memory error occurred belongs to a big page memory address; if the memory address does not belong to a big page memory address, then controlling the cloud scheduler to schedule the first scheduling service from the scheduling service library; or if the memory address belongs to a big page memory address, then allocating big page memory to the original host machine to which the virtual machine belongs, and after allocating big page memory, controlling the cloud scheduler to schedule the first scheduling service from the scheduling service library.

[0053] It should be understood that if a memory error occurs in the massive page memory, the capacity of the massive page memory will decrease. In this case, directly scheduling the first scheduling service and restarting the virtual machine on the original host machine will result in insufficient memory. Therefore, in this embodiment of the disclosure, it is first determined whether the memory address where the memory error occurred belongs to the massive page memory address.

[0054] If the memory address does not belong to a big page memory address, the first scheduling service can be scheduled to restart the virtual machine on its original host machine. If the memory address belongs to a big page memory address, big page memory can be allocated to the original host machine first, meaning the big page memory capacity of the original host machine is consistent with the big page memory capacity before the memory error occurred. Then, the first scheduling service can be scheduled on the original host machine after allocating big page memory to restart the virtual machine. This avoids the problem of insufficient memory during the virtual machine restart process, ensuring the normal operation of the virtual machine after restarting.

[0055] The cloud service control method provided in this disclosure will now be described through another exemplary embodiment. (Refer to...) Figure 2 The cloud service control methods include:

[0056] Step 201: Determine that the virtual machine monitor has detected a memory error.

[0057] Step 202: If the memory address falls within the memory address range corresponding to the virtual machine monitor, determine whether the memory error can be repaired by memory erasure coding. If yes, proceed to step 203; otherwise, proceed to step 204.

[0058] Step 203: Determine that the memory error is not a crash error, and proceed to step 206.

[0059] Step 204: Determine the memory error as a crash error and proceed to step 210.

[0060] Step 205: If the memory address belongs to the memory address range corresponding to the virtual machine, determine whether a memory error can be injected into the virtual machine. If yes, proceed to step 203; otherwise, proceed to step 204.

[0061] Step 206: Determine whether the total number of memory errors has reached the preset threshold. If yes, proceed to step 207; otherwise, proceed to step 208.

[0062] Step 207: Use the cloud scheduler to schedule the second scheduling service to hot migrate the virtual machine to another host.

[0063] Step 208: Continue recording the total number of memory errors.

[0064] Step 209: Determine whether the total number of memory errors has reached the preset threshold. If yes, proceed to step 210; otherwise, proceed to step 211.

[0065] Step 210: Determine whether the memory address where the memory error occurred belongs to a big page memory address. If so, proceed to step 212; otherwise, proceed to step 213.

[0066] Step 211: The second scheduling service is scheduled through the cloud scheduler to cold migrate the virtual machine to another host.

[0067] Step 212: Allocate large page memory to the original host machine to which the virtual machine belongs, and schedule the first service after allocating the large page memory, and restart the virtual machine on the original host machine to which the virtual machine belongs.

[0068] Step 213: Schedule the first service to restart the virtual machine on its original host machine.

[0069] The specific implementation methods for each of the above steps have been described in detail above and will not be repeated here. It should also be understood that, for the sake of simplicity, the above method embodiments are described as a series of actions; however, those skilled in the art should understand that this disclosure is not limited to the order of actions described above. Furthermore, those skilled in the art should also understand that the embodiments described above are preferred embodiments, and the steps involved are not necessarily essential to this disclosure.

[0070] It should be understood that in the management of large-scale cloud servers, manual handling of memory errors is inefficient and untimely, and directly restarting virtual machines cannot effectively prevent memory errors. Therefore, this disclosure proposes to automatically schedule corresponding scheduling services for different types of memory errors. While resolving memory errors in a timely and appropriate manner, it records the number of memory errors that occur on the host machine to monitor the host machine's hardware health status, and can migrate virtual machines as early as possible to prevent impact on business operations.

[0071] Based on the same concept, this disclosure also provides a cloud service control device, which can be part or all of an electronic device through software, hardware, or a combination of both. (See also...) Figure 3 The cloud service control device 300 may include:

[0072] The first determining module 301 is used to determine the memory address where the memory error occurred when a memory error is detected, and to determine the error type of the memory error based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash.

[0073] The second determining module 302 is used to determine the total number of detected memory errors;

[0074] The third determining module 303 is used to control the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library according to the error type of the memory error and the total number of errors. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from the original host machine to another host machine.

[0075] Optionally, the third determining module 303 is used to:

[0076] When the error type of the memory error indicates that the memory error will not cause the virtual machine to crash, the cloud scheduler is controlled to schedule the second scheduling service from the scheduling service library when the total number of errors reaches a preset threshold.

[0077] Optionally, the third determining module 303 is used to:

[0078] When the error type of the memory error indicates that the memory error will cause the virtual machine to crash, if the total number of errors has not reached the preset threshold, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library, or if the total number of errors has reached the preset threshold, the cloud scheduler is controlled to schedule the second scheduling service from the scheduling service library.

[0079] Optionally, the third determining module 303 is used to:

[0080] Determine whether the memory address where the memory error occurred belongs to a big page memory address;

[0081] When the memory address does not belong to a big page memory address, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library. Alternatively, when the memory address belongs to a big page memory address, big page memory is allocated to the original host machine to which the virtual machine belongs, and after allocating the big page memory, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library.

[0082] Optionally, the first determining module 301 is used to:

[0083] If the memory address falls within the memory address range corresponding to the virtual machine monitor, determine whether the memory error can be repaired using memory erasure coding.

[0084] When the memory error can be repaired by the memory erasure coding, the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. When the memory error cannot be repaired by the memory erasure coding, the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

[0085] Optionally, the first determining module 301 is used to:

[0086] If the memory address belongs to the memory address range corresponding to the virtual machine, determine whether the memory error can be injected into the virtual machine;

[0087] When the memory error is successfully injected into the virtual machine, the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. When the memory error is not successfully injected into the virtual machine, the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

[0088] Optionally, the device 300 further includes:

[0089] The sending module is used to send memory error notifications to the virtual machine monitor via the cloud server's operating system kernel;

[0090] The fourth determining module is used to determine that a memory error has been detected when the virtual machine monitor receives the memory error notification.

[0091] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0092] Based on the same concept, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of any of the cloud service control methods described above.

[0093] Based on the same concept, this disclosure also provides an electronic device, including:

[0094] A storage device on which computer programs are stored;

[0095] A processing device for executing the computer program in the storage device to implement the steps of any of the cloud service control methods described above.

[0096] The following is for reference. Figure 4 This diagram illustrates a structural schematic of an electronic device 400 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0097] like Figure 4 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0098] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0099] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.

[0100] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0101] In some implementations, communication can be conducted using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0102] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0103] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following to occur: upon detecting a memory error, determine the memory address where the memory error occurred, and determine the error type of the memory error based on the memory address, wherein the error type is used to identify whether the memory error will cause the virtual machine to crash; determine the total number of detected memory errors; and, based on the error type of the memory error and the total number of errors, control the cloud scheduler to schedule a corresponding scheduling service from the scheduling service library, wherein the scheduling service library stores a first scheduling service and a second scheduling service, wherein the first scheduling service is used to control the restart of the virtual machine on its original host machine, and the second scheduling service is used to control the migration of the virtual machine from its original host machine to another host machine.

[0104] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0106] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.

[0107] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0108] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0109] According to one or more embodiments of this disclosure, Example 1 provides a cloud service control method, including:

[0110] When a memory error is detected, the memory address where the memory error occurred is determined, and the error type of the memory error is determined based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash.

[0111] Determine the total number of memory errors detected;

[0112] Based on the error type of the memory error and the total number of errors, the cloud scheduler is controlled to schedule the corresponding scheduling service from the scheduling service library. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from its original host machine to another host machine.

[0113] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein controlling the cloud scheduler to schedule a corresponding scheduling service from the scheduling service library based on the error type of the memory error and the total number of errors includes:

[0114] If the error type of the memory error indicates that the memory error will not cause the virtual machine to crash, then when the total number of errors reaches a preset threshold, the cloud scheduler is controlled to schedule the second scheduling service from the scheduling service library.

[0115] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, wherein controlling the cloud scheduler to schedule a corresponding scheduling service from the scheduling service library based on the error type of the memory error and the total number of errors includes:

[0116] If the memory error type indicates that the memory error will cause the virtual machine to crash, then if the total number of errors does not reach the preset threshold, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library; or if the total number of errors reaches the preset threshold, the cloud scheduler is controlled to schedule the second scheduling service from the scheduling service library.

[0117] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, wherein controlling the cloud scheduler to schedule the first scheduling service from the scheduling service library includes:

[0118] Determine whether the memory address where the memory error occurred belongs to a big page memory address;

[0119] If the memory address does not belong to a big page memory address, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library; or, if the memory address belongs to a big page memory address, big page memory is allocated to the original host machine to which the virtual machine belongs, and after allocating the big page memory, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library.

[0120] According to one or more embodiments of this disclosure, Example 5 provides a method of any one of Examples 1-4, wherein determining the error type of the memory error based on the memory address includes:

[0121] If the memory address falls within the memory address range corresponding to the virtual machine monitor, determine whether the memory error can be repaired using memory erasure coding.

[0122] If the memory error can be repaired by the memory erasure coding, then the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. If the memory error cannot be repaired by the memory erasure coding, then the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

[0123] According to one or more embodiments of this disclosure, Example 6 provides a method of any one of Examples 1-4, wherein determining the error type of the memory error based on the memory address includes:

[0124] If the memory address belongs to the memory address range corresponding to the virtual machine, determine whether the memory error can be injected into the virtual machine;

[0125] If the memory error is successfully injected into the virtual machine, the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. If the memory error is not successfully injected into the virtual machine, the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

[0126] According to one or more embodiments of this disclosure, Example 7 provides a method of any one of Examples 1-4, the method further comprising:

[0127] Send memory error notifications to the virtual machine monitor via the cloud server's operating system kernel;

[0128] When the virtual machine monitor receives the memory error notification, it determines that a memory error has been detected.

[0129] According to one or more embodiments of this disclosure, Example 8 provides a cloud service control device, the device comprising:

[0130] The first determining module is used to determine the memory address where the memory error occurred when a memory error is detected, and to determine the error type of the memory error based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash.

[0131] The second determination module is used to determine the total number of memory errors detected;

[0132] The third determining module is used to control the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library according to the error type of the memory error and the total number of errors. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from the original host machine to another host machine.

[0133] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-7.

[0134] According to one or more embodiments of this disclosure, Example 10 provides an electronic device, including:

[0135] A storage device on which computer programs are stored;

[0136] A processing device for executing the computer program in the storage device to implement the steps of any one of the methods in Examples 1-7.

[0137] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0138] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0139] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A cloud service control method, characterized in that, The method includes: When a memory error is detected, the memory address where the memory error occurred is determined, and the error type of the memory error is determined based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash; wherein, the error type includes crash error or non-crash error, the crash error indicates that the memory error will cause the virtual machine to crash; the non-crash error indicates that the memory error will not cause the virtual machine to crash; Determine the total number of memory errors detected; Based on the error type of the memory error and the total number of errors, the cloud scheduler is controlled to schedule the corresponding scheduling service from the scheduling service library. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from its original host machine to another host machine.

2. The method according to claim 1, characterized in that, The step of controlling the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library based on the error type of the memory error and the total number of errors includes: If the error type of the memory error indicates that the memory error will not cause the virtual machine to crash, then when the total number of errors reaches a preset threshold, the cloud scheduler is controlled to schedule the second scheduling service from the scheduling service library.

3. The method according to claim 1, characterized in that, The step of controlling the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library based on the error type of the memory error and the total number of errors includes: If the memory error type indicates that the memory error will cause the virtual machine to crash, then if the total number of errors does not reach a preset threshold, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library; or if the total number of errors reaches the preset threshold, the cloud scheduler is controlled to schedule the second scheduling service from the scheduling service library.

4. The method according to claim 3, characterized in that, The control of the cloud scheduler to schedule the first scheduling service from the scheduling service library includes: Determine whether the memory address where the memory error occurred belongs to a big page memory address; If the memory address does not belong to a big page memory address, then control the cloud scheduler to schedule the first scheduling service from the scheduling service library; or If the memory address is a big page memory address, then big page memory is allocated to the original host machine to which the virtual machine belongs, and after allocating the big page memory, the cloud scheduler is controlled to schedule the first scheduling service from the scheduling service library.

5. The method according to any one of claims 1-4, characterized in that, Determining the error type of the memory error based on the memory address includes: If the memory address falls within the memory address range corresponding to the virtual machine monitor, determine whether the memory error can be repaired using memory erasure coding. If the memory error can be repaired by the memory erasure coding, then the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. If the memory error cannot be repaired by the memory erasure coding, then the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

6. The method according to any one of claims 1-4, characterized in that, Determining the error type of the memory error based on the memory address includes: If the memory address belongs to the memory address range corresponding to the virtual machine, determine whether the memory error can be injected into the virtual machine; If the memory error is successfully injected into the virtual machine, the error type of the memory error is determined to be a non-crash error that will not cause the virtual machine to crash. If the memory error is not successfully injected into the virtual machine, the error type of the memory error is determined to be a crash error that will cause the virtual machine to crash.

7. The method according to any one of claims 1-4, characterized in that, The method further includes: Send memory error notifications to the virtual machine monitor via the cloud server's operating system kernel; When the virtual machine monitor receives the memory error notification, it determines that a memory error has been detected.

8. A cloud service control device, characterized in that, The device includes: The first determining module is used to determine the memory address where the memory error occurred when a memory error is detected, and to determine the error type of the memory error based on the memory address. The error type is used to identify whether the memory error will cause the virtual machine to crash. The error type includes a crash error or a non-crash error. A crash error indicates that the memory error will cause the virtual machine to crash. A non-crash error indicates that the memory error will not cause the virtual machine to crash. The second determination module is used to determine the total number of memory errors detected; The third determining module is used to control the cloud scheduler to schedule the corresponding scheduling service from the scheduling service library according to the error type of the memory error and the total number of errors. The scheduling service library stores a first scheduling service and a second scheduling service. The first scheduling service is used to control the restart of the virtual machine on the original host machine to which the virtual machine belongs, and the second scheduling service is used to control the migration of the virtual machine from the original host machine to another host machine.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method described in any one of claims 1-7.

10. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Memory error processing method and device and server

    CN111625387A