Service instance management method based on cloud technology, cloud management platform, and cluster

Receiving event information through the cloud management platform and migrating service instances in advance, solving the impact of resource node operations on service instance QoS, and achieving the guarantee of service quality and improvement of user experience.

WO2025140269A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/142182
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2024-12-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In cloud computing, the separation of management of resource nodes and management of service instances causes service instances to be unable to perceive the operation of resource nodes, resulting in the impact of quality of service (QoS). Especially during migration operations, the QoS of service instances is lower than the requirements.

Method used

The cloud management platform receives event information sent by the physical machine, predicts the execution time of the operation, and migrates the service instance to a healthy physical machine resource node before the operation to avoid QoS being affected.

Benefits of technology

It effectively avoids the QoS of service instances being affected by resource node operations, ensures that the service quality is not lower than the requirements, improves user experience, and saves computing resources of the cloud management platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024142182_03072025_PF_FP_ABST
    Figure CN2024142182_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a service instance management method based on cloud technology, a cloud management platform, and a cluster. The method comprises: a cloud management platform receiving event information sent by a first physical machine among a plurality of physical machines, wherein the event information comprises an expected execution time of an operation, and the operation affects the state of a first resource node of the first physical machine; and before the expected execution time, the cloud management platform migrating a service instance in the first resource node to a second resource node of a second physical machine among the plurality of physical machines on the basis of a QoS requirement of the service instance in the first resource node. By means of the method, the QoS of a service instance can be prevented from being affected by an operation for a resource node where the service instance is located.
Need to check novelty before this filing date? Find Prior Art

Description

Cloud technology-based service instance management method, cloud management platform and cluster

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 25, 2023, with application number 202311797508.4 and application name “A Fault Management System”, and the Chinese patent application filed with the State Intellectual Property Office of China on March 6, 2024, with application number 202410256871.3 and application name “Service Instance Management Method, Cloud Management Platform and Cluster Based on Cloud Technology”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of cloud computing, and in particular to a cloud technology service instance management method, a cloud management platform, and a cluster. Background Art

[0003] In cloud computing, virtualization technology is used to create service instances such as containers on resource nodes, such as physical machines or virtual machines (VMs). Resource node management and service instance management are separate. For example, VMs, which serve as resource nodes, are managed by the Infrastructure as a Service (IaaS) layer, while containers, which serve as service instances, are managed by the Kubernetes (K8s) control plane.

[0004] Because resource node management and service instance management are separate, service instance management is unaware of operations targeting resource nodes. This causes service instances in resource nodes to passively adapt to resource node operations. For example, when a resource node is migrated, the service instances in that resource node passively migrate along with the resource node. This can affect the service instance's quality of service (QoS), causing the QoS of the service instance to fall below the required QoS. Summary of the Invention

[0005] The present application provides a cloud-based service instance management method, a cloud management platform, and a cluster, which can prevent the QoS of a service instance from being affected by the operation of the resource node where the service instance is located.

[0006] In a first aspect, a cloud-based service instance management method is provided. The method is applied to a cloud management platform connected to a cloud infrastructure, the cloud management platform being used to manage multiple service instances deployed in the cloud infrastructure, the cloud infrastructure comprising multiple physical machines, each of the multiple physical machines being used to deploy one or more resource nodes, each of the multiple service instances being deployed in at least one resource node in the cloud infrastructure, and the quality of service (QoS) of the service instance being affected by the state of the resource node where the service instance is located; the method comprising: the cloud management platform receiving event information emitted by a first physical machine among the multiple physical machines, the event information including an expected execution time of an operation, wherein the operation affects the state of a first resource node of the first physical machine; and before the expected execution time, the cloud management platform migrating the service instance in the first resource node to a second resource node of a second physical machine among the multiple physical machines based on the QoS requirements of the service instance in the first resource node. The second physical machine is a healthy physical machine, wherein a healthy physical machine refers to a physical machine that has not experienced any failures or undergone any changes.

[0007] When a physical machine is about to perform an operation that affects the resource node of the physical machine, it can issue an event message. The event message includes the expected execution time of the operation. The cloud management platform can receive the event message and obtain the expected execution time of the operation from the event message. The cloud management platform can migrate the service instance to the resource node of another physical machine based on the QoS requirements of the service instance in the resource node. In this way, the execution of the operation will not affect the QoS of the service instance, that is, the QoS of the service instance is avoided from being affected by the operation of the resource node where the service instance is located.

[0008] In one possible implementation, the event information also includes the type of operation; the cloud management platform migrates the service instance in the first resource node to the second resource node of the second physical machine among the multiple physical machines based on the QoS requirements of the service instance in the first resource node, including: the cloud management platform confirms that the operation will cause the QoS of the service instance in the first resource node to be lower than the QoS requirements based on the type of operation; the cloud management platform migrates the service instance in the first resource node to the second resource node.

[0009] When it is confirmed that the operation on the resource node will cause the QoS of the service instance to be lower than the QoS requirement, the service instance is migrated to the resource node of another physical machine. In this way, when the operation is performed, the QoS of the service instance is guaranteed to be no lower than the QoS requirement, thereby improving the user's service experience.

[0010] In one possible implementation, the cloud management platform migrates the service instance in the first resource node to the second resource node of the second physical machine among multiple physical machines based on the QoS requirement of the service instance in the first resource node, including: the cloud management platform confirms that the QoS requirement is greater than or equal to a preset QoS threshold; the cloud management platform migrates the service instance in the first resource node to the second resource node.

[0011] This QoS threshold is the threshold for judging whether the service instance can tolerate operations that affect resource nodes. If the QoS requirement of the service instance is greater than or equal to the QoS threshold, it means that the QoS requirement of service instance A1 cannot tolerate any operations that affect resource nodes. In this case, operations that affect resource nodes will cause the QoS of the service instance to be no lower than the QoS requirement by default. Therefore, there is no need to judge whether the operations that affect resource nodes will cause the QoS of the service instance to be no lower than the QoS requirement, and the service instance can be directly migrated to the resource node of another physical machine. In this way, while ensuring that the QoS of the service instance is no lower than the QoS requirement and improving the user's service experience, it can also save the computing resources of the cloud management platform.

[0012] In a possible implementation, the method further includes: the cloud management platform prohibiting deployment of new service instances in the first resource node.

[0013] The cloud management platform can confirm through the event information that an operation affecting the first resource node is about to be performed on the first resource node. In this case, the cloud management platform does not deploy a new service instance to the first resource node to avoid the impact of the operation affecting the first resource node on the QoS of the newly deployed service instance.

[0014] In one possible implementation, the operation is used to respond to a failure or change in hardware, where the hardware is used for a first physical machine to run a first resource node; or, the operation is used to respond to a failure or change in software, where the software includes at least one of an operating system of the first physical machine and management software of the first resource node.

[0015] In other words, a failure or change in the hardware or software associated with a resource node on a physical machine triggers the physical machine to perform an action on the resource node to address the failure or change, and issues an event message. The cloud management platform can use this event message to migrate services on the resource node to resource nodes on other physical machines before the action is executed, thereby minimizing the impact of the action on the resource node on the QoS of the service instance.

[0016] In a possible implementation, the service instance in the first resource node includes any one or more of a container, a database instance, and a data warehouse instance.

[0017] In a possible implementation, the first resource node is a virtual machine or a bare metal server, or the first resource node is the first physical machine itself.

[0018] On the second aspect, a cloud management platform is provided, which is connected to a cloud infrastructure. The cloud management platform is used to manage multiple service instances deployed in the cloud infrastructure. The cloud infrastructure includes multiple physical machines, each of the multiple physical machines is used to deploy one or more resource nodes, and each of the multiple service instances is deployed in at least one resource node in the cloud infrastructure. The service quality (QoS) of the service instance is affected by the state of the resource node where the service instance is located; the cloud management platform includes: a receiving module for receiving event information issued by a first physical machine among the multiple physical machines, the event information including an expected execution time of an operation, wherein the operation affects the state of a first resource node of the first physical machine; a migration module for migrating the service instance in the first resource node to a second resource node of a second physical machine among the multiple physical machines before the expected execution time based on the QoS requirements of the service instance in the first resource node.

[0019] In one possible implementation, the event information also includes the type of operation; the migration module is used to: based on the type of operation, confirm that the operation will cause the QoS of the service instance in the first resource node to be lower than the QoS requirement; and migrate the service instance in the first resource node to the second resource node.

[0020] In a possible implementation, the migration module is configured to: confirm that the QoS requirement is greater than or equal to a preset QoS threshold; and migrate the service instance in the first resource node to the second resource node.

[0021] In a possible implementation, the cloud management platform further includes: a prohibition module, configured to prohibit deployment of new service instances in the first resource node.

[0022] In one possible implementation, the operation is used to respond to a failure or change in hardware, where the hardware is used for a first physical machine to run a first resource node; or, the operation is used to respond to a failure or change in software, where the software includes at least one of an operating system of the first physical machine and management software of the first resource node.

[0023] In a possible implementation, the service instance in the first resource node includes any one or more of a container, a database instance, and a data warehouse instance.

[0024] In a possible implementation, the first resource node is a virtual machine or a bare metal server, or the first resource node is the first physical machine itself.

[0025] In a third aspect, a computing device cluster is provided, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method provided in the first aspect.

[0026] In a fourth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method provided in the first aspect.

[0027] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computer device cluster, the computer device cluster executes the method provided in the first aspect.

[0028] The beneficial effects of the second to fifth aspects can be referred to the above introduction to the beneficial effects of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] FIG1 is a schematic diagram of a system architecture provided in an embodiment of the present application;

[0030] FIG2 is a schematic diagram of the structure of an abnormality perception example provided in an embodiment of the present application;

[0031] FIG3 is a schematic diagram of the structure of an event generation example provided in an embodiment of the present application;

[0032] FIG4 is a functional diagram of an event sensing channel provided in an embodiment of the present application;

[0033] FIG5 is a schematic diagram of the structure of an event sensing example provided in an embodiment of the present application;

[0034] FIG6 is a flowchart of a service instance management method provided in an embodiment of the present application;

[0035] FIG7 is a schematic diagram of a service instance management method provided in an embodiment of the present application;

[0036] FIG8 is a schematic diagram of a service instance management method provided in an embodiment of the present application;

[0037] FIG9 is a schematic diagram of the structure of a cloud management platform provided in an embodiment of the present application;

[0038] FIG10 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0039] FIG11 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0040] FIG12 is a schematic diagram of a structure of a computing device cluster connected via a network provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The following describes the solutions provided by the embodiments of the present application in conjunction with the accompanying drawings. In the embodiments of the present application, "plurality" refers to two or more, and "multiple" refers to two or more. Terms such as "first" and "second" are used only to distinguish similar objects and do not necessarily describe a specific order or quantity of objects.

[0042] To facilitate understanding of the solutions provided by the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application are first introduced.

[0043] Cloud technology refers to a hosting service that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0044] Cloud infrastructure, also known as infrastructure, is the facility that supports cloud computing services and includes at least one data center, each of which includes multiple physical machines (e.g., servers). Physical machines can serve as resource nodes or be deployed with resource nodes. Service instances such as containers, database instances, and data warehouse service instances can be run within these resource nodes to implement elastic cloud computing services. For example, if the cloud infrastructure includes multiple data centers, these data centers can be distributed across different geographic regions, with remote connections between the data centers achieved via a backbone network.

[0045] Resource nodes: Nodes with computing and storage resources that can be used to deploy service instances. Resource nodes can be physical machines, such as servers. Resource nodes can also be virtual machines deployed on physical machines or bare metal servers (BMS).

[0046] Service instance: A collection of programs or processes that provide services such as data storage and data computation. Common service instances include database instances, data warehouse service instances, and containers. A service instance can be deployed on at least one resource node.

[0047] A cloud management platform is a platform provided by cloud service providers for managing service instances, such as creating, shutting down, and migrating service instances. Cloud service providers offer cloud services to users through service instances. The cloud management platform interacts with users, allowing them to register accounts on the platform and subscribe to cloud services provided by one or more service instances, thereby becoming tenants of cloud services.

[0048] Quality of service (QoS) refers to the quality of the service provided by a service instance. QoS can be expressed through metrics such as latency and throughput. QoS is negatively correlated with latency, while it is positively correlated with throughput. QoS reflects the service instance's ability to provide services.

[0049] QoS requirements: These are the QoS requirements that a cloud provider or tenant places on a service instance. For example, if a service instance's QoS includes latency, the QoS requirement specifies that latency must not exceed a preset latency threshold. Another example is if a service instance's QoS includes throughput, the QoS requirement specifies that throughput must not fall below a preset throughput threshold. QoS requirements for a service instance can be set by the cloud provider or the tenant to which the service instance belongs.

[0050] Planned events, also known as planned maintenance events or maintenance operations, are proactively initiated to improve the reliability and security of resource nodes. Typically, planned events are generated based on proactive predictions of physical machine hardware and software failure risks. These typically involve migrating resource nodes within a physical machine to mitigate the impact of hardware or software failures on resource nodes.

[0051] Unplanned events, also known as planned operational events or maintenance operations, are triggered by hardware or software failures on physical machines. Unlike planned events, unplanned events are generated passively in response to existing hardware or software failures. Unplanned events often involve migrating resource nodes within a physical machine to mitigate the impact of hardware or software failures on resource nodes.

[0052] Live migration, also known as dynamic migration or real-time migration, involves preserving the complete operating state of an entire resource node on a physical machine and quickly restoring that state to another physical machine. The restored resource node continues to operate smoothly. During live migration, the resource node's central processing unit (CPU) and memory frequency must be reduced to increase the time interval between CPU and memory execution and data processing. The live migration is then completed within this time interval.

[0053] Cold migration, also known as static migration, involves shutting down a resource node and migrating it from its current physical machine to another. Cold migration requires shutting down the resource node and is time-consuming. Compared to hot migration, cold migration has a greater impact on the resource node's services.

[0054] In cloud technology, cloud services are provided through service instances, which are deployed on resource nodes. Therefore, the QoS of service instances is affected by the status of resource nodes. Specifically, the status of a resource node can be its operational state. When the status of a resource node deteriorates, the QoS of the service instance degrades. Operations on resource nodes often affect their status. For example, to address hardware or software failures on a physical machine, operating system upgrades on a physical machine, or sub-healthy hardware resources on a physical machine, hot or cold migration of a resource node is often triggered. Hot migration of a resource node can cause CPU and memory throttling, resulting in the resource node operating at low CPU and memory frequencies. This impacts the QoS of the service instance. If the service instance has high QoS requirements, CPU and memory throttling on the resource node can cause the service instance's QoS to fall below the required level. Furthermore, cold migration requires shutting down the resource node, further impacting the service instance's QoS.

[0055] The cloud management platform used to manage service instances is separate from the underlying management platform used to manage resource nodes. This makes it impossible for the cloud management platform to predict operations on resource nodes (such as planned or unplanned events). As a result, the cloud management platform often reschedules service instances, for example, by migrating them to other resource nodes, only after operations on resource nodes have already begun and their status has already affected the QoS of the service instances. This impacts the reliability and resilience of cloud services, and even affects the user experience of cloud services.

[0056] An embodiment of the present application provides a service instance management method based on cloud technology. In this method, before an operation is performed on a resource node in a physical machine, the physical machine may issue an event message. The event message includes the expected execution time of the operation, which affects the resource node in the physical machine. The cloud management platform may receive the event message and obtain the expected execution time of the operation from the event message. Then, the cloud management platform may migrate the service instance to a resource node in another physical machine before the expected execution time of the operation based on the QoS requirements of the service instance in the resource node.

[0057] In this method, the cloud management platform can use event information to obtain the expected execution time of an operation affecting a resource node on a physical machine. Before the expected execution time, the cloud management platform can migrate the service instance in the resource node to a resource node on another physical machine. This prevents the service quality of the service instance in the resource node from being affected by the operation of the resource node.

[0058] Next, the service instance management method provided in the embodiment of the present application is introduced in detail.

[0059] Figure 1 illustrates a system architecture that can be used to implement the method. As shown in Figure 1 , the system architecture includes a cloud infrastructure 100 and a cloud management platform 200. Cloud infrastructure 100 includes multiple physical machines, such as physical machine 110 and physical machine 120. For example, the physical machines can be servers.

[0060] In some embodiments, cloud infrastructure 100 may include multiple data centers, where each data center may include at least one of the multiple physical machines. In other words, different physical machines in the multiple physical machines may be deployed in different data centers. For example, physical machine 110 and physical machine 120 may be deployed in different data centers.

[0061] Each physical machine may be deployed with one or more resource nodes. For example, physical machine 110 is deployed with resource node 111, and physical machine 120 is deployed with node 121. In some embodiments, the resource node deployed in the physical machine is the physical machine itself. In some embodiments, the resource node deployed in the physical machine may be a virtual machine or a bare metal server, etc. The cloud management platform 200 may create one or more service instances in the resource nodes in the cloud infrastructure 100, such as creating service instance A1 in resource node 111, creating service instance B1 in resource node 112, and so on. Among them, a service instance may be deployed in at least one resource node, and the QoS of the service instance is affected by the state of the resource node where the service instance is located. The cloud management platform 200 may manage service instances in resource nodes, for example, migrating service instances from current resource nodes to other resource nodes.

[0062] Referring to Figure 1 , each physical machine includes an anomaly detection instance and an event generation instance. The anomaly detection instance detects anomalies or changes in the physical machine and issues alerts. The event generation instance generates events based on alerts and sends these events to the cloud management platform 200. The anomaly detection instance and event generation instance run directly on the physical machine, for example, as system services or applications running directly within the physical machine's operating system.

[0063] The anomaly of the physical machine may include a hardware failure of the physical machine. Referring to FIG2 , the anomaly perception instance may include a hardware fault detection module. The hardware fault detection module is used to detect the status of the hardware of the physical machine. The hardware may include one or more of a processor, a platform controller hub (PCH), a data card, a redundant array of independent disks (RAID) card, a fan, a memory, a disk, a network card, a motherboard, a power supply, etc. The processor may be a CPU, a graphics processing unit (GPU), etc. The data card may be a high-speed serial computer expansion bus standard (Peripheral Component Interconnect Express, PCIE) card, etc.

[0064] In some embodiments, the hardware fault detection module can detect hardware operating indicators, such as CPU temperature, CPU frequency, accumulated memory dirty pages, etc. In some embodiments, the hardware fault detection module can detect kernel events of the physical machine operating system, such as CPU internal error (CPU IERR) events, system-level memory initialization error events, and hard disk drive fault events generated in the kernel log. The hardware fault detection module can periodically detect hardware operating indicators or kernel events, for example, once every 5 seconds.

[0065] As shown in Figure 2, the anomaly perception example includes a fault identification module. The fault identification module can obtain hardware operating indicators and / or kernel events detected by the hardware fault detection module and identify whether a hardware fault has occurred based on the hardware operating indicators and / or kernel events.

[0066] As shown in Figure 2, the anomaly detection instance has a fault identification rule base. The fault identification rule base includes fault identification rules. These rules can be pre-set. The fault identification rules can be used to identify hardware failures based on hardware operating indicators and / or kernel events detected by the faulty hardware fault detection module. In one example, the fault identification rules are shown in Table 1.

[0067] Table 1

[0068] As shown in Figure 2, the anomaly detection instance also includes an alarm generation module. The alarm generation module is configured to generate an alarm based on the hardware fault identified by the fault identification module. When a hardware fault occurs, and the faulty hardware is required for a physical machine to run a resource node, the fault identification module can trigger the alarm generation module to generate an alarm. In one example, the alarm generation module can generate an alarm based on the alarm format shown in Table 2.

[0069] Table 2

[0070] The alarm generation module can generate an alarm based on the hardware fault, referring to the alarm format shown in Table 2. The alarm corresponds to the fault type, and different fault types generate different alarms.

[0071] In some embodiments, as shown in FIG2 , the anomaly perception example includes a software fault detection module. The software fault detection module is used to detect the running status of the software, such as the response time of the software, the frequency of software crashes or freezes, thread leaks, etc. The software is the operating system of the physical machine or the management software of the resource node. For example, when the resource node is a virtual machine, the software can be a virtual machine monitor (VMM). The fault identification module can identify whether the software has failed based on the running status of the software detected by the software fault detection module. For example, when the response time of the software is greater than a preset time threshold, it can be confirmed that the software has failed. In addition, the fault can also include a sub-healthy state. If the software has a thread leak, it means that the software is in a sub-healthy state. When the fault identification module confirms that the software has failed, it can trigger the alarm generation module to generate an alarm.

[0072] In some embodiments, as shown in FIG2 , the anomaly detection instance includes a change detection module. A change can be a planned maintenance operation, such as replacing one or more hardware components of a physical machine, upgrading the operating system or other software on the physical machine, or installing a patch. Maintenance personnel can use the change detection module to input the change time and change type. The alarm generation module can then generate an alarm based on the change time and change type.

[0073] The event generation instance can generate events based on alarms. As shown in Figure 3, the event generation instance may include an alarm receiving module, an event generation module, and an event information sending module. Among them, the alarm receiving module can receive alarms from the abnormality perception instance and send the received alarms to the event generation module. The event generation module can generate events based on alarms. An event refers to the execution of a certain operation on a resource node at the expected execution time. In other words, an event includes an operation and the expected execution time of the operation. Operations can be of various types. Common operation types include resource node migration, physical machine restart, etc. Among them, the migration of resource nodes can be divided into hot migration and cold migration. Among them, compared with cold migration, hot migration of resource nodes takes less time and has less impact on the QoS of service instances in the resource nodes.

[0074] The event generation module can generate different events based on different alarms. The difference in events can refer to the different types of operations in the events. For example, if the alarm is an alarm corresponding to a CPU failure, the type of operation in the generated event is resource node cold migration. For another example, if the alarm is an alarm corresponding to a memory configuration error, the type of operation in the generated event is physical machine restart, and so on. In some embodiments, the event generation instance has an event generation rule base. The event generation rule base includes a pre-set correspondence between alarms and events. The event generation module can generate corresponding events based on the correspondence and the alarm.

[0075] The actions in an event are the actions that are performed to address the failure or change corresponding to the alarm. In other words, the actions that a physical machine will perform to address hardware or software failures or hardware or software changes. For example, if a hardware failure on a physical machine affects the operation of a resource node, the node may be migrated to ensure its operation. This is similar to the actions that are performed in the previous example.

[0076] An event can be represented by event information. Event information includes the expected execution time of the operation, the type of operation, etc. In some embodiments, event information may also include an event identifier (ID), an event description, and the execution status of the operation. The execution status of the operation may include pending, executing, completed, or canceled.

[0077] In an example, the format of event information may be as shown in Table 3.

[0078] Table 3

[0079] Whenever the event generation module generates an event, the event information of the event can be sent to the event information issuing module. The event information issuing module can issue the event information.

[0080] The cloud management platform 200 can receive the event information and, based on the event information, schedule a service instance in the resource node of the physical machine. Referring to FIG1 , an event-aware instance is deployed in the resource node of the physical machine. An event-aware instance is a service instance, specifically a service instance for sensing event information. The event-aware instance in the physical machine can obtain event information emitted by the event-generating instance in the physical machine. The event-aware instance can send the obtained event information to the cloud management platform 200.

[0081] In some embodiments of cloud computing, for security reasons, the virtual network environment where service instances reside is isolated from the physical network environment where physical machines reside. Service instances and physical machines cannot directly access each other. In other words, event-aware instances cannot obtain event information by directly accessing physical machines. Referring to Figure 4, this embodiment provides an event-aware channel. Through this channel, event-aware instances can obtain event information.

[0082] The event perception channel can be an interface through which the event perception instance can obtain event information. Specifically, the interface is associated with a storage space of the physical machine. The storage space is used to store the event information generated by the physical machine. That is, the event generation instance can store the event information in the storage space. The physical machine can open the interface to or provide the event perception instance in the resource node of the physical machine. The event perception instance can access the storage space associated with the interface through the interface, so that the event information can be read from the storage space. The event perception instance can periodically access the storage space through the interface to detect whether new event information is generated. For example, the event perception instance can access the storage space through the interface every 1 second. Only the storage space associated with the interface can be accessed through the interface, and other storage spaces cannot be accessed. In this way, the information security of the physical machine is guaranteed.

[0083] In some embodiments, as shown in FIG5 , an event-sensing instance is pre-installed with a sensing rule library. The sensing rule library includes a preset correspondence between event-sensing instances and event-sensing channels. Based on this correspondence, the event-sensing instance can obtain the event-sensing channel corresponding to the event-sensing instance, thereby obtaining event information through the event-sensing channel.

[0084] In some embodiments, the event perception instance can filter the acquired event information. Specifically, the operations corresponding to some event information may have been canceled or completed. In this case, there is no need to send the event information to the cloud management platform 200. The event perception instance can filter the event information corresponding to the operation to be executed from the acquired event information, and send the filtered event information to the cloud management platform. As mentioned above, the event information includes the execution status of the operation, and the event perception instance filters the event information based on the execution status of the operation in the event information.

[0085] In some embodiments, as shown in FIG5 , the event-aware instance is pre-installed with a QoS requirement library. The QoS requirement library includes the QoS requirements of the service instances in the resource node where the event-aware instance is located. Based on the QoS requirement library, the event-aware instance can also send the QoS requirements of the service instances to the cloud management platform 200. For example, when the event-aware instance sends event information to the cloud management platform 200, it also sends the QoS requirements of the service instances to the cloud management platform 200, so that the cloud management platform can schedule the service instances based on the event information and the QoS requirements of the service instances.

[0086] The above examples introduce the system architecture provided by the embodiment of the present application. Next, in conjunction with the system architecture, the service instance management method provided by the embodiment of the present application is introduced.

[0087] The method is executed by the cloud management platform 200. As shown in FIG6 , the method includes the following steps.

[0088] In step 601 , the cloud management platform 200 receives event information from a physical machine 110 among multiple physical machines. The event information includes an expected execution time of an operation C1 , where the operation C1 affects the status of a resource node 111 of the physical machine 110 .

[0089] The cloud management platform 200 may receive event information from the physical machine 110, for example, from an event-aware instance in the resource node 111 of the physical machine 110. Please refer to the above description for details, which will not be repeated here.

[0090] The expected execution time is later than the current time, and the expected execution time is later than the execution time of step 601 .

[0091] Operation C1 is an operation that the physical machine 110 will begin to execute at the expected execution time, and operation C1 affects the state of the resource node 111. The state of the resource node can be the running state of the resource node. That is, when the physical machine 111 executes operation C1, the state of the resource node 111 will change. Specifically, the state of the resource node 111 will deteriorate. As described above, the deterioration of the state of the resource node causes the QoS of the service instance in the resource node to decrease. In other words, the execution of operation C1 will cause the QoS of the service instance in the resource node 111 to decrease.

[0092] In some embodiments, operation C1 is an operation used to address hardware failures or changes in physical machine 110, where the hardware is required for physical machine 110 to run resource node 111. That is, the hardware provides the hardware resources for physical machine 1110 to run resource node 111. In some embodiments, operation C1 is an operation used to address software failures or changes; where the software is the operating system of physical machine 111 or resource node management software. The generation and function of operation C1 can be found in the above description of the embodiments shown in Figures 2 and 3 and will not be further elaborated here.

[0093] In step 602, before the expected execution time, cloud management platform 200 migrates the service instance in resource node 111 to resource node 121 of physical machine 120 among the multiple physical machines based on the QoS requirements of the service instance in resource node 111. Physical machine 120 is a healthy physical machine, and cloud management platform 200 has not received any event information from physical machine 120.

[0094] A service instance has QoS requirements. These requirements can be set by the tenant to which the service instance belongs or by the cloud provider.

[0095] In some embodiments, the cloud management platform 200 stores the QoS requirements of each service instance. The cloud management platform 200 may execute step 602 based on the stored QoS requirements of the service instances in the resource nodes 111 .

[0096] In some embodiments, the cloud management platform 200 may receive the QoS requirements of the service instances in the resource node 111 from the resource node 111. For example, as described above, the event-aware instance in the resource node 111 may send the QoS requirements of the service instances in the resource node 111 to the cloud management platform 200.

[0097] In some embodiments, when the resource node is a virtual machine or a bare metal server, the types of operations on the resource node may include hot migration of the resource node, cold migration of the resource node, and physical machine restart. That is, operation C1 belongs to one of hot migration of the resource node, cold migration of the resource node, and physical machine restart. Among them, hot migration of the resource node, cold migration of the resource node, and physical machine restart will all cause the state of the resource node to deteriorate, that is, hot migration of the resource node, cold migration of the resource node, and physical machine restart will all reduce the QoS of the service instance in the resource node. Among them, the decline in the QoS of the service instance caused by hot migration of the resource node, cold migration of the resource node, and physical machine restart increases in sequence.

[0098] The QoS requirements of a service instance can have multiple QoS levels, and different QoS levels tolerate different types of operations. The multiple QoS levels can be set to include QoS level D1, QoS level D2, QoS level D3, and QoS level D4, with QoS decreasing in sequence. Among them, the QoS requirement of a service instance that provides game backend computing services is usually QoS level D1, the QoS requirement of a service instance that provides e-commerce services is usually QoS level D2, and the QoS requirement of a service instance that provides offline AI model training services is usually QoS level D3. The QoS requirement of a service instance that provides periodic dialing services is usually QoS level D4.

[0099] As shown in Table 4, QoS level D1 cannot tolerate any type of operation. Specifically, QoS level D1 cannot tolerate hot migration of resource nodes. A QoS level that cannot tolerate hot migration of resource nodes also cannot tolerate cold migration of resource nodes or physical machine restarts. In other words, hot migration of resource nodes causes the QoS of service instances in that resource node to fall below QoS level D1.

[0100] QoS level D2 can tolerate hot migration of resource nodes, but cannot tolerate cold migration of resource nodes or physical machine restarts. In other words, the QoS of a service instance affected by hot migration of resource nodes is not lower than QoS level D2, but cold migration of resource nodes or physical machine restarts cause the QoS of the service instance to fall below QoS level D2.

[0101] QoS level D3 can tolerate cold migration of resource nodes, but not physical machine restarts. QoS levels that tolerate cold migration of resource nodes can also tolerate hot migration of resource nodes. In other words, the QoS of a service instance affected by a cold or hot migration of a resource node is no less than QoS level D3, but a physical machine restart can cause the QoS of the service instance to drop below QoS level D3.

[0102] QoS level D4 can tolerate operations of any operation type. That is, no matter what type of operation is performed on the resource node, the QoS of the service instance in the resource node is higher than QoS level D4.

[0103] Table 4

[0104] In some embodiments, in step 602, the cloud management platform 200 determines whether the QoS requirement of the service instance A1 in the resource node 111 is greater than or equal to a preset QoS threshold. The preset QoS threshold may be QoS level D1. If the QoS requirement of the service instance A1 is greater than or equal to the preset QoS threshold, in step 602, the cloud management platform 200 migrates the service instance A1 to the resource node 121. That is, in step 602, the cloud management platform 200 migrates the service instance A1 to the resource node 121 after confirming that the QoS requirement of the service instance A1 in the resource node 111 is greater than or equal to the preset QoS threshold.

[0105] Specifically, the QoS requirement of service instance A1 is greater than or equal to the preset QoS threshold, indicating that the QoS requirement of service instance A1 cannot tolerate any type of operation. In this case, regardless of the type of operation C1, the QoS requirement of service instance A1 cannot tolerate it. In other words, operation C1 defaults to causing the QoS of service instance A1 to be lower than the QoS requirement of service instance A1. Therefore, there is no need to identify the type of operation C1 and the service instance can be migrated directly.

[0106] Service instance migration can be performed as follows.

[0107] Referring to Figure 7, first, create a service instance A1' in the resource node 121. Service instance A1' is a copy of service instance A1 and can provide the services provided by service instance A1. When service instance A1' is created successfully, that is, when service instance A1' can provide the services provided by service instance A1, service instance A1' is started, and service instance A1 is isolated, and new service requests are sent to service instance A1', and no new service requests are sent to service instance A1. Among them, service instance A1 still keeps running to process existing service requests (that is, service requests sent to service instance A1 before starting service instance A1'). Exemplarily, when service instance A1 has processed the existing service requests, the cloud management platform 200 can delete service instance A1 from the resource node 111.

[0108] Through the above-mentioned service instance migration method, lossless migration of the service instance is achieved, so that the QoS of the service provided by the service instance A1 is not lower than the QoS level D1.

[0109] When the expected execution time of operation C1 arrives, physical machine 110 executes operation C1 on resource node 111. At this point, service instance A1 has already been migrated from resource node 111, and operation C1 executed on resource node 111 has no impact on service instance A1. In other words, the service instance providing the service provided by service instance A1 has already been migrated from resource node 111, and operation C1 executed on resource node 111 has no impact on the QoS of that service.

[0110] In some embodiments, the event information received by the cloud management platform 200 also includes the type of operation C1. In step 602, based on the type of operation C1, it can be determined whether operation C1 will cause the QoS of the service instance in the resource node 111 to be lower than the QoS requirement of the service instance. It can be determined whether the QoS requirement of the service instance can tolerate the type of operation C1. If it cannot be tolerated, it is confirmed that operation C1 will cause the QoS of the service instance in the resource node 111 to be lower than the QoS requirement of the service instance. If it can be tolerated, it is confirmed that operation C1 will not cause the QoS of the service instance in the resource node 111 to be lower than the QoS requirement of the service instance.

[0111] For example, as shown in Table 4, it can be assumed that service instance A2 is also deployed in resource node 111, where the QoS requirement of service instance A2 is QoS level D2. As described above, QoS level D2 can only tolerate hot migration of resource nodes. If operation C1 is a cold migration of a resource node or a physical machine restart, then confirm that operation C1 will cause the QoS of service instance A2 to be lower than the QoS requirement of service instance A2. If operation C1 is a hot migration of a resource node, then confirm that operation C1 will not cause the QoS of service instance A2 to be lower than the QoS requirement of service instance A2.

[0112] It is also possible to configure resource node 111 to further deploy service instance A3, where the QoS requirement for service instance A3 is QoS level D3. As described above, QoS level D3 can tolerate both cold and hot migrations of resource nodes. If operation C1 is a physical machine restart, then confirm that operation C1 will cause the QoS of service instance A3 to fall below the QoS requirement for service instance A3. If operation C1 is a hot or cold migration of a resource node, then confirm that operation C1 will not cause the QoS of service instance A3 to fall below the QoS requirement for service instance A3.

[0113] In step 602 , when it is confirmed that operation C1 will cause the QoS of the service instance in resource node 111 to be lower than the QoS requirement of the service instance, the cloud management platform 200 migrates the service instance in resource node 111 to resource node 121 .

[0114] For example, as shown in Figure 8 , operation C1 causes the QoS of service instance A2 in resource node 111 to fall below the QoS requirement of service instance A2. Cloud management platform 200 then migrates service instance A2 from resource node 111 to resource node 121. This creates service instance A2' in resource node 121. Service instance A2' is a copy of service instance A2 and can take over the service instance A2 in resource node 111 to provide the corresponding service. For details, please refer to the above description and will not be repeated here.

[0115] In step 602, when it is confirmed that C1 will not cause the QoS of the service instance in resource node 111 to be lower than the QoS requirement of the service instance, the cloud management platform 200 may not migrate the service instance. When performing operation C1 on resource node 111, the service instance is scheduled in a traditional manner.

[0116] 8 , operation C1 does not cause the service instance A3 in the resource node 111 to fall below the QoS requirement of the service instance A3 . In this case, the cloud management platform may not migrate the service instance A3 from the resource node 111 .

[0117] In some embodiments, when the cloud management platform 200 receives event information, the cloud management platform 200 prohibits deploying new service instances on the resource node corresponding to the event information (i.e., resource node 111). This prevents new service instances from being affected by operation C1 and reduces the scheduling pressure on the cloud management platform 200.

[0118] To summarize, when a physical machine is about to perform an operation that affects a resource node on that physical machine, it can send an event message to the cloud management platform. This event message includes the expected execution time of the operation. Upon receiving this event message, the cloud management platform can obtain the expected execution time and, before the expected execution event, migrate the service instance on that resource node to a resource node on another physical machine. This prevents the QoS of the service instance on the resource node from being affected by the operation.

[0119] Referring to FIG9 , an embodiment of the present application provides a cloud management platform 900. The cloud management platform 900 is connected to a cloud infrastructure and is used to manage multiple service instances deployed in the cloud infrastructure. The cloud infrastructure includes multiple physical machines, each of which is used to deploy one or more resource nodes. Each of the multiple service instances is deployed in at least one resource node in the cloud infrastructure. The QoS of the service instance is affected by the status of the resource node where the service instance is located. As shown in FIG9 , the cloud management platform 900 includes:

[0120] A receiving module 910 is configured to receive event information sent by a first physical machine among the multiple physical machines, the event information including an expected execution time of an operation, wherein the operation affects a state of a first resource node of the first physical machine;

[0121] The migration module 920 is used to migrate the service instance in the first resource node to the second resource node of the second physical machine among the multiple physical machines based on the QoS requirements of the service instance in the first resource node before the expected execution time.

[0122] In some embodiments, the event information also includes the type of the operation; the migration module 920 is used to: based on the type of the operation, confirm that the operation will cause the QoS of the service instance in the first resource node to be lower than the QoS requirement; and migrate the service instance in the first resource node to the second resource node.

[0123] In some embodiments, the migration module 920 is configured to: confirm that the QoS requirement is greater than or equal to a preset QoS threshold; and migrate the service instance in the first resource node to the second resource node.

[0124] In some embodiments, the cloud management platform 900 further includes: a prohibition module, configured to prohibit deployment of new service instances in the first resource node.

[0125] In some embodiments, the operation is used to respond to hardware failures or changes, wherein the hardware is used for the first physical machine to run the first resource node; or, the operation is used to respond to software failures or changes; wherein the software includes at least one of the operating system of the first physical machine and the management software of the first resource node.

[0126] In some embodiments, the service instance in the first resource node includes any one or more of a container, a database instance, and a data warehouse instance.

[0127] In some embodiments, the first resource node is a virtual machine or a bare metal server, or the first resource node is the first physical machine itself.

[0128] The receiving module 910 and the migration module 920 can be implemented in software or hardware. For example, the implementation of the receiving module 910 will be described below using the receiving module 910 as an example. Similarly, the implementation of the migration module 920 can refer to the implementation of the receiving module 910.

[0129] As an example of a software functional unit, the receiving module 910 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the receiving module 910 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone AZ or in different AZs, and each AZ includes one data center or multiple geographically close data centers. Generally, a region may include multiple AZs.

[0130] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0131] As an example of a hardware functional unit, the receiving module 910 may include at least one computing device, such as a server. Alternatively, the receiving module 910 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0132] The multiple computing devices included in the receiving module 910 can be distributed in the same region or in different regions. The multiple computing devices included in the receiving module 910 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the receiving module 910 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0133] It should be noted that, in other embodiments, the receiving module 910 can be used to execute any step in the method shown in FIG6 , and the migration module 920 can be used to execute any step in the method shown in FIG6 . The steps that the receiving module 910 and the migration module 920 are responsible for implementing can be specified as needed. By having the receiving module 910 and the migration module 920 respectively implement different steps in the method shown in FIG6 , the full functionality of the cloud management platform 900 is realized.

[0134] This application also provides a computing device 1000. As shown in Figure 10, computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1000.

[0135] Bus 1002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG10 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1002 may include a path for transmitting information between various components of computing device 1000 (e.g., memory 1006, processor 1004, and communication interface 1008).

[0136] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0137] The memory 1006 may include volatile memory, such as random access memory (RAM). The memory 1006 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0138] The memory 1006 stores executable program codes, and the processor 1004 executes the executable program codes to implement the functions of the aforementioned receiving module 910 and the migration module 920, thereby implementing the method shown in Figure 6. That is, the memory 1006 stores instructions for executing the method shown in Figure 6.

[0139] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.

[0140] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0141] As shown in Figure 11, the computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the method shown in Figure 6.

[0142] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also respectively store some instructions for executing the method shown in Figure 6. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the method shown in Figure 6.

[0143] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions, each for executing part of the functions of the cloud management platform 900. In other words, the instructions stored in the memory 1006 in different computing devices 1000 can implement the functions of one or more modules in the receiving module 910 and the migration module 920.

[0144] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG12 illustrates a possible implementation. As shown in FIG12 , two computing devices 1000A and 1000B are connected via a network. Specifically, the connection to the network is made via a communication interface in each computing device. In this type of possible implementation, the memory 1006 in the computing device 1000A stores instructions for executing the functions of the receiving module 910. Simultaneously, the memory 1006 in the computing device 1000B stores instructions for executing the functions of the migration module 920.

[0145] It should be understood that the functionality of the computing device 1000A shown in FIG12 may also be accomplished by multiple computing devices 1000. Similarly, the functionality of the computing device 1000B may also be accomplished by multiple computing devices 1000.

[0146] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 11 and 12. However, the memory 1006 in one or more computing devices 1000 in this computing device cluster can store the same instructions for executing the method shown in Figure 6.

[0147] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also respectively store some instructions for executing the method shown in Figure 6. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the method shown in Figure 6.

[0148] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the method shown in FIG6 .

[0149] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a host migration device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method shown in FIG6 .

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A service instance management method based on cloud technology, characterized in that The method is applied to a cloud management platform connected to a cloud infrastructure. The cloud management platform is used to manage multiple service instances deployed in the cloud infrastructure. The cloud infrastructure includes multiple physical machines, and each physical machine in the multiple physical machines is used to deploy one or more resource nodes. Each service instance in the multiple service instances is deployed in at least one resource node in the cloud infrastructure, and the quality of service (QoS) of the service instance is affected by the state of the resource node where the service instance is located. The method includes: The cloud management platform receives event information sent by a first physical machine among the multiple physical machines. The event information includes the expected execution time of an operation, where the operation affects the state of a first resource node of the first physical machine. Before the expected execution time, the cloud management platform migrates the service instance in the first resource node to a second resource node of a second physical machine among the multiple physical machines based on the QoS requirements of the service instance in the first resource node.

2. The method according to claim 1, wherein The event information further includes the type of the operation. The cloud management platform migrates the service instance in the first resource node to a second resource node of a second physical machine among the multiple physical machines based on the QoS requirements of the service instance in the first resource node, including: The cloud management platform confirms, based on the type of the operation, that the operation will cause the QoS of the service instance in the first resource node to be lower than the QoS requirements. The cloud management platform migrates the service instance in the first resource node to the second resource node.

3. The method according to claim 1, characterized in that, The cloud management platform migrates the service instance in the first resource node to a second resource node of a second physical machine among the multiple physical machines based on the QoS requirements of the service instance in the first resource node, including: The cloud management platform confirms that the QoS requirements are greater than or equal to a preset QoS threshold. The cloud management platform migrates the service instance in the first resource node to the second resource node.

4. The method according to any one of claims 1 to 3, characterized in that The method further includes: The cloud management platform prohibits deploying new service instances in the first resource node.

5. The method according to any one of claims 1-4, wherein The operation is used to cope with a hardware failure or change, where the hardware is used for the first physical machine to run the first resource node; or The operation is used to cope with a software failure or change; where the software includes at least one of the operating system of the first physical machine and the management software of the first resource node.

6. The method according to any one of claims 1-5, characterized in that, The service instance in the first resource node includes any one or more of a container, a database instance, and a data warehouse instance.

7. The method according to any one of claims 1-6, characterized in that, The first resource node is a virtual machine or a bare metal server, or the first resource node is the first physical machine itself.

8. A cloud management platform, characterized in that, The cloud management platform is connected to the cloud infrastructure. The cloud management platform is used to manage multiple service instances deployed in the cloud infrastructure. The cloud infrastructure includes multiple physical machines. Each physical machine among the multiple physical machines is used to deploy one or more resource nodes. Each service instance among the multiple service instances is deployed in at least one resource node in the cloud infrastructure. The quality of service (QoS) of the service instance is affected by the state of the resource node where the service instance is located. The cloud management platform includes: A receiving module, configured to receive event information sent by a first physical machine among the multiple physical machines. The event information includes an expected execution time of an operation, where the operation affects the state of a first resource node of the first physical machine. A migration module, configured to, before the expected execution time, based on the QoS requirements of the service instance in the first resource node, migrate the service instance in the first resource node to a second resource node of a second physical machine among the multiple physical machines.

9. The cloud management platform according to claim 8, wherein The event information further includes the type of the operation. The migration module is configured to: Based on the type of the operation, confirm that the operation will cause the QoS of the service instance in the first resource node to be lower than the QoS requirements. Migrate the service instance in the first resource node to the second resource node.

10. The cloud management platform according to claim 8, wherein The migration module is configured to: Confirm that the QoS requirements are greater than or equal to a preset QoS threshold. Migrate the service instance in the first resource node to the second resource node.

11. The cloud management platform according to any one of claims 8-10, characterized in that, The cloud management platform further includes: A prohibition module, configured to prohibit deploying a new service instance in the first resource node.

12. The cloud management platform according to any one of claims 8-11, wherein The operation is used to cope with a hardware failure or change, where the hardware is used for the first physical machine to run the first resource node; or The operation is used to cope with a software failure or change; where the software includes at least one of an operating system of the first physical machine and management software of the first resource node.

13. The cloud management platform according to any one of claims 8-12, characterized in that, The service instance in the first resource node includes any one or more of a container, a database instance, and a data warehouse instance.

14. The cloud management platform according to any one of claims 8-13, characterized in that, The first resource node is a virtual machine or a bare metal server, or the first resource node is the first physical machine itself.

15. A cluster of computing devices, characterized in that, Including at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1-7.

16. A computer-readable storage medium, characterized in that, Including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1-7.

17. A computer program product comprising instructions, characterized in that, When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Service instance management method based on cloud technology, cloud management platform and cluster

    CN120216091A

  • Method and system for deploying services

    CN103428241A

  • Fault recovery method and device of cloud computing server and management system

    CN108632057A

  • Virtual instance setting method and device

    CN114489922A

  • Implementing a private network isolated from a user network for virtual machine deployment and migration and for monitoring and managing the cloud environment

    US20140201365A1