Container instance migration method and device, electronic equipment and storage medium
By identifying and migrating risky nodes in the computing power cluster, the problem of delayed fault handling of computing power nodes was solved, thereby improving the stability and service reliability of the cluster.
Patent Information
- Application Number
- CN202511313006.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, the handling of computing node failures suffers from lag and an inability to cope with gradual performance degradation, resulting in poor cluster stability and low service availability and reliability.
By acquiring time-series data of operational metrics from computing cluster nodes, risky computing nodes can be identified, and container instances can be proactively migrated to healthy nodes to prevent failures.
It improved cluster stability and service availability, reduced business interruptions, and enhanced service reliability.
Smart Images

Figure CN120803619A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and in particular relates to a container instance migration method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of cloud computing, big data and artificial intelligence technologies, large-scale heterogeneous computing clusters have become the core infrastructure supporting modern digital businesses, and currently mainly include Kubernetes clusters that manage and schedule long-running microservice applications using container orchestration systems and Slurm clusters that are oriented to high-performance computing (HPC) or batch processing tasks. With the expansion of the scale of heterogeneous computing clusters and the increase in hardware complexity, the failure of computing nodes has become a major challenge to business continuity and stability.
[0003] However, the cluster management system in the related art starts the recovery process only after the computing node has already failed seriously (such as being down or lost), which belongs to a passive and reactive failure handling mode, thereby existing problems of reaction lag and being unable to cope with the gradual performance degradation of computing nodes, and completely relying on after-the-fact remedies, resulting in poor cluster stability, and low availability and reliability of services. SUMMARY
[0004] To solve the problems of the prior art, the embodiments of the present application provide a container instance migration method and device, an electronic device and a storage medium. The technical solution is as follows: In one aspect, a container instance migration method is provided, and the method comprises: obtaining at least one running index time series data corresponding to each computing node in a computing cluster; the at least one running index time series data corresponds to at least one running index dimension one by one, and the at least one running index dimension is used to represent the health status of the computing node; based on the matching of the at least one running index time series data corresponding to each computing node and a preset risk condition, identifying a risk computing node to obtain a risk computing node identification result; when the risk computing node identification result indicates that there is a risk computing node, determining a to-be-migrated container instance from the container instances running on the risk computing node; migrating the to-be-migrated container instance to a non-risk computing node of the computing cluster for running.
[0005] In some example embodiments, the preset risk condition comprises at least one preset risk identification rule corresponding to the at least one running index dimension; the risk computing node identification based on the matching of the at least one running index time series data corresponding to each of the computing nodes with the preset risk condition comprises: traversing the at least one preset risk identification rule, for the current preset risk identification rule traversed, matching the target running index time series data of each of the computing nodes with the current preset risk identification rule, if matched, determining the computing node corresponding to the target running index time series data matched as a risk computing node, obtaining a risk computing node identification sub-result corresponding to the current preset risk identification rule; the target running index time series data is the running index time series data of the running index dimension corresponding to the current preset risk identification rule; determining the risk computing node identification result based on the risk computing node identification sub-result corresponding to each of the preset risk identification rules.
[0006] In some example embodiments, when the risk computing node identification result indicates that there is a risk computing node, determining a to-be-migrated container instance from the container instances running on the risk computing node comprises: when the risk computing node identification result indicates that there is a risk computing node, calling an application programming interface service of the computing cluster to determine a container instance running on the risk computing node, obtaining a container instance set of the risk computing node; selecting a to-be-migrated container instance from the container instance set of the risk computing node according to a preset screening strategy, obtaining a to-be-migrated container instance list; wherein, the preset screening strategy comprises: excluding container instances marked as prohibited migration in annotations; excluding container instances belonging to key system components in a preset key system component list; the remaining available replica number of the application to which each of the container instances belongs meets the minimum replica number configuration.
[0007] In some example embodiments, the preset screening strategy further comprises a migration sorting strategy, and the to-be-migrated container instances in the to-be-migrated container instance list are sorted according to the migration sorting strategy; wherein, the migration sorting strategy comprises: using priority information marked in annotations to determine the migration order of the corresponding to-be-migrated container instance; the to-be-migrated container instance belonging to a stateless application is prior to the to-be-migrated container instance belonging to a stateful application.
[0008] In some example embodiments, the migration of the to-be-migrated container instance to the non-risk computing node of the computing cluster for running comprises: constructing a banishment object according to the container instance to be migrated; sending a banishment request to an application programming interface service of the computing power cluster based on the banishment object, so that the computing power cluster, in response to the banishment request, migrates the container instance to be migrated from the risk computing power node to a non-risk computing power node of the computing power cluster for running.
[0009] In some exemplary embodiments, the method further comprises: monitoring a migration process of the container instance to be migrated, and generating a migration event record corresponding to the container instance to be migrated; the migration event record comprises timestamp information, source computing power node identification, target computing power node identification, migrated container instance identification, and migration result.
[0010] In some exemplary embodiments, after the risk computing power node identification result is obtained, the method further comprises: adding a stain label to the risk computing power node indicated by the risk computing power node identification result; After the container instance to be migrated is migrated to a non-risk computing power node of the computing power cluster for running, the method further comprises: When the risk computing power node is determined to be a non-risk computing power node, removing the stain label on the risk computing power node.
[0011] In another aspect, a container instance migration apparatus is provided, the apparatus comprising: a running index data acquisition module configured to acquire at least one running index time series data corresponding to each computing power node in a computing power cluster respectively; the at least one running index time series data corresponds to at least one running index dimension one by one, and the at least one running index dimension is used to represent the health status of the computing power node; a risk computing power node identification module configured to perform risk computing power node identification based on the matching of the at least one running index time series data corresponding to each computing power node respectively and a preset risk condition, and obtain a risk computing power node identification result; a container instance to be migrated determination module configured to determine a container instance to be migrated from the container instances running on the risk computing power node when the risk computing power node identification result indicates that there is a risk computing power node; a migration module configured to migrate the container instance to be migrated to a non-risk computing power node of the computing power cluster for running.
[0012] In some example embodiments, the preset risk condition comprises at least one preset risk identification rule corresponding to the at least one operation index dimension; the risk computing power node identification module is specifically configured to: traverse the at least one preset risk identification rule, for a current preset risk identification rule traversed, match the target operation index time series data of each computing power node with the current preset risk identification rule respectively, if matched, determine that the computing power node corresponding to the target operation index time series data matched is a risk computing power node, and obtain a risk computing power node identification sub-result corresponding to the current preset risk identification rule; the target operation index time series data is operation index time series data of an operation index dimension corresponding to the current preset risk identification rule; and determine the risk computing power node identification result based on the risk computing power node identification sub-result corresponding to each preset risk identification rule.
[0013] In some example embodiments, the to-be-migrated container instance determination module comprises: a container instance set determination module configured to, when the risk computing power node identification result indicates that there is a risk computing power node, call an application programming interface service of the computing power cluster, determine a container instance running on the risk computing power node, and obtain a container instance set of the risk computing power node; a screening module configured to select a to-be-migrated container instance from the container instance set of the risk computing power node according to a preset screening strategy, and obtain a to-be-migrated container instance list; wherein the preset screening strategy comprises: excluding a container instance marked as prohibited from migration in an annotation; excluding a container instance belonging to a key system component in a preset key system component list; and each container instance belonging to an application whose remaining available replica number in the computing power cluster meets a minimum replica number configuration.
[0014] In some example embodiments, the preset screening strategy further comprises a migration sorting strategy, and the to-be-migrated container instances in the to-be-migrated container instance list are sorted according to the migration sorting strategy; wherein the migration sorting strategy comprises: using priority information marked in an annotation to determine a migration order of a corresponding to-be-migrated container instance; and a to-be-migrated container instance belonging to a stateless application is prior to a to-be-migrated container instance belonging to a stateful application.
[0015] In some example embodiments, the migration module comprises: an object construction module configured to construct an eviction object according to the to-be-migrated container instance; The eviction request sending module is configured to send an eviction request to an application programming interface service of the computing power cluster based on the eviction object, so that the computing power cluster migrates the to-be-migrated container instance from the risk computing power node to a non-risk computing power node of the computing power cluster in response to the eviction request.
[0016] In some example embodiments, the migration module further includes: The migration process monitoring module is configured to monitor a migration process of the to-be-migrated container instance and generate a migration event record corresponding to the to-be-migrated container instance. The migration event record includes timestamp information, a source computing power node identifier, a target computing power node identifier, a migrated container instance identifier, and a migration result.
[0017] In some example embodiments, the device further includes: The tainted label adding module is configured to add a tainted label to the risk computing power node indicated by the risk computing power node identification result after obtaining the risk computing power node identification result. The tainted label removing module is configured to remove the tainted label from the risk computing power node when the risk computing power node is determined to be a non-risk computing power node after the to-be-migrated container instance is migrated to a non-risk computing power node of the computing power cluster.
[0018] In another aspect, an electronic device is provided, including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the container instance migration method of any of the above aspects.
[0019] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the container instance migration method of any of the above aspects.
[0020] In another aspect, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the container instance migration method of any of the above aspects.
[0021] The embodiment of the application obtains at least one running index time series data corresponding to each computing power node in the computing power cluster, the at least one running index time series data corresponds to at least one running index dimension one by one, and the at least one running index dimension is used to represent the health status of the computing power node. Then, based on the matching condition of the at least one running index time series data corresponding to each computing power node and the preset risk condition, the risk computing power node is identified, and the risk computing power node identification result is obtained, so that the risk computing power node with overload, overheating or potential failure risk can be identified in advance. When the risk computing power node identification result indicates that there is a risk computing power node, the to-be-migrated container instance is determined from the container instance running on the risk computing power node, and the to-be-migrated container instance is migrated to the non-risk computing power node of the computing power cluster for running, so that after the risk computing power node is identified, the to-be-migrated container instance running on the risk computing power node is automatically and actively migrated to the healthy computing power node for running before the risk computing power node completely fails. This predictive migration effectively avoids business interruption caused by computing power node hardware or performance problems, improves the stability of the cluster, and significantly improves the overall availability and reliability of the service. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0023] Figure 1 is a schematic diagram of an application environment of a container instance migration method provided by the embodiment of the present application; Figure 2 is a flowchart of a container instance migration method provided by the embodiment of the present application; Figure 3 is a flowchart of another container instance migration method provided by the embodiment of the present application; Figure 4 is a schematic diagram of an architecture of a container instance migration system provided by the embodiment of the present application; Figure 5 is a structural block diagram of a container instance migration device provided by the embodiment of the present application; Figure 6 is a hardware structural block diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0024] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of the present application.
[0025] It should be noted that the terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0027] It can be understood that in the specific embodiments of the present application, data related to user information and the like is involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in countries and regions.
[0028] In the prior art, the failure forms of computing power nodes in a computing power cluster are various, including but not limited to hardware layer failures such as central processing unit (CPU) or graphics processing unit (GPU) overheating, hard disk damage, disk I / O performance sharp decline, network interface card (NIC) failure, etc.; and also including software or configuration layer problems such as GPU driver unable to load normally, core system process crash, network storage (such as NFS, Ceph) mounting failure leading to that a persistent volume (PV) of a container cannot be accessed, etc.
[0029] For the fault of computing nodes, the Kubernetes cluster (also known as K8s) in the related art provides a built-in fault tolerance mechanism. Specifically, the core component kubelet periodically reports the state of the computing node to the control plane (Control Plane). When a computing node cannot report the heartbeat for a long time due to network partition, downtime or kubelet process crash, the control plane will mark it as NotReady. After a configurable waiting time (pod-eviction-timeout), the pods (container instances) running on the computing node will be automatically evicted and rebuilt on other healthy computing nodes by kube-scheduler. This built-in fault tolerance mechanism is a passive and reactive fault handling mode, and its main disadvantages are: 1) processing lag: the mechanism starts the recovery process after the computing node has already failed (such as downtime, disconnection). From the occurrence of the fault to the detection of the system, and then to the completion of the pod migration and reconstruction, there is a time window of minutes or even hours, during which the service running on the faulty computing node is completely unavailable, which may cause business interruption. 2) Unable to cope with performance degradation: For those "sub-healthy" nodes that are not completely down but have serious performance degradation (for example, CPU / GPU with high temperature and reduced frequency, or I / O congested disk), kubelet may still be able to report the heartbeat normally, resulting in the node state still being Ready. The default mechanism of Kubernetes cannot identify this performance degradation, so it will not trigger migration, causing the performance of the application running on it to be damaged and affecting user experience. 3) Lack of predictability: The mechanism relies entirely on after-the-fact remedies and cannot predict impending failures based on the evolving trends of node state (such as continuously rising temperature and load lingering at high levels), thus missing the valuable opportunity to actively avoid risks before the failure occurs.
[0030] For traditional HPC scheduling systems such as Slurm, their fault tolerance mechanisms are usually for batch jobs. For example, Slurm allows jobs to be automatically requeued after failure, but this is also a reactive mechanism and mainly targets stateless and reentrant computing tasks, which is not suitable for stateful online services that need to run continuously 7x24 hours.
[0031] Therefore, the related art has problems of reaction lag, inability to handle "sub-healthy" state (gradual performance degradation of computing nodes) and lack of predictability in handling computing node faults, resulting in poor cluster stability, low service availability and reliability.
[0032] In view of this, the embodiment of the present application provides a container instance migration method, by obtaining at least one running index time series data corresponding to each computing power node in the computing power cluster, the at least one running index time series data corresponds to at least one running index dimension one by one, and the at least one running index dimension is used to represent the health status of the computing power node, and then based on the matching condition of the at least one running index time series data corresponding to each computing power node and the preset risk condition, the risk computing power node is identified, and the risk computing power node identification result is obtained, so as to identify the risk computing power node with overload, overheating or potential fault risk in advance, and when the risk computing power node identification result indicates that there is a risk computing power node, the to-be-migrated container instance is determined from the container instance running on the risk computing power node, and the to-be-migrated container instance is migrated to the non-risk computing power node of the computing power cluster for running, so that after the risk computing power node is identified, the to-be-migrated container instance running on it is automatically and actively migrated to a healthy computing power node for running before the risk computing power node completely fails, which effectively avoids business interruption caused by computing power node hardware or performance problems, improves the stability of the cluster, and significantly improves the overall availability and reliability of the service.
[0033] Please refer to Figure 1 , which shows a system application environment diagram of a container instance migration method provided by the embodiment of the present application, including a server 110 and a computing power cluster 120, the computing power cluster 120 includes a plurality of computing power nodes (computing power node A, computing power node B, …, computing power node X), the server 110 can be connected and communicated with each computing power node in the computing power cluster 120, and one or more container instances (Pod) are running on each computing power node.
[0034] The cluster type of the computing power cluster 120 can include but is not limited to a Kubernetes cluster and a Slurm cluster, and the computing power cluster 120 can be a heterogeneous computing power cluster. The heterogeneous computing power in the heterogeneous computing power cluster can include but is not limited to a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), an FPGA (Field Programmable Gate Array) hardware, etc.
[0035] Each computing node in the computing cluster 120 is deployed with an operating indicator data collection component (Exporter). Specifically, this deployment can be implemented through DaemonSet to ensure that each computing node automatically runs the corresponding collector instance. The operating indicator data collection component (Exporter) collects real-time operating indicator data from the corresponding computing node across multiple operating indicator dimensions. These dimensions can reflect the health of the computing node and may include CPU usage, average load over a preset time period (e.g., 1 minute, 5 minutes, or 15 minutes), memory usage and pressure, I / O usage, and the number of bytes sent and received on the network. For heterogeneous computing nodes that include GPUs, these multiple operating indicator dimensions may also include the current GPU temperature, GPU utilization, memory bandwidth utilization, and a count of serious GPU errors. In a specific implementation, the operating indicator data collection component (Exporter) for each computing node may include a node operating indicator data collection component (Node Exporter) for exposing hardware and operating system kernel-related metrics of the computing node, and a GPU operating indicator data collection component (GPU Exporter) for specifically collecting GPU operating metrics.
[0036] Server 110 is used to automatically discover and periodically pull the operating indicator data from the operating indicator data collection component (Exporter) deployed on each computing power node, and store it in the form of a time series in a time series database. Then, based on the operating indicator time series data of each computing power node in each operating indicator dimension in the time series database, the container instance migration method of the embodiment of the present application is executed to actively and imperceptibly migrate the container instance on the computing power node to a healthy computing power node before a computing power node failure occurs.
[0037] It should be noted that the nodes / servers involved in the embodiments of this application can be independent physical servers, or server clusters or distributed systems composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal devices involved in the embodiments of this application include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals.
[0038] In one exemplary embodiment, each computing node of server 110 and computing cluster 120 may be a node device in a blockchain system, capable of sharing acquired and generated information with other node devices in the blockchain system, thereby enabling information sharing among multiple node devices. Multiple node devices in a blockchain system may be configured with the same blockchain, which is composed of multiple blocks, and adjacent blocks are associated with each other. This allows any tampering of data in any block to be detected by the next block, thereby preventing tampering of data in the blockchain and ensuring the security and reliability of the data in the blockchain.
[0039] See also Figure 2 , which shows a flow chart of a container instance migration method provided by an embodiment of the present application, which can be applied to Figure 1 The server 110 in the embodiment. It should be noted that this specification provides the method operation steps as described in the embodiment or flowchart, but based on conventional or non-creative work, more or fewer operation steps may be included. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the drawings. Specifically, Figure 2 As shown, the container instance migration method of the embodiment of the present application may include: S201, obtaining at least one operation indicator time series data corresponding to each computing node in the computing power cluster.
[0040] The at least one operational indicator time series data corresponds one-to-one to at least one operational indicator dimension, and the at least one operational indicator dimension is used to characterize the health status of the computing power node. The operational indicator time series data of the computing power node in each operational indicator dimension includes operational indicator data for that operational indicator dimension collected within a preset time period before the current time, arranged in chronological order of collection time.
[0041] For example, at least one operating indicator dimension may include but is not limited to CPU usage, average load within a preset time period (such as 1 minute, 5 minutes, 15 minutes), memory usage and pressure, I / O usage, number of bytes sent and received on the network, GPU temperature, GPU utilization, video memory bandwidth utilization, and count of serious errors occurring on the GPU.
[0042] In a specific implementation, at least one operating indicator time series data corresponding to each computing node in the computing cluster can be obtained according to a preset time interval. The preset time interval can be set based on actual needs. For example, it can be obtained once every hour, and the container migration method of the embodiment of the present application can be executed once every hour.
[0043] S203, performing risk computing power node identification based on matching of the at least one running index time series data corresponding to each of the computing power nodes and the preset risk condition, to obtain a risk computing power node identification result.
[0044] The risk computing power node identification result indicates whether there is a risk computing power node at present. The risk computing power node refers to a "sub-healthy" computing power node with risks of overload, overheating or potential failure. It can be understood that the computing power nodes other than the risk computing power nodes in the computing power cluster can be considered as non-risk computing power nodes, or healthy computing power nodes.
[0045] The preset risk condition is a risk identification rule pre-set for indicating the existence of risks. The risk identification rule can be associated with a change trend of the data in the running index dimension within a period of time.
[0046] Specifically, the preset risk condition includes at least one preset risk identification rule corresponding to at least one running index dimension one by one. The above step S203 can include the following steps when implemented: traversing the at least one preset risk identification rule, for the current preset risk identification rule traversed, matching the target running index time series data of each of the computing power nodes with the current preset risk identification rule, if matched, determining that the computing power node corresponding to the matched target running index time series data is a risk computing power node, obtaining a risk computing power node identification sub-result corresponding to the current preset risk identification rule; the target running index time series data is the running index time series data of the running index dimension corresponding to the current preset risk identification rule; determining the risk computing power node identification result based on the risk computing power node identification sub-result corresponding to each of the preset risk identification rules.
[0047] Specifically, the risk computing power node identification sub-result corresponding to each of the preset risk identification rules is aggregated, and the risk computing power node identification result is obtained. For example, the risk computing power node identification sub-result of the preset risk identification rule corresponding to the GPU temperature indicates that the computing power node A is a risk computing power node, and the risk computing power node identification sub-result of the preset risk identification rule corresponding to the CPU usage rate indicates that the computing power node B and the computing power node A are both risk computing power nodes. Therefore, the finally determined risk computing power node identification result is {computing power node A computing power node B}.
[0048] Taking the running index dimension corresponding to the current preset risk identification rule as an example, the GPU temperature, in a specific implementation, a PromQL query can be constructed: avg_over_time((DCGM_FI_DEV_GPU_TEMP{node_name=~".*"}>85)[5m:30s]), which is used to return all computing power nodes with a temperature higher than 85 degrees for the past 5 minutes. The computing power node list returned by executing the query can be obtained. Then, based on the GPU temperature time series data of each computing power node in the computing power node list in the last 1 hour, the average GPU temperature of each computing power node in 1 hour is calculated, and the computing power node with an average GPU temperature greater than or equal to 85 degrees Celsius is determined as a risk computing power node, obtaining the risk computing power node identification sub-result of the current preset risk identification rule.
[0049] The preset risk identification rule corresponding to the multiple dimensions and fine-grained running index in the above embodiment can comprehensively and accurately identify the risk computing power node in the computing power cluster, which is beneficial to further improve the overall availability and reliability of the service.
[0050] S205, when the risk computing power node identification result indicates that there is a risk computing power node, determining a to-be-migrated container instance from the container instances running on the risk computing power node.
[0051] Specifically, when the risk computing power node identification result indicates that there is a risk computing power node, an application programming interface service of the computing power cluster can be called to determine all container instances running on the risk computing power node through the application programming interface service, obtaining a container instance set of the risk computing power node; and according to a preset filtering strategy, a to-be-migrated container instance is selected from the container instance set of the risk computing power node, obtaining a to-be-migrated container instance list. The preset filtering strategy includes: excluding container instances marked as prohibited migration in the annotation, i.e. container instances marked as prohibited migration in the annotation are not to-be-migrated container instances; excluding container instances belonging to a preset key system component list, i.e. container instances belonging to key system components are not to-be-migrated container instances; and the remaining available copy number of the application to which each container instance belongs in the computing power cluster meets the minimum copy number configuration.
[0052] Taking a Kubernetes cluster as an example, the preset key system component list can include the Pods of system components in core namespaces such as kube-system, which are not used as container instances to be migrated, so as to avoid affecting the stability of the cluster itself during the migration operation. The container instances marked as prohibited migration in the annotations are not used as container instances to be migrated, for example, if the annotations of a container instance contain operator.example.com / allow-migration:"false", the container instance is not used as a container instance to be migrated.
[0053] By checking the PDB (Pod Disruption Budget) of the application to which the container instance belongs, the minimum replica number configuration of the application can be determined to ensure that this migration does not violate the PDB agreement, thereby avoiding service replica number shortage caused by migration and service interruption, for example, if the PDB of the application app1 to which the Pod1 on the risk computing power node A belongs defines the minimum replica number as 4, the computing power cluster currently has 5 Pods belonging to the application app1, and the container instance set of the risk computing power node A has two Pods belonging to the application app1, only one of the two Pods can be selected as a container instance to be migrated, at this time, the remaining available replica number of the application app1 in the computing power cluster is 4, which meets the minimum replica number configuration (4) of the application, if both of the two Pods are used as container instances to be migrated, the remaining available replica number of the application app1 in the computing power cluster is 3, which does not meet the minimum replica number configuration (4) of the application.
[0054] By using the preset screening strategy to screen the container instances to be migrated, the Pods can be automatically evacuated from the risk computing power node to the healthy computing power node under the premise of minimizing service interruption risk, complying with application constraints (PDB), and ensuring cluster stability (protecting system components).
[0055] In an exemplary embodiment, the preset screening strategy further includes a migration sorting strategy, and the container instances to be migrated in the list of container instances to be migrated are sorted according to the migration sorting strategy, wherein the migration sorting strategy includes: using the priority information marked in the annotations to determine the migration order of the corresponding container instances to be migrated; and the container instances to be migrated belonging to stateless applications are given priority over the container instances to be migrated belonging to stateful applications.
[0056] Specifically, if the annotation of the to-be-migrated container instance includes operator.example.com / migration-priority: "high", it can be determined that the to-be-migrated container instance is high priority and needs to be migrated first. In addition, stateful applications and stateless applications are distinguished. Since the reconstruction process of the to-be-migrated container instance belonging to the stateless application is simple, the to-be-migrated container instance (such as a Pod controlled by Deployment or ReplicaSet) belonging to the stateless application is migrated first, and for the to-be-migrated container instance (such as a Pod of StatefulSet) belonging to the stateful application, migration needs to ensure that the associated persistent storage (Persistent Volume, PV) can be remounted on the new computing node. Through the migration sorting strategy, the subsequent migration process of the to-be-migrated container instance can be ensured to be orderly controllable.
[0057] S207, migrating the to-be-migrated container instance to run on a non-risk computing node of the computing cluster.
[0058] Specifically, according to the order of the to-be-migrated container instance list, each to-be-migrated container instance is migrated to run on a non-risk computing node of the computing cluster in turn.
[0059] Exemplarily, when migrating the to-be-migrated container instance to run on a non-risk computing node of the computing cluster can include: constructing an eviction object according to the to-be-migrated container instance; sending an eviction request to an application programming interface service of the computing cluster based on the eviction object, so that the computing cluster responds to the eviction request and migrates the to-be-migrated container instance from the risk computing node to run on a non-risk computing node of the computing cluster.
[0060] Embodiments of the present application do not directly delete the to-be-migrated container instance on the risk computing node, but use the API provided by the computing cluster for providing the eviction service (Eviction). The eviction object, such as the policy / v1 / Eviction object, can be first constructed according to the to-be-migrated container instance, and the eviction object is submitted to the application programming interface service of the computing cluster, such as the Kubernetes API Server, through the POST request, so that the Kubernetes will automatically complete the rescheduling and migrate the to-be-migrated container instance from the risk computing node to run on a non-risk computing node of the computing cluster. Compared with direct deletion, it can be safer and more orderly, follow the standard eviction process of the computing cluster, and ensure that the application has a cleaning opportunity.
[0061] In some exemplary embodiments, as shown in Figure 3 The method can further include: S301, monitoring the migration process of the to-be-migrated container instance, and generating a migration event record corresponding to the to-be-migrated container instance.
[0062] The migration event record includes timestamp information, source computing power node identifier, target computing power node identifier, migrated container instance identifier, and migration result.
[0063] Specifically, the server can monitor the migration process and record the migration event through the migration status field (status) to generate the corresponding migration event record. The timestamp information in the migration event record represents the start time of the migration, the source computing power node identifier indicates the risk computing power node running the to-be-migrated container instance before migration, for example, node-1, the target computing power node identifier indicates the non-risk computing power node running the to-be-migrated container instance after migration, for example, node-3, the migrated container instance identifier represents the migrated Pod, and the migration result represents whether the migration is successful. By monitoring the migration process and recording the migration event in detail, it is convenient for operation and maintenance audit and troubleshooting.
[0064] In some exemplary embodiments, after obtaining the risk computing power node identification result, the foregoing step S203 can further include: adding a taint label to the risk computing power node indicated by the risk computing power node identification result. Correspondingly, after migrating the to-be-migrated container instance to run on the non-risk computing power node of the computing power cluster, the method further includes: removing the taint label on the risk computing power node when it is determined that the risk computing power node is a non-risk computing power node.
[0065] Specifically, a NoSchedule taint (Taint) label can be added to these risk computing power nodes by calling the application programming interface service of the computing power cluster, so that subsequent container instance pods can be prevented from being scheduled to these risk computing power nodes. After the migration of the to-be-migrated container instance on the risk computing power node is completed, if it is determined that the risk computing power node recovers to be a non-risk computing power node, for example, in the next risk computing power node identification, the risk computing power node identification result indicates that it is a non-risk computing power node, then the taint label added to the risk computing power node is removed, so that the computing power node can receive new Pod scheduling again.
[0066] It can be understood that in the embodiments of the present application, when the risk computing power node identification result indicates that there is no risk computing power node at each time, no container instance migration is performed.
[0067] The technical scheme of the embodiment of the application can identify risk computing power nodes with risks of overload, overheating or potential failure in advance. After identifying the risk computing power nodes, the application automatically and actively migrates the to-be-migrated container instances running on the risk computing power nodes to healthy computing power nodes before the risk computing power nodes completely fail. The predictive migration effectively avoids service interruption caused by hardware or performance problems of the computing power nodes, improves the stability of the cluster, and significantly improves the overall availability and reliability of the service. For critical services, this means reducing the downtime from minutes to nearly zero non-sense migration, greatly improving the service SLA (service level agreement), and completely automating the process without human intervention. This greatly liberates the operation and maintenance personnel from tedious and repetitive computing power node monitoring and manual migration work, reduces the possibility of human error, and achieves 7x24-hour uninterrupted intelligent cluster self-healing capability.
[0068] To facilitate understanding of the technical scheme of the embodiments of the application, the Kubernetes computing power cluster is taken as an example and the specific system example is combined with Figure 4 for detailed description. As shown in Figure 4 , the container instance migration method of the embodiments of the application can be implemented through the cooperative work among the node state monitoring and index collection module, the self-defined controller and predictive analysis module, and the active migration execution module. The three modules are described in detail as follows.
[0069] Node state monitoring and index collection module: The node state monitoring and index collection module can comprehensively, real-timely and accurately obtain the running state of each computing power node in the computing power cluster.
[0070] As shown in Figure 4 , the index collection component (Exporter) is deployed in each computing power node of the Kubernetes computing power cluster through the DaemonSet mode. The DaemonSet ensures that each node will automatically run an instance of the collector. The index collection component (Exporter) includes the Node Exporter and the GPU Exporter.
[0071] Node Exporter: It is the most basic collector in the Prometheus ecosystem, used to expose node's hardware and operating system kernel-related metrics. The core metrics collected by it include: 1) node_cpu_seconds_total: CPU usage time, which can calculate CPU usage rate. 2) node_load1, node_load5, node_load15: System 1, 5, 15-minute average load. 3) node_memory_MemAvailable_bytes: Available memory size, which can calculate memory usage rate and pressure. 4) node_disk_io_time_seconds_total: Disk I / O time spent, which can calculate I / O usage rate. 5) node_network_receive_bytes_total, node_network_transmit_bytes_total: Network receive and transmit byte count.
[0072] GPU Exporter: A specialized GPU metrics collector for heterogeneous computing nodes containing GPUs. This component can provide GPU state data, including: 1) DCGM_FI_DEV_GPU_TEMP: Current GPU temperature (in Celsius), which is a key indicator for predicting overheating failures. 2) DCGM_FI_DEV_GPU_UTIL: GPU utilization. 3) DCGM_FI_DEV_MEM_COPY_UTIL: Memory bandwidth utilization. 4) DCGM_FI_DEV_XID_ERRORS: Serious error count of GPU, which is a strong signal of hardware failure.
[0073] Deploy the Prometheus service in the cluster and configure its scrape_configs to automatically discover and periodically pull (scrape) running metrics data from the above-mentioned Exporters deployed on each computing node. All collected data is labeled with labels identifying the source computing node and specific metrics, and stored in the Prometheus built-in time series database (TSDB) in the form of time series, providing a real-time and reliable data source for subsequent analysis and decision-making.
[0074] Customized Controller and Predictive Analysis Module: Custom controllers and predictive analysis modules can be implemented using custom controllers running in the Kubernetes control plane. A KubernetesOperator can be developed using the Kubebuilder framework as a custom controller. This controller follows the Operator pattern, monitoring changes in specific resources by watching the Kubernetes API server and executing the corresponding reconciliation logic (reconciliation loop). To flexibly configure migration policies, one or more custom resource definitions (CRDs) can be defined, creating a CRD named PredictiveMigrationPolicy, which represents a predictive migration policy. In its reconciliation loop, the controller first calls the Prometheus HTTP API through an HTTP client to pull data from the node status monitoring and metrics collection modules, based on the scrapeInterval period (data pull frequency) defined in the PredictiveMigrationPolicy.
[0075] When performing analysis and prediction, the controller iterates over all rules defined in the PredictiveMigrationPolicy CRD. For each rule, it performs the following operations: 1) constructs a complete PromQL query; 2) executes the query and retrieves a list of compute nodes. 3) identifies risky compute nodes based on whether the time series data of each compute node's operating metrics matches the rule. To prevent subsequent pods from being scheduled to these nodes, the controller immediately adds a "NoSchedule" taint label to these risky compute nodes through the Kubernetes API.
[0076] Active migration execution module: When the controller identifies a risky computing node, the active migration execution module is responsible for executing the Pod migration operation safely and orderly.
[0077] First, select the Pods to be migrated: The controller calls the Kubernetes API to list all the Pods running on the at-risk compute node. The controller follows the following preset filtering strategies to decide which Pods to migrate and the order of migration: 1) Follow PodDisruptionBudget (PDB): First, check if the application to which the target Pod belongs has a PDB configured. PDB defines the minimum number of replicas that an application must maintain at any point in time. The controller checks the PDB status before initiating eviction to ensure that this eviction does not violate the PDB agreement, thereby avoiding service replica shortages caused by migration. 2) Based on priority and annotations: Administrators can set specific annotations (Annotations) for Pods to guide migration decisions, such as operator.example.com / migration-priority: "high" or operator.example.com / allow-migration: "true". The controller will prefer to migrate high-priority Pods and skip Pods marked as non-migratable. 3) Differentiate between stateful and stateless: Generally, stateless applications (such as Deployment or ReplicaSet controlled Pods) are preferred for migration because their reconstruction process is simple. For stateful applications (StatefulSet), migration needs to ensure that their associated persistent storage (PV) can be re-mounted on the new node. 4) Exclude critical system components: The controller will configure an exclusion list to ignore system component Pods in core namespaces such as kube-system, avoiding the impact of migration operations on the stability of the cluster itself.
[0078] Then initiate Pod eviction (Eviction): For the selected target Pods, the controller does not directly delete the Pod object, but uses the Eviction API provided by Kubernetes. It constructs a policy / v1 / Eviction object and submits it to the Kubernetes API Server through a POST request. Kubernetes will automatically complete the rescheduling.
[0079] The controller monitors the migration process and can record migration events in the status field of its CRD, including timestamp, source compute node, target compute node, migrated Pod name, and whether the migration was successful. After migration is complete, if the health indicators of the at-risk compute node return to normal, the controller will remove the previously added taint, allowing the compute node to receive new Pod scheduling again.
[0080] The above implementation mode realizes predictive fault avoidance and active container instance migration, minimizes service interruption time caused by node hardware or software problems through a "prediction-migration" closed-loop mechanism, minimizes the impact of faults on business, and significantly improves the overall availability and reliability of services; avoids that the container application continuously runs on a performance-degraded computing power node, ensures that it can obtain stable and efficient computing resources, and also provides an effective technical means for realizing green intelligent calculation and load balancing, optimizes cluster resource performance; and reduces the artificial monitoring and intervention burden of the operation and maintenance personnel on the node state, and improves the cluster management efficiency.
[0081] Corresponding to the container instance migration method provided by the above several embodiments, the embodiment of the application also provides a container instance migration device. Since the container instance migration device provided by the embodiment of the application corresponds to the container instance migration method provided by the above several embodiments, the implementation mode of the foregoing container instance migration method is also applicable to the container instance migration device provided by the embodiment of the application, which will not be described in detail in this embodiment.
[0082] Please refer to Figure 5 which is a structural schematic diagram of a container instance migration device provided by the embodiment of the application. The device has the function of realizing the container instance migration method in the above method embodiment, which can be realized by hardware or corresponding software executed by hardware. As Figure 5 shown, the container instance migration device 500 can include: The running index data acquisition module 510 is configured to acquire at least one running index time series data corresponding to each computing power node in the computing power cluster; the at least one running index time series data corresponds to at least one running index dimension one by one, and the at least one running index dimension is used to represent the health state of the computing power node; The risk computing power node identification module 520 is configured to identify the risk computing power node based on the matching of the at least one running index time series data corresponding to each computing power node and the preset risk condition, and obtain a risk computing power node identification result; The to-be-migrated container instance determination module 530 is configured to determine a to-be-migrated container instance from the container instance running on the risk computing power node when the risk computing power node identification result indicates that there is a risk computing power node; The migration module 540 is configured to migrate the to-be-migrated container instance to run on a non-risk computing power node of the computing power cluster.
[0083] In some example embodiments, the preset risk condition comprises at least one preset risk identification rule corresponding to the at least one operation index dimension; the risk computing power node identification module 520 is specifically configured to: traverse the at least one preset risk identification rule, for a current preset risk identification rule traversed, match the target operation index time series data of each computing power node with the current preset risk identification rule respectively, if matched, determine that the computing power node corresponding to the target operation index time series data matched is a risk computing power node, obtain a risk computing power node identification sub-result corresponding to the current preset risk identification rule; the target operation index time series data is the operation index time series data of the operation index dimension corresponding to the current preset risk identification rule; determine the risk computing power node identification result based on the risk computing power node identification sub-result corresponding to each preset risk identification rule.
[0084] In some example embodiments, the to-be-migrated container instance determination module 530 comprises: a container instance set determination module configured to, when the risk computing power node identification result indicates that there is a risk computing power node, call an application programming interface service of the computing power cluster, determine a container instance running on the risk computing power node, and obtain a container instance set of the risk computing power node; a screening module configured to select a to-be-migrated container instance from the container instance set of the risk computing power node according to a preset screening strategy, and obtain a to-be-migrated container instance list; wherein, the preset screening strategy comprises: excluding a container instance marked as prohibited migration in an annotation; excluding a container instance belonging to a key system component in a preset key system component list; the remaining available replica number of an application to which each container instance belongs in the computing power cluster meets a minimum replica number configuration.
[0085] In some example embodiments, the preset screening strategy further comprises a migration sorting strategy, and the to-be-migrated container instances in the to-be-migrated container instance list are sorted according to the migration sorting strategy; wherein, the migration sorting strategy comprises: using priority information marked in an annotation to determine the migration order of the corresponding to-be-migrated container instance; a to-be-migrated container instance belonging to a stateless application is prior to a to-be-migrated container instance belonging to a stateful application.
[0086] In some example embodiments, the migration module 540 comprises: an object construction module configured to construct an eviction object according to the to-be-migrated container instance; The eviction request sending module is configured to send an eviction request to an application programming interface service of the computing power cluster based on the eviction object, so that the computing power cluster migrates the to-be-migrated container instance from the risk computing power node to a non-risk computing power node of the computing power cluster in response to the eviction request.
[0087] In some example embodiments, the migration module 540 further includes: A migration process monitoring module is configured to monitor a migration process of the to-be-migrated container instance, and generate a migration event record corresponding to the to-be-migrated container instance. The migration event record includes timestamp information, a source computing power node identifier, a target computing power node identifier, a migrated container instance identifier, and a migration result.
[0088] In some example embodiments, the device 500 further includes: A tainted label adding module is configured to add a tainted label to the risk computing power node indicated by the risk computing power node identification result after obtaining the risk computing power node identification result. A tainted label removing module is configured to remove the tainted label on the risk computing power node when the risk computing power node is determined to be a non-risk computing power node after the to-be-migrated container instance is migrated to a non-risk computing power node of the computing power cluster for running.
[0089] It should be noted that the device provided in the above embodiments, in realizing its functions, is only exemplified by the above division of each functional module. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules, such as the node state monitoring and index collecting module, the self-defined controller and prediction analysis module, and the active migration execution module shown in the above embodiments, to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here. Figure 4
[0090] The electronic device provided in the embodiments of the present application includes a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or the at least one program is loaded and executed by the processor to implement any one of the container instance migration methods provided in the embodiments of the present application.
[0091] The memory can be used to store software programs and modules, and the processor performs various function applications and data processing by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required for functions, etc.; and the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory can also include a memory controller to provide access for the processor to the memory.
[0092] The method embodiments provided by the embodiments of the application can be executed in a computer terminal, a server or a similar computing device, that is, the above-mentioned electronic device can include a computer terminal, a server or a similar computing device. Taking the case of running on a server as an example, Figure 6 is a hardware structure block diagram of a server running a container instance migration method provided by the embodiments of the application, as Figure 6 shown, the server 600 can have a large difference due to different configurations or performances, and can include one or more central processing units (CPU) 610 (the central processing unit 610 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 630 for storing data, one or more storage media 620 (such as one or more mass storage devices) for storing application programs 623 or data 622. Among them, the memory 630 and the storage medium 620 can be temporary storage or persistent storage. The programs stored in the storage medium 620 can include one or more modules, and each module can include a series of instruction operations in the server. Further, the central processing unit 610 can be configured to communicate with the storage medium 620 and execute a series of instruction operations in the storage medium 620 on the server 600. The server 600 can also include one or more power supplies 660, one or more wired or wireless network interfaces 650, one or more input and output interfaces 640, and / or one or more operating systems 621, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0093] The input / output interface 640 can be configured to receive or transmit data via a network. Examples of the network can include a wireless network provided by a communication provider of the server 600. In an example, the input / output interface 640 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the input / output interface 640 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.
[0094] Those skilled in the art can understand that, Figure 6 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 600 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 6 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 600 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 6 The structure shown is only schematic, and does not limit the structure of the electronic device. For example, the server 600 can further include more or fewer components than those shown, or have a different configuration of components than those shown.
[0095] Embodiments of the present application also provide a computer readable storage medium, which can be arranged in an electronic device to save at least one instruction or at least one program for implementing a container instance migration method. The at least one instruction or the at least one program is loaded and executed by the processor to implement any container instance migration method provided by the above method embodiments.
[0096] Embodiments of the present application also provide a computer program product or computer program, which includes computer instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the electronic device execute any container instance migration method provided by the above method embodiments.
[0097] Optionally, in the present embodiment, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0098] It should be noted that the above-mentioned order of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. And the above describes the specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0099] Each of the embodiments in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0100] A person of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk.
[0101] The above only describes the preferred embodiments of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A container instance migration method, characterized in that: The method comprises: Obtain at least one operating indicator time series data corresponding to each computing node in the computing power cluster; the at least one operating indicator time series data corresponds one-to-one to at least one operating indicator dimension, and the at least one operating indicator dimension is used to represent the health status of the computing power node; Based on the matching of at least one operation indicator time series data corresponding to each of the computing power nodes and the preset risk condition, risk computing power node identification is performed to obtain a risk computing power node identification result; When the risky computing power node identification result indicates that a risky computing power node exists, determining a container instance to be migrated from container instances running on the risky computing power node; Migrate the container instance to be migrated to a non-risk computing node in the computing cluster for operation.
2. The method according to claim 1, characterized in that The preset risk condition includes at least one preset risk identification rule corresponding to the at least one operating indicator dimension; the risk computing power node identification is performed based on the matching of the at least one operating indicator time series data corresponding to each computing power node and the preset risk condition, and the risk computing power node identification result obtained includes: Traversing the at least one preset risk identification rule, for the traversed current preset risk identification rule, matching the target operating indicator time series data of each computing power node with the current preset risk identification rule respectively; if there is a match, determining that the computing power node corresponding to the matched target operating indicator time series data is a risk computing power node, and obtaining the risk computing power node identification sub-result corresponding to the current preset risk identification rule; the target operating indicator time series data is the operating indicator time series data of the operating indicator dimension corresponding to the current preset risk identification rule; The risk computing power node identification result is determined based on the risk computing power node identification sub-results corresponding to each of the preset risk identification rules.
3. The method according to claim 1, characterized in that When the risky computing power node identification result indicates that a risky computing power node exists, determining the container instance to be migrated from the container instance running on the risky computing power node includes: When the risk computing power node identification result indicates that a risk computing power node exists, calling the application programming interface service of the computing power cluster, determining the container instance running on the risk computing power node, and obtaining the container instance set of the risk computing power node; Selecting a container instance to be migrated from the container instance set of the risk computing power node according to a preset screening strategy to obtain a list of container instances to be migrated; Among them, the preset screening strategy includes: excluding container instances marked as prohibited from migration in the annotation; excluding container instances that belong to key system components in the preset key system component list; the number of remaining available copies of the application to which each container instance belongs in the computing power cluster meets the minimum copy number configuration.
4. The method according to claim 3, characterized in that The preset screening strategy further includes a migration sorting strategy, and the container instances to be migrated in the list of container instances to be migrated are sorted according to the migration sorting strategy; The migration ordering strategy includes: using priority information marked in the annotation to determine the migration order of the corresponding container instances to be migrated; and giving priority to container instances to be migrated belonging to stateless applications over container instances to be migrated belonging to stateful applications.
5. The method according to claim 1, wherein Migrating the container instance to be migrated to a non-risk computing node in the computing cluster for operation includes: Constructing an eviction object according to the container instance to be migrated; An eviction request is sent to the application programming interface service of the computing power cluster based on the eviction object, so that the computing power cluster migrates the container instance to be migrated from the risky computing power node to a non-risk computing power node of the computing power cluster for operation in response to the eviction request.
6. The method according to claim 5, characterized in that The method further comprises: The migration process of the container instance to be migrated is monitored to generate a migration event record corresponding to the container instance to be migrated; the migration event record includes timestamp information, source computing node identifier, target computing node identifier, migrated container instance identifier and migration result.
7. The method according to claim 1, characterized in that After obtaining the risk computing power node identification result, the method further includes: Adding a taint label to the risky computing power node indicated by the risky computing power node identification result; After migrating the container instance to be migrated to a non-risk computing node in the computing cluster for operation, the method further includes: When it is determined that the risky computing power node is a non-risky computing power node, the taint label on the risky computing power node is removed.
8. A container instance migration device, characterized in that: The device comprises: An operating indicator data acquisition module is used to acquire at least one operating indicator time series data corresponding to each computing power node in the computing power cluster; the at least one operating indicator time series data corresponds one-to-one with at least one operating indicator dimension, and the at least one operating indicator dimension is used to represent the health status of the computing power node; A risk computing power node identification module is used to identify risk computing power nodes based on the matching of at least one operating indicator time series data corresponding to each computing power node with a preset risk condition, and obtain a risk computing power node identification result; a module for determining container instances to be migrated, configured to determine, when the risky computing power node identification result indicates the presence of a risky computing power node, a container instance to be migrated from container instances running on the risky computing power node; The migration module is used to migrate the container instance to be migrated to a non-risk computing node in the computing cluster for operation.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the container instance migration method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the container instance migration method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Resource coordination method and coordination device for container cluster and storage medium
CN114064226A
Container group scheduling method and system, electronic equipment, cluster and readable storage medium
CN117667358A
Intelligent computing power scheduling method and device, electronic equipment and storage medium
CN118897723A
Container rescheduling method, system and equipment, storage medium and program product
CN120560776A