Node regression method, device and equipment and computer readable storage medium
By evacuating the service nodes and selecting the target service nodes for container group migration, the problem of inability to evaluate the service nodes suitable for vacancies in the existing technology is solved, and the vacancies success rate and resource recycling efficiency are improved.
Patent Information
- Application Number
- CN202311519365.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art cannot evaluate which service node is more suitable for vacancies when migrating container groups in service nodes, resulting in migration failure and affecting resource recycling efficiency.
By determining the set of service nodes for the resources to be recycled and the vacancies, each service node in the service node collection is vacancies simulation, candidate service nodes that can be vacancies are selected, and node scores are determined based on resource occupation information and resource categories, thereby selecting the target service node for container group migration and resource recycling.
This improves the vacancies success rate of service nodes, ensures resource recycling efficiency, and reduces the cases of vacancies failure.
Smart Images

Figure CN120017659A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a node evacuation method, device, equipment and computer-readable storage medium. Background Art
[0002] Cloud services provide computing resources and data storage through the Internet. They provide computing power, storage space, databases, applications and other resources to users, enterprises and other objects through the Internet. Taking computing power as an example, pod is a container group deployed in a service node, which is the smallest business operation unit. Cloud service providers can migrate the container group in the service node to vacate the service node and recycle the vacated service node.
[0003] When migrating a container group in a service node, the related technology first performs simulated scheduling on the container group in the service node, and then performs migration processing on the container group in the service node after the simulated scheduling passes, so as to free up service node resources.
[0004] During the research and practice of related technologies, the inventors of the present application discovered that when migrating container groups in service nodes, the related technologies are unable to evaluate which current service node is more suitable for vacating, which can easily lead to the failure of migration of container groups in service nodes, thereby affecting the recovery efficiency of service node resources. Summary of the invention
[0005] The embodiments of the present application provide a node vacating method, apparatus, device and computer-readable storage medium, which can select a target service node for vacating from the service nodes that have passed the simulated scheduling, thereby improving the success rate of vacating the service nodes and ensuring the resource recovery efficiency of the service nodes.
[0006] The present application provides a node evacuation method, including:
[0007] Determine a service node set of resources to be reclaimed, and determine a vacating speed, wherein the service node set includes at least one service node;
[0008] Performing a vacancy simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated, the set of candidate service nodes including at least one candidate service node;
[0009] Determine the node score of each candidate service node based on the resource occupancy information and resource category of each candidate service node;
[0010] Selecting a target service node from the set of candidate service nodes according to the evacuation speed and the node score;
[0011] A migration process is performed on the running container group on the target service node, so as to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0012] Accordingly, an embodiment of the present application provides a node evacuation device, including:
[0013] A first determining unit is used to determine a service node set of resources to be reclaimed and determine a vacating speed, wherein the service node set includes at least one service node;
[0014] A prediction unit, configured to perform a vacancy simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated, wherein the set of candidate service nodes includes at least one candidate service node;
[0015] A second determining unit, configured to determine a node score of each candidate service node according to resource occupancy information and resource category of each candidate service node;
[0016] A selection unit, configured to select a target service node from the set of candidate service nodes according to the evacuation speed and the node score;
[0017] The recycling unit is used to perform migration processing on the running container group on the target service node, so as to recycle the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0018] In some implementations, the second determining unit is further configured to:
[0019] Determine, according to the resource occupancy information of each candidate service node, a resource occupancy score for the deployed target service in each candidate service node;
[0020] Determine a resource category score corresponding to the resource category of each candidate service node;
[0021] A node score of each candidate service node is determined according to the resource occupancy score and the resource category score of each candidate service node.
[0022] In some implementations, the resource occupancy information includes a resource allocation rate of the candidate service node for the deployed target service and the number of container groups of the running container groups, and the second determining unit is further configured to:
[0023] Determine a resource allocation score of each candidate service node according to the resource allocation rate of each candidate service node; the resource allocation score is negatively correlated with the resource allocation rate;
[0024] Determine, according to the number of container groups running on each candidate service node, a service impact score when each candidate service node performs migration processing on the container group occupied by the target service, wherein the service impact score is negatively correlated with the number of container groups;
[0025] The resource occupancy score of each candidate service node is determined according to the resource allocation score and the business impact score of each candidate service node.
[0026] In some implementations, the second determining unit is further configured to:
[0027] Query resource categories of other service nodes outside the service node set;
[0028] Determine the number of service nodes corresponding to each resource category in the other service nodes;
[0029] Determine the weight value of the service node of each resource category in the other service nodes based on the number of service nodes corresponding to each resource category;
[0030] According to the weight value of the service node of each resource category, the priority coefficient of the resource category of each candidate service node is determined, and the resource category score of each candidate service node is determined according to the priority coefficient of each candidate service node.
[0031] In some implementations, the selection unit is further configured to:
[0032] Sorting the candidate service nodes in the candidate service node set in descending order of the node scores to obtain a candidate service node ranking;
[0033] Determine the number of nodes to be vacated per unit time according to the vacating speed;
[0034] According to the number of nodes, a target service node ranked first is selected from the candidate service nodes.
[0035] In some implementations, the prediction unit is further configured to:
[0036] Identify a service type of a target service deployed at each service node in the service node set;
[0037] Filtering the service nodes whose service types are preset service types in the service node set to obtain a filtered service node set;
[0038] Querying the historical evacuation record of each service node in the filtered service node set, and removing the service nodes that have failed to be evacuated in the historical evacuation record, to obtain an initial service node set;
[0039] Compatibility deployment prediction is performed on the container group on each service node in the initial service node set to predict a set of candidate service nodes that can perform evacuation.
[0040] In some implementations, the prediction unit is further configured to:
[0041] For each service node in the initial service node set, determine a target resource category compatible with each container group deployed on the service node;
[0042] Querying the service node set for the to-be-processed service nodes corresponding to the target resource category, and determining the total number of the to-be-processed service nodes;
[0043] If the total number of the service nodes to be processed is greater than a preset threshold, and it is detected that no other container group incompatible with the container group is deployed on the service node to be processed, the service node is determined as a candidate service node that can be vacated to obtain a set of candidate service nodes.
[0044] In some embodiments, the recovery unit is further used for:
[0045] Performing migration processing on each container group in the target service node to obtain a migration processing result;
[0046] If the migration process result is that the migration is successful, determining to recycle the computing resources of the target service node after the migration process;
[0047] If the migration processing result is migration failure, it is determined not to recycle the computing resources of the target service node.
[0048] In some embodiments, the recovery unit is further used for:
[0049] Filtering out a to-be-processed service node having the same resource category as the target service node from the service node set;
[0050] Adjust the configuration state of the target service node to an unschedulable state, and request the to-be-processed service node to create a target container group corresponding to the container group, to obtain a creation result;
[0051] If the creation result is that each target container group is successfully created, the corresponding container group in the target service node in the unschedulable state is deleted, and the migration processing result is determined to be a successful migration;
[0052] If the unit creation result is that at least one target container group fails to be created, then the migration processing result is determined to be a migration failure.
[0053] In some embodiments, the recovery unit is further used for:
[0054] receiving a deregistration request for a target service node after the migration process, wherein the deregistration request carries target account information for initiating the deregistration request;
[0055] Determine the target address information of the target service node after the migration process;
[0056] The permission binding relationship between the target account information and the target address information is released, and the target address information whose permission binding relationship has been released is added to the available resource node list.
[0057] In addition, an embodiment of the present application also provides a computer device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in any node evacuation method provided in the embodiment of the present application.
[0058] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps in any node evacuation method provided in the embodiment of the present application.
[0059] In addition, an embodiment of the present application also provides a computer program product, including computer instructions, which, when executed, implement the steps of any node vacancy method provided in the embodiment of the present application.
[0060] The embodiment of the present application can first set the service node set (cluster) that needs to be resource recovered and the service node withdrawal speed, and simulate the withdrawal of each service node in the service node set, so as to simulate the scheduling of the container group running in the service node, and screen out the candidate service nodes that can be withdrawn. Then, each candidate service node is scored according to its own resource occupancy and resource type, so as to select at least one target service node from all candidate service nodes according to the withdrawal speed and the score size of the candidate service node. Finally, the container group on the target service node is migrated, and the computing resources of each successfully migrated target service node are recovered. In this way, the target service node can be selected from the service nodes that have passed the simulation scheduling for withdrawal, the withdrawal success rate of the service node is improved, and the resource recovery efficiency of the service node is ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0062] Figure 1 It is a scene schematic diagram of the node evacuation system provided in an embodiment of the present application;
[0063] Figure 2 It is a schematic diagram of the steps of the node evacuation method provided in an embodiment of the present application;
[0064] Figure 3 It is another step flow chart of the node evacuation method provided in an embodiment of the present application;
[0065] Figure 4 It is a schematic diagram of an interface for a single node evacuation scenario provided in an embodiment of the present application;
[0066] Figure 5 This is a schematic diagram of a detailed interface of node evacuation results after node evacuation provided in an embodiment of the present application;
[0067] Figure 6 It is a schematic diagram of an interface for a scenario in which multiple nodes in a service node set are vacated provided in an embodiment of the present application;
[0068] Figure 7 This is a schematic diagram of a node evacuation process provided in an embodiment of the present application;
[0069] Figure 8 It is a flow chart of a node vacating quantity controller in a node vacating process provided in an embodiment of the present application;
[0070] Fig. 9 It is a flowchart of a solver in a node vacating process provided in an embodiment of the present application;
[0071] Fig.10 It is a flow chart of the executor in the node vacating process provided by an embodiment of the present application;
[0072] Fig.11 It is a schematic diagram of the result post-processing process in the node evacuation process provided in an embodiment of the present application;
[0073] Fig.12 It is a structural schematic diagram of a node evacuation device provided in an embodiment of the present application;
[0074] Fig.13 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0075] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0076] In some processes described in the specification, claims and the above drawings, multiple steps appearing in a specific order are included, but it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, the descriptions such as "first" and "second" in this document are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0077] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0078] The embodiments of the present application provide a node vacating method, device, equipment and computer-readable storage medium. Specifically, the embodiments of the present application will be described from the dimension of the node vacating device, and the node vacating device can be specifically integrated in a computer device, which can be a server or a user terminal and other devices. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Among them, the user terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart home appliance, a car terminal, an intelligent voice interaction device, an aircraft, etc., but is not limited to this.
[0079] It should be noted that in the specific implementation of this application, data related to user information, user usage records, user status, etc. (such as information related to the "account information" below) is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0080] It should be noted that the node vacating method provided in the embodiment of the present application can be applied to resource recovery scenarios of cloud services, such as storage, computing power and other resource recovery scenarios. These scenarios are not limited to being implemented through cloud technology and other methods, and are specifically described through the following embodiments:
[0081] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing. Specifically, cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.
[0082] Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. The network that provides resources is called a "cloud". From the user's point of view, the resources in the "cloud" can be infinitely expanded, and can be obtained at any time, used on demand, expanded at any time, and paid for by use. As a basic capability provider for cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0083] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer. SaaS can also be deployed directly on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0084] Among them, the node vacancy service involved in the embodiments of the present application can be implemented through cloud computing. The node vacancy service can be applied to cloud technology, smart transportation, assisted driving, and map fields, and is specifically described by the following embodiments:
[0085] For example, see Figure 1 , which is a scenario diagram of the node vacancy system provided in an embodiment of the present application. The device in the scenario system may include a server. The server can directly execute the node vacancy method of the embodiment of the present application to vacate the managed service nodes and reclaim the resources of the managed service nodes, such as reclaiming storage resources, computing resources, etc.
[0086] Specifically, when a server in the system can determine a set of service nodes for resources to be reclaimed and determine a evacuation speed, wherein the service node set includes at least one service node; a evacuation simulation is performed on each service node in the service node set to predict a set of candidate service nodes that can be evacuated, wherein the candidate service node set includes at least one candidate service node; a node score of each candidate service node is determined based on resource occupancy information and resource category of each candidate service node; a target service node is selected from the candidate service node set based on the evacuation speed and the node score; and a migration process is performed on the running container group on the target service node to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0087] It should be noted that the service node can be a managed physical server or virtual server, and the "server" in the above system can be understood as a management server, which is used to manage resources of the managed service nodes in the system, such as responding to the application of service nodes, vacating and recycling resources, etc. Exemplarily, a resource orchestration / management system (kubernetes, k8s) is deployed on the management server to manage resource recycling of each service node.
[0088] For example, taking the leasing scenario of cloud computing resources as an example, suppose that enterprise A leases a certain number of service nodes from cloud computing resource provider B to carry and run the application programs of enterprise A's target business. Enterprise A can access the cloud through network connection and enjoy various services such as computing resources. When the target business is reduced or due to other reasons, enterprise A needs to reduce the number of service nodes that it has currently leased. At this time, it needs to apply to cloud computing resource provider B to reclaim the computing resources of the target number of service nodes. At this time, cloud computing resource provider B can first determine the service node cluster (set) leased to enterprise A, and determine the service node evacuation speed during resource recovery according to actual conditions, such as determining the evacuation speed according to historical evacuation speed or a pre-set evacuation speed, which is not limited here; then, simulate scheduling of the container group in each service node in the service node cluster to preliminarily screen out a set of candidate service nodes that can be evacuated; further, score each candidate service node according to its own resource occupancy and resource type, and select the target service node based on the score of each candidate service node; finally, migrate the container group in the target service node to migrate each container group in the target service node to other service nodes in the service node cluster for operation, and recycle computing resources for the successfully migrated target service node.
[0089] It should be noted that the above is only an example and can also be applied to other node evacuation scenarios, which will not be described one by one here.
[0090] For ease of understanding, each step of the node evacuation method will be described in detail below. It should be noted that the order of the following embodiments is not intended to limit the preferred order of the embodiments.
[0091] In the embodiment of the present application, the node vacating device will be described from the perspective of the node vacating device, and the node vacating device can be specifically integrated in a computer device, such as a server. Figure 2 , Figure 2 This is a schematic diagram of the steps of the node evacuation method provided in the embodiment of the present application. In the embodiment of the present application, the node evacuation device is specifically integrated on the server as an example. When the processor on the server executes the program instructions corresponding to the node evacuation method, the specific process is as follows:
[0092] 101. Determine the service node set of resources to be recycled and determine the evacuation speed.
[0093] In the embodiment of the present application, the node vacating method can be applied to resource recovery scenarios of cloud services, such as computing resources, storage resources and other recovery scenarios in cloud services. It should be noted that in the resource recovery scenario of the service node, the premise of resource recovery is to vacate the service node, that is, to migrate the container group deployed in the service node. Therefore, it is necessary to first determine the set of service nodes that need to be processed for resource recovery, so as to recover resources for one or more service nodes in the set of service nodes; in addition, it is also necessary to determine the vacating speed of the service node according to the actual situation, so as to improve the vacating efficiency of the server while utilizing system resources.
[0094] It should be noted that cloud service providers can build a distributed system based on regional physical service machines and networking as cloud services, and the distributed system includes multiple service nodes (Node) and a management server for managing each service node. The management server is the execution subject of the embodiment of this application. The management server can be installed with a resource orchestration / management system (kubernetes, k8s) to execute the node vacating method proposed in the embodiment of this application. The above is only an example and is not a specific implementation method of this application.
[0095] The resource may be computing power resources and storage resources provided by the cloud service provider. Resource recycling means migrating the container groups originally deployed on the service node one by one, and obtaining the vacated service node after all container groups are successfully migrated, so as to recycle the resources of the vacated service node, such as recycling the computing resources of the service node to the management of the cloud service provider. For example, cloud computing resources are recycled to the public cloud, and the computing resources are not limited to the central processing unit (CPU), cache, etc.
[0096] The service node set may be a service node cluster occupied by the target business of the target object when it is deployed, and the service node set includes at least one service node. Specifically, it can be understood as a service node cluster used by a cloud service provider when providing cloud services to the target object. The target object can deploy the target business in the service node of the service node cluster. When deploying the target business, a container group will be created in the service node of the service node cluster. Each container group is used to carry the application and / or related code running the target business to support the operation of the target business.
[0097] In some embodiments, the target object may actively send a service node vacancy request (or resource recovery request) to the execution subject server of an embodiment of the present application through a terminal. The target object may be an individual user or an enterprise. The request may include a target object identifier or a target business identifier or a list of currently occupied service node identifiers. The request may also include the number of nodes to be vacated (or resources recovered). The execution subject server may determine the set of service nodes occupied by the target object when deploying the target business based on the target object identifier, the target business identifier or the list of currently occupied service node identifiers.
[0098] Among them, the vacating speed can be the speed when vacating the service node, and the vacating speed can reflect the number of service nodes vacated per unit time. Therefore, it is mainly based on the service node, specifically how many service nodes are vacated per unit time. For example, how many service nodes are vacated in 10 minutes, how many service nodes are vacated in one hour, the above are just examples. It should be noted that when vacating a service node, one unit time (such as 1 hour, 30 minutes, etc.) can be used as a vacating round, and within this vacating round, multiple service nodes can be vacated in parallel, or each service node can be vacated one by one, which is not limited here.
[0099] In some implementations, the current evacuation speed can be determined based on the evacuation record of historical time or a pre-set preset evacuation speed. Specifically, taking the evacuation record of historical time as an example, the server can query the evacuation speed of the node in the same time period as the current time information in the evacuation record according to the current time information, and determine the node evacuation speed as the current evacuation speed. Alternatively, a pre-built evacuation speed list is obtained, and the evacuation speed list includes multiple preset evacuation speeds, such as the corresponding relationship between each preset evacuation speed and the corresponding time period. The corresponding target preset evacuation speed can be queried from the evacuation speed list according to the current time information, and the target preset evacuation speed is used as the evacuation speed of the current time period.
[0100] In some implementations, the evacuation speed can be determined based on the resource occupancy of the managed server. Specifically, local resource occupancy information, such as the number of CPU cores, cache usage capacity, etc., is obtained to determine the evacuation speed based on the occupancy rate of the target resource item in the local resource occupancy information. In addition, local resource occupancy information can also be obtained to determine the remaining resources based on the resource occupancy information, and then, the target number of resources required for evacuation is predicted to determine the target ratio of the target number of resources to the remaining resources, and the evacuation speed is determined according to the size of the target ratio. For example, the target ratio is compared with a preset ratio threshold, and the preset ratio threshold is used as a demarcation point for the evacuation speed. If the target ratio is less than the preset ratio threshold, the evacuation speed is determined according to the required evacuation number. For example, the number of nodes to be evacuated is 100, and the evacuation speed is 100 / hour, or 100 / 10 minutes, depending on the actual situation. Conversely, if the target ratio is greater than or equal to the preset ratio threshold, the evacuation speed will be determined according to a multiple of the required evacuation number. For example, if the number of nodes to be evacuated is 100, the evacuation speed is set to 0.5 times the number of nodes to be evacuated, such as 50 / hour. The multiples and unit time can be determined according to the actual situation and are not used as a specific implementation method here.
[0101] By the above method, the service node set that needs to be managed can be determined, and the evacuation speed when performing evacuation management on the service node set can be determined, so that the target service node for evacuation can be selected from the service node set according to the evacuation speed.
[0102] 102. Perform a vacancy simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated.
[0103] In an embodiment of the present application, before selecting a service node from a service node set for vacating, the service node may also be simulated for vacating. Specifically, simulated scheduling is performed on the container group on each service node in the service node set, and candidate service nodes that can be vacated are preliminarily screened out according to the simulation results. In this way, service nodes that cannot be vacated can be excluded, so that the candidate service nodes that can be vacated can be vacated subsequently, thereby reducing cases of vacating failures and ensuring reliability.
[0104] The candidate service node set includes at least one candidate service node, each candidate service node is a service node that has passed the evacuation simulation, and the container group on the candidate service node meets the actual migration process.
[0105] It should be noted that since some service nodes themselves carry container groups of special business types, or there are cases of service node evacuation failure in historical evacuation, or there is no suitable service node to receive the container group that needs to be migrated on the current service, this will cause service nodes to fail to evacuate during the evacuation process. Therefore, in the simulation of the evacuation process, the main focus is on excluding service nodes that cannot be evacuated in order to predict the set of candidate service nodes that can be evacuated.
[0106] In some implementations, the evacuation simulation process may include simulation of the business type of the target business deployed on each service node, the actual evacuation case of each service node, and the subsequent compatibility deployment of the container group on each service node. For example, step 102 may include:
[0107] (102.1) Identify the service type of the target service deployed on each service node in the service node set;
[0108] (102.2) filtering the service nodes whose service types are preset service types in the service node set to obtain a filtered service node set;
[0109] (102.3) querying the historical evacuation record of each service node in the filtered service node set, and removing the service nodes that have failed to be evacuated in the historical evacuation record, to obtain an initial service node set;
[0110] (102.4) Compatibility deployment prediction is performed on the container group on each service node in the initial service node set to predict a set of candidate service nodes that can perform evacuation.
[0111] Among them, the target business refers to the business service deployed on the service node, which can be an application, logic component, code, function, etc., which is not limited here. It should be noted that the target business is deployed on the service node and needs to be carried and run by a container group, and a target business may require one or more container groups to carry and run after deployment. For example, assuming that the target business contains multiple business functions, and a container group contains a container, each business function is carried and run by a container group. For example, the container group contains multiple containers, and each business function is carried and run by a container. Exemplarily, assuming that an instant chat service includes chat management, friend addition, instant video chat, instant voice chat, friend list, short video and other business functions, when the instant chat service is deployed on the service node, each business function can occupy a container group, or multiple business functions occupy a container, which is not limited here.
[0112] Among them, each target business has a corresponding business type. According to the business type, the target business can be divided into node system business and non-node system business, and the node system business can be a business that guarantees the basic business operation of the service node. For example, common system components, AB testing components and other businesses are all node system businesses, which guarantee the normal operation of the service node itself or other service nodes. Non-node system businesses are usually third-party businesses, such as the "target object" business mentioned above, for example, instant messaging business, online shopping mall business, video application business, online ticket booking business, etc.; illustratively, cloud service providers can lease out some computing resources, and some enterprises or individual users can lease computing resources from cloud service providers through the network to deploy the target business in the server system of the cloud service provider. It should be noted that the preset business type can be a node system business, which is used to ensure that the service node itself or other service nodes can operate normally.
[0113] Among them, the historical evacuation record can be a evacuation record of the corresponding service node in the historical time, which can reflect the evacuation status of the current service node in the historical time, and can specifically be a evacuation record in the historical time period, such as the evacuation record in the past 1 minute, the past 10 minutes, the past 1 hour, the past 12 hours, and the past 24 hours, which is not limited here. It should be noted that when a service node has a evacuation case in the historical time, whether the evacuation is successful or failed, a evacuation case will be generated in the historical evacuation record. Each evacuation case can include the evacuation time, evacuation result, the time spent on evacuation, the number of evacuations, etc. When there is no evacuation record for the service node in the historical time, the historical evacuation record does not contain a evacuation case.
[0114] Among them, the container group can be the smallest computing unit on the service node, which can be understood as a business computing unit, used to support the operation of the target business. Exemplarily, a resource orchestration / management system (k8s) is installed on the server used to manage the application and withdrawal of service nodes, and the container group is a pod, which is the smallest computing unit created, managed and deployed in the resource orchestration / management system, used to carry the operation of the target business.
[0115] Specifically, for each service node in the determined service node set, first, the business type of the target business deployed in each service node can be determined. It should be noted that when multiple target businesses are deployed on the service node, the business type of each target business needs to be determined; then, each service node is filtered according to the business type. Since the service node with the node system class business cannot be vacated, otherwise it may cause the entire server system to crash, the preset business type can be set to the node system class, and the service nodes with the target business of the node system class are filtered to obtain the filtered service node set. Then, for each service node in the filtered service node set, the historical vacated record of the service node can be queried to determine whether there is a vacated failure case in the historical time length of the current service node. For example, the service node is queried for a vacated failure case in the past short time (such as 1 minute, 1 hour). At this time, if the service node continues to be vacated, a vacated failure case may still occur again, or a vacated failure case may occur with a high probability. In order to avoid affecting the subsequent resource recovery efficiency, the service node with a vacated failure case in the historical time can be removed to obtain the initial service node set. Finally, the compatibility deployment of the container group on each service node in the initial service node set is predicted. The compatibility deployment prediction process is mainly to detect whether there is a service node in the original service node set that is suitable for the deployment of the container group on the current service node. It should be noted that the service node withdrawal process is mainly to migrate the container group on the service node to other service nodes, such as migrating to other service nodes in the service node set. Therefore, in order to ensure that the normal operation of the target business of the original deployment is not affected after the service node is vacated, it is necessary to test the compatibility deployment of the container group on the service node. After completing the compatibility deployment test of the container group on each service node in the initial service node set, the service nodes that pass the compatibility deployment test can be further screened out from the initial service node set as candidate service nodes, thereby obtaining a candidate service node set.
[0116] In some implementations, when predicting the compatibility of the container group on the service node, the resource category of the service node that is compatible with the container group when it is deployed can be measured, and whether the current container group is compatible with the container group on other service nodes can be measured to filter out candidate service nodes suitable for deployment from the original service node set. For example, step (102.4) may include:
[0117] (102.4.1) for each service node in the initial set of service nodes, determining a target resource category that is compatible with each container group deployed on the service node;
[0118] (102.4.2) querying the service node set for the service nodes to be processed corresponding to the target resource category, and determining the total number of the service nodes to be processed;
[0119] (102.4.3) If the total number of service nodes to be processed is greater than a preset threshold, and it is detected that no other container groups incompatible with the container group are deployed on the service nodes to be processed, the service nodes are determined as candidate service nodes that can be vacated to obtain a set of candidate service nodes.
[0120] The resource category may be a category of a certain operating resource of the corresponding service node. It is understandable that the service node may include multiple operating resources, such as operating resources of the central processing unit (CPU), cache memory, random access memory (RAM), storage, etc. The target resource category is the category of the target resource on the corresponding service node. For example, taking the central processing unit (CPU) as an operating resource, some container groups require a central processing unit of the target model / category. When predicting the compatibility deployment of the container group, it is possible to detect whether the original service node combination contains a service node with a central processing unit (CPU) of the target model / category.
[0121] Specifically, in order to facilitate the understanding of the compatibility deployment prediction, the following will be described using one of the service nodes in the service node set. The specific process is: when predicting the compatibility deployment, for each service node in the initial service node set, first, determine the target resource category that each container group deployed on the service node is compatible with. The target resource category can be the resource category of the service node where the container group is currently deployed. The resource categories compatible with multiple container groups deployed on the same service node are consistent, but there are also some container groups that can be compatible with multiple resources, that is, the target resource category can also be other resource categories in addition to the resource category of the service node where the container group is currently deployed; then, determine the service node to be processed corresponding to the target resource category in the original service node set (that is, the service node set before the evacuation simulation is performed). The service node to be processed is the service node other than the service node currently deployed with the container group in the original service node set. service nodes, and determine the total number of the service nodes to be processed; finally, the minimum number of service nodes required for deploying the target business is used as the preset threshold, or the number of service nodes currently occupied by the target business is used as the preset threshold, and the total number of service nodes to be processed is compared with the preset threshold to determine whether there are enough service nodes to be processed for the subsequent deployment of the target business. When the total number of service nodes to be processed is greater than or equal to the preset threshold, it is determined that there are enough service nodes to be processed for the subsequent deployment of the target business. At the same time, it is also necessary to determine whether other container groups incompatible with the current container group have been deployed on each service node to be processed. If other container groups incompatible with the container group are not deployed on each service node to be processed, then the service node can be determined as a candidate service node. In this way, according to the above method, the compatibility deployment prediction of the container group on each service node in the service node set can be performed to determine whether the service node is suitable for performing the vacated task, thereby screening out at least one candidate service node to obtain a candidate service node set.
[0122] Exemplarily, taking the resource category of CPU computing resources as an example, for each service node in the service node cluster, determine the target resource category compatible with the container group deployed by it, or directly determine the target resource category of the current service node, thereby determining the total number of other service nodes (to-be-processed service nodes) corresponding to the target resource category contained in the service node cluster, or determine the total number of available computing resources (the number of available CPU cores) commonly contained in all other service nodes corresponding to the target resource category. When the total number is greater than a preset threshold (the preset threshold may be the number of service nodes required to deploy the target business, or the number of computing resources required to deploy the target business), When the number of containers that are incompatible with the container group corresponding to the current target business is determined, it is further determined whether other container groups that are incompatible with the container group corresponding to the current target business are deployed on each service node to be processed. For example, the target business and a certain business service are mutually exclusive, and it is not desired that the container group of the target business be deployed on the same service node as the container group of the business service after migration, otherwise the services of the target business will be mutually restricted or affected. In order to avoid the above phenomenon, when a service node contains other container groups that are incompatible with the container group on the currently predicted service node, the currently predicted service node is excluded, and thus each remaining service node is used as a candidate service node to obtain a set of candidate service nodes.
[0123] Through the above method, the service nodes can be simulated for evacuation, and the candidate service nodes that can be evacuated can be preliminarily screened according to the simulation results. In this way, the service nodes that cannot be evacuated can be excluded, so that the candidate service nodes that can be evacuated can be evacuated later, reducing the cases of evacuation failure and ensuring reliability.
[0124] 103. Determine a node score of each candidate service node according to the resource occupancy information and resource category of each candidate service node.
[0125] In the embodiment of the present application, after simulating the evacuation of the candidate service nodes that can be evacuated, a score can be given according to the resource situation of each candidate service node to indicate the difficulty of each candidate service node in the subsequent evacuation. The higher the score of the candidate service node, the easier it is to evacuate the candidate service node, and the smaller the impact on the deployed target business during the evacuation process. On the contrary, the lower the score of the candidate service node, the more difficult it is to evacuate the candidate service node, and the greater the impact on the deployed target business during the evacuation process. In this way, the target service node can be selected for evacuation according to the size of the node score.
[0126] The resource occupancy information may be information indicating resource usage on the corresponding candidate service node, which is not limited to including the resource allocation rate of the candidate service node and the number of container groups deployed under the resource allocation rate. It should be noted that the resource allocation rate may represent the ratio between the resource occupancy of one or more target services in the candidate service node and the total amount of resources of the candidate service node, and the number of container groups deployed under the resource allocation rate may be understood as the total number of container groups deployed in the candidate service node for one or more target services.
[0127] The resource category may be any type of running resource on the service node, for example, the resource category is not limited to running resources including central processing unit (CPU), cache memory, random access memory (RAM), storage, etc. The resource category may be represented by a resource category tag, such as a key-value pair. For example, taking the CPU as an example, assuming that the CPU of the service node is a type A processor, it is represented as "CPU = type A processor" or "CPU = A". The above are only examples, and other resource category tags may also be included.
[0128] Among them, the node score can refer to the comprehensive score of the candidate service node under multiple factors. For example, the node score can be determined by the resource occupancy score and resource category score of the candidate service node. The resource occupancy score indicates the resource usage within the current candidate service node; and the resource category score reflects the importance, inventory or importance of the resources of the current candidate service node. The resource category score can be used to decide the priority of vacating the corresponding candidate service node in terms of resource category.
[0129] In some implementations, for each candidate service node, a resource occupancy score may be determined based on the resource occupancy, and a resource category score may be determined based on the resource category, and the node score of the current candidate service node may be determined in combination with the resource occupancy score and the resource category score. For example, step 103 may include:
[0130] (103.1) determining a resource occupancy score for a deployed target service in each candidate service node according to the resource occupancy information of each candidate service node;
[0131] (103.2) Determine a resource category score corresponding to the resource category of each candidate service node;
[0132] (103.3) Determine a node score of each candidate service node based on the resource occupancy score and resource category score of each candidate service node.
[0133] Among them, the resource occupancy score can reflect the resource usage of the candidate service node. When the resource occupancy score is higher, it means that the resource occupancy of the target service deployed in the candidate service node is smaller. In addition, it can also indicate that the number of container groups deployed under the resource occupancy is smaller. Conversely, when the resource occupancy score is lower, it means that the resource occupancy of the target service deployed in the candidate service node is larger. In addition, it can also indicate that the number of container groups deployed under the resource occupancy is larger.
[0134] Among them, the resource category score can indicate the importance of the candidate service node of the current resource category in the selection of vacating. It can be understood that each candidate service node generally only selects one resource category for each operating resource. The resource category score can reflect the importance of the current candidate service node in the idle resources currently managed by the cloud service provider. Assuming that the idle service nodes include the first resource category and the second resource category, when the number of service nodes of the first resource category in the idle service nodes is less than that of the service nodes of the second resource category, the weight of the first resource category is higher, and the resource category score is larger. Exemplarily, taking the central processing unit as an operating resource as an example, the resource categories can be distinguished according to different manufacturers or different product models. Assuming that the first manufacturer resource category and the second manufacturer resource category are included, when the number of idle service nodes of the first manufacturer resource category in the idle service nodes currently managed by the cloud service provider is small, the weight of the candidate service node belonging to the first manufacturer resource category is higher, and the resource category score of the candidate service node of the first manufacturer resource category is larger. The above is only an example and is not a specific limitation of implementation.
[0135] Specifically, after obtaining a set of candidate service nodes that can be vacated, each candidate service node can be scored so that the service node can be selected for vacating according to the node score of the candidate service node. When scoring each candidate service node, the resource occupancy of each candidate service node can be scored first to determine the resource occupancy score of each candidate service node. The resource occupancy score can represent the resource occupancy of the corresponding candidate service node. For example, when the resource occupancy score is higher, the resource occupancy rate is lower, and vice versa, the resource occupancy rate is higher; in addition, the higher the resource occupancy score, the fewer the number of container groups deployed, and vice versa, the more the number of container groups deployed; it should be noted that the greater the resource occupancy rate and / or the more the number of container groups deployed, the greater the difficulty of vacating the candidate service node, and the greater the impact on the operation of the originally deployed target business during the vacating process. Further, the resource category score of the candidate service node is determined according to the resource category of the candidate service node. When the proportion of the resource category in the idle resources is smaller, the weight of the resource category is greater, and the subsequent resource category score is greater. Finally, the resource occupancy score and resource category score can be integrated, such as weighted or summed, to obtain the node score of each candidate service node. In this way, the target service node with a high score can be selected for vacating according to the size of the node score.
[0136] In some implementations, the resource occupancy score is determined by the resource allocation score of the resource allocation rate and the business impact score of the number of container groups deployed under the resource allocation rate. For example, the resource occupancy information includes the resource allocation rate of the candidate service node for the deployed target business and the number of container groups of the running container groups. Step (103.1) may include:
[0137] (103.1.1) determining a resource allocation score for each candidate service node based on the resource allocation rate of each candidate service node;
[0138] (103.1.2) Determine, based on the number of container groups running on each candidate service node, a service impact score when each candidate service node performs migration processing on the container groups occupied by the target service;
[0139] (103.1.3) Determine a resource occupancy score for each candidate service node based on the resource allocation score and business impact score of each candidate service node.
[0140] Among them, the resource allocation rate can be the resource occupancy rate of the target service deployed under the corresponding candidate service node. For example, taking the central processing unit as an operating resource as an example, the central processing unit includes multiple CPU cores. Assuming that the candidate service node includes 10 CPU cores, and 1 CPU core is occupied when the target service is run after the target service is deployed, the resource allocation rate for the central processing unit is 1 / 10; for another example, taking the memory (RAM) as an example, assuming that the memory (RAM) of the candidate service node is 50G, and 10G is occupied when the target service is run after the target service is deployed, the resource allocation rate is 1 / 5. The above is only an example and is not a limited implementation method of this application.
[0141] The resource allocation score may be a score reflecting the resource usage of the target service deployed on the candidate service node during operation. The resource allocation score is negatively correlated with the resource allocation rate. Specifically, the higher the resource allocation rate, the higher the resource allocation score, indicating that the service node is less difficult to vacate. Conversely, the lower the resource allocation rate, the lower the resource allocation score, indicating that the service node is more difficult to vacate.
[0142] It should be noted that, since the target business may include multiple business functions, in order to avoid mutual influence between different business functions, when deploying the target business on the candidate service node, a separate container group can be created for each business function, or multiple sub-functions under different parent objects can be used as a container group. Exemplarily, taking an instant messaging application as the target business as an example, the instant messaging application can include business functions such as chat sessions, short videos, friend management, virtual resource transfer, and online shopping malls. A container group can be deployed separately on the candidate service node for each of the above business functions; further, a container group can contain multiple containers, each of which can represent a subdivided sub-function. Taking the business function of friend management as an example, the business function of friend management can also include business sub-functions such as "friend query", "friend addition", "friend list loading", "face-to-face friend addition", and "shrink nearby friends". A separate container can be created for each of the above business sub-functions, and the containers of these multiple business sub-functions are included in the container group of friend management. The above is only an example, and a separate container group can also be created for each business sub-function.
[0143] Among them, the number of container groups can reflect the complexity of the target business. It can be understood that the more container groups deployed for the target business, the more functions the target business has. If the candidate service node is vacated, it will have a greater impact on the normal operation of the target business, such as the phenomenon of business function disorder. Therefore, the number of container groups can be used as the basis for the business impact score of the candidate service node. The business impact score can indicate the degree of influence of the current candidate service node when performing vacating. The business impact score is negatively correlated with the number of container groups. For example, if the number of container groups deployed on the candidate service node is more, the business impact score is higher. Conversely, if the number of container groups deployed on the candidate service node is smaller, the business impact score is lower.
[0144] Specifically, in order to determine the resource occupancy score of each candidate service node, first, the resource allocation score can be determined according to the resource allocation rate of each candidate service node according to the negative correlation between the resource allocation score and the resource allocation rate. For example, for the negative correlation between the resource allocation score and the resource allocation rate, a list of corresponding relationships between the resource allocation rate and the resource allocation score can be pre-built. After determining the resource allocation rate of the candidate service node, the resource allocation score corresponding to the resource allocation rate can be queried from the list; at the same time, according to the negative correlation between the business impact score and the number of container groups, the corresponding business impact score can be determined according to the number of container groups deployed in the candidate service node. For example, for the negative correlation between the business impact score and the number of container groups, a list of corresponding relationships between the number of container groups and the business impact score can be pre-built. After determining the number of container groups deployed in the candidate service node, the business impact score corresponding to the number of container groups can be queried from the list. Finally, for the resource allocation score and business impact score of each candidate service node, its resource allocation score and business impact score are integrated. Specifically, the resource occupancy score of each candidate service node can be calculated by weighting, summing, or the like. In this way, when selecting candidate servers for decommissioning, the resource usage in each candidate service node can be fully considered to ensure the efficiency of the candidate service node when decommissioning, and avoid the impact of the service node decommissioning process on the operation of the target business, thereby maximizing the protection of the stable operation of the target business.
[0145] In some embodiments, for each resource category, the weight value of the resource category when executing the vacated service node may be determined according to the number of service nodes in an idle state or an allocatable state, so as to determine the resource category score of the candidate service nodes under the resource category according to the weight value. For example, step (103.2) may include: querying the resource categories of other service nodes outside the service node set; determining the number of service nodes corresponding to each resource category in other service nodes; determining the weight value of service nodes of each resource category in other service nodes based on the number of service nodes corresponding to each resource category; determining the priority coefficient of the resource category of each candidate service node according to the weight value of the service node of each resource category, and determining the resource category score of each candidate service node according to the priority coefficient of each candidate service node.
[0146] Among them, the other service nodes can be understood as any service nodes other than the set of service nodes for resources to be reclaimed. For example, assume that a cloud service provider builds a distributed system including multiple physical service machines, including 30 service nodes and a management server for managing the 30 service nodes, and the management server is used to execute the node vacating method of the embodiment of the present application, wherein the cloud service provider allocates 10 service nodes as a resource pool according to the needs of the target enterprise, and the resource pool provides computing resources or storage resources for the target enterprise, and there are 20 service nodes in an idle state or an allocatable state, and these 20 service nodes are "other service nodes" outside the resource pool.
[0147] The inventory can be the number of service nodes corresponding to each resource category in other service nodes. For example, among the service nodes managed by the cloud service provider, only 20 service nodes are idle or available for allocation, including 8 service nodes of resource category A and 12 service nodes of resource category B. Then the inventory of other service nodes of resource category A is 8, and the inventory of other service nodes of resource category B is 12.
[0148] Among them, the resource category can be represented by a label, and the label can be in the form of a key-value pair. Exemplarily, taking the CPU as an example, assuming that the CPU of the service node is a type A processor, it is expressed as "CPU = A type processor" or "CPU = A". It should be noted that each resource category has a corresponding weight value, that is, the label weight value, which can be dynamically determined according to the number of other idle service nodes under the resource category, that is, when the number of other idle service nodes under a certain resource category changes, the weight value of the resource category will also change. It should be noted that the weight value corresponding to each resource category is negatively correlated with its holdings, or the weight value corresponding to each resource category is negatively correlated with the proportion of its holdings. For example, assume that a cloud service provider builds a distributed system consisting of multiple physical service machines, including 30 service nodes. According to the needs of the target enterprise, 10 service nodes are allocated as a resource pool, which provides computing resources or storage resources for the target enterprise. There are 20 service nodes in an idle or allocatable state. These 20 service nodes are "other service nodes" outside the resource pool. The number of other service nodes in resource category A is 8, and the number of other service nodes in resource category B is 12. The weight value of the service node in resource category A is 0.4, that is, the weight of label A is 0.4, and the weight value of the service node in resource category B is 0.6, that is, the weight of label B is 0.6.
[0149] Among them, the priority coefficient can be used to represent the priority of different resource categories when they are vacated. The priority coefficient is related to the size of the weight value. When the size of the priority coefficient is positively correlated with the weight value, the larger the weight value of the resource category, the larger the priority coefficient, and the smaller the weight value of the resource category, the smaller the priority coefficient. It should be noted that the weight value of the corresponding resource category can be directly used as the priority coefficient; in addition, it can also be determined according to a functional relationship (such as a sigmoid function, a linear function) so that the priority coefficient of a resource category with a larger weight value is much larger than the priority coefficient of a resource category with a smaller weight value, which is conducive to the accurate execution of the subsequent evacuation optimization strategy, which is not limited here.
[0150] It is understandable that each service node has a corresponding resource category, and the number of service nodes of different resource categories may be different. Due to the different service deployment requirements between different target businesses, the actual occupancy of service nodes of different resource categories is also different. Therefore, the remaining (idle or allocable) "other service nodes" of different resource categories have different holdings. In order to ensure that the cloud service supply needs can be met in the future, when making the optimal decision to vacate, it is necessary to balance the holdings of "other service nodes" of different resource categories as much as possible, so as to determine the resource category score of each resource category based on the holdings of service nodes of each resource category. Specifically, the (management) server may query the resource category of each other service node except the set of service nodes for resources to be recycled, and determine the inventory of service nodes corresponding to each resource category in multiple other service nodes. When the inventory of "other service nodes" of a certain resource category is small, the weight value of the resource category is larger. Conversely, when the inventory of "other service nodes" of a certain resource category is large, the weight value of the resource category is smaller. Then, according to the weight value of each resource category, the priority coefficient of the candidate service node is determined. It can be understood that in the set of candidate service nodes, the priority coefficients of candidate service nodes of the same resource category are consistent. Finally, according to the priority coefficient of each candidate service node, the resource category score of each candidate service node is calculated. The resource category score can be calculated by multiplying the priority coefficient by the weight value, or by determining it according to a functional relationship (such as a sigmoid function or a linear function). In addition, the priority coefficient can also be directly used as the resource category score, which is not limited here.
[0151] For example, assume that a cloud service provider builds a distributed system including multiple physical service machines, including 30 service nodes. According to the needs of the target enterprise, 6 service nodes of resource category A and 4 service nodes of resource category B are allocated to the target enterprise as a resource pool for deploying the target business. The remaining 20 service nodes are idle or allocatable. These 20 service nodes are "other service nodes" outside the resource pool. The number of other service nodes of resource category A is 8, and the number of other service nodes of resource category B is 12. The weight value of the service node of resource category A is 0.4, that is, the weight of label A is 0.4, and the weight value of the service node of resource category B is 0.6, that is, the weight of label B is 0.6. Furthermore, assuming that the weight value is used as the priority coefficient, in the resource pool allocated to the target enterprise, the priority coefficient of the candidate service node of resource category A is 0.4 when it is preferred for evacuation, and the priority coefficient of the candidate service node of resource category B is 0.6 when it is preferred for evacuation. Furthermore, using the priority coefficient as the resource category score, the resource category score of the candidate service node of resource category A is 0.4, and the resource category score of the candidate service node of resource category B is 0.6. The above is only an example, and the resource category score can also be calculated in other ways. For example, when the resource category score is a percentage system, it can also be calculated through a linear relationship to increase the difference between the scores of different resource categories, which is conducive to considering the resource category factor to a large extent when executing the evacuation optimization strategy. It is not limited here.
[0152] Through the above method, each candidate service node can be scored according to its resource occupancy and resource category, so as to fully consider various factors such as resource allocation on each candidate service node, the number of deployed container groups and resource category, so that the subsequent evacuation optimization strategy can be executed according to the score size, and the favorable service nodes can be prioritized for evacuation, which is reliable.
[0153] 104. Select a target service node from the set of candidate service nodes according to the evacuation speed and the node score.
[0154] In the embodiment of the present application, after determining the node score of each candidate service node, the vacating optimization stage can be entered. Since the higher the score of the candidate service node, the smaller the impact on the deployed target business, and the easier it is to vacate the candidate service node, in the vacating optimization stage, the target service node with a large node score can be preferentially selected from the candidate service node set according to the size of the node score to perform vacating. In addition, one or more service nodes can be vacated in each vacating round. When multiple service nodes need to be vacated, the number of service nodes that can be vacated per unit time can be determined according to the vacating speed, and the target service node with a larger score can be selected according to the number of service nodes that can be vacated per unit time to perform vacating in the current round.
[0155] It should be noted that each evacuation round can be divided according to a unit time, such as 30 minutes, 1 hour, etc., or can be divided according to the end of the evacuation of all target service nodes selected in the round as the end point of the evacuation round. In each evacuation round, if multiple target service nodes are evacuated, the evacuation can be performed on multiple target service nodes at the same time, or the evacuation can be performed on each selected target service node step by step in a queue form, which is not limited here.
[0156] In some embodiments, when making a decision on the preferred evacuation, all candidate service nodes in the candidate service node set may be sorted according to the node score size relationship, and the number of nodes that can be evacuated within a unit time may be determined based on the evacuation speed, so as to preferentially select the target service node with a higher ranking from the candidate service node ranking according to the number of nodes. For example, step 104 may include: sorting the candidate service nodes in the candidate service node set according to the order of node scores from large to small to obtain the candidate service node ranking; determining the number of nodes to be evacuated within a unit time based on the evacuation speed; and selecting the target service node with a higher ranking from the candidate service node ranking based on the number of nodes.
[0157] Exemplarily, assume that the candidate service node set contains 6 candidate service nodes, which are represented as node 1, node 2, node 3, node 4, node 5 and node 6 respectively. After scoring processing, the node scores are 8, 7, 5, 5, 6, and 9 respectively. The 6 candidate service nodes are sorted from large to small according to the scores. The sorting relationship of these 6 candidate service nodes is "node 6, node 1, node 2, node 5, node 3 and node 4"; assuming that the evacuation speed is 3 nodes / hour, the number of nodes that can be evacuated within 1 hour (unit time) is 3, and then, the 3 candidate service nodes with the first sorting are selected as the target service nodes in the sorting, that is, the preferred target service nodes in the current evacuation round are "node 6", "node 1" and "node 2" respectively, so that the three target service nodes can be evacuated subsequently.
[0158] Through the above method, the target service node in each withdrawal round can be selected according to the size of the node score and the withdrawal speed, so as to reduce the impact on the operation of the deployed target business during the subsequent withdrawal, and improve the success rate and efficiency of the withdrawal of the service node, with reliability.
[0159] 105. Perform migration processing on the running container group on the target service node to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0160] In the embodiment of the present application, the evacuation process mainly migrates the container groups in the target service node to the corresponding service node. When all container groups on the target service node are successfully migrated, the target service node is in a vacated state. At this time, the operating resources of the vacated target service node can be recycled for subsequent allocation and supply of cloud services.
[0161] It should be noted that computing resources can be understood as computing power resources provided by cloud services, which are used to run various programs and logic codes of the target business. Reclaiming the computing resources of the target service node means releasing the computing power of the target service node and recycling it to the public cloud service of the cloud service provider for subsequent allocation and use.
[0162] In some implementations, when migrating the container groups in each target service node, each migration may result in a successful migration or a failed migration. If and only if all container groups in a target service node are successfully migrated, the migration processing result for the target service node is a successful migration. Otherwise, the migration processing result for the target service node is a failed migration. Then, different resource recovery strategies are executed for different migration processing results. For example, step 105 may include:
[0163] (105.1) Perform migration processing on each container group in the target service node to obtain a migration processing result;
[0164] (105.2) If the migration processing result is successful, determine to recycle the computing resources of the target service node after the migration processing;
[0165] (105.3) If the migration processing result is a migration failure, it is determined not to recycle the computing resources of the target service node.
[0166] Among them, the migration processing result represents the migration results of all container groups in the corresponding target service node, and the migration processing result can record the migration results of each container group in the form of a list. For example, assuming that the target service node includes 10 container groups, after the container groups in the target service node are migrated, the list records 10 migration results. It should be noted that the migration processing result is determined to be a successful migration if and only if all container groups in a target service node are successfully migrated. If all 10 migration results are successful, it can be decided to recycle resources for the successfully migrated target service node; when there is at least one migration result that is a migration failure, it means that the migration processing result is a migration failure, and it can be decided not to recycle the resources of the target service node that failed to migrate.
[0167] In some implementations, when the container group in the target service node is migrated, the container group may be migrated to another service node in the set of service nodes whose resources are to be reclaimed, so as to ensure that the deployed target service can continue to run within the limited service node resources, so as to achieve migration. For example, step (105.1) may include:
[0168] (105.1.1) Filtering out the to-be-processed service nodes having the same resource category as the target service node from the service node set;
[0169] (105.1.2) Adjust the configuration state of the target service node to an unschedulable state, and request the service node to be processed to create a target container group corresponding to the container group, and obtain a creation result;
[0170] (105.1.3) If the creation result is that each target container group is successfully created, the corresponding container group in the target service node in the unschedulable state is deleted, and the migration processing result is determined to be a successful migration;
[0171] (105.1.4) If the unit creation result is that at least one target container group fails to be created, the migration processing result is determined to be a migration failure.
[0172] The service node to be processed may be any service node in the service node set of resources to be reclaimed except the current target service node, or may be any service node in the service node set of resources to be reclaimed except the candidate service node. It should be noted that the resource category of the service node to be processed needs to be consistent with the resource category of the target service node currently performing the vacating process. In addition, the resource category of the service node to be processed may also be a resource category compatible with the container group in the target service node. For example, when the container group in the target service node is compatible with multiple resource categories, the service node to be processed may be a service node of any compatible resource category in the service node set.
[0173] Among them, the configuration state can be understood as the resource allocation state, which may include a schedulable state and an unschedulable state. The schedulable state may be a state indicating that the target service node can be migrated or a container group can be created, and the unschedulable state may be a state indicating that the current target service node cannot be migrated or a container group cannot be created. In the unschedulable state, no container group is allowed to migrate to the target service node that is currently performing evacuation, so as to effectively avoid the phenomenon of business function disorder.
[0174] It should be noted that when a service node set recycles resources, it may mean vacating some service nodes in the service node set to release the computing resources of some service nodes, that is, shrinking the service node set to reduce the number of service nodes occupied by the target business. Therefore, during the node vacancy process, it is necessary to migrate the container groups in a part of the target service nodes to any compatible service node to be processed in the original service node set. Specifically, in the original set of service nodes to be shrunk, the service nodes to be processed that are consistent with the resource category of the target service node are screened out, or the service nodes to be processed are screened out according to the resource category compatible with the container group in the target service node. The service node to be processed can be any service node in the service node set except the target service node or the candidate service node; then, the configuration state of the target service node is adjusted to an unschedulable state to intercept the migration of the container group of other service nodes into the current target service node, or intercept any target business to create a container group in the current target service node. At the same time, the selected service node to be processed is requested to create a target container group corresponding to each container group to be migrated, and the creation result of the target container group in the service node to be processed is obtained, and the creation result includes the creation of each target container group in the service node to be processed. Further, when the creation result is that each target container is successfully created, the corresponding container group originally deployed in the target service node in the unschedulable state is deleted, and the migration processing result is set to migration success. When the creation result is that at least one target container group fails to be created in the service node to be processed, the migration processing result is regarded as migration failure.
[0175] It should be noted that when selecting a service node to be processed, it can also be determined according to the specified address information. Specifically, a node selection request corresponding to the target object is received, and the node selection request includes node address information and node resource category. The node address information can be the address information of a service node, such as IP address information or a string in the IP address information, or it can be the address range information of multiple service nodes, that is, multiple IP address information. The service node to be processed can be screened from the original service node set according to the node address information and node resource category. The above is only an example, which can be combined with the aforementioned description of "migration processing of container groups in target service nodes".
[0176] In some embodiments, the resource recovery process involves an unbinding operation of the address information of the service node, and the target address information of the unbound service node is added to the node list of the public cloud. For example, the step (105.2) of "determining to recycle the computing resources of the target service node after the migration process" may include: receiving a deregistration request for the target service node after the migration process, the deregistration request carries the target account information for initiating the deregistration request; determining the target address information of the target service node after the migration process; releasing the permission binding relationship between the target account information and the target address information, and adding the target address information after the permission binding relationship is released to the available resource node list.
[0177] The target account information may be the account information registered by the target object on the cloud service provider platform, or the object identification information assigned by the cloud service provider to the target object, which may be in the form of a contact address, custom character information, contact number, email address, etc.
[0178] The target address information may be IP address information, physical address information, node label information, globally unique identifier (GUID), etc. It should be noted that there is a permission binding relationship between the target account information and the target address information, indicating that the target object corresponding to the target account information has the right to use the service node corresponding to the target address information.
[0179] Among them, the list of available resource nodes can be an information list of service nodes in an allocatable state, which can be understood as an information list of service nodes that can currently allocate cloud services to the cloud service provider. The list is not limited to including multiple information such as address information, node identification, node resource category, and node resource occupancy of the service nodes in an allocatable state.
[0180] Specifically, when the computing resources of the successfully migrated target service node are recovered, a logout request for the successfully vacated target service node (all container groups are successfully migrated) can be received. In response to the logout request, the target address information of the currently successfully migrated target service node is determined, and the permission binding relationship between the target account information and the target address information is deleted based on the target account information carried in the logout request to remove the permission binding relationship between the target account information and the target address information, and the target address information after the permission binding relationship is released is added to the list of available resource nodes to indicate that the computing resources of the target service node after the successful migration of the container group are recovered.
[0181] Through the above method, the container group in the selected target service node can be migrated, and the service node evaluated to be suitable for vacating can be vacated. Furthermore, a decision is made based on the migration result whether to recycle the computing resources of the current service node, ensuring that only the computing resources of the target service node that has been successfully migrated are recycled, which is reliable.
[0182] As can be seen from the above, the embodiment of the present application can first set the service node set (cluster) that needs to be resource recovered and the service node withdrawal speed, and perform a withdrawal simulation on each service node in the service node set, so as to simulate the scheduling of the container group running in the service node, and screen out the candidate service nodes that can be withdrawn. Then, each candidate service node is scored according to its own resource occupancy and resource type, so as to select at least one target service node from all candidate service nodes according to the withdrawal speed and the score size of the candidate service node, and finally, the container group on the target service node is migrated, and the computing resources of each successfully migrated target service node are recovered. In this way, it is possible to select the target service node for withdrawal from the service nodes that have passed the simulation scheduling, improve the withdrawal success rate of the service node, and ensure the resource recovery efficiency of the service node.
[0183] According to the method described in the above embodiment, the following is further described in detail with examples.
[0184] The embodiment of the present application takes the test question image processing as an example to further describe the test question image processing method provided in the embodiment of the present application.
[0185] Figure 3 This is another step flow chart of the test image processing method provided in the embodiment of the present application. Figure 3 Give a description.
[0186] In the embodiment of the present application, the test image processing device will be described from the perspective of the test image processing device, which can be integrated into a computer device such as a terminal or a server. For example, when the processor on the computer device executes the program corresponding to the test image processing method, the specific process of the test image processing method is as follows:
[0187] 201. Determine a set of service nodes for resources to be recycled and determine a evacuation speed.
[0188] In the embodiment of the present application, the node vacating method can be applied to resource recovery scenarios of cloud services, such as computing resources, storage resources and other recovery scenarios in cloud services. It should be noted that in the resource recovery scenario of the service node, the premise of resource recovery is to vacate the service node, that is, to migrate the container group deployed in the service node. Therefore, it is necessary to first determine the set of service nodes that need to be processed for resource recovery, so as to recover resources for one or more service nodes in the set of service nodes; in addition, it is also necessary to determine the vacating speed of the service node according to the actual situation, so as to improve the vacating efficiency of the server while utilizing system resources.
[0189] It should be noted that cloud service providers can build a distributed system based on regional physical service machines and networking as cloud services, and the distributed system includes multiple service nodes (Node) and a management server for managing each service node. The management server is the execution subject of the embodiment of this application. The management server can be installed with a resource orchestration / management system (kubernetes, k8s) to execute the node vacating method proposed in the embodiment of this application. The above is only an example and is not a specific implementation method of this application.
[0190] The service node set may be a service node cluster occupied when a target service of a target object is deployed, and the service node set includes at least one service node.
[0191] Among them, the vacating speed can be the speed when vacating the service node, and the vacating speed can reflect the number of service nodes vacated per unit time. Therefore, it is mainly based on the service node, specifically how many service nodes are vacated per unit time. For example, how many service nodes are vacated in 10 minutes, how many service nodes are vacated in one hour, the above are just examples. It should be noted that when vacating a service node, one unit time (such as 1 hour, 30 minutes, etc.) can be used as a vacating round, and within this vacating round, multiple service nodes can be vacated in parallel, or each service node can be vacated one by one, which is not limited here.
[0192] 202. Perform a vacancy simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated.
[0193] In an embodiment of the present application, before selecting a service node from a service node set for vacating, the service node may also be simulated for vacating. Specifically, simulated scheduling is performed on the container group on each service node in the service node set, and candidate service nodes that can be vacated are preliminarily screened out according to the simulation results. In this way, service nodes that cannot be vacated can be excluded, so that the candidate service nodes that can be vacated can be vacated subsequently, thereby reducing cases of vacating failures and ensuring reliability.
[0194] The candidate service node set includes at least one candidate service node, each candidate service node is a service node that has passed the evacuation simulation, and the container group on the candidate service node meets the actual migration process.
[0195] Specifically, for each service node in the determined service node set, first, the business type of the target business deployed in each service node can be determined. It should be noted that when multiple target businesses are deployed on the service node, the business type of each target business needs to be determined; then, each service node is filtered according to the business type. Since the service nodes with node system-type businesses deployed cannot be vacated, otherwise the entire server system may crash. Therefore, the preset business type can be set to the node system class, and the service nodes with the target business of the node system class deployed can be filtered to obtain a filtered service node set.
[0196] Then, for each service node in the filtered service node set, the historical evacuation record of the service node can be queried to determine whether there are any evacuation failure cases for the current service node within the historical time period. For example, the service node can be queried to see if there are any evacuation failure cases in the past short period of time (such as 1 minute, 1 hour). At this time, if the evacuation of the service node continues, there may still be evacuation failure cases again, or there is a high probability of evacuation failure cases. In order to avoid affecting the subsequent resource recovery efficiency, the service nodes that have evacuation failure cases in the historical time period can be removed to obtain the initial service node set.
[0197] Finally, the compatibility deployment of the container group on each service node in the initial service node set is predicted. The compatibility deployment prediction process is mainly to detect whether there is a service node in the original service node set that is suitable for the deployment of the container group on the current service node. It should be noted that the service node withdrawal process is mainly to migrate the container group on the service node to other service nodes, such as migrating to other service nodes in the service node set. Therefore, in order to ensure that the normal operation of the target business of the original deployment is not affected after the service node is vacated, it is necessary to test the compatibility deployment of the container group on the service node. After completing the compatibility deployment test of the container group on each service node in the initial service node set, the service nodes that pass the compatibility deployment test can be further screened out from the initial service node set as candidate service nodes, thereby obtaining a candidate service node set.
[0198] It should be noted that when predicting the compatibility deployment of container groups on service nodes, the resource categories of the service nodes that are compatible when deploying the container group can be measured, as well as whether the current container group is compatible with container groups on other service nodes, so as to screen out candidate service nodes suitable for deployment from the original service node set.
[0199] 203. Determine a resource occupancy score for the deployed target service in each candidate service node according to the resource occupancy information of each candidate service node.
[0200] In an embodiment of the present application, the resource occupancy information includes the resource allocation rate of the candidate service node for the deployed target service and the number of container groups of the running container groups, and the resource occupancy score is determined by the resource allocation score of the resource allocation rate and the business impact score of the number of container groups deployed under the resource allocation rate.
[0201] The resource occupancy information may be information indicating resource usage on the corresponding candidate service node, which is not limited to including the resource allocation rate of the candidate service node and the number of container groups deployed under the resource allocation rate. It should be noted that the resource allocation rate may represent the ratio between the resource occupancy of one or more target services in the candidate service node and the total amount of resources of the candidate service node, and the number of container groups deployed under the resource allocation rate may be understood as the total number of container groups deployed in the candidate service node for one or more target services.
[0202] Specifically, in order to determine the resource occupancy score of each candidate service node, first, the resource allocation score can be determined according to the resource allocation rate of each candidate service node according to the negative correlation between the resource allocation score and the resource allocation rate. For example, for the negative correlation between the resource allocation score and the resource allocation rate, a list of corresponding relationships between the resource allocation rate and the resource allocation score can be pre-built. After determining the resource allocation rate of the candidate service node, the resource allocation score corresponding to the resource allocation rate can be queried from the list; at the same time, according to the negative correlation between the business impact score and the number of container groups, the corresponding business impact score can be determined according to the number of container groups deployed in the candidate service node. For example, for the negative correlation between the business impact score and the number of container groups, a list of corresponding relationships between the number of container groups and the business impact score can be pre-built. After determining the number of container groups deployed in the candidate service node, the business impact score corresponding to the number of container groups can be queried from the list. Finally, for the resource allocation score and business impact score of each candidate service node, its resource allocation score and business impact score are integrated. Specifically, the resource occupancy score of each candidate service node can be calculated by weighting, summing, or the like. In this way, when selecting candidate servers for decommissioning, the resource usage in each candidate service node can be fully considered to ensure the efficiency of the candidate service node when decommissioning, and avoid the impact of the service node decommissioning process on the operation of the target business, thereby maximizing the protection of the stable operation of the target business.
[0203] 204. Determine a resource category score corresponding to the resource category of each candidate service node.
[0204] In an embodiment of the present application, for each resource category, the weight value of the resource category when executing evacuation can be determined according to the number of service nodes in idle state or allocatable state, so as to determine the resource category score of the candidate service nodes under the resource category according to the weight value.
[0205] It should be noted that each service node has a corresponding resource category, and the number of service nodes of different resource categories may be different. Due to the different service deployment requirements between different target businesses, the actual occupancy of service nodes of different resource categories is also different. Therefore, the remaining (idle or allocable) "other service nodes" of different resource categories have different holdings. In order to ensure that the cloud service supply needs can be met in the future, when making the optimal decision to vacate, it is necessary to balance the holdings of "other service nodes" of different resource categories as much as possible, so as to determine the resource category score of each resource category based on the holdings of service nodes of each resource category.
[0206] Specifically, the (management) server may query the resource category of each other service node except the set of service nodes for resources to be recycled, and determine the inventory of service nodes corresponding to each resource category in multiple other service nodes. When the inventory of "other service nodes" of a certain resource category is small, the weight value of the resource category is larger. Conversely, when the inventory of "other service nodes" of a certain resource category is large, the weight value of the resource category is smaller. Then, according to the weight value of each resource category, the priority coefficient of the candidate service node is determined. It can be understood that in the set of candidate service nodes, the priority coefficients of candidate service nodes of the same resource category are consistent. Finally, according to the priority coefficient of each candidate service node, the resource category score of each candidate service node is calculated. The resource category score can be calculated by multiplying the priority coefficient by the weight value, or by determining it according to a functional relationship (such as a sigmoid function or a linear function). In addition, the priority coefficient can also be directly used as the resource category score, which is not limited here.
[0207] 205. Determine a node score of each candidate service node according to the resource occupation score and resource category score of each candidate service node.
[0208] Among them, the node score can refer to the comprehensive score of the candidate service node under multiple factors. For example, the node score can be determined by the resource occupancy score and resource category score of the candidate service node. The resource occupancy score indicates the resource usage within the current candidate service node; and the resource category score reflects the importance, inventory or importance of the resources of the current candidate service node. The resource category score can be used to decide the priority of vacating the corresponding candidate service node in terms of resource category.
[0209] Specifically, the resource occupancy score and the resource category score are integrated, such as weighted or summed calculation, to obtain the node score of each candidate service node, so that the target service node with a high score can be selected for vacating according to the size of the node score.
[0210] 206. Select a target service node from the set of candidate service nodes according to the evacuation speed and the node score.
[0211] In the embodiment of the present application, after determining the node score of each candidate service node, the vacating optimization stage can be entered. Since the higher the score of the candidate service node, the smaller the impact on the deployed target business, and the easier it is to vacate the candidate service node, in the vacating optimization stage, the target service node with a large node score can be preferentially selected from the candidate service node set according to the size of the node score to perform vacating. In addition, in each vacating round, one or more service nodes can be vacated. When multiple service nodes need to be vacated, the number of service nodes that can be vacated per unit time can be determined according to the vacating speed, and the target service node with a larger score can be selected according to the number of service nodes that can be vacated per unit time to perform vacating in the current round.
[0212] 207. Perform migration processing on each container group in the target service node to obtain a migration processing result.
[0213] In the embodiment of the present application, the evacuation process mainly migrates the container group in the target service node to other service nodes in the service node set of the resources to be reclaimed. It can be understood that when migrating the container group in each target service node, each migration may be successful or failed. If and only if all container groups in a target service node are successfully migrated, the migration processing result for the target service node is a successful migration. Otherwise, the migration processing result of the target service node is a migration failure.
[0214] The migration processing result indicates the migration results of all container groups in the corresponding target service node, which includes the migration results of each container group in the corresponding target service node. The migration result can be a migration success or a migration failure. If and only if all container groups in a target service node are successfully migrated, the migration processing result for the target service node is a migration success. Otherwise, the migration processing result of the target service node is a migration failure.
[0215] Specifically, in the original set of service nodes to be shrunk, the service nodes to be processed that are consistent with the resource category of the target service node are screened out, or the service nodes to be processed are screened out according to the resource category compatible with the container group in the target service node. The service node to be processed can be any service node in the service node set except the target service node or the candidate service node; then, the configuration state of the target service node is adjusted to an unschedulable state to intercept the migration of the container group of other service nodes into the current target service node, or intercept any target business to create a container group in the current target service node. At the same time, the selected service node to be processed is requested to create a target container group corresponding to each container group to be migrated, and the creation result of the target container group in the service node to be processed is obtained, and the creation result includes the creation of each target container group in the service node to be processed. Further, when the creation result is that each target container is successfully created, the corresponding container group originally deployed in the target service node in the unschedulable state is deleted, and the migration processing result is set to migration success. When the creation result is that at least one target container group fails to be created in the service node to be processed, the migration processing result is regarded as migration failure.
[0216] It should be noted that whether to perform resource recovery on the target service node after the migration process can be determined according to the migration process result. For example, when the migration process result is migration success, step 208 is executed; when the migration process result is migration failure, step 209 is executed.
[0217] 208. If the migration processing result is that the migration is successful, determine to recycle the computing resources of the target service node after the migration processing.
[0218] In an embodiment of the present application, when it is determined according to the migration processing result that all container groups in a target service node have been successfully migrated, it is determined to recycle the computing resources of the target service node after the migration processing. The resource recycling process involves the unbinding operation of the address information of the service node, and the target address information of the unbound service node is added to the node list of the public cloud.
[0219] The computing resources may be computing power resources provided by cloud services, which are used to run various programs, logic codes, etc. of the target business.
[0220] Specifically, when the computing resources of the successfully migrated target service node are recovered, a logout request for the successfully vacated target service node (all container groups are successfully migrated) can be received. In response to the logout request, the target address information of the currently successfully migrated target service node is determined, and the permission binding relationship between the target account information and the target address information is deleted based on the target account information carried in the logout request to remove the permission binding relationship between the target account information and the target address information, and the target address information after the permission binding relationship is released is added to the list of available resource nodes to indicate that the computing resources of the target service node after the successful migration of the container group are recovered.
[0221] 209. If the migration processing result is a migration failure, determine not to recycle the computing resources of the target service node, and display a prompt message indicating the migration failure.
[0222] In the embodiment of the present application, for any container group in each target service node, when at least one migration failure occurs, it is decided not to reclaim the resources of the target service node where the migration failed.
[0223] The prompt information includes but is not limited to the reason for the migration failure, the identifier of the container group that failed to migrate, the identifier and address information of the target service node that failed to be vacated, etc., which are not limited here.
[0224] In order to facilitate the understanding of the embodiment of the present application, the embodiment of the present application will be described with a specific application scenario example. Specifically, the application scenario example is described by executing the above steps 201-209.
[0225] It should be noted that the node vacancy method is applicable to node vacancy instances in cloud technology, smart transportation, assisted driving, map field, online shopping and other scenarios. For example, the specifics of the scenario instance are as follows:
[0226] 1. The node evacuation scenario example is briefly introduced as follows:
[0227] This node evacuation scenario example can find some key terms, which are explained as follows:
[0228] k8s: short for kubernete, a resource orchestration strategy in a cloud environment.
[0229] Pod: The smallest computing unit that can be created and managed in k8s.
[0230] Node: is the logical working node in k8s, that is, the service node mentioned above. It can be a physical machine or a virtual machine. Pods can be deployed on the node.
[0231] Label: It is a key-value pair attached to a k8s object (for example, a Pod or a node), representing a logical characteristic. These labels can be labels such as the CPU model and the location of the computer room.
[0232] Resource pool: is a set of nodes with one or more labels.
[0233] Pod migration: Migrate pods to other service nodes, with features such as expansion before deletion and the number of replicas remaining unchanged before and after migration.
[0234] Machine evacuation / node evacuation: Node evacuation refers to the process of evacuating and migrating the pods deployed on a node so that they can be returned to the public cloud or other machine providers.
[0235] This scenario instance performs the "whole machine migration" process for the node according to the Kubernetes standard. Specifically, the node is first marked as unschedulable through taints, and then the pods deployed on the node are migrated one by one. At this point, when the process is completed, the machine (node) is in an empty state, and no pods related to the target business are deployed on the node. The empty state can be understood as a state that can be recycled to the public cloud vendor at any time.
[0236] This scenario instance simulates and obtains the pre-scheduling results of pods in batches in advance; based on the pre-scheduling results, it avoids machines that cannot be moved in the large disk. In this way, while improving the success rate and efficiency of the entire machine migration, it minimizes the impact of the migration process on the business and ensures that the number of business replicas remains consistent before and after the machine is moved.
[0237] In addition, this scenario instance can be vacated through the preferred target service node. Specifically, based on the current machine allocation rate, the previous machine migration results, the custom label weight and other processes, after a comprehensive multi-dimensional evaluation in the global scope (across multiple k8s clusters), the appropriate node is selected for vacating, so that the cluster resource dimension can be as "elastic" as a pod.
[0238] For ease of understanding, this scenario is introduced through the following examples.
[0239] Example 1, combined with Figure 4As shown, taking the daily single-machine operation and maintenance requirements of the Site Reliability Engineer (SRE) as an example, the SRE needs to shield and isolate the business on a specific machine for operation and maintenance and troubleshooting. At this time, k8s can perform targeted evacuation by specifying the node address information (such as IP). Furthermore, after the evacuation task is completed, the background system will write the task status of the evacuation task and the specific reason for the migration failure (if it is due to a pod migration failure, it will include specific task details). If necessary, you can confirm the details. For details, please refer to Figure 5 shown.
[0240] Example 2, see Figure 6 As shown in the figure, taking the resource water level operation requirements of various "resource pools" as an example, by enabling the automatic withdrawal configuration of a specific city, the water level of the city (the number of allocated nodes or the number of used nodes) is adjusted to the expected value. Then, when automatic withdrawal is enabled, the background service will automatically select and withdraw nodes that meet the requirements in the city, and finally recycle the computing resources of the machines that have been evacuated successfully to the public cloud.
[0241] Second, this scenario instance can be combined with multiple "resource pools" to achieve elastic scaling. For ease of understanding, take the application scenario of "configuring automatic evacuation for a specific city" as an example. Figure 7 As shown in the figure, the process of selecting nodes in multiple clusters, evacuating them, and finally returning them to the public cloud is fully described.
[0242] Among them, Figure 7 As shown, after setting the resource pool (such as the logical pool range: the computer room is in area A) and specifying the automatic evacuation speed, the prerequisites for this function are met. Figure 7 The module is simply classified according to its responsibilities. These modules are not limited to the evacuation quantity controller, solver, actuator and result post-processing module. The implementation of each module is explained separately below.
[0243] (1) The evacuation quantity controller is used to dynamically calculate the number of nodes that need to be evacuated in this round according to the set evacuation speed and the length of the current evacuation queue. It should be noted that in order to reduce the potential risks of the evacuation process to the business, the entire evacuation workflow can be designed to be serially executed and blocked. When there are no extra vacancies in the evacuation queue, no node will be selected for evacuation in this round.
[0244] Specifically, its implementation logic is as follows Figure 8As shown, the logical flow is: if it is detected that the evacuation speed is greater than 0, it is allowed to enter the evacuation process. In this evacuation process, it is determined whether the number of node nodes currently being evacuated in the queue is online. If so, the upper limit is reached. At this time, it is necessary to avoid trying to start the evacuation process again within the interval time. If not, the upper limit has not been reached, and the number of nodes that need to be evacuated in the current evacuation round (round) can be output.
[0245] (2) Solver, which is used to select a suitable node from the resource pool to perform evacuation according to the evacuation quantity dynamically calculated by the evacuation quantity controller. Its implementation logic is as follows: Fig. 9 Specifically, the solver selects nodes that can be vacated from the resource pool through two processes: (2.1) simulation (filter) stage and (2.2) scoring (score) stage, so as to perform subsequent evacuation.
[0246] (2.1) Simulation (filter) stage: This stage is mainly responsible for calculating the feasibility of evacuation. It can be specifically designed through plug-in. The current system specifically includes the following plug-ins (Filter plugin):
[0247] 1). Pod expansion prediction plug-in, used to confirm that the container group (pod) can be rebuilt on other service nodes to ensure that the number of pod copies of the business-side workload is consistent before and after the evacuation; 2) Whitelist exemption plug-in, the node itself is not marked as a evacuation whitelist, and some nodes may have special uses, such as deploying system components, doing abtest, etc. The whitelist exemption plug-in is used to exempt such nodes; 3) Node evacuation task portrait plug-in, which is used to directly skip the node that failed to evacuate within a short period of time when the node has an execution record of evacuation failure within the short period of time.
[0248] (2.2) Score stage: used to select the best option from several executable options for evacuation. The specific implementation also uses a plug-in design. The current system contains the following default plug-ins: 1) Allocation rate plug-in, used to determine the resource allocation score based on the node's allocation rate. The lower the allocation rate, the higher the resource allocation score, indicating that this type of node is easiest to evacuate in terms of time; 2) Business impact assessment plug-in, used to determine the business impact score based on the number of pods deployed on the node. The fewer the number of pods, the higher the business impact score, indicating that evacuating this type of node has the least impact on the business side; 3) Label dynamic weight plug-in, used to dynamically calculate the label score (i.e. the aforementioned resource category score) based on the node's label weight.
[0249] (3) Executor, used to remove all pods deployed on the preferred node (specify the corresponding node by IP) by migration. Fig.10 As shown in the figure, the logical process is as follows: first, the executor will mark the node as unschedulable, that is, set the current node as a "tainted" mark to prevent incremental pods from being rescheduled to the node during the migration process. Then, the executor will initiate migration of the pods on the node one by one, and query the migration results regularly until all pod migration tasks are completed. Finally, the executor will update the task status according to the migration results of these pods. The task status is successful only when and only when the migration of all pods is successfully completed. The rest of the statuses are judged as failures, and the detailed information is written to the disk.
[0250] (4) The result post-processing module is used to process the nodes after the evacuation is completed. Fig.11 As shown, the post-processing logic flow of this scenario example is: start post-processing, query and load the current post-processing plug-in, the plug-in can be used to determine whether the node evacuation task status is successful, when the node evacuation is successful, the computing resources of the machine (node) are directly returned to the public cloud; when the node evacuation fails, the node node unschedulable state flag is removed, and the node is released in place, and the next batch (next round) of evacuation processing is carried out. It should be noted that the result post-processing module can also provide extensibility through plug-in design, which can include two slots, OnFailed and OnSucceed.
[0251] It should be noted that the default plug-ins in the current system and their corresponding slot implementations are: 1) Virtual machine recycling plug-in, used to roll back excess resources, which are recycled to the public cloud when successful, and released on the spot when failed. 2) Single machine troubleshooting plug-in, used to troubleshoot single machine scenarios. In addition, regardless of whether the final result is successful or not, the machine's unschedulable flag is retained and waits for manual intervention.
[0252] Through the above scenario examples, the following examples can be implemented: Example 1, SRE performs daily maintenance and troubleshooting on machines, which is more efficient than the original SRE manual vacancy. Example 2, in the resource return scenario for the city's service nodes, taking CPU core resources as an example, a single city can automatically complete the resource return of more than 20,000 cores per day, and all cities can return an average of 80,000 cores per day, which greatly improves the speed, and the impact of the entire process of vacating on the business is completely controllable.
[0253] By executing the following scenario steps, the following effects can be achieved: after the node is vacated, there must be enough resources globally for the deleted pod to be rebuilt; secondly, the entire reconstruction compensation process can span multiple clusters, and the scope of action is no longer limited to a single k8s cluster. In other words, the pod deleted due to node evacuation in a cluster can be rebuilt in another cluster, which is more flexible; finally, we introduced the "optimal selection" process, which can dynamically score and evacuate the node according to the label weight when selecting the node for evacuation, which is more scalable.
[0254] As can be seen from the above, when the target scene data to be rendered is obtained, the embodiment of the present application can first obtain the vertex data of each vertex on the strip element (such as line elements) that are prone to intermittent display abnormalities after rendering, and perform geometric processing to convert each vertex on the strip element into the screen space; then, according to the target primitive to which each vertex belongs, fragment configuration is performed in the screen space to determine the target fragment that intersects with the primitive of the strip element; then, the weight of each target fragment is determined according to the overlap ratio between each target fragment and the target primitive to which it belongs, so as to perform fragment coloring on the target primitive belonging to the strip element according to the weight of each target fragment, and obtain a first color rendering image, so as to realize the separate node vacancy of the strip elements that are prone to intermittent display effects in the target scene; finally, according to the weight of each target fragment, the first color rendering image is merged with the color rendering images of other elements in the target scene to obtain the target image corresponding to the target scene. In this way, the strip elements in the scene that are prone to intermittent display abnormalities can be separated from other elements in the scene and vacated as independent nodes. The weight of each fragment with intersection is combined in the independent node vacancy to ensure the normal display of the strip elements. The image of the independent node vacancy is merged with the images of other elements to ensure that the line elements in the final rendered image can be displayed normally, thereby improving the image quality.
[0255] In order to better implement the above method, the embodiment of the present application also provides a node vacating device. Fig.12 As shown, the node vacating device may include a determining unit 401 , a matching unit 402 , a determining unit 403 , a coloring unit 404 and a blending unit 405 .
[0256] A first determining unit 401 is used to determine a service node set of resources to be reclaimed and determine a vacating speed, wherein the service node set includes at least one service node;
[0257] A prediction unit 402 is configured to perform a vacancy simulation on each service node in the service node set to predict a candidate service node set that can be vacated, the candidate service node set including at least one candidate service node;
[0258] A second determining unit 403 is used to determine a node score of each candidate service node according to resource occupancy information and resource category of each candidate service node;
[0259] A selection unit 404, configured to select a target service node from a set of candidate service nodes according to a vacating speed and a node score;
[0260] The recycling unit 405 is used to perform migration processing on the running container group on the target service node, so as to recycle the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0261] In some embodiments, the second determination unit 403 is further used to: determine the resource occupancy score for the deployed target service in each candidate service node based on the resource occupancy information of each candidate service node; determine the resource category score corresponding to the resource category of each candidate service node; and determine the node score of each candidate service node based on the resource occupancy score and resource category score of each candidate service node.
[0262] In some embodiments, the resource occupancy information includes the resource allocation rate of the candidate service node for the deployed target business and the number of container groups of the running container groups. The second determination unit 403 is further used to: determine the resource allocation score of each candidate service node according to the resource allocation rate of each candidate service node; the resource allocation score is negatively correlated with the resource allocation rate; determine the business impact score of each candidate service node when migrating the container group occupied by the target business according to the number of container groups of the container groups running on each candidate service node, and the business impact score is negatively correlated with the number of container groups; determine the resource occupancy score of each candidate service node according to the resource allocation score and the business impact score of each candidate service node.
[0263] In some embodiments, the second determination unit 403 is further used to: query the resource categories of other service nodes outside the service node set; determine the number of service nodes corresponding to each resource category in other service nodes; determine the weight value of the service nodes of each resource category in other service nodes based on the number of service nodes corresponding to each resource category; determine the priority coefficient of the resource category of each candidate service node according to the weight value of the service node of each resource category, and determine the resource category score of each candidate service node according to the priority coefficient of each candidate service node.
[0264] In some embodiments, the selection unit 404 is further used to: sort the candidate service nodes in the candidate service node set in descending order of the node scores to obtain a candidate service node ranking; determine the number of nodes to be vacated per unit time based on the vacancy speed; and select a target service node that is ranked first from the candidate service node ranking based on the number of nodes.
[0265] In some embodiments, the prediction unit 402 is further used to: identify the business type of the target business deployed on each service node in the service node set; filter the service nodes in the service node set whose business type is a preset business type to obtain a filtered service node set; query the historical evacuation records of each service node in the filtered service node set, and remove the service nodes with failed evacuation in the historical evacuation records to obtain an initial service node set; predict the compatibility deployment of the container group on each service node in the initial service node set to predict a set of candidate service nodes that can perform evacuation.
[0266] In some embodiments, the prediction unit 402 is further used to: determine, for each service node in the initial service node set, a target resource category that is compatible with each container group deployed on the service node; query the service node set for service nodes to be processed that correspond to the target resource category, and determine the total number of service nodes to be processed; if the total number of service nodes to be processed is greater than a preset threshold, and it is detected that no other container group incompatible with the container group is deployed on the service node to be processed, determine the service node as a candidate service node that can be vacated, so as to obtain a set of candidate service nodes.
[0267] In some embodiments, the recycling unit 405 is also used to: perform migration processing on each container group in the target service node to obtain a migration processing result; if the migration processing result is a successful migration, determine to recycle the computing resources of the target service node after the migration processing; if the migration processing result is a failed migration, determine not to recycle the computing resources of the target service node.
[0268] In some implementations, the recycling unit 405 is further used to: filter out a to-be-processed service node having the same resource category as the target service node from the service node set; adjust the configuration state of the target service node to an unschedulable state, and request the to-be-processed service node to create a target container group corresponding to the container group to obtain a creation result; if the creation result is that each target container group is successfully created, delete the corresponding container group in the target service node in the unschedulable state, and determine that the migration processing result is a successful migration; if the unit creation result is that at least one target container group fails to be created, determine that the migration processing result is a migration failure.
[0269] In some embodiments, the recycling unit 405 is also used to: receive a logout request for the target service node after the migration process, the logout request carries the target account information that initiates the logout request; determine the target address information of the target service node after the migration process; release the permission binding relationship between the target account information and the target address information, and add the target address information after the permission binding relationship is released to the list of available resource nodes.
[0270] As can be seen from the above, the embodiment of the present application can first set the service node set (cluster) that needs to be resource recovered and the service node withdrawal speed, and perform a withdrawal simulation on each service node in the service node set, so as to simulate the scheduling of the container group running in the service node, and screen out the candidate service nodes that can be withdrawn. Then, each candidate service node is scored according to its own resource occupancy and resource type, so as to select at least one target service node from all candidate service nodes according to the withdrawal speed and the score size of the candidate service node. Finally, the container group on the target service node is migrated, and the computing resources of each successfully migrated target service node are recovered. In this way, the target service node can be selected from the service nodes that have passed the simulation scheduling for withdrawal, the withdrawal success rate of the service node is improved, and the resource recovery efficiency of the service node is ensured.
[0271] The present application also provides a computer device, such as Fig.13 As shown, it shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0272] The computer device may include one or more processing core processors 501, one or more computer-readable storage media memories 502, a power supply 503, an input unit 504 and other components. Those skilled in the art will appreciate that Fig.13 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0273] The processor 501 is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 502, and calling data stored in the memory 502, the processor 501 performs various functions of the computer device and processes data. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 501.
[0274] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and node vacancy processes by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0275] The computer device also includes a power supply 503 for supplying power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 503 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.
[0276] The computer device may further include an input unit 504, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0277] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in the embodiment of the present application, the processor 501 in the computer device will load the executable file corresponding to the process of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application program stored in the memory 502, thereby realizing various functions, as follows:
[0278] Determine a set of service nodes for resources to be reclaimed, and determine a evacuation speed, wherein the service node set includes at least one service node; perform evacuation simulation on each service node in the service node set to predict a set of candidate service nodes that can be evacuated, wherein the candidate service node set includes at least one candidate service node; determine a node score for each candidate service node based on resource occupancy information and resource category of each candidate service node; select a target service node from the candidate service node set based on the evacuation speed and the node score; perform migration processing on a running container group on the target service node to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0279] The specific implementation of the above operations can be found in the previous embodiments and will not be described in detail here.
[0280] It can be concluded that this solution can first set the service node set (cluster) that needs to be resource recovered and the service node withdrawal speed, and simulate the withdrawal of each service node in the service node set, so as to simulate the scheduling of the container group running in the service node, and screen out the candidate service nodes that can be withdrawn. Then, each candidate service node is scored according to its own resource occupancy and resource type, so as to select at least one target service node from all candidate service nodes according to the withdrawal speed and the score size of the candidate service node. Finally, the container group on the target service node is migrated, and the computing resources of each successfully migrated target service node are recovered. In this way, the target service node can be selected from the service nodes that have passed the simulation scheduling for withdrawal, the withdrawal success rate of the service node is improved, and the resource recovery efficiency of the service node is ensured.
[0281] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any of the node vacating methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0282] Determine a set of service nodes for resources to be reclaimed, and determine a evacuation speed, wherein the service node set includes at least one service node; perform evacuation simulation on each service node in the service node set to predict a set of candidate service nodes that can be evacuated, wherein the candidate service node set includes at least one candidate service node; determine a node score for each candidate service node based on resource occupancy information and resource category of each candidate service node; select a target service node from the candidate service node set based on the evacuation speed and the node score; perform migration processing on a running container group on the target service node to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated.
[0283] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0284] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0285] Since the instructions stored in the computer-readable storage medium can execute the steps in any node evacuation method provided in the embodiments of the present application, the beneficial effects that can be achieved by any node evacuation method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0286] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional implementations provided in the above embodiments.
[0287] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0288] The above is a detailed introduction to a node vacancy method, device, equipment and computer-readable storage medium provided in an embodiment of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A node evacuation method, characterized in that: include: Determine a service node set of resources to be reclaimed, and determine a vacating speed, wherein the service node set includes at least one service node; Performing a vacancy simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated, the set of candidate service nodes including at least one candidate service node; Determine the node score of each candidate service node based on the resource occupancy information and resource category of each candidate service node; Selecting a target service node from the set of candidate service nodes according to the evacuation speed and the node score; A migration process is performed on the running container group on the target service node, so as to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated.
2. The method according to claim 1, characterized in that Determining the node score of each candidate service node according to the resource occupancy information and resource category of each candidate service node includes: Determine, according to the resource occupancy information of each candidate service node, a resource occupancy score for the deployed target service in each candidate service node; Determine a resource category score corresponding to the resource category of each candidate service node; A node score of each candidate service node is determined according to the resource occupancy score and the resource category score of each candidate service node.
3. The method according to claim 2, characterized in that The resource occupancy information includes a resource allocation rate of the candidate service node for the deployed target service and the number of container groups of the running container groups. The determining, according to the resource occupancy information of each candidate service node, a resource occupancy score for the deployed target service in each candidate service node includes: Determine a resource allocation score of each candidate service node according to the resource allocation rate of each candidate service node; the resource allocation score is negatively correlated with the resource allocation rate; Determine, according to the number of container groups running on each candidate service node, a service impact score when each candidate service node performs migration processing on the container group occupied by the target service, wherein the service impact score is negatively correlated with the number of container groups; The resource occupancy score of each candidate service node is determined according to the resource allocation score and the business impact score of each candidate service node.
4. The method according to claim 2, characterized in that: The determining of the resource category score corresponding to the resource category of each candidate service node includes: Query resource categories of other service nodes outside the service node set; Determine the number of service nodes corresponding to each resource category in the other service nodes; Determine the weight value of the service node of each resource category in the other service nodes based on the number of service nodes corresponding to each resource category; According to the weight value of the service node of each resource category, the priority coefficient of the resource category of each candidate service node is determined, and the resource category score of each candidate service node is determined according to the priority coefficient of each candidate service node.
5. The method according to any one of claims 1 to 4, characterized in that: The selecting a target service node from the set of candidate service nodes according to the evacuation speed and the node score includes: Sorting the candidate service nodes in the candidate service node set in descending order of the node scores to obtain a candidate service node ranking; Determine the number of nodes to be vacated per unit time according to the vacating speed; According to the number of nodes, a target service node ranked first is selected from the candidate service nodes.
6. The method according to claim 1, characterized in that The performing a vacating simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated includes: Identify a service type of a target service deployed at each service node in the service node set; Filtering the service nodes whose service types are preset service types in the service node set to obtain a filtered service node set; Querying the historical evacuation record of each service node in the filtered service node set, and removing the service nodes that have failed to be evacuated in the historical evacuation record, to obtain an initial service node set; Compatibility deployment prediction is performed on the container group on each service node in the initial service node set to predict a set of candidate service nodes that can perform evacuation.
7. The method according to claim 6, characterized in that The predicting of the compatibility deployment of the container group on each service node in the initial service node set to predict a set of candidate service nodes that can be vacated includes: For each service node in the initial service node set, determine a target resource category compatible with each container group deployed on the service node; Querying the service nodes to be processed corresponding to the target resource category in the service node set, and determining the total number of the service nodes to be processed; If the total number of the service nodes to be processed is greater than a preset threshold, and it is detected that no other container group incompatible with the container group is deployed on the service node to be processed, the service node is determined as a candidate service node that can be vacated to obtain a set of candidate service nodes.
8. The method according to claim 1, characterized in that The migrating the container group running on the target service node to reclaim the computing resources of the target service node after the container group on the target service node is successfully migrated includes: Performing migration processing on each container group in the target service node to obtain a migration processing result; If the migration process result is that the migration is successful, determining to recycle the computing resources of the target service node after the migration process; If the migration processing result is migration failure, it is determined not to recycle the computing resources of the target service node.
9. The method according to claim 8, characterized in that The performing migration processing on each container group in the target service node to obtain a migration processing result includes: Filter out a to-be-processed service node having the same resource category as the target service node from the service node set; Adjust the configuration state of the target service node to an unschedulable state, and request the to-be-processed service node to create a target container group corresponding to the container group, to obtain a creation result; If the creation result is that each target container group is successfully created, the corresponding container group in the target service node in the unschedulable state is deleted, and the migration processing result is determined to be a successful migration; If the unit creation result is that at least one target container group fails to be created, then the migration processing result is determined to be a migration failure.
10. The method according to claim 8, characterized in that The determining to recycle the computing resources of the target service node after the migration process includes: receiving a deregistration request for a target service node after the migration process, wherein the deregistration request carries target account information for initiating the deregistration request; Determine the target address information of the target service node after the migration process; The permission binding relationship between the target account information and the target address information is released, and the target address information whose permission binding relationship has been released is added to the list of available resource nodes.
11. A resource recovery device, characterized in that: include: A first determining unit is used to determine a service node set of resources to be reclaimed and determine a vacating speed, wherein the service node set includes at least one service node; A prediction unit, configured to perform a vacancy simulation on each service node in the service node set to predict a set of candidate service nodes that can be vacated, wherein the set of candidate service nodes includes at least one candidate service node; A second determining unit, configured to determine a node score of each candidate service node according to resource occupancy information and resource category of each candidate service node; A selection unit, configured to select a target service node from the set of candidate service nodes according to the evacuation speed and the node score; The recycling unit is used to perform migration processing on the running container group on the target service node, so as to recycle the computing resources of the target service node after the container group on the target service node is successfully migrated.
12. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the resource recovery method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the resource recovery method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The method comprises computer instructions, which, when executed, implement the steps of the resource recovery method according to any one of claims 1 to 10.
Citation Information
Cited By
Container deployment method and system
CN121455666A
Container deployment method and system
CN121455666B