Method, system, medium and computer program for reducing placement conflicts
By avoiding collisions and adjusting the dynamic deployment strategy of the system through randomized allocation, the problem of deployment conflicts in traditional cloud computing resource allocation systems is solved, thereby improving resource utilization and allocation efficiency.
Patent Information
- Application Number
- CN202511382849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-14
- Filing Date
- 2021-09-14
- Publication Date
- 2025-12-12
AI Technical Summary
Traditional cloud computing resource allocation systems are prone to delays and deployment conflicts when faced with a large number of deployment requests, and cannot effectively adapt to the growth in computing resource demand, resulting in low resource allocation efficiency.
A collision avoidance system is adopted, which reduces conflicts between allocation agents by maintaining and dynamically updating the deployment strategy, including partially randomizing the allocation of computational resources and relaxing or adjusting the allocation rules based on observed conflict situations.
It effectively reduces deployment conflicts between distribution agents, improves the utilization and allocation efficiency of computing resources, and optimizes resource management of cloud computing systems.
Smart Images

Figure CN121125496A_ABST
Abstract
Description
[0001] This application is a divisional application of the application patent application with the application date of September 14, 2021, the application number of 202180062763.7, and the invention name of “Method, system, medium, and computer program for reducing placement conflicts”. BACKGROUND
[0002] A cloud computing system refers to a collection of computing devices that are capable of providing remote services and resources. For example, modern cloud computing infrastructures often include a collection of physical server devices organized in a hierarchical structure, including compute zones, virtual local area networks (VLANs), racks, fault domains, and the like. For example, many cloud computing services are divided into clusters of nodes (e.g., node clusters). Cloud computing systems often use different types of virtual services (e.g., compute containers, virtual machines) that provide remote storage and computing functionality to various clients or customers. These virtual services can be hosted by server nodes on the cloud computing system.
[0003] As cloud computing continues to proliferate, it becomes increasingly difficult to manage different types of services and provide sufficient cloud-based resources to customers. For example, as the demand for cloud computing resources continues to grow, more and more customers and tenants request to deploy cloud computing resources at a higher rate. However, as the demand for computing resources increases, traditional systems for allocating computing resources to accommodate resource requests have many problems and shortcomings.
[0004] For example, traditional allocation systems are often limited by the ability of the allocation system to handle a large number of incoming requests. For example, allocation agents are often implemented on server devices (e.g., server nodes) that have limited computing capabilities. Thus, while traditional allocation systems can often allocate hundreds of discrete resources per minute, in the case of an allocation agent that receives thousands of resource requests in a short period of time, it can cause delays in allocating resources to accommodate sudden spikes in received resource requests. Moreover, as the processing of resources continues to improve, the demand for cloud computing resources continues to increase, and modern server devices are unable to provide sufficient throughput to accommodate a large number of time periods of deployment requests.
[0005] To accommodate a larger number of service requests, some traditional systems operate multiple server devices that provide multiple allocation agents that are capable of operating in parallel. For example, traditional systems can handle received placement requests concurrently, which enables a larger number of resource requests to be placed on available computing resources. However, these parallel allocation agents often encounter placement conflicts as one or more allocation agents attempt to allocate overlapping computing resources for two or more resource requests. These placement conflicts can cause significant delays and often result in multiple processing of placement requests before a resource is successfully placed on the cloud computing system.
[0006] These and other problems exist with allocating computing resources in response to receiving a large number of placement requests. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 An example environment of a computing region including a collision avoidance system is shown in accordance with one or more embodiments.
[0008] Figure 2 An example implementation of a collision avoidance system in maintaining and updating a placement policy is shown in accordance with one or more embodiments.
[0009] Figures 3A-3B An example implementation of a placement policy being modified to accommodate a large number of placement requests is shown in accordance with one or more embodiments.
[0010] Figure 4 An example implementation of a collision avoidance system in selectively modifying a placement policy is shown in accordance with one or more embodiments.
[0011] Figures 5A-5B An example implementation of a placement policy being selectively modified to accommodate a large number of placement requests is shown in accordance with one or more embodiments.
[0012] Figure 6 An example series of actions for modifying and implementing a placement policy to reduce allocation conflicts is shown in accordance with one or more embodiments.
[0013] Figure 7 Certain components that can be included within a computer system are shown. DETAILED DESCRIPTION
[0014] The present disclosure generally relates to systems and methods for reducing conflicts in placement (e.g., allocation, deployment) of services (e.g., virtual machines) on server nodes of a cloud computing system. In particular, the present disclosure relates to a collision avoidance system that prevents or otherwise reduces conflicts resulting from allocation agents operating in parallel and attempting to allocate overlapping computing resources to accommodate a large number of placement requests (e.g., container or virtual machine placement requests). The collision avoidance system is able to reduce placement conflicts by maintaining and dynamically updating a placement policy that causes allocation agents to allocate computing resources so as to maximize or otherwise optimize utilization of computing resources on a particular computing region.
[0015] As will be discussed in further detail herein, the collision avoidance system is able to reduce placement conflicts in a variety of ways. For example, the collision avoidance system is able to implement a placement policy that results in a distribution agent, in response to an incoming placement request, pseudo-randomly assigning a compute resource. In one or more embodiments, the collision avoidance system reduces placement conflicts by modifying the policy in a manner that reduces the probability that an assignment implemented by a first distribution agent will collide with a concurrent assignment implemented by a second distribution agent. Further, the collision avoidance system can selectively modify the manner and location of assignment for certain types of placement requests to further reduce placement conflicts between distribution agents.
[0016] As an example, and as will be discussed in further detail below, the collision avoidance system can maintain a placement store including records of deployed services (e.g., containers, virtual machines) on a plurality of compute nodes of a compute region, where the compute region has a plurality of agent distributors implemented thereon. The collision avoidance system can determine that a number of placement conflicts between the plurality of agent distributors with respect to incoming placement requests exceeds a threshold number of placement conflicts in a recent period of time. Based on detecting the threshold number of placement conflicts, the collision avoidance system can modify the placement policy by relaxing one or more restrictions from the placement policy associated with assigning resources on the plurality of compute nodes of the compute region. Additional information related to one or more examples will be discussed herein.
[0017] The present disclosure includes a number of practical applications that provide benefits and / or solve problems associated with reducing or otherwise preventing the occurrence of placement conflicts due to multiple distribution agents operating in parallel with one another on a particular compute region. Some examples of these applications and advantages will be discussed in further detail below.
[0018] For example, in one or more embodiments, the collision avoidance system partially randomizes the assignment of compute resources according to a placement policy. In particular, where a placement policy having a same set of assignment rules would likely cause concurrently operating distribution agents to conflict with one another, one or more embodiments described herein involve randomly assigning resources from a set of compute nodes by using a partially random placement policy. For example, the collision avoidance system reduces the probability that two or more concurrently operating distribution agents assign overlapping sets of resources by selecting a sufficiently large set of eligible compute nodes to reduce the probability that the two or more concurrently operating distribution agents assign overlapping sets of resources to reduce placement conflicts.
[0019] In addition to partially randomizing the placement of services, the collision avoidance system can also dynamically modify placement policies based on observed conflicts for compute regions. In particular, in cases where a high volume of placement requests begin to cause placement conflicts between two or more allocation agents, the collision avoidance system can relax or otherwise reduce one or more allocation rules to increase the number of eligible compute nodes that are able to accommodate allocation of compute resources. This temporary expansion of eligible compute nodes allows the collision avoidance system to reduce placement conflicts while continuing to enable allocation agents to pursue valuable placement goals for compute regions (e.g., reduce fragmentation, optimize resources).
[0020] Further, in one or more embodiments, the collision avoidance system enables selective randomization of resource allocation for particular types of resources. For example, when the collision avoidance system observes that only certain types of virtual machines are associated with a small uptick in observed placement conflicts, the collision avoidance system can selectively update placement policies to affect the first type of virtual machine without modifying placement policies to affect other types of virtual machines. In this way, the collision avoidance can reduce placement conflicts while allowing allocation agents to continue to optimize allocation of cloud computing resources as much as possible.
[0021] As indicated by the foregoing discussion, the present disclosure utilizes various terminology to describe features and advantages of the systems described herein. Additional details regarding the meaning of some example terminology are now provided.
[0022] For example, as used herein, a "cloud computing system" refers to a network of connected computing devices that provide various services to client devices (e.g., client devices, network devices). For example, as described above, a distributed computing system can include a collection of physical server devices (e.g., server nodes) organized in a hierarchical structure, including clusters, compute regions, virtual local area networks (VLANs), racks, fault domains, etc. A cloud computing system can refer to a private or public cloud computing system.
[0023] As used herein, a "compute region" or "region" relates to any grouping or set of multiple compute nodes on a cloud computing system. For example, a compute region can relate to a cluster of nodes, a cluster of clusters of nodes, a data center, multiple data centers (e.g., a region of multiple data centers), a server rack, a row or multiple rows of server racks, a group of nodes powered by a common power supply, or other hierarchical structure in which network devices are grouped together physically or virtually. In one or more embodiments, a compute region refers to a cloud computing system. Alternatively, a compute region can relate to any subset of network devices of a cloud computing system. A compute region can include any number of nodes thereon. As an example and not a limitation, in one or more embodiments, a compute region can include anywhere between 1000 nodes to 250,000 nodes.
[0024] In one or more embodiments described herein, a compute region may include multiple allocation agents. As used herein, "allocation agent" refers to an application, routine, executable instruction, or other software and / or hardware mechanism used by a server device to allocate resources on server nodes of a compute region. For example, an allocation agent may refer to a service deployed on a cloud computing system on a server device, configured to execute instructions for allocating compute resources in response to a received deployment request. In one or more embodiments, the allocation agent allocates resources for the deployment of virtual machines, compute containers, or any other type of service that may be deployed on server nodes of a cloud computing system. Indeed, while one or more embodiments described herein specifically refer to allocation agents for deploying virtual machines on nodes of a compute region, similar features can be applied to deploying any type of service, such as compute containers or other cloud-based services. As will be discussed in more detail herein, a compute region may include multiple allocation agents operating in parallel to allocate resources according to one or more deployment strategies.
[0025] As used herein, a “deployment request” refers to any request for deploying a service on a cloud computing system. For example, a deployment request may refer to a request for deploying or launching a virtual machine according to a virtual machine specification. A deployment request may include the customer’s identity, specifications for the virtual machine (or other resource), such as the size of the virtual machine (e.g., the number of compute cores), the type of the virtual machine, or any other information that may be used by the allocation agent to determine the location (e.g., a server node, a set of cores on a server node) on which to deploy (multiple) virtual machines. A deployment request may include a request to deploy a single instance of compute resource (e.g., a single virtual machine). Alternatively, a deployment request may include a request to deploy multiple service instances (e.g., multiple virtual machines) for one or more customers.
[0026] As used herein, a “placement conflict” or “placement collision” refers to a conflict between two or more allocation agents when attempting to allocate computing resources in a compute region. Specifically, in one or more embodiments described herein, a placement conflict refers to a situation where a first allocation agent attempts to allocate compute resources to a second virtual machine (or other type of service) in the same location (e.g., using the same set of compute resources). For example, a placement conflict may exist where a first allocation agent allocates a first set of resources to a first service (e.g., a first virtual machine), while a second allocation agent attempts to allocate the same set of resources (or some overlapping sets of resources) to a second service (e.g., a second virtual machine). In practice, a placement conflict can also refer to any situation where an allocation agent is unable to place a virtual machine or other service on compute resources due to a prior action by another allocation agent.
[0027] As used herein, a “deployment policy” refers to any instruction or rule associated with allocating resources on a compute region. For example, a deployment policy may include a hierarchy of rules followed by an allocation agent when allocating resources for deploying services on compute resources within a compute region. A deployment policy may include information such as the number of compute cores, the type or generation of server nodes(s), the maximum fragmentation of candidate compute nodes, or other characteristics of the hardware(s) on which they are allocated for an associated request. A deployment policy may include objectives associated with the desired state of the compute region, such as the desired fragmentation of compute nodes, a target or minimum number of healthy empty nodes on which no virtual machines are deployed, or other preferences associated with the compute region. In one or more embodiments described herein, a deployment policy includes a list of rules governing the deployment of virtual machines (and other services) on a compute region. For example, in one or more implementations, a deployment policy includes a list of rules ordered by importance.
[0028] As used herein, “deployment,” “service deployment,” or “deployment” can be used interchangeably to refer to one or more associated services and allocations provided by a cloud computing system via a computing region. For example, an deployment can refer to the deployment of one or more service instances (e.g., virtual machines, containers) on one or more server nodes that are capable of providing computing resources according to the specifications of one or more deployment requests. An deployment can involve one or more services provided based on a single deployment request. In one or more implementations, a deployment refers to one or more service instances provided via server nodes.
[0029] As used herein, the terms "core," "computing core," or "node core" can be used interchangeably to refer to computing resources or units of computing resources provided via computing nodes (e.g., server nodes) of a cloud computing system. A computing core can refer to a virtual core that uses the same processor without interfering with other virtual cores operating with that processor. Alternatively, a computing core can refer to a physical core that is physically separate from other computing cores. A computing core implemented on one or more server nodes can refer to various different cores with different sizes and capabilities. A server node may include one or more computing cores implemented thereon. Furthermore, multi-core sets can be allocated for hosting one or more virtual machines or other cloud-based services.
[0030] Additional details about the collision avoidance system will now be provided in conjunction with illustrative diagrams depicting the example implementation. For example, Figure 1 An example environment 100 including a computing area 102 is shown. The computing area 102 may include any number of devices. For example, in one or more embodiments, the computing area 102 refers to a cloud computing system or a portion of a cloud computing system having any number of networked devices. Figure 1As shown, the computing area 102 includes (multiple) server devices 104, on which a collision avoidance system 106 is implemented.
[0031] like Figure 1 As shown, the collision avoidance system 106 includes a collision detector 108, an allocation area manager 110, and a data storage device 112. The data storage device 112 may include layout strategy data 114 and layout status data 116. As will be discussed in further detail below, according to one or more embodiments described herein, the collision avoidance system 106 performs features and functions related to maintaining layout storage, identifying layout conflicts, and modifying layout strategies implemented by the allocation agent. Additional details relating to each component 108-116 of the collision avoidance system 106 will be described in more detail below.
[0032] like Figure 1 As shown, the computing region 102 includes multiple allocation agents 118a-118b. Specifically, Figure 1 The example compute region 102 shown includes a first allocation agent 118a and a second allocation agent 118b. Compute region 102 may include any number of allocation agents communicating with (or incorporated into) collision avoidance system 106. In one or more embodiments, allocation agents 118a-118b are implemented on a single server device. Alternatively, in one or more implementations, environment 100 includes one or more agents implemented across multiple server devices. As described above, allocation agents 118a-118b may include applications, routines, software, and / or hardware implemented on one or more server devices configured to allocate compute resources to compute region 102 to enable the deployment of virtual machines (or other services) on the allocated resources.
[0033] In one or more embodiments, allocation agents 118a-118b may include a placement policy 120, which includes rules and instructions for allocating computing resources on compute region 102. As will be discussed in further detail below, in addition to a determined number of placement conflicts occurring between allocation agents 118a-118b (e.g., due to a large number of incoming placement requests), placement policy 120 may include updated or modified placement policies for allocation or deployment targets of compute region 102. Additional information relating to resource allocation and virtual machine placement will be discussed below.
[0034] like Figure 1As shown, computing region 102 may further include multiple node clusters 122a-122n. Node clusters 122a-122n can be grouped by geographical location (e.g., the region of the node cluster). Node clusters 122a-122n can be implemented across multiple geographical regions (e.g., in different data centers and / or on different server racks). Note that although one or more embodiments described herein specifically relate to the grouping of server nodes within a given node cluster, other device groupings can be similarly used in resource allocation and deployment strategies according to one or more embodiments described herein.
[0035] Each node cluster in node clusters 122a-122n can include various server nodes 124a-124n with multiple and different computing cores. For example... Figure 1 As shown, server nodes 124a-124n may include virtual machines 126a-126n implemented thereon. For illustration, the first node cluster 122a may include a first server node set 124a having a first virtual machine set 126a. Specifically, the first server node set 124a may have computing cores capable of hosting multiple virtual machines 126a for clients of a cloud computing system. Each of the additional node clusters 122b-122n having server nodes 124b-124n and virtual machines 126b-126n may have similar features and functions to the first node cluster 122a, the first server node set 124a, and the associated virtual machines 126a.
[0036] Multiple virtual machines 126a-126n can occupy a portion of the computing resources (e.g., computing cores) of server nodes 124a-124n to varying degrees of fragmentation. Specifically, although... Figure 1 Not shown, but server nodes 124a-124n may include a combination of occupied nodes, empty nodes and fragmented nodes.
[0037] As used herein, an occupied node can refer to a server node where each compute core is occupied by a virtual machine or other cloud-based service (e.g., no compute cores on such a server node are available for new allocation of cloud-based resources on it). An empty node can refer to a server node on which no virtual machines are deployed, nor are any compute cores allocated to prevent virtual machines or other cloud-based services from being used. An empty node can refer to a server node that can be used to deploy a variety of virtual machines, or can act as a recovery node in the event of a downtime of another server node with virtual machines on it, thus contributing to the greater overall health of the associated node cluster. As used herein, a fragmented node can refer to a server node where one or more compute cores are occupied by virtual machines or other cloud-based services, and one or more compute cores are available for allocation. Fragmented nodes can have an associated degree of fragmentation based on the ratio of empty cores to occupied cores (or the number of empty cores available for allocation).
[0038] like Figure 1 As shown, environment 100 includes multiple client devices 128a-128n communicating with computing region 102 via network 130 (e.g., communicating with different server nodes 124a-124n). Client devices 128a-128n can refer to various types of computing devices, including mobile devices, desktop computers, server devices, or other types of computing devices. Network 130 can include one or more networks that use one or more communication platforms or technologies to transmit data. For example, network 130 can include the Internet or other data links that enable the transport of electronic data between the respective client devices 128a-128n and devices in computing region 102. In one or more embodiments described herein, client devices 128a-128n can provide deployment requests to allocation agents 118a-118b, requesting the allocation of resources on node cluster 122a-122n and / or the deployment of virtual machines 126a-126n. Furthermore, although one or more embodiments described herein relate to client devices 128a-128n that provide deployment requests, other types of clients (e.g., internal cloud clients) may act as the source for deployment requests.
[0039] In one or more embodiments, the collision avoidance system 106 and allocation agents 118a-118b cooperate to allocate computing resources of node clusters 122a-122n in response to an incoming deployment request. For example, client devices 128a-128n may provide deployment requests, including requests for the deployment of virtual machines, containers, or other cloud computing resources on available resources of node clusters 122a-122n. As described above, the deployment request may indicate a specific type of virtual machine family(s) and an indicated number of virtual machine instances to be deployed on compute region 102. In one or more embodiments, the deployment request is provided to one of the allocation agents 118a-118b to determine the location for deployment of the requested resources. In one or more embodiments, the deployment request(s) are provided to one of the allocation agents 118a-118b via a load balancer or other mechanism for routing the deployment request(s) to any allocation agent(s) available to receive the request.
[0040] The allocation agents 118a-118n can determine the allocation of the requested service(s) based on the allocation strategy 120 implemented thereon. As described above, the allocation strategy 120 may include a set of instructions and / or rules that affect the allocation of the service(s) on the computing resources of the node clusters 122a-122n. The allocation strategy 120 may include any number of rules to optimize the allocation of services on computing resources in order to maximize the utilization of computing resources and reduce fragmentation of the selective server nodes and / or node clusters as a whole.
[0041] As an example, in one or more embodiments, the deployment strategy 120 includes rules for prioritizing the deployment of services on server nodes based on the fragmentation of server nodes on the node cluster and / or the fragmentation of server nodes on the node cluster. For example, in response to receiving a deployment request, the first allocation agent 118a may identify the first node cluster 122a based on the overall fragmentation of the first node cluster 122a relative to the attached node clusters. Within the first node cluster 122a, the first allocation agent 118a may also selectively identify one or more server nodes with a sufficient number of compute cores (e.g., fragmented server nodes) capable of hosting the virtual machine indicated by the deployment request.
[0042] In one or more embodiments, the placement strategy 120 includes a series of rules that indicate the criteria by which the allocation agent should place virtual machines on server nodes. The allocation agent may iterate through each rule until it finds the optimal location for the virtual machines on a particular node cluster and / or server node. Where the allocation agent does not necessarily identify the most likely placement (e.g., according to each placement rule), the allocation agent may identify the next best or acceptable placement of the virtual machines on any computing resources indicated by the placement rules.
[0043] As described above, in one or more embodiments, the placement rules may be overly specific, causing allocation agents 118a-118b to attempt to place virtual machines on the same set of available compute resources based on the same set of placement rules from the same placement policy 120. Therefore, in one or more embodiments, the allocation agents perform semi-random placement of virtual machines (and other services) according to one or more embodiments described herein. For example, instead of identifying specific servers, the allocation agents may apply the placement rules of placement policy 120 to identify placement regions with multiple possible server nodes that conform to (or largely conform to) the criteria indicated by the placement rules. A placement region may represent a subset of a cluster or a subset of server nodes from a larger set of capable nodes used to host (multiple) virtual machines. After identifying the placement region, the allocation agents (multiple) may randomly allocate compute resources for the placement of virtual machines in response to a placement request.
[0044] While randomizing resource placement can significantly reduce placement conflicts between allocation agents 118a-118b, allocation agents 118a-118b can still attempt to allocate the same or overlapping computing resources for the placement of two or more services. In particular, when the volume of incoming placement requests is particularly high, allocation agents 118a-118b may begin to experience a large number of placement conflicts, which cause the allocation agents to reprocess incoming placement requests, potentially leading to a slowdown in resource deployment on compute region 102.
[0045] As described above, the collision avoidance system 106 can implement the features and functions described herein to enable the allocation agents 118a-118b to experience fewer arrangement conflicts with each other. For example, as Figure 1 As shown, the collision avoidance system 106 includes a collision detector 108. The collision detector 108 can detect conflicts between allocation agents 118a-118b within a predetermined time period. For example, in one or more embodiments, the collision detector 108 monitors instances of conflicts to determine whether placement conflicts are occurring at an increasing rate. More specifically, in one or more implementations, the collision detector 108 determines whether the number of placement conflicts within a predetermined time period is greater than or equal to a threshold number of placement conflicts. This may involve determining whether the number of placement conflicts exceeds a predetermined number and / or percentage of conflicts (e.g., relative to the number of incoming placement requests within the predetermined time period).
[0046] Collision detector 108 can detect placement conflicts in a variety of ways. For example, in one or more embodiments, collision detector 108 queries placement storage (e.g., placement state data 116) for the current state of computing resources that one of the allocation agents 118a-118b is attempting to place. If placement storage indicates that the computing resources(s) have already been allocated and used by another, collision detector 108 can determine that a conflict exists and provide an indication of the conflict to the allocation agent.
[0047] In one or more embodiments, the allocation agent provides an indication of an identified set of computing resources (e.g., an identified set of compute cores and / or an identified set of server nodes) to enable the collision detector 108 to locally determine whether a placement conflict exists. For example, upon receiving the identification of computing resources, the collision detector 108 can compare the identified computing resources with information from the placement store to determine whether one or more allocation agents 118a-118b have previously allocated the same set of computing resources for placement of another service. In one or more embodiments, the collision detector 108 provides an indication of a placement conflict to the allocation agents 118a-118b.
[0048] Collision detector 108 can track these collisions to determine the number of collisions within a predetermined time period. For example, collision detector 108 can maintain the total number of detected collisions over a time period of 1-2 minutes (or any other time interval) to determine whether placement collisions are occurring at an increasing frequency. As described above, collision detector 108 can determine whether the number of tracked placement collisions exceeds a threshold number or percentage for that predetermined time period, which can be used to determine whether adjustments are needed to how computing resources are allocated in response to incoming placement requests.
[0049] like Figure 1 As further shown, the collision avoidance system 106 includes an allocation region manager 110. As described above, allocation agents 118a-118b can perform partially random placement of virtual machines in an attempt to reduce placement conflicts between them. In one or more embodiments described herein, the allocation region manager 110 can manage the degree of randomness in which virtual machines are placed on appropriate resources in compute region 102.
[0050] For example, in one or more embodiments, allocation region manager 110 identifies a set of candidate server nodes, which may include a subset of server nodes within compute region 102. The candidate nodes can act as a target node set, on which allocation agents 118a-118b can allocate compute resources in response to an incoming deployment request. For example, in one or more embodiments, allocation agents 118a-118b randomly deploy virtual machines onto server nodes based on the set of candidate server nodes identified by allocation region manager 110.
[0051] In one or more embodiments, the allocation region manager 110 modifies the placement region based on the number of identified placement conflicts. For example, in one or more embodiments, the allocation region manager 110 expands the placement region to include a larger number of server nodes based on a determination that the number of identified placement conflicts (e.g., within a predetermined time period) exceeds a threshold number or percentage of placement conflicts. In one or more embodiments, the allocation region manager 110 may further expand the placement region based on a determination that the number of placement conflicts continues to exceed a threshold number of placement conflicts. Alternatively, in one or more embodiments, the allocation region manager 110 may reduce or shrink the placement region to a more targeted set of candidate server nodes based on a determination that the number of identified placement conflicts has decreased by a certain threshold amount.
[0052] Therefore, the allocation region manager 110 can modify the degree of randomness associated with the allocation of computing resources and the placement of virtual machines in response to an incoming placement request. Specifically, the allocation region manager 110 can increase or decrease the size of the placement region (e.g., increase or decrease the number of candidate server nodes) based on the number of identifications of placement conflicts experienced by the allocation agents 118a-118b, and randomly place virtual machines on one or more server nodes within the placement region.
[0053] It will be understood that the collision detector 108 and the allocation area manager 110 can perform actions associated with identifying layout conflicts and modifying layout areas in various ways. Additional details relating to each of these features and functions of the collision avoidance system 106 will be discussed further in detail, and through... Figure 1 The example configurations and implementations shown are discussed in relation to the examples.
[0054] As stated above, and as Figures 2-5BAs further shown, the collision avoidance system 106 includes a data storage device 112 on which deployment policy data 114 is stored. Deployment policy data 114 may include any information used by allocation agents 118a-118b in determining the deployment of virtual machines or other cloud-based services. For example, deployment policy data 114 may include a list of allocation or deployment rules followed by allocation agents 118a-118b in response to an incoming deployment request to determine where to deploy a virtual machine. In one or more embodiments, the rules of deployment policy data 114 may also include a hierarchy of rules corresponding to allocation objectives, such as optimizing allocations or resources to optimize fragmentation of cloud computing resources on compute region 102.
[0055] Although Figure 1 An example is shown where deployment strategy data 114 is maintained within data storage device 112 on collision avoidance system 106; however, in one or more embodiments, deployment strategy data is represented as a series of hard-coded rules as part of allocation agents 118a-118b. In one or more embodiments, the behavior of allocation agents 118a-118b can be controlled by changing one or more configuration settings.
[0056] like Figure 1 As further shown, the data storage device 112 may include arrangement status data 116. Arrangement status data 116 may include any information related to the current state of resource allocation on compute region 102. For example, arrangement status data 116 may include a record indicating which server nodes are currently occupied by virtual machines. Arrangement status data 116 may include information identifying any number of compute cores that are occupied or available for allocation, and may also include an indication of fragmentation for each server node on compute region 102. In one or more embodiments, arrangement status data 116 includes arrangement storage having key-value pairs representing the arrangement of services on the corresponding server nodes of compute region 102. Specifically, arrangement status data 116 may include stored virtual machine identifier pairs and identifiers of the server nodes (or specific cores of the corresponding server nodes) occupied by the identified virtual machines.
[0057] Although Figure 1 Not shown, but in one or more embodiments, allocation agents 118a-118b may be coupled to collision avoidance system 106 via a communication reverse channel (or simply "reverse channel"). Specifically, allocation agents 118a-118b and collision avoidance system 106 may maintain communication. This reverse channel may be used in a variety of ways.
[0058] For example, in one or more embodiments, the reverse channel enables allocation agents 118a-118b to transmit information about newly allocated resources to collision avoidance system 106 for storage in the deployment storage. For instance, after identifying compute resources (e.g., server nodes, compute core sets), the allocation agents may provide the identifier of the compute resource to collision avoidance system 106 for verification against the current version of the deployment storage. Collision avoidance system 106 can provide confirmation of availability or an indication that the identified compute resource has recently been allocated to a deployment for another virtual machine.
[0059] In one or more embodiments, the collision avoidance system 106 may provide periodic updates to allocation agents 118a-118b every few seconds or minutes via a reverse channel to provide an updated view of the current allocation status on compute region 102. In this way, the collision avoidance system 106 can maintain a current allocation record on the compute region while providing allocation agents 118a-118b with a semi-current version of the layout status data 116. While this may not eliminate layout conflicts for recent layout requests, it can still reduce layout conflicts because the allocation agents attempt to allocate compute resources previously allocated before receiving the most recent layout store update.
[0060] Now we will combine Figure 1 We will discuss additional information using example implementations and workflows. For example, Figures 2-5B An example implementation of a collision avoidance system 106 according to one or more embodiments described herein is shown. Specifically, Figures 2-3B An example workflow is shown, illustrating the actions that can be performed by combining the example. Figure 2 This illustrates how to implement a multi-node cluster in conjunction with an example. Figures 3A-3B The example visualization of the workflow shown is shown below.
[0061] Specifically, Figure 2 An example workflow 200 is shown, which includes a series of actions that can be performed by the collision avoidance system 106 to reduce placement conflicts between two or more allocation agents due to receiving a large number of placement requests in a short period of time. Figure 2 Each action shown can be performed by the collision avoidance system 106 and / or by one or more allocation agents of the associated computing region.
[0062] like Figure 2 As shown, the collision avoidance system 106 can perform action 202 to implement the initial deployment strategy. The initial deployment strategy may include... Figure 2The placement strategy 120 shown relates to any of the features discussed above. In one or more embodiments, the initial placement strategy includes a default placement strategy implemented by each of a plurality of allocation agents on the compute region. For example, the initial placement strategy may refer to a strategy that considers each of a plurality of placement rules when determining the location of a virtual machine (or other cloud-based service) on the compute region.
[0063] In one or more embodiments, the initial placement strategy refers to the most restrictive or most optimistic version of a placement strategy. For example, as described above, the initial placement strategy may include instructions to consider each of any number of placement rules affecting the placement of virtual machines on server nodes in a compute region. To illustrate, where the placement strategy includes instructions to randomly allocate compute resources across a candidate set of server nodes, the initial placement strategy may refer to a smaller set of candidate nodes for random allocation, rather than a set of rules representing other versions or potential modifications of the placement strategy. As another example, where the placement strategy includes a hierarchy of placement rules (e.g., rules ordered by importance), the initial placement strategy may include instructions to consider each placement rule in order of importance when determining which compute resources to allocate for the placement of virtual machines.
[0064] like Figure 1 As further illustrated, the collision avoidance system 106 can perform the action 204 of tracking placement conflicts. Specifically, in cases where a computing region includes multiple allocation agents that allocate computing resources according to a placement policy, the collision avoidance system 106 can observe whether one or more attempts to place virtual machines result in placement conflicts between the multiple allocation agents. As described above, the collision avoidance system 106 can track placement conflicts by using a placement store that identifies server nodes and the associated virtual machine identifiers (or other identifiers of services) deployed on them. Specifically, the collision avoidance system 106 can receive allocation information from an allocation agent and determine whether the indicated allocation attempt conflicts with an allocation previously performed by another allocation agent in the same computing region.
[0065] In addition to typically tracking placement conflicts, the collision avoidance system 106 can also perform an action 206 to determine whether placement conflicts exceed a threshold. For example, the collision avoidance system 106 can determine whether the number of placement conflicts exceeds a threshold number and / or percentage. In one or more embodiments, the collision avoidance system 106 determines whether the identified placement conflicts exceed the threshold within a predetermined time period (e.g., 1-2 minutes). Figure 2 As shown, if the arrangement conflict does not exceed the threshold, the collision avoidance system 106 can perform action 204 and continue to track the arrangement conflict.
[0066] Alternatively, if the collision avoidance system 106 observes that the number of recently identified placement conflicts exceeds a threshold, the collision avoidance system 106 may perform an action to modify the placement strategy to identify a larger placement area. The collision avoidance system 106 may modify the placement strategy in a variety of ways. For example, in one or more implementations, the collision avoidance system 106 may modify the placement strategy by reducing, discarding, or otherwise relaxing one or more restrictions of the placement strategy, thereby making the previously applicable placement area (e.g., the placement area based on the initial placement strategy) larger.
[0067] In one or more embodiments, the collision avoidance system 106 modifies the placement strategy by providing a modification instruction to one or more of a plurality of allocation agents. In one or more implementations, the collision avoidance system 106 provides the modified placement strategy instruction to each of the plurality of allocation agents, causing each of the plurality of allocation agents to begin allocating computing resources across a wider range of server nodes (e.g., relative to a placement area based on the initial placement strategy).
[0068] like Figure 2 As shown, after modifying the placement strategy, the collision avoidance system 106 can re-execute actions 204-206 to determine whether placement conflicts exceed a threshold. Then, the collision avoidance system 106 can further modify the placement strategy accordingly. For example, if the number of observed conflicts continues to exceed the threshold number or percentage of placement attempts, the collision avoidance system 106 can further modify the placement strategy by further relaxing one or more placement rules and / or further expanding the placement area on which the allocation agent can allocate resources. Alternatively, if the number of observed conflicts decreases predictably, and the reduction in placement conflicts remains low over a predetermined time period, the collision avoidance system 106 can modify the placement strategy by reverting the modified placement strategy to the initial or default placement strategy.
[0069] In one or more embodiments, modifying the placement strategy involves an adaptive approach based on the observed collision rate. For example, with a collision rate of 30% or other proportional values, collision avoidance system 106 can modify the placement strategy by applying a relaxed strategy to the corresponding number or percentage of incoming requests. In this example, in response to a collision rate of 30%, 30% (or other proportional values) of requests will be placed using a relaxed strategy, while the remaining number or percentage of requests will be placed using a default strategy (e.g., without relaxing one or more rules). As another example, with a collision rate of 60%, 60% (or other proportional values) of requests can be placed using a relaxed strategy, while the remaining number or percentage of requests will be placed using a default strategy. Thus, as discussed in conjunction with one or more example implementations, collision avoidance system 106 can modify the placement strategy using a self-adjusting or adaptive approach based on different observed levels of placement conflict.
[0070] In this example, the threshold number or percentage of placement conflicts can refer to any non-zero proportion of placement conflicts. Furthermore, modifying the placement strategy in response to observed placement conflicts can involve selectively modifying or relaxing the placement strategy for a corresponding proportion of incoming placement requests while using the default placement strategy for other incoming placement requests. Therefore, according to a non-limiting example, the collision avoidance system 106 can utilize a sliding scale or dynamic modification of the placement strategy based on the observed proportion of placement conflicts relative to the total number of incoming placement requests. The collision avoidance system 106 can then dynamically modify the proportion of placement requests for which a relaxed placement strategy is applied (e.g., selectively applied) based on the real-time collision rate of the incoming placement requests.
[0071] Figure 2 An example visualization of an implementation of a collision avoidance system 106 according to one or more embodiments described herein is shown. For example, Figures 3A-3B A collision avoidance system 106 communicating with multiple allocation agents 118 is shown. (Example) Figure 3A As shown, the collision avoidance system 106 can enable the initial deployment strategy 301a to be implemented on each allocation agent 118. Therefore, in Figure 3A In the example shown, allocation agent 118 can receive incoming virtual machine requests (e.g., from multiple client devices 128). Allocation agent 118 can then allocate computing resources according to the initial deployment strategy 301a.
[0072] like Figure 3AAs shown, the allocation agent 118 can deploy virtual machines on a compute region 300 comprising multiple node clusters 302a-302f. Each node cluster 302a-f can include any number of server nodes. Furthermore, one or more node clusters can be located in one or more data centers. For example, node clusters can be located in a single regional data center or in different geographical locations within different data centers.
[0073] like Figure 3A As shown, and according to the initial deployment strategy 301a, the allocation agent 118 can selectively allocate computing resources on one of a plurality of server nodes within the identified deployment area 304. In this example, the allocation agent 118 can identify the deployment area 304 comprising a first node cluster 302a of a plurality of node clusters 302a-302f. Therefore, the allocation agent 118 can selectively allocate computing resources in response to incoming virtual machine requests on one or more server nodes of the first node cluster 302a.
[0074] More specifically, in this example, allocation agent 118 can identify the placement region 304, which includes the first node cluster 302a, based on the determination that the first node cluster 302a meets one or more placement rules from the initial placement strategy 301a. For example, allocation agent 118 can determine the threshold number of empty nodes or the degree of fragmentation of the first node cluster 302a, which would enable the placement of specific types of virtual machines on the server nodes of the first node cluster 302a, resulting in higher resource utilization efficiency on the computing region 300 compared to identifying one or more additional node clusters 302b-302f as candidate node clusters.
[0075] According to one or more embodiments described herein, after identifying the deployment area 304, the allocation agent 118 may randomly allocate computing resources within the deployment area 304. For example, in response to each virtual machine request in an incoming virtual machine request, the allocation agent 118 may randomly identify a server node within the first node cluster 302a and allocate computing resources for the deployment of virtual machines(multiple) on the randomly identified server nodes.
[0076] As described above Figure 3A As discussed, if allocation agent 118 successfully responds to virtual machine requests and deploys virtual machines without exceeding a deployment conflict threshold, allocation agent 118 can continue to allocate resources and deploy virtual machines on the first node cluster 302a. This process can continue until the first node cluster 302a is full or no longer meets the criteria for being designated as deployment area 304. However, in one or more embodiments, allocation agents 118 may begin to conflict with each other due to receiving a large number of virtual machine requests within a short period of time.
[0077] As described above, the collision avoidance system 106 can identify when the number of arrangement conflicts exceeds a threshold and modify the arrangement strategy accordingly. For example, such as Figure 2 As shown, the collision avoidance system 106 can provide a modified placement strategy 301b to the allocation agent 118. This modified placement strategy 301b allows the allocation agent 118 to identify an updated placement area 306 that includes a larger group of node clusters. In this example, as a result of reducing one or more restrictions on the initial placement strategy 301a to achieve the modified placement strategy 301b, the allocation agent 118 can identify a new placement area 306 that includes a first node cluster 302a, a second node cluster 302b, and a third node cluster 302c.
[0078] Although Figure 3B The example shown illustrates a tripling of the deployment area size (e.g., from a single-node cluster to a three-node cluster), but the deployment area can be increased with any number of server nodes. For example, with... Figure 3B Compared to the previous method, the placement area can be expanded with fewer server nodes (e.g., some or all of the node cluster) or more server nodes. Furthermore, in one or more embodiments, the change in the size of the placement area can be based on a metric of placement conflicts relative to a threshold. For example, if the collision avoidance system 106 identifies significantly more placement conflicts than the threshold (e.g., due to a large influx of placement requests), the placement area can be expanded with more server nodes compared to when the collision avoidance system 106 determines that the number of placement conflicts exceeds the threshold by a smaller amount (e.g., a lower percentage).
[0079] Alternative locations, as mentioned above Figure 3B The placement area discussed can be incrementally modified until the number of placement conflicts no longer exceeds a threshold. For example, the collision avoidance system 106 can modify the initial placement policy 301a multiple times as it reaches a modified placement policy 301b by gradually relaxing one or more placement rules from the original placement policy. In this example, the placement area can be increased by the incremental number of server nodes, for example, by increasing the size of the placement area through a single node cluster, to reach... Figure 2 The arrangement area 306 for the three node clusters shown. Alternatively, the arrangement area can be increased by a fixed number of server nodes and / or a percentage of server nodes relative to the initial arrangement area 304.
[0080] In each of the examples above, the allocation agent 118 can randomly allocate computing resources from the identified layout area. For example, in the identified... Figure 3BAfter the modified placement area 306 shown, the allocation agent 118 can randomly allocate computing resources and place virtual machines on random server nodes from any of the three node clusters 302a-302c based on the modified placement policy 301b. Although placement area 306 includes a larger number of node clusters 302a-302c, the allocation agent 118 can experience fewer placement conflicts due to the larger number of server nodes on which computing resources can be randomly allocated.
[0081] Figure 3B Another example implementation is shown, in which the collision avoidance system 106 and multiple allocation agents 118 can collaboratively reduce placement conflicts on server nodes in a computing region. Specifically, Figures 4-5B Example workflow 400 is shown, illustrating a series of actions that can be performed by collision avoidance system 106 in reducing layout conflicts between allocation agents. Figure 4 One or more actions shown may be performed by the collision avoidance system 106 and / or by one or more assignment agents targeting the associated computing region.
[0082] like Figure 4 As shown, workflow 400 includes elements combined with the above. Figure 4 Several actions are similar to the actions discussed. For example, actions 402-406 may include features and functions similar to the corresponding actions 202-206 discussed above, which combine the implementation of the initial placement strategy, tracking placement conflicts, and determining whether the number of placement conflicts observed within a predetermined time period is greater than or equal to a threshold (e.g., a percentage threshold number).
[0083] In this example, and in other embodiments, a deployment strategy may include different rules or instructions applicable to deploying different types of virtual machines (or other cloud-based services). For example, a virtual machine of a first type associated with a first set of characteristics may be associated with a different deployment region or a different set of deployment rules than a virtual machine of a second type associated with a second set of characteristics. In fact, different characteristics of virtual machines may make them more suitable for deployment on different clusters, rather than one location within a cluster. This could be the result of different characteristics of the servers themselves, different sizes of the virtual machines, differences in the applications used by the virtual machines, or other factors.
[0084] Therefore, as Figure 2As shown, in addition to typically determining whether a placement conflict exceeds a threshold, the collision avoidance system 106 can also perform an action 408 to determine whether the identified placement conflict is specific to a particular machine type. For example, the collision avoidance system 106 may determine that the identified placement conflict occurred for the placement of a first type of virtual machine, while a placement request for a second type of virtual machine did not cause any placement conflict (or at least was less than the threshold).
[0085] Then, the collision avoidance system 106 can selectively modify the placement policy based on whether the placement conflict is located for a specific type of virtual machine. For example, if the collision avoidance system 106 determines that the placement conflict is not unique for a particular virtual machine, the collision avoidance system 106 can perform an action 410 to modify the placement policy to identify a larger placement area. This action 410 may include actions related to the above. Figure 4 The discussion corresponds to action 208 and has similar features.
[0086] Alternatively, if the collision avoidance system 106 determines that a placement conflict is unique for a specific virtual machine type, the collision avoidance system 106 may perform an action 412 to selectively modify the placement policy for that specific virtual machine type. Specifically, similar to one or more embodiments described herein, the collision avoidance system 106 may selectively modify the placement policy to identify a larger placement area for virtual machine types experiencing placement conflicts at a higher rate.
[0087] In either case (e.g., whether the collision avoidance system 106 determines that the placement conflict is specific to the virtual machine type), the collision avoidance system 106 can return to action 404 and continue tracking placement conflicts between allocation agents in the compute regions. Furthermore, based on observed changes in the rate at which placement conflicts occur on the compute regions, the collision avoidance system 106 can continue to modify or restore placement policies or portions of placement policies applicable to different virtual machine types.
[0088] continue, Figure 2 An example visualization of an implementation of a collision avoidance system 106 according to one or more embodiments described herein is shown. Similar to the above combination Figures 5A-5B Examples discussed, Figures 3A-3B A collision avoidance system 106 communicating with multiple allocation agents 118 is shown. Similar to... Figure 5A The collision avoidance system 106 enables the initial deployment strategy 501a to be implemented on each allocation agent 118. Similarly, allocation agents 118 can receive virtual machine deployment requests from multiple client devices 128. Figure 3A As shown, a deployment request may include a first set of deployment requests associated with a first type of virtual machine (represented as a VM-A request) and a second type of virtual machine (represented as a VM-B request).
[0089] like Figure 5A As shown, the allocation agent 118 can deploy virtual machines on a compute region 500 comprising multiple node clusters 502a-502f. Each node cluster in node clusters 502a-502f may include characteristics similar to the other node clusters discussed herein.
[0090] According to the initial deployment strategy 501a, the allocation agent 118 can selectively allocate computing resources on server nodes based on the identified deployment regions 504a-504b. In this example, the allocation agent 118 can allocate resources for a first type of virtual machine on server nodes in the first deployment region 504a (e.g., in response to a first set of deployment requests). Similarly, the allocation agent 118 can allocate resources for a second type of virtual machine on server nodes in the second deployment region 504b (e.g., in response to a second set of deployment requests). In this example, the first deployment region 504a includes a first node cluster 502a, while the second deployment region 504b includes a third node cluster 502c. Deployment regions may include overlapping server nodes shared between deployment regions. Alternatively, deployment regions may include, for example, Figure 5A The server nodes shown are non-overlapping groups.
[0091] The allocation agent 118 can allocate resources within a corresponding area according to the initial deployment strategy 501a. For example, similar to one or more embodiments described herein, the allocation agent 118 can randomly select server nodes within the corresponding deployment areas 504a-504b.
[0092] Based on the number of placement conflicts observed by the collision avoidance system 106, the placement strategy can be modified by reducing one or more restrictions on how the allocation agent 118 is instructed to allocate computing resources. In one or more embodiments, the initial placement strategy 501a is modified similarly to one or more examples discussed above. In this example, the initial placement strategy 501a can be selectively modified relative to the rules associated with a particular type of virtual machine.
[0093] Specifically, such as Figure 5A As shown, the collision avoidance system 106 can selectively relax one or more placement rules associated with the placement of virtual machines of a first type associated with the first placement region 504a. Specifically, when the collision avoidance system 106 determines that placement conflicts occur at a higher frequency relative to virtual machines of the first type (e.g., associated with the first placement request set), the collision avoidance system 106 can selectively modify the rules associated with the placement of the first virtual machine type. This can be performed without modifying the rules associated with the placement of the second virtual machine type.
[0094] like Figure 5B As shown, the collision avoidance system 106 can implement a modified layout policy 501b on the allocation agent 118. As further shown, the modified layout policy 501b may include one or more relaxed layout rules, thereby generating a first modified layout region 506a associated with a first virtual machine type and an original layout region 504b associated with a second virtual machine type. Figure 5B As shown, the updated first arrangement area 506a includes a first node cluster 502a and a fourth node cluster 502b, while the second arrangement area 504b again includes a third node cluster 502c.
[0095] Now go to Figure 5B The diagram illustrates an example flowchart that includes a series of actions to reduce placement conflicts between distribution agents when deploying cloud-based services across compute regions. While Figure 6 Actions according to one or more embodiments are shown, but alternative embodiments may omit, add, reorder, and / or modify. Figure 6 Any action shown. Figure 6 The action can be performed as part of a method. Alternatively, a non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause a computing device (e.g., a server device) to perform. Figure 6 The system can perform the following actions. In a further embodiment, the system can execute... Figure 6 The action.
[0096] Figure 6 A series of example actions 600 for reducing placement conflicts between allocation agents are illustrated. For example, the series of actions 600 includes action 610 of maintaining a placement storage device that includes records of services allocated across multiple compute nodes in a compute region. In one or more embodiments, action 610 relates to maintaining a placement storage that includes records of compute resources allocated across multiple compute nodes in a compute region, wherein the compute region includes multiple agent allocators for allocating resources according to a placement policy in response to an incoming placement request. In one or more embodiments, maintaining the placement storage includes pairing storage service identifiers with node identifiers that indicate the placement of one or more services on corresponding compute nodes in the compute region.
[0097] As further shown, the series of actions 600 includes action 620 determining that the number of arrangement conflicts among multiple agent allocators regarding incoming arrangement requests is greater than a threshold. For example, action 620 may include determining, based on information from records of allocated computing resources, that the number of arrangement conflicts among multiple agent allocators regarding incoming arrangement requests is greater than or equal to a threshold number of arrangement conflicts within a predetermined time period.
[0098] In one or more embodiments, a series of actions 600 includes identifying placement conflicts based on detected conflicts between service placements attempted by one or more of a plurality of proxy allocators and services previously placed as indicated in the record of allocated computing resources. Furthermore, in one or more embodiments, determining that the number of placement conflicts is greater than or equal to a threshold number of placement conflicts includes detecting a threshold percentage of failed submissions by the plurality of proxy allocators for incoming placement requests.
[0099] As further illustrated, a series of actions 600 includes action 630 modifying a placement policy for multiple agent allocators by reducing restrictions from a placement policy associated with allocating resources on a compute region, based on the number of placement conflicts exceeding a threshold. For example, action 630 may include modifying a placement policy for multiple agent allocators, associated with allocating resources on multiple compute nodes in a compute region, by reducing one or more restrictions from a placement policy determined to be greater than or equal to a threshold number of placement conflicts.
[0100] In one or more embodiments, the deployment strategy includes a set of rules executable by a plurality of proxy allocators to identify candidate nodes for resource allocation by identifying a subset of compute nodes from a plurality of compute nodes in a compute region, and to randomly allocate resources for an incoming deployment request on the identified candidate nodes. In one or more embodiments, modifying the deployment strategy includes modifying the set of rules to expand the candidate nodes to include a subset of compute nodes and additional compute nodes from the plurality of compute nodes in the compute region. In one or more embodiments, modifying the set of rules includes ignoring one or more rules from the set of rules to expand the candidate nodes eligible for resource allocation.
[0101] In one or more embodiments, a series of actions 600 includes identifying a first set of arrangement requests for a first type of resource and a second set of arrangement requests for a second type of resource from incoming arrangement requests. The series of actions 600 also includes determining a number of arrangement conflicts greater than or equal to a threshold number associated with the first set of arrangement requests. In one or more implementations, modifying the arrangement policy includes selectively reducing one or more restrictions from the arrangement policy for the first set of arrangement requests, while not reducing one or more restrictions from the arrangement policy for the second set of arrangement requests.
[0102] In one or more embodiments, a series of actions 600 includes determining that the number of updates to placement conflicts among a plurality of agent allocators continues to be greater than or equal to a threshold number of placement conflicts under a modified placement policy. In one or more implementations, a series of actions 600 includes further modifying the placement policy by reducing one or more additional constraints in the placement policy based on the determination that the number of updates to placement conflicts continues to be greater than or equal to the threshold number of placement conflicts.
[0103] In one or more embodiments, a series of actions 600 includes determining, under a modified placement policy, that the number of updates to placement conflicts among multiple agent allocators has decreased by a threshold amount. The series of actions 600 may further include, based on the determination that the number of updates to placement conflicts has decreased by the threshold amount, causing the multiple agent allocators to revert to the placement policy.
[0104] Figure 6 Some components that may be included within computer system 700 are shown. One or more computer systems 700 may be used to implement the various devices, components, and systems described herein.
[0105] Computer system 700 includes processor 701. Processor 701 can be a general-purpose single-chip or multi-chip microprocessor (e.g., an advanced RISC (Reduced Instruction Set Computer) machine (ARM)), a special-purpose microprocessor (e.g., a digital signal processor (DSP)), a microcontroller, a programmable gate array, etc. Processor 701 can be referred to as a central processing unit (CPU). Although in... Figure 7 The computer system 700 shows only a single processor 701, but in alternative configurations, a combination of processors (e.g., ARM and DSP) can be used.
[0106] The computer system 700 also includes a memory 703 that is in electronic communication with the processor 701. The memory 703 can be any electronic component capable of storing electronic information. For example, the memory 703 can be embodied as random access memory (RAM), read-only memory (ROM), magnetic disk storage medium, optical storage medium, flash memory in RAM, on-board memory included in the processor, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, etc., including combinations thereof.
[0107] Instruction 705 and data 707 may be stored in memory 703. Instruction 705 may be executed by processor 701 to implement some or all of the functions disclosed herein. Executing instruction 705 may involve using data 707 stored in memory 703. Any of the various examples of modules and components described herein may be implemented in part or in whole as instruction 705 stored in memory 703 and executed by processor 701. Any of the various examples of data described herein may be data in data 707 stored in memory 703 and used by processor 701 during the execution of instruction 705.
[0108] The computer system 700 may also include one or more communication interfaces 709 for communicating with other electronic devices. The communication interfaces 709 may be based on wired communication technology, wireless communication technology, or both. Some examples of communication interfaces 709 include Universal Serial Bus (USB), Ethernet adapters, and wireless adapters operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 wireless communication protocol. Wireless communication adapter and infrared (IR) communication port.
[0109] Computer system 700 may also include one or more input devices 711 and one or more output devices 713. Some examples of input devices 711 include keyboards, mice, microphones, remote control devices, buttons, joysticks, trackballs, touchpads, and light pens. Some examples of output devices 713 include speakers and printers. A particular type of output device typically included in computer system 700 is a display device 715. Display devices 715 used with the embodiments disclosed herein can utilize any suitable image projection technology, such as liquid crystal displays (LCDs), light-emitting diodes (LEDs), gas plasma, electroluminescence, etc. A display controller 717 may also be provided for converting data 707 stored in memory 703 into text, graphics, and / or moving images (as applicable) displayed on display device 715.
[0110] The various components of the computer system 700 can be coupled together via one or more buses, which may include power buses, control signal buses, status signal buses, data buses, etc. For clarity, the various buses are... Figure 7 Figure 7 It is shown as bus system 719.
[0111] The techniques described herein can be implemented in hardware, software, firmware, or any combination thereof, unless specifically described as being implemented in a particular manner. Any features described as modules, components, etc., can also be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a non-transitory processor-readable storage medium comprising instructions that, when executed by at least one processor, perform one or more methods described herein. Instructions can be organized into routines, programs, objects, components, data structures, etc., which can perform specific tasks and / or implement specific data types, and can be combined or distributed as needed in various embodiments.
[0112] As used herein, a non-transitory computer-readable storage medium (device) may include RAM, ROM, EEPROM, CD-ROM, solid-state drive (“SSD”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and is accessible by a general-purpose or special-purpose computer.
[0113] The steps and / or actions of the methods described herein may be interchanged without departing from the scope of the claims. In other words, unless the correct operation of the described methods requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0114] The term "determine" encompasses a wide variety of actions; therefore, "determine" can include accounting, calculation, processing, deriving, investigating, searching (e.g., looking in a table, database, or other data structure), ascertaining, etc. Furthermore, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in storage), etc. Additionally, "determine" can include resolving, selecting, choosing, establishing, etc.
[0115] The terms “comprising,” “including,” and “having” are intended to be inclusive, meaning that other elements besides those listed may be present. Furthermore, it should be understood that references to “one embodiment” or “an embodiment” in this disclosure are not intended to exclude the existence of additional embodiments that also include the described features. For example, where compatible, any element or feature described with respect to embodiments herein may be combined with any element or feature of any other embodiment described herein.
[0116] This disclosure may be embodied in other specific forms without departing from its spirit or characteristics. The described embodiments are to be considered illustrative rather than restrictive. Therefore, the scope of this disclosure is indicated by the appended claims rather than by the foregoing description. Modifications within the meaning and equivalent scope of the claims should be included within its scope.
Claims
1. A method comprising: Maintain a record of computing resources allocated on multiple computing nodes in a computing region, the computing region including multiple proxy allocators configured to allocate resources in parallel according to a first allocation strategy in response to an incoming allocation request, the first allocation strategy being associated with allocating a first type of resources on computing nodes in the computing region; Based on the information from the records of the allocated computing resources, it is determined that the number of arrangement conflicts among the plurality of agent allocators regarding the allocation of the first type of resources is greater than or equal to a threshold number of arrangement conflicts. Based on the determination that the number of placement conflicts is greater than or equal to the threshold number, the first placement strategy is modified by reducing one or more restrictions of the first placement strategy associated with allocating the first type of resources in the computing region. as well as The plurality of agent allocators allocate the first type of resources in parallel across the plurality of computing nodes according to the modified first deployment strategy.
2. The method of claim 1, wherein the first arrangement strategy applies to each of the plurality of agent allocators.
3. The method according to claim 1, wherein modifying the first arrangement strategy includes: Selectively, one or more restrictions from the first deployment strategy are reduced for the first deployment request set, while one or more restrictions from the second deployment strategy are not reduced for the second deployment request set, which is associated with allocating resources of a different type than the first type.
4. The method according to claim 1, further comprising: The first arrangement strategy is modified without modifying the second arrangement strategy associated with allocating the second type of resources on computing nodes within the computing region, based on the fact that the second number of arrangement conflicts among the plurality of agent allocators is less than the second threshold number of arrangement conflicts.
5. The method according to claim 1, further comprising: It was determined that, under the modified first placement strategy, the updated number of placement conflicts among the plurality of agent allocators had been reduced by a threshold amount. as well as The threshold amount has been reduced based on the updated number of determined placement conflicts, causing the plurality of agent allocators to revert to the first placement strategy.
6. The method of claim 1, wherein maintaining the record of the allocated computing resources comprises: A pairing of storage service identifiers and node identifiers, the pairing indicating the arrangement of one or more services on corresponding computing nodes in the computing region.
7. The method according to claim 1, further comprising: The arrangement conflict is identified based on a detected conflict between the arrangement of services attempted by one or more of the plurality of agent allocators and the previously arranged services indicated in the record of the allocated computing resources.
8. The method of claim 1, wherein determining that the number of arrangement conflicts is greater than or equal to the threshold number of arrangement conflicts comprises: Detect the threshold percentage of submission failures of the multiple proxy allocators for incoming placement requests.
9. The method of claim 1, wherein the first placement strategy comprises a set of rules, the set of rules being executable by each of the plurality of proxy allocators to: Candidate nodes for resource allocation are identified by identifying a subset of computing nodes from the plurality of computing nodes in the computing region; and Resources for the incoming deployment request are randomly allocated on the identified candidate nodes.
10. The method of claim 9, wherein modifying the first arrangement strategy comprises: The set of rules is modified so that the plurality of agent allocators expand the candidate nodes to include a subset of the compute nodes and additional compute nodes from the plurality of compute nodes in the compute region.
11. The method of claim 10, wherein modifying the set of rules comprises: One or more rules are ignored from the set of rules to expand the candidate nodes that are eligible for resource allocation.
12. The method of claim 1, wherein causing the plurality of agent allocators to allocate the first type of resources in parallel comprises: Each of the plurality of proxy allocators processes the placement request in parallel and allocates resources on the plurality of computing nodes in the computing region according to the modified first placement strategy.
13. A system comprising: At least one processor; A memory, which is in electronic communication with the at least one processor; as well as Instructions, the instructions stored in the memory, the instructions being executable by the at least one processor to: Maintain a record of computing resources allocated on multiple computing nodes in a computing region, the computing region including multiple proxy allocators configured to allocate resources in parallel according to a first allocation strategy in response to an incoming allocation request, the first allocation strategy being associated with allocating a first type of resources on computing nodes in the computing region; Based on the information from the records of the allocated computing resources, it is determined that the number of arrangement conflicts among the plurality of agent allocators regarding the allocation of the first type of resources is greater than or equal to a threshold number of arrangement conflicts. Based on the determination that the number of placement conflicts is greater than or equal to the threshold number, the first placement strategy is modified by reducing one or more restrictions of the first placement strategy associated with allocating the first type of resources in the computing region. as well as The plurality of agent allocators are configured to: follow the modified first arrangement strategy. The first type of resources are allocated in parallel across the plurality of computing nodes.
14. The system of claim 13, wherein the first arrangement strategy applies to each of the plurality of agent allocators.
15. The system of claim 13, wherein modifying the first arrangement strategy comprises: Selectively, one or more restrictions from the first deployment strategy are reduced for the first deployment request set, while one or more restrictions from the second deployment strategy are not reduced for the second deployment request set, which is associated with allocating resources of a different type than the first type.
16. The system of claim 13, wherein determining that the number of arrangement conflicts is greater than or equal to the threshold number of arrangement conflicts comprises: Detect the threshold percentage of submission failures of the multiple proxy allocators for incoming placement requests.
17. The system of claim 13, wherein the first placement strategy comprises a set of rules, the set of rules being executable by each of the plurality of agent allocators to: Candidate nodes for resource allocation are identified by identifying a subset of computing nodes from the plurality of computing nodes in the computing region; and Resources for the incoming deployment request are randomly allocated on the identified candidate nodes.
18. The system of claim 17, wherein modifying the first arrangement strategy comprises: The set of rules is modified so that the plurality of agent allocators expand the candidate nodes to include a subset of the compute nodes and additional compute nodes from the plurality of compute nodes in the compute region, wherein modifying the set of rules includes ignoring one or more rules from the set of rules to expand the candidate nodes that are eligible for resource allocation.
19. The system of claim 13, wherein causing the plurality of agent allocators to allocate the first type of resources in parallel comprises: Each of the plurality of proxy allocators processes the placement request in parallel and allocates resources on the plurality of computing nodes in the computing region according to the modified first placement strategy.
20. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: Maintain a record of computing resources allocated on multiple computing nodes in a computing region, the computing region including multiple proxy allocators configured to allocate resources in parallel according to a first allocation strategy in response to an incoming allocation request, the first allocation strategy being associated with allocating a first type of resources on computing nodes in the computing region; Based on the information from the records of the allocated computing resources, it is determined that the number of arrangement conflicts among the plurality of agent allocators regarding the allocation of the first type of resources is greater than or equal to a threshold number of arrangement conflicts. Based on the determination that the number of placement conflicts is greater than or equal to the threshold number, the first placement strategy is modified by reducing one or more restrictions of the first placement strategy associated with allocating the first type of resources in the computing region. as well as The plurality of agent allocators allocate the first type of resources in parallel across the plurality of computing nodes according to the modified first deployment strategy.