A resource matching method and device
By finding the resources required for jobs in multi-layer resources, the problem of low resource matching efficiency in the existing technology is solved, and faster and more efficient resource allocation is achieved.
Patent Information
- Application Number
- CN202011551356.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-12-24
AI Technical Summary
In the prior art, when the scheduler allocates resources for a job, it needs to traverse all nodes, resulting in high time complexity and low resource matching efficiency.
By finding the resources required for the job in a multi-layer resource, the specific steps include obtaining the request for the job scheduling resource, determining the resources of each layer based on the attributes, and finding the corresponding resource level in the resource collection, and determining to run the job on the corresponding node.
It improves the efficiency of resource matching, reduces time complexity, and makes resource allocation faster and more efficient.
Smart Images

Figure CN114679448B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computing, and in particular, to a resource matching method and apparatus. Background Art
[0002] A cluster can be understood as a virtual computer cluster formed by connecting a large number of homogeneous or heterogeneous computers through a fast local area network. Clusters are used to provide a model for solving large-scale computing problems. When running an application in a cluster, a scheduler allocates resources for the application (usually called a job). In the prior art, the scheduler allocates resources for a job, and the specific implementation can be as follows: The scheduler selects a batch of nodes from all nodes of the computer cluster, and compares them one by one among these nodes according to the amount or attributes of the resources required by the job to find a suitable node. However, the above resource matching method needs to traverse all nodes, with a high time complexity, resulting in low efficiency of resource matching. Summary of the Invention
[0003] A resource matching method and apparatus provided by embodiments of the present application can reduce the time complexity and improve the efficiency of resource matching.
[0004] In a first aspect, a resource matching method provided by embodiments of the present application is applied to a computer cluster. The computer cluster includes multiple layers of resources, each layer of resources corresponding to an attribute, and each layer of resources including one or more levels of resources, each level of resources corresponding to an attribute value. The method may include: obtaining a request for job scheduling resources; the request includes a first attribute and a second attribute of the resources. Determining a first layer of resources according to the first attribute, and finding a first-level resource corresponding to the first attribute value in the first layer of resources. Determining a second layer of resources according to the second attribute, and determining a resource set associated with the first-level resource in the second layer of resources, and finding a second-level resource corresponding to the second attribute value in the resource set. Determining to run the job on the node corresponding to the second-level resource. In this way, embodiments of the present application can improve the efficiency of resource matching and reduce the time complexity by searching for resources required by a job in multiple layers of resources.
[0005] In a feasible implementation, a resource matching method provided by embodiments of the present application further includes a historical resource list, where the historical resource list includes one or more released resources, and the released resources are resources released after the job runs; before determining the first layer of resources according to the first attribute and finding the first-level resource corresponding to the first attribute value in the first layer of resources, it includes: determining that there is no first-level resource among the one or more released resources in the historical resource list. In embodiments of the present application, when the scheduling node determines that the resources required by the job do not exist in the historical resource list, the scheduling node searches for the resources required by the job in the above hierarchical resource tree, so as to quickly match the resources required by the job.
[0006] In one implementable manner, a resource matching method provided by an embodiment of the present application further includes: when it is determined that there is a first-level resource among one or more released resources in the historical resource list, search for the first-level resource corresponding to the first attribute value among one or more released resources in the historical resource list. In the embodiment of the present application, when the scheduling node determines that there are resources required by a job in the historical resource list, the scheduling node searches for the resources required by the job in the historical resource list, so as to achieve the purpose of quickly matching the resources required by the job.
[0007] In one implementable manner, after determining to run a job on a node corresponding to a second-level resource, it further includes: obtaining information about the released resources. Insert the information about the released resources into the historical resource list. In the embodiment of the present application, the scheduling node stores the information about the released resources in the historical resource list, so that when matching resources subsequently, the resources in the historical resource list can be searched first, further improving the resource matching efficiency.
[0008] In one implementable manner, after inserting the information about the released resources into the historical resource list, it further includes: updating the node information of the node corresponding to the released resources.
[0009] In one implementable manner, in the above embodiment, the multi-level resources include a pointer layer, and the pointer layer includes a first pointer, and the first pointer is used to point to a first node among multiple nodes corresponding to a resource. Determining to run a job on a node corresponding to a second-level resource includes: determining to run a job on a first node corresponding to a second-level resource found according to the first pointer.
[0010] In one implementable manner, the pointer layer further includes a second pointer, and the second pointer is used to point to a second node among multiple nodes corresponding to a resource, and the second node is different from the first node. Specifically, determining to run a job on a node corresponding to a second-level resource is: determining to run a job on a second node corresponding to a second-level resource found according to the second pointer.
[0011] Second aspect, a resource matching device provided by an embodiment of the present application is applied to a computer cluster. The computer cluster includes multiple layers of resources, each layer of resources corresponds to an attribute, each layer of resources includes one or multiple levels of resources, and each level of resources corresponds to an attribute value. The device includes: a first acquisition unit, configured to acquire a request for job scheduling resources; the request includes a first attribute and a second attribute of the resources; a first determination unit, configured to determine the first layer of resources according to the first attribute, and find the first-level resources corresponding to the first attribute value in the first layer of resources; a second determination unit, configured to determine the second layer of resources according to the second attribute, and determine a resource set associated with the first-level resources in the second layer of resources, and find the second-level resources corresponding to the second attribute value in the resource set; a third determination unit, configured to determine to run the job on the node corresponding to the second-level resources. In this way, the embodiment of the present application can improve the efficiency of resource matching and reduce the time complexity by finding the resources required by the job in multiple layers of resources.
[0012] In an implementable manner, the resource matching device provided by the embodiment of the present application further includes a historical resource list, and the historical resource list includes one or more released resources, and the released resources are resources released after the job runs. The device further includes: a fourth determination unit, configured to determine that the first-level resources do not exist in one or more released resources in the historical resource list. In the embodiment of the present application, when the scheduling node determines that the resources required by the job do not exist in the historical resource list, the scheduling node finds the resources required by the job in the above hierarchical resource tree, so as to achieve the purpose of quickly matching the resources required by the job.
[0013] In an implementable manner, the resource configuration device provided by the embodiment of the present application further includes: a search unit, configured to find the first-level resources corresponding to the first attribute value in one or more released resources in the historical resource list when it is determined that the first-level resources exist in one or more released resources in the historical resource list. In the embodiment of the present application, when the scheduling node determines that the resources required by the job exist in the historical resource list, the scheduling node finds the resources required by the job in the historical resource list, so as to achieve the purpose of quickly matching the resources required by the job.
[0014] In an implementable manner, the resource configuration device provided by the embodiment of the present application further includes: a second acquisition unit, configured to acquire information about the released resources. An insertion unit, configured to insert the information about the released resources into the historical resource list. In the embodiment of the present application, the scheduling node stores the information about the released resources in the historical resource list, so that the resources in the historical resource list can be searched first when matching resources subsequently, further improving the resource matching efficiency.
[0015] In an implementable manner, the resource configuration device provided by the embodiment of the present application further includes: an update unit, configured to update the node information of the node corresponding to the released resources.
[0016] In one possible implementation, the multi-layer resources described above include a pointer layer, and the pointer layer includes a first pointer, where the first pointer is used to point to a first node among multiple nodes corresponding to the resources; the third determination unit is further configured to determine to run a job on the first node corresponding to the second-level resources found according to the first pointer.
[0017] In one possible implementation, the pointer layer further includes a second pointer, where the second pointer is used to point to a second node among multiple nodes corresponding to the resources, and the second node is different from the first node; the third determination unit is further configured to determine to run a job on the second node corresponding to the second-level resources found according to the second pointer.
[0018] In a third aspect, there is provided a computer-readable storage medium including computer instructions, which when running on a terminal, cause the terminal to execute the method described in the above aspect and any one of its possible implementation manners.
[0019] In a fourth aspect, there is provided a computer program product, which when running on a computer, causes the computer to execute the method described in the above aspect and any one of its possible implementation manners.
[0020] In a fifth aspect, there is provided a chip system including a processor, which when executing instructions, executes the method described in the above aspect and any one of its possible implementation manners.
[0021] Among them, for the specific implementation manners and corresponding technical effects of each embodiment in the above second aspect to fifth aspect, reference may be made to the specific implementation manners and technical effects of the above first aspect. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic diagram of the architecture of the resource matching system provided by the embodiment of the present application;
[0024] Figure 2 It is a schematic structural diagram of a device provided by the embodiment of the present application;
[0025] Figure 3 It is a schematic flowchart of a resource matching method provided by the embodiment of the present application;
[0026] Figure 4A structural schematic diagram of a hierarchical resource tree provided by an embodiment of the present application;
[0027] Figure 5 A structural schematic diagram of a node link list provided by an embodiment of the present application;
[0028] Figure 6 A flowchart of another resource matching method provided by an embodiment of the present application;
[0029] Figure 7 A structural schematic diagram of a historical resource list provided by an embodiment of the present application;
[0030] Figure 8 A flowchart of another resource matching method provided by an embodiment of the present application;
[0031] Figure 9 An overall data architecture diagram provided by an embodiment of the present application;
[0032] Figure 10 A structural schematic diagram of a resource matching device provided by an embodiment of the present application. Detailed implementation manners
[0033] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0034] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0035] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0036] Figure 1 An architecture diagram of a resource matching system provided by an embodiment of the present application. As Figure 1As shown in the figure, when the cluster runs an application, the terminal 11 can create an application (usually called a job), which is a set of program instances required to complete a specific computing task, and usually corresponds to a set of processes, containers, or other runtime entities on one or more computers. The terminal 11 submits the job to the scheduling node 12, which can be a distributed cluster management system software installed on a computer and is responsible for the management of computing resources and the scheduling of batch jobs; of course, the scheduling node 12 can also be a hardware device, such as a scheduler. The scheduling node 12 determines the target node in the computing nodes 13 according to the resource information required by the job, and determines to run the above job on the target node. Among them, the computing nodes 13 can be understood as multiple computers in the computer cluster.
[0037] In the existing resource matching method, the following matching methods may exist: The scheduling node selects a batch of nodes from all the computing nodes in the computer cluster, and makes a one-by-one comparison among these nodes according to the resource amount or attributes required by the job to find a suitable node. However, the above resource matching method needs to traverse all nodes, and the time complexity is o(n), resulting in low resource matching efficiency.
[0038] Therefore, to solve the above technical problems, in the embodiments of the present application, a resource matching method is proposed. This method is applied to a computer cluster, which includes multiple layers of resources, each layer of resources corresponds to an attribute, each layer of resources includes one or more levels of resources, and each level of resources corresponds to an attribute value. This method obtains a request for job scheduling resources, and the request includes the first attribute and the second attribute of the resources. Determine the first layer of resources according to the first attribute, and find the first-level resources corresponding to the first attribute value in the first layer of resources. Determine the second layer of resources according to the second attribute, and determine the resource set associated with the first-level resources in the second layer of resources, and find the second-level resources corresponding to the second attribute value in the resource combination. Determining to run the job on the node corresponding to the second-level resources can improve the efficiency of resource matching and reduce the time complexity.
[0039] The following describes the scheduling node 12, the computing node 13, and the storage node 14 in the resource matching system shown in Figure 1 the accompanying drawings in the embodiments of the present application.
[0040] Figure 2Schematic diagram of the composition of a device 200 provided by an embodiment of the present application. The device 200 may include a processor 201 and a memory 204. Further, the device 200 may also include a communication line 202 and a communication interface 203. Among them, the processor 201, the memory 204, and the communication interface 203 may be connected through the communication line 202. Among them, the device 200 may be a scheduling node 12. The device 200 may also be a computing node 13. The device 200 may also be a storage node 14. Of course, the memory 204 in the device 200 may be the storage node 14.
[0041] The processor 201 may be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 201 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.
[0042] The communication line 202 is used to transmit information between the components included in the communication device 300.
[0043] The communication interface 203 is used to communicate with other devices or other communication networks. The other communication network may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 203 may be a module, a circuit, a transceiver, or any device capable of implementing communication.
[0044] The memory 204 is used to store instructions. Among them, the instructions may be computer programs.
[0045] Among them, the memory 204 can be a read-only memory (ROM) or other types of static storage devices that can store static information and / or instructions, or it can be a random access memory (RAM) or other types of dynamic storage devices that can store information and / or instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs), magnetic disk storage media, and other magnetic storage devices, without limitation.
[0046] It should be noted that the memory 204 can exist independently of the processor 201 or be integrated with the processor 201. The memory 204 can be used to store instructions, program codes, or some data, etc. The memory 204 can be located inside the device 200 or outside the device 200, without limitation.
[0047] The processor 201 is used to execute the instructions stored in the memory 204 to implement the service switching method provided in the following embodiments of this application. For example, when the device 200 is a chip or a system-on-chip in a network device, the processor 201 executes the instructions stored in the memory 204 to implement the steps performed by the network device in the following embodiments of this application.
[0048] In one example, the processor 201 may include one or more CPUs, such as Figure 2 CPU0 and CPU1 in
[0049] As an alternative implementation, the device 200 includes multiple processors. For example, in addition to Figure 2 the processor 201 in
[0050] As an alternative implementation, the device 200 further includes an output device 205 and an input device 206. Exemplarily, the input device 206 is a device such as a keyboard, a mouse, a microphone, or a joystick, and the output device 205 is a device such as a display screen or a speaker.
[0051] It should be noted that the device 200 can be a desktop computer, a laptop computer, a network server, a mobile phone, a tablet computer, a wireless terminal, an embedded device, a chip system, or a device with a Figure 2 similar structure in Figure 2The constituent structure shown does not constitute a limitation on the communication device. Except Figure 2 for the components shown, the communication device may include more or fewer components than those shown, or combine certain components, or have different component arrangements.
[0052] In the embodiments of the present application, the chip system may be composed of chips or may include chips and other discrete devices.
[0053] In addition, actions, terms, etc. involved among the embodiments of the present application can be referred to each other without limitation. The message names or parameter names in the messages exchanged between devices in the embodiments of the present application are only examples, and other names can also be used in specific implementations without limitation.
[0054] Next, taking Figure 2 the architecture shown as an example, the resource matching method provided by the embodiments of the present application will be described. Each network element in the following embodiments may have Figure 2 the components shown and will not be elaborated. It should be noted that the message names or parameter names in the messages exchanged between devices in the embodiments of the present application are only examples, and other names can also be used in specific implementations. The determination in the embodiments of the present application can also be understood as create or generate, and "include" in the embodiments of the present application can also be understood as "carry". This is uniformly stated here, and the embodiments of the present application do not make specific limitations on this.
[0055] Next, in combination with the accompanying drawings in the embodiments of the present application, the resource matching method provided by the embodiments of the present application will be described.
[0056] A resource matching method provided by the embodiments of the present application can be applied to Figure 1 the resource matching system shown. Figure 1 The scheduling node 12, computing node 13, and storage node 14 shown in may each be one or more computers, and these computers can form a computer cluster. When running a job in this computer cluster, the terminal 11 submits the job to the scheduling node 12, and the scheduling node 12 determines a target node in the computing node 13 according to the resource information required by the job and determines to run the above job on the target node. The realizable manner for the scheduling node 13 to match resources for the job is as follows:
[0057] Figure 3 is a schematic flowchart of a resource matching method provided by the embodiments of the present application. As Figure 3 shown, the method may include:
[0058] S301. The scheduling node reads the configuration information
[0059] The configuration information may be sent by the terminal to the scheduling node or may be stored by the scheduling node in advance, and no specific limitation is made here. The configuration information may include: information on layer resources, information on level resources, the association relationships between layer resources, the association relationships between layer resources and level resources, and so on.
[0060] S302. The scheduling node generates a hierarchical resource tree according to the configuration information
[0061] The scheduling node constructs a hierarchical resource tree from the multi-layer resources included in the computer cluster according to the configuration information. Among them, each layer of resources corresponds to an attribute, each layer of resources includes one or more levels of resources, and each level of resources corresponds to an attribute value.
[0062] Exemplarily, the multi-layer resources include a first layer of resources and a second layer of resources. The first layer of resources may correspond to a first attribute, which may be represented as resource1. The first layer of resources may include a first-level resource and a second-level resource. The first-level resource may correspond to a first attribute value, which may be represented as bucket1. The second-level resource may correspond to a second attribute value, which may be represented as bucket2. The second layer of resources may correspond to a second attribute, which may be represented as resource2. The second layer of resources may include a first-level resource, a second-level resource, and a third-level resource. The first-level resource may correspond to a first attribute value, which may be represented as bucket1. The second-level resource may correspond to a second attribute value, which may be represented as bucket2. The third-level resource may correspond to a third attribute value, which may be represented as bucket3.
[0063] The expression of the hierarchical resource tree may be:
[0064] resource1:
[0065] -bucket1
[0066] -bucket2
[0067] resource2:
[0068] -bucket1
[0069] -bucket2
[0070] -bucket3
[0071] The multi-layer resource may further include a pointer layer, which may include a first pointer for pointing to a first node to which the resources required by a job belong. The first node can be understood as the first node suitable for allocating the resources required by the job. Of course, the pointer layer may further include a second pointer for pointing to a second node to which the resources required by the job belong. The second node can be understood as a node suitable for allocating the resources required by the job other than the first node. The second node is different from the first node. Therefore, after resource matching is completed, the second pointer points to the next node to which the resources required by the job belong to achieve load balancing. The pointer layer may further include a quantity pointer for recording the number of nodes falling into the resource tree at this level. Therefore, the current pointer can quickly determine whether there are sufficient resources for allocation.
[0072] Of course, the multi-layer resource may include a third-layer resource. The multi-layer resource may further include a fourth-layer resource. In the embodiments of the present application, the specific number of layers of the multi-layer resource is not limited, and it needs to be selected according to actual requirements during specific implementation.
[0073] Among them, there is an association relationship between the resources of each layer in the multi-layer resource. This association relationship can be artificially defined, randomly formed, or established according to the attributes of each layer of resources, or formed by arranging in a specified order (for example, arranged in ascending order of serial number).
[0074] Example 1 Figure 4 is a schematic structural diagram of a hierarchical resource tree provided by an embodiment of the present application. As Figure 4 shown, assume that the first-layer resource is a central processing unit (CPU), and the first attribute corresponding to the first-layer resource is the CPU architecture. The architecture of this CPU may include the X86 architecture (the first attribute value corresponding to the first-level resource) and the Aarch64 architecture (the second attribute value corresponding to the second-level resource). Therefore, the expression of the first-layer resource (Res1) can be:
[0075] CPU Type:
[0076] -X86
[0077] -aarch64
[0078] The second-layer resource is the number of CPU cores, and the second attribute corresponding to the second-layer resource is the available number of CPU cores (FREE_CPU). The available number of CPU cores may include 0 (the first attribute value corresponding to the first-level resource), 8 (the second attribute value corresponding to the second-level resource), and 16 (the third attribute value corresponding to the third-level resource). Therefore, the expression of the second-layer resource (Res2) can be:
[0079] FREE_CPU: 0 8 16
[0083] The third - layer resource is memory, and the third attribute corresponding to the third - layer resource can be available memory (FREE_MEM). This available memory can include 0 (the first - attribute value corresponding to the first - level resource), 16 (the second - attribute value corresponding to the second - level resource), and 32 (the third - attribute value corresponding to the third - level resource). Therefore, the expression for the third - layer resource (Res3) can be:
[0084] FREE_MEM: 0 16 32
[0088] The pointer layer can include a first pointer (head pointer or header pointer), a second pointer (current pointer or current - position pointer), and a quantity pointer (#nodes). Among them, the first pointer can be the first - level resource, the second pointer can be the second - level resource, and the third pointer can be the third - level resource. Therefore, the expression for the pointer layer (node head) can be:
[0089] node head:
[0090] head
[0091] current
[0092] #nodes
[0093] Among them, the association relationships between the first - layer resource, the second - layer resource, the third - layer resource, and the pointer layer can vary according to the different arrangement orders of the first - layer resource, the second - layer resource, and the third - layer resource. No specific limitation is made here, and it can be set according to actual requirements during specific implementation.
[0094] For example, as Figure 4 shown, the first - level resource in the first - layer resource is associated with the resource set composed of the first - level resource, the second - level resource, and the third - level resource in the second - layer resource. The second - level resource in the first - layer resource is associated with the resource set composed of the first - level resource, the second - level resource, and the third - level resource in the second - layer resource. The first - level resource in the second - layer resource is associated with the resource set composed of the first - level resource, the second - level resource, and the third - level resource in the third - layer resource. The first - level resource in the third - layer resource is associated with the resource set composed of the first - level resource, the second - level resource, and the third - level resource in the fourth - layer resource.
[0095] In summary, through the organization of the first - level resource, the second - level resource, the third - level resource, and the pointer layer in the embodiments of the present application, it is generated as Figure 4A hierarchical resource tree as shown, which can be similar to B + tree. It should be explained here that B + tree is a tree data structure, and its characteristic is that it can keep the data stable and ordered. Therefore, in the embodiments of the present application, the generation of the hierarchical resource tree can realize the pre-grouping of resources, and can solve the problem of resource fragmentation to a certain extent, so as to facilitate the rapid selection of resources.
[0096] S303. The scheduling node initializes the hierarchical resource tree so that each resource finds its belonging node
[0097] Specifically, the scheduling node initializes the above-mentioned hierarchical resource tree so that each resource finds its belonging node, and forms a node link list with the nodes to which multiple resources belong. This node link list is maintained by a doubly linked list between nodes. This doubly linked list is a type of linked list, and each of its data has two pointers, which respectively point to the direct successor and the direct predecessor.
[0098] That is to say, the computer cluster is grouped according to the hierarchical resource tree to obtain a node link list. The node table in this node link list is a hash table. A hash table is a data structure that can be directly accessed according to the key value. The key of the hash table is the identification information of the node, such as name, IP address, etc. The value of the hash table can be node information, such as the available number of cores (free_core), available memory (free_mem), available disk (free_disk), next pointer (next), and previous pointer (pre), etc.
[0099] Figure 5 is a schematic structural diagram of the node link list provided by the embodiments of the present application. As Figure 5 shown, the node table (nodetable) includes node 1 (node 1), node 2 (node 2), node 3 (node 3), and so on. Node 1 (node 1), node 2 (node 2), and node 3 (node 3) can all include the available number of cores (free_core), available memory (free_mem), available disk (free_disk), next pointer (next), and previous pointer (pre), etc.
[0100] Therefore, the node link list in the embodiments of the present application adopts a hash table and a doubly linked list to realize the maintenance of node information, so as to reduce the time complexity of node information modification, node information addition, and node information removal, and can make the time complexity close to O(1).
[0101] In summary, the hierarchical resource tree can be pre-generated or generated in real time during the operation of computer cluster jobs. The embodiments of the present application do not make specific limitations in this regard.
[0102] It should be added here that the resource matching method provided by the embodiments of the present application may further include: when the attributes of the newly added resources do not meet the existing attributes in the hierarchical resource tree, updating the hierarchical resource tree.
[0103] It should be understood that when the newly added resources can cause changes to the layer resources (or level resources) of the hierarchical resource tree, the hierarchical resource tree is updated. Exemplarily, assume that the existing attributes are the available number of cores and the available storage capacity. The attribute of the newly added resource is a disk, and the available disk does not exist in the existing attributes. Therefore, the hierarchical resource tree is updated so that the new hierarchical resource tree includes the layer resources (or level resources) of the available disk.
[0104] Figure 6 It is a schematic flowchart of another resource matching method provided by the embodiments of the present application. As Figure 6 shown, after the hierarchical resource tree is generated at the scheduling node, the resource matching method provided by the embodiments of the present application may further include:
[0105] S601. The scheduling node obtains a request for job scheduling resources.
[0106] The scheduling node can be a distributed cluster management system software installed on one or more computers in the computer cluster and is used to be responsible for the management of computing resources and the scheduling of batch jobs.
[0107] The resources may include hardware resources such as a central processing unit (CPU), memory, network, and network storage, etc. The resources may also include other non-hardware resources such as software resources, and the software resources may include available network services, software licenses, etc.
[0108] The request may include a simple resource request and a complex resource request. Among them, the simple resource request may have a fixed request type for resources and / or be enumerable, etc. The complex resource request may have a dynamically changing request type for resources and may involve custom resource types.
[0109] Since each layer of resources corresponds to an attribute, the request includes the first attribute and the second attribute of the resources, and it can be understood that the request includes the first attribute of the first layer of resources and the second attribute of the second layer of resources.
[0110] Exemplarily, assume that the first-layer resource is the CPU. Correspondingly, the first attribute can be the type of the CPU. The second-layer resource is the memory. Correspondingly, the second attribute can be the available memory. Of course, for different layer resources, the corresponding attributes are also different, which will not be listed one by one here.
[0111] This step can be specifically implemented as follows. Figure 1 As shown, the terminal 11 sends a request for job scheduling resources to the scheduling node 12. The scheduling node 12 obtains the request for job scheduling resources.
[0112] S602. The scheduling node determines the first-layer resources according to the first attribute, and searches for the first-level resources corresponding to the first attribute value in the first-layer resources.
[0113] The attribute value can refer to the quantifiable value of the attribute. Exemplarily, continuing with the above example, if the attribute is the type of the CPU, correspondingly, the attribute value can be X86 or aarch64. If the attribute is the available memory, correspondingly, the attribute value can be 32G, 64G, or 128G.
[0114] Example 1. As Figure 1 shown, the terminal 11 sends a request for job scheduling resources to the scheduling node 12. The request can include the type of the CPU. The scheduling node 12 determines that the first-layer resource is the CPU according to the type of the CPU, and searches for the first-level resources with the first attribute value of X86 in the layer where the CPU is located.
[0115] Example 2. As Figure 1 shown, the terminal 11 sends a request for job scheduling resources to the scheduling node 12. The request can include the available memory. The scheduling node 12 determines that the first-layer resource is the memory according to the available memory, and searches for the first-level resources with the first attribute value of 32G in the layer where the memory is located.
[0116] S603. The scheduling node determines the second-layer resources according to the second attribute, determines the resource set associated with the first-level resources in the second-layer resources, and searches for the second-level resources corresponding to the second attribute value in the resource set.
[0117] Continuing with the above Example 1, the request for job scheduling resources can also include the available number of CPU cores. The scheduling node 12 determines that the second-layer resource is the number of CPU cores according to the available number of CPU cores, and searches for the second-level resources with the CPU type of X86 and the available number of CPU cores less than 16 in the layer where the number of CPU cores is located.
[0118] Continuing with the above Example 2, the request for job scheduling resources can also include the available number of CPU cores. The scheduling node 12 determines that the second-layer resource is the number of CPU cores according to the available number of CPU cores, and searches for the second-level resources with the available memory of 32G and the available number of CPU cores less than 16 in the layer where the number of CPU cores is located.
[0119] S604. The scheduling node determines to run a job on the node corresponding to the second-level resource.
[0120] This step can be specifically implemented as follows:
[0121] If the multi-level resources can include a pointer layer, and the pointer layer can include a first pointer that is used to point to a first node among the multiple nodes corresponding to the resource, then S604 can be specifically implemented as: the scheduling node determines to run a job on the first node corresponding to the second-level resource found according to the first pointer.
[0122] If the pointer layer further includes a second pointer that is used to point to a second node among the multiple nodes corresponding to the resource, and the second node is different from the first node, then S604 can be specifically implemented as: the scheduling node determines to run a job on the second node corresponding to the second-level resource found according to the second pointer.
[0123] That is to say, if the first node corresponding to the second-level resource pointed to by the first pointer is occupied, the scheduling node determines to run a job on the second node corresponding to the second-level resource pointed to by the second pointer. Therefore, while ensuring the rapid allocation of resources, each node that meets the requirements can be used to achieve the purpose of load balancing and effectively alleviate the problem of load skew.
[0124] In some embodiments, the resource matching method provided by the embodiments of the present application may further include a historical resource list, and the historical resource list may include one or more released resources, and the released resources are the resources released after the job runs.
[0125] Figure 7 It is a schematic structural diagram of the historical resource list provided by the embodiments of the present application, as Figure 7As shown, the historical resource list (free list) is a hash table, whose key value (Res_key1) is the resource requirement of the job, and the value of this hash table is a series of available resource information. The historical resource list includes a first pointer layer (RES1) and a second pointer layer (RES2). Both the first pointer layer and the second pointer layer can include a next pointer, which is used to point to the node to which the next available resource belongs. Both the first pointer layer and the second pointer layer can also include a Node_ptr, which is used to point to the node (node) to which this resource block belongs. Both the first pointer layer and the second pointer layer can also include Free_res 1, Free_res 2... Free_res n. Among them, Free_res 1 indicates that res1 in the first pointer layer is idle. Free_res 2 indicates that res2 in the second pointer layer is idle. Free_res n indicates that res n in the nth pointer layer is idle. Therefore, the historical resource list in the embodiments of the present application is essentially to store the historically used resources, form a hash table with the historically used resources, and it is a single-layer data structure rather than a multi-layer data structure, which can achieve precise matching of the historically used resources.
[0126] In summary, the historical resource list is a single-layer data structure, and the hierarchical resource tree is a multi-layer data structure. Therefore, in order to facilitate finding the resources required by the job more quickly, before executing S602, the resource matching method provided by the embodiments of the present application may further include:
[0127] S605. The scheduling node determines whether there is a first-level resource among one or more released resources in the historical resource list. If not, execute S607; if so, execute S606.
[0128] It should be understood that the scheduling node first determines whether there are resources required by the job in the historical resource list. If not, the scheduling node searches for the resources required by the job in the hierarchical resource tree; if so, the scheduling node searches for the resources required by the job in the historical resource list.
[0129] S606. When the scheduling node determines that there is a first-level resource among one or more released resources in the historical resource list, it searches for the first-level resource corresponding to the first attribute value among one or more released resources in the historical resource list.
[0130] It should be understood that the historical resource list stores the complete information of the resources required by the job. For example, the resources required by the job are X86, the available number of cores is 10, and the available memory is 20G. Then, the historical resource list stores the resource information as X86, the available number of cores is 10, the available memory is 20G, and the node information to which this resource information belongs.
[0131] S607. The scheduling node determines whether there is a first-level resource among the multiple resources of the hierarchical resource tree. If there is, S602 is executed; if not, S608 is executed.
[0132] S608. Resource matching fails.
[0133] Optionally, after executing S604, the resource matching method provided by the embodiments of the present application may further include: S609. The scheduling node updates the information of the node corresponding to the second resource.
[0134] In the embodiments of the present application, when the scheduling node determines that the resources required by the job do not exist in the historical resource list, the scheduling node searches for the resources required by the job in the above hierarchical resource tree. When the scheduling node determines that the resources required by the job exist in the historical resource list, the scheduling node searches for the resources required by the job in the historical resource list, so as to achieve the purpose of quickly matching the resources required by the job.
[0135] After the computer cluster runs a job, the computer cluster needs to release resources. The specific operations are as follows:
[0136] Figure 8 It is a schematic flowchart of another resource matching method provided by the embodiments of the present application. As Figure 8 shown, the method may further include:
[0137] S801. The scheduling node sends a resource release request.
[0138] The scheduling node sends a resource release request to the computer cluster. The computer cluster releases resources according to the resource release request and feeds back the information of the released resources to the scheduling node.
[0139] S802. The scheduling node obtains the information of the released resources.
[0140] S803. The scheduling node inserts the information of the released resources into the historical resource list.
[0141] This step may be specifically implemented as: S8031. The scheduling node searches the historical resource list according to the information of the released resources to check if there are released resources. If so, S8032 is executed; if not, S8033 is executed. S8032. The scheduling node inserts the information of the released resources into the head of the historical resource list. S8033. The scheduling node inserts the information of the released resources into the historical resource list.
[0142] Preferably, after executing S803, it further includes:
[0143] S804. The scheduling node updates the node information of the node corresponding to the released resources.
[0144] It should be understood that the scheduling node updates the information of the node to which the released resources belong in the historical resource list. That is to say, in the historical resource list, when the scheduling node updates the information of the released resources, it also updates the information of the node to which the released resources belong, so that the node to which the released resources belong can be found when the released resources are found.
[0145] In the embodiment of the present application, the scheduling node stores the information of the released resources in the historical resource list, so that when matching resources subsequently, the resources in the historical resource list can be searched first, further improving the resource matching efficiency.
[0146] In practical applications, Figure 9 FIG. is a data overall architecture diagram provided by the embodiment of the present application. Figure 9 It is a combined diagram of the above-mentioned hierarchy resource bucketing tree, free list, and node link list. As Figure 9 shown, assume that the hierarchy resource tree includes first-level resources and second-level resources. Among them,
[0147] The first-level resources are the number of cores, and the first attribute corresponding to the first-level resources is the free core. The free core can include 0 (the first attribute value corresponding to the first-level resources), 8 (the second attribute value corresponding to the second-level resources), 16 (the third attribute value corresponding to the third-level resources), 32 (the fourth attribute value corresponding to the fourth-level resources), and 64 (the fifth attribute value corresponding to the fifth-level resources). Therefore, the expression of the first-level resources (free core) can be:
[0148] Free Core: 0 8 16 32 64
[0154] The second-level resources are memory, and the second attribute corresponding to the second-level resources is Free Mem. The free core can include 0 (the first attribute value corresponding to the first-level resources), 16 (the second attribute value corresponding to the second-level resources), 32 (the third attribute value corresponding to the third-level resources), 64 (the fourth attribute value corresponding to the fourth-level resources), and 128 (the fifth attribute value corresponding to the fifth-level resources). Therefore, the expression of the second-level resources (Free Mem) can be:
[0155] Free Mem: 0 16 32 64 128
[0161] The node link list includes node (node 1), node 2 (node 2), node 3 (node 3), node N (nodeN), and node K (node K). Each node may include: the number of available cores (free_core), available memory (free_mem), available disk (free_disk), next pointer (next), and previous pointer (pre), etc.
[0162] The historical resource list (free list) includes key values res_key1 and res_key2. The key value res_key1 corresponds to the first released resource, and the key value res_key2 corresponds to the second released resource. The historical resource list may also include a next pointer, which is used to point to the node to which the next available resource belongs. The historical resource list may also include a Node_ptr, which is used to point to the node (node) to which this resource block belongs.
[0163] Taking the association between the fourth-level resource in the first-level resource and the resource set in the second-level resource in the hierarchical resource tree as an example, when the computer cluster runs a job, the terminal submits the job to the scheduling node, and the scheduling node determines the target node in the computer cluster according to the resources required by the job.
[0164] Suppose the resources required by the job are 50 available cores and 64G of available memory. The resources required by the job exist in the historical resource list. For example, the key value res_key1 in the historical resource list corresponds to the resources required by the job. The scheduling node receives a request for job scheduling resources. The scheduling node determines that the resources required by the job exist in the historical resource list. The scheduling node locates the node to which the resource belongs in the historical resource list, which is node N. The scheduling node determines to run the above job on node N.
[0165] Suppose again that the resources required by the job are 50 available cores and 128G of available memory. The resources required by the job do not exist in the historical resource list. The scheduling node receives a request for job scheduling resources. The scheduling node determines that the resources required by the job do not exist in the historical resource list. The scheduling node locates the node to which the resource belongs in the hierarchical resource tree, which is node 1. The scheduling node determines to run the above job on node 1.
[0166] In the embodiments of the present application, by searching the hierarchical resource tree or the historical resource list, the resources matching the job are found, and the problem of repeated failures in the application operation due to hardware problems frequently occurring in the HPC (High Performance Computing Cluster) scenario is solved. The embodiments of the present application try to actively and timely discover the fault points and actively avoid them by continuously analyzing the operation results, so as to avoid the current mode of only passive discovery and manual intervention. Furthermore, the operation time of the application is reduced, and the utilization rate of the HPC cluster is improved.
[0167] Of course, in addition to being applied to the HPC scenario, the embodiments of the present application can also be applied to the public / private cloud virtualization scenario. The requirements for resources in public / private cloud virtualization are relatively simple and enumerable, such as 8U64G, 16U128G. The resources applied for are applied and recycled as a whole, and resource matching can be achieved more efficiently. For the public / private cloud virtualization scenario, the specific principle is as described above and will not be elaborated here.
[0168] Figure 10 FIG. is a schematic structural diagram of a resource matching device provided by an embodiment of the present application. The resource matching device is applied to a computer cluster. The computer cluster includes multiple layers of resources. Each layer of resources corresponds to an attribute. Each layer of resources includes one or more levels of resources. Each level of resources corresponds to an attribute value. The device 1000 may include:
[0169] A first obtaining unit 1001, configured to obtain a request for job scheduling resources; the request includes a first attribute and a second attribute of the resources;
[0170] A first determining unit 1002, configured to determine the first layer of resources according to the first attribute, and search for the first-level resources corresponding to the first attribute value in the first layer of resources;
[0171] A second determining unit 1003, configured to determine the second layer of resources according to the second attribute, and determine a resource set associated with the first-level resources in the second layer of resources, and search for the second-level resources corresponding to the second attribute value in the resource set;
[0172] A third determining unit 1004, configured to determine to run the job on the node corresponding to the second-level resources.
[0173] In some embodiments, the device 1000 may further include a historical resource list, and the historical resource list includes one or more released resources, and the released resources are resources released after the job runs;
[0174] The device 1000 may further include:
[0175] A fourth determining unit 1005, configured to determine that the first-level resources do not exist in one or more released resources in the historical resource list.
[0176] In some embodiments, the device 1000 further includes:
[0177] A search unit 1006, configured to search for a first-level resource corresponding to a first attribute value among one or more released resources in a historical resource list when it is determined that there is a first-level resource among the one or more released resources in the historical resource list.
[0178] In some embodiments, the apparatus 1000 may further include:
[0179] A second acquisition unit 1007, configured to acquire information of a released resource;
[0180] An insertion unit 1008, configured to insert the information of the released resource into the historical resource list.
[0181] In some embodiments, the apparatus 1000 may further include:
[0182] An update unit 1009, configured to update node information of a node corresponding to a released resource.
[0183] In some embodiments, the apparatus may further include a pointer layer, and the pointer layer includes a first pointer, where the first pointer is used to point to a first node among a plurality of nodes corresponding to a resource;
[0184] The third determination unit 1004 is further configured to determine to run a job on a first node corresponding to a second-level resource found according to the first pointer.
[0185] In some embodiments, the pointer layer further includes a second pointer, where the second pointer is used to point to a second node among a plurality of nodes corresponding to a resource, and the second node is different from the first node;
[0186] The third determination unit 1004 is further configured to determine to run a job on a second node corresponding to a second-level resource found according to the second pointer.
[0187] Optionally, in this possible design, all relevant contents of each step of the electronic device involved in the foregoing Figures 1 to 9 method embodiments may be cited to the function descriptions of the corresponding functional modules, and will not be elaborated herein. The electronic device described in this possible design is used to execute Figures 1 to 9 the functions of the electronic device in the real-time communication method shown, and thus can achieve the same effect as the foregoing resource matching method.
[0188] An embodiment of the present application further provides a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected through a line. For example, the interface circuit can be used to receive signals from other devices (such as a memory). For another example, the interface circuit can be used to send signals to other devices (such as a processor). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can execute each step in the above embodiment. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations thereon.
[0189] An embodiment of the present application further provides a device, which is included in an electronic device and has a function of implementing the behavior of the electronic device in any of the above methods. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, a detection module or unit, a determination module or unit, etc.
[0190] An embodiment of the present application further provides a computer-readable storage medium, including computer instructions, which cause the electronic device to execute any of the above methods when the computer instructions run on the electronic device.
[0191] An embodiment of the present application further provides a computer program product, which causes the computer to execute any of the above methods when the computer program product runs on the computer.
[0192] It can be understood that, in order to implement the above functions, the above terminal, etc. includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraint conditions of the technical solution. Those skilled in the art can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.
[0193] The embodiments of the present application can perform functional module division on the above terminal, etc. according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present invention is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0194] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above function modules is used as an example. In actual applications, the above functions can be allocated to different function modules as needed, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. For the specific working processes of the system, device, and unit described above, reference can be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0195] In each embodiment of this application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0196] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of this application. The foregoing storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc.
[0197] The above is only the specific embodiment of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
Claims
1. A resource matching method, characterized in that, it is applied to a computer cluster, the computer cluster includes multiple layers of resources, each layer of resources corresponds to an attribute, each layer of resources includes one or more levels of resources, and each level of resources corresponds to an attribute value. The method includes: Obtaining a request for job scheduling resources; the request includes a first attribute and a second attribute of the resources; Determining the first layer of resources according to the first attribute, and searching for the first-level resources corresponding to the first attribute value in the first layer of resources; Determining the second layer of resources according to the second attribute, and determining a resource set associated with the first-level resources in the second layer of resources, and searching for the second-level resources corresponding to the second attribute value in the resource set; Determining to run the job on the node corresponding to the second-level resources; The multiple layers of resources include a pointer layer, and the pointer layer includes a first pointer, and the first pointer is used to point to the first node among the multiple nodes corresponding to the resources; Determining to run the job on the node corresponding to the second-level resources includes: Determining to run the job on the first node corresponding to the second-level resources found according to the first pointer; The pointer layer further includes a second pointer, and the second pointer is used to point to the second node among the multiple nodes corresponding to the resources, and the second node is different from the first node; Determining to run the job on the node corresponding to the second-level resources includes: Determining to run the job on the second node corresponding to the second-level resources found according to the second pointer; The computer cluster further includes a historical resource list, and the historical resource list includes one or more released resources, and the released resources are the resources released after the job runs; Before determining the first layer of resources according to the first attribute and searching for the first-level resources corresponding to the first attribute value in the first layer of resources, it includes: Determining that the first-level resources do not exist in one or more of the released resources in the historical resource list.
2. The method according to claim 1, characterized in that, it further includes: When it is determined that the first-level resources exist in one or more of the released resources in the historical resource list, searching for the first-level resources corresponding to the first attribute value in one or more of the released resources in the historical resource list.
3. The method according to claim 1 or 2, characterized in that, After determining to run the job on the node corresponding to the second-level resources, it further includes: Obtaining information of the released resources; Inserting the information of the released resources into the historical resource list.
4. The method according to claim 3, characterized in that, After inserting the information of the released resources into the historical resource list, it further includes: Updating the node information of the node corresponding to the released resources.
5. A resource matching device, characterized in that, it is applied to a computer cluster, the computer cluster includes multiple layers of resources, each layer of resources corresponds to an attribute, each layer of resources includes one or more levels of resources, and each level of resources corresponds to an attribute value. The device includes: A first acquisition unit, configured to acquire a request for job scheduling resources; the request includes a first attribute and a second attribute of the resources; A first determination unit, configured to determine a first layer of resources according to the first attribute, and search for a first-level resource corresponding to the first attribute value among the first layer of resources; A second determination unit, configured to determine a second layer of resources according to the second attribute, determine a resource set associated with the first-level resource in the second layer of resources, and search for a second-level resource corresponding to the second attribute value in the resource set; A third determination unit, configured to determine to run the job on a node corresponding to the second-level resource; The multi-layer resources include a pointer layer, and the pointer layer includes a first pointer, and the first pointer is used to point to a first node among multiple nodes corresponding to the resource; The third determination unit is further configured to determine to run the job on the first node corresponding to the second-level resource found according to the first pointer; The pointer layer further includes a second pointer, and the second pointer is used to point to a second node among multiple nodes corresponding to the resource, and the second node is different from the first node; The third determination unit is further configured to determine to run the job on the second node corresponding to the second-level resource found according to the second pointer; The computer cluster further includes a historical resource list, and the historical resource list includes one or more released resources, and the released resources are resources released after the job runs; The apparatus further includes: A fourth determination unit, configured to determine that the first-level resource does not exist in one or more of the released resources in the historical resource list.
6. The apparatus according to claim 5, wherein, it further includes: A search unit, configured to search for a first-level resource corresponding to the first attribute value in one or more of the released resources in the historical resource list when it is determined that the first-level resource exists in one or more of the released resources in the historical resource list.
7. The apparatus according to claim 5 or 6, wherein, it further includes: A second acquisition unit, configured to acquire information of the released resources; An insertion unit, configured to insert the information of the released resources into the historical resource list.
8. The apparatus according to claim 7, wherein, it further includes: An update unit, configured to update the node information of the node corresponding to the released resource.
9. A computer-readable storage medium, wherein, the computer-readable storage medium includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the resource matching method according to any one of claims 1-4.
10. A computer program product, wherein, when the computer program product runs on a computer, the computer is caused to execute the resource matching method according to any one of claims 1 to 4.
11. A chip system, wherein, it includes one or more processors, and when the one or more processors execute instructions, the one or more processors execute the resource matching method according to any one of claims 1-4.