Node resource processing method and device, equipment, medium and program product

By deploying proxy components and scheduling components in the NUMA architecture cloud computing platform, obtaining node resource topology information and analyzing and processing it, the problem of unreasonable node resource scheduling is solved, and reasonable resource allocation and improved scheduling effect are achieved.

CN120675955APending Publication Date: 2025-09-19SHENZHEN TENCENT COMP SYST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410316379.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In cloud computing platforms with NUMA architecture, existing solutions fail to effectively perceive the NUMA topology of nodes, resulting in pod resources not being able to run normally after being scheduled to the target node, and the scheduling effect does not meet expectations.

Method used

By deploying proxy components and scheduling components in the cloud computing platform, node resource topology information is obtained, resource analysis and processing are performed based on the topology information to determine the target node, and resource annotation information is generated to bind resource nodes to achieve reasonable resource allocation and scheduling.

Benefits of technology

It improves the rationality and accuracy of resource scheduling, ensures that resource allocation meets user needs, and improves resource scheduling effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675955A_ABST
    Figure CN120675955A_ABST
Patent Text Reader

Abstract

The invention provides a node resource processing method and apparatus, a device, a medium and a program product, which are applied to a cloud computing platform architecture comprising N nodes, one node is deployed with an agent component, and one node comprises one or more resource nodes. The method comprises the following steps: when a resource scheduling request is received, acquiring a node resource topology information set; based on the resource scheduling request and the node resource topology information set, performing resource analysis processing on the N nodes, and determining a target node; resources in the target node are pre-allocated based on the resource scheduling request, resource annotation information is generated based on a resource allocation result of the target node, and the resource annotation information is used for indicating that the allocated resource node is bound with the resource object in the target node. According to the method, the node resources can be allocated and scheduled according to the node resource topological structure and the resource scheduling request, so that the node resources can be reasonably allocated, the scheduling result meets user requirements, and the resource scheduling effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a node resource processing method, a node resource processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In cloud computing platform architectures such as NUMA (Non-Uniform Memory Access), the primary concern is allocating and scheduling pod resources. Existing solutions typically schedule pod resources directly on a node-by-node basis. However, during resource scheduling, because the scheduling component is unaware of the node's NUMA topology, the target node selected by the scheduling component may not meet the pod's required NUMA topology (Topology Affinity Error). Consequently, after the pod is scheduled to the target node, it may not run properly, resulting in unsatisfactory scheduling results. Summary of the Invention

[0003] The embodiments of the present application propose a node resource processing method, device, equipment, medium and program product, which can allocate and schedule node resources according to the node resource topology structure and resource scheduling requests, so that node resources can be reasonably allocated and the scheduling results meet user needs, thereby improving the resource scheduling effect.

[0004] In one aspect, an embodiment of the present application provides a node resource processing method, which is applied to a cloud computing platform architecture. The cloud computing platform architecture includes N nodes, wherein an agent component is deployed in each node, and each node includes one or more resource nodes. The cloud computing platform architecture also includes a scheduling component, which is used to schedule resources for the N nodes, where N is a positive integer. The method includes:

[0005] When a resource scheduling request is received, a node resource topology information set is obtained; wherein the node resource topology information set includes node resource topology information of N nodes, and the node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component;

[0006] Based on the resource scheduling request and the node resource topology information set, perform resource analysis on N nodes and determine the target node from the N nodes; the target node meets the resource requirements of the resource object requested by the resource scheduling request;

[0007] Pre-allocating resources in the target node based on the resource scheduling request, and generating resource annotation information of the resource object based on the resource allocation result of the target node; the resource annotation information is used to instruct the agent component in the target node to bind the allocated resource node with the resource object in the resource scheduling request in the target node;

[0008] The bound resource node is used to execute the data processing service provided by the resource object.

[0009] In one aspect, an embodiment of the present application provides a node resource processing device, which is applied to a cloud computing platform architecture. The cloud computing platform architecture includes N nodes, wherein an agent component is deployed in each node, and each node includes one or more resource nodes. The cloud computing platform architecture also includes a scheduling component, which is used to schedule resources for the N nodes, where N is a positive integer. The device includes:

[0010] An acquiring unit is configured to acquire a node resource topology information set upon receiving a resource scheduling request; wherein the node resource topology information set includes node resource topology information of N nodes, and the node resource topology information of any node is collected by an agent component deployed in the corresponding node and stored in the scheduling component;

[0011] The processing unit is configured to perform resource analysis and processing on N nodes based on the resource scheduling request and the node resource topology information set, and determine a target node from the N nodes; the target node satisfies the resource requirements of the resource object requested for scheduling by the resource scheduling request;

[0012] The processing unit is further configured to pre-allocate resources in the target node based on the resource scheduling request, and generate resource annotation information of the resource object based on the resource allocation result of the target node; the resource annotation information is used to instruct the agent component in the target node to bind the allocated resource node with the resource object in the resource scheduling request in the target node;

[0013] The bound resource node is used to execute the data processing service provided by the resource object.

[0014] In a possible implementation, during the process of the agent component collecting node resource topology information of any node, the processing unit is configured to perform the following operations:

[0015] Determine the startup parameters of the agent component and configure the collection cycle of the agent component in the startup parameters of the agent component;

[0016] According to the configured collection cycle, the agent component is used to collect and process the resource information of the node to obtain the node collection information of the node; wherein the node collection information includes: the system resources and reserved resources of the node;

[0017] The difference between the system resources and the reserved resources of the node is calculated to obtain the node resource topology information of the node.

[0018] In one possible implementation, in a process of using an agent component to collect and process resource information of a node to obtain reserved resources of the node, the processing unit is configured to perform the following operations:

[0019] Obtaining startup parameters of the cloud computing platform architecture, parsing the startup parameters of the cloud computing platform architecture to obtain parsed parameters, and calculating the parsed parameters according to priority rules to obtain reserved resources of the node; or,

[0020] Call the remote procedure call service to collect and process information on the node and obtain the reserved resources of the node.

[0021] In one possible implementation, the processing unit performs resource analysis on N nodes based on the resource scheduling request and the node resource topology information set, and determines a target node from the N nodes to perform the following operations:

[0022] Based on the resource scheduling request and the node resource topology information set, N nodes are filtered according to the node resource topology management strategy to obtain K candidate nodes that meet the affinity constraints; K is a positive integer and K≤N;

[0023] Scoring the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes; the scoring node of any candidate node includes the score value of the corresponding candidate node;

[0024] Based on the scores of K candidate nodes, the target node is determined.

[0025] In one possible implementation, the resource objects requested for scheduling by the resource scheduling request include: X container resources, where X is a positive integer; when the node resource topology management policy is the first policy, the first policy is used to indicate that a resource node in a node needs to be used for resource scheduling; any one of the N nodes is represented as node i;

[0026] Based on the resource scheduling request and the node resource topology information set, the processing unit filters N nodes according to the node resource topology management policy to obtain K candidate nodes that meet the affinity constraints, which are used to perform the following operations:

[0027] Based on the node resource topology information set, analyze and obtain the available resource amounts corresponding to each resource node in the node i, and determine the maximum value of each available resource amount as the maximum available resource amount Y of the node;

[0028] Among the N nodes, the nodes whose maximum available resource amount Y is less than the required resource amount X are determined as nodes to be filtered;

[0029] One or more nodes to be filtered are determined from the N nodes and filtered to obtain K candidate nodes that meet the affinity constraint.

[0030] In one possible implementation, when the node resource topology management strategy is the second strategy, the second strategy is used to indicate that one or more resource nodes in a node need to be used for resource scheduling; any candidate node among the K candidate nodes is represented as candidate node j; and the processing unit is further used to perform the following operations:

[0031] Based on the node resource topology information of the candidate node j, the resource amount X indicated by the resource scheduling request is allocated to obtain at least one resource node combination included in the candidate node j; a resource node combination includes one or more resource nodes;

[0032] According to the preset selection conditions, the target resource node combination corresponding to the candidate node j is selected from each resource node combination.

[0033] In one possible implementation, the scoring strategy includes a first scoring substratum and a second scoring substratum, and the scoring result of any candidate node includes a first scoring value and a second scoring value. The processing unit scores the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes, which are used to perform the following operations:

[0034] Get K target resource node combinations of K candidate nodes; one candidate node corresponds to one target resource node combination;

[0035] According to the first scoring strategy, the number of resource nodes in the K target resource node combinations is calculated to obtain the first scoring value of each candidate node;

[0036] According to the second scoring sub-strategy, the resource node distances in the K target resource node combinations are calculated respectively to obtain the second scoring value of each candidate node.

[0037] In one possible implementation, the processing unit determines a target node based on the scores of the K candidate nodes, and performs the following operations:

[0038] Based on the first scores of the K candidate nodes, determining one or more target candidate nodes with the largest first scores from the K candidate nodes;

[0039] If the number of target candidate nodes is 1, the target candidate node is determined as the target node;

[0040] If the number of target candidate nodes is greater than 1, based on the second scores of the respective target candidate nodes, a target candidate node with the largest second score is determined as the target node.

[0041] In one possible implementation, the node resource topology information of any node includes the available resource amount of each resource node in the node; the resource object requested for scheduling by the resource scheduling request includes: X container resources, any container resource among the X container resources is represented as container resource p; the resource allocation result includes the allocated quantity of container resource p in one or more resource nodes;

[0042] The processing unit pre-allocates resources in the target node based on the resource scheduling request to perform the following operations:

[0043] Determine a target resource node combination corresponding to the target node, where the target resource node combination includes one or more target resource nodes to be used;

[0044] According to the X container resources to be scheduled in the resource scheduling request and the available resources of each target resource node, the corresponding allocation quantity of the container resource p in each target resource node is determined.

[0045] In one possible implementation, the resource annotation information includes multiple fields, including a policy field and a resource allocation field. The processing unit generates resource annotation information of the resource object based on the resource allocation result of the target node, and is used to perform the following operations:

[0046] Parse the resource scheduling request, obtain the resource binding policy of the container resource p, and write the resource binding policy of the container resource p into the policy field of the resource annotation information;

[0047] The allocated quantity of the container resource p in each target resource node is written into the resource allocation field of the resource annotation information.

[0048] In a possible implementation, the processing unit is further configured to perform the following operations:

[0049] Calling the proxy component in the target node to obtain the target node resource topology information of the target node;

[0050] Verify each target resource node based on the allocated quantity of the container resource p indicated by the resource allocation field in the resource annotation information and the available resources of each target resource node;

[0051] If the verification passes, the container resource p is bound in the operating system of the target node according to the resource binding policy and allocation quantity of the container resource p indicated by the policy field in the resource annotation information, and the container runtime interface is called.

[0052] In a possible implementation, the target node includes one or more target resource nodes, and any target resource node includes at least one multi-core processor; the resource binding strategy includes: a first binding strategy, a second binding strategy, a third binding strategy, and a fourth binding strategy; wherein,

[0053] The first binding policy is used to indicate that the container resource p is bound to all multi-core processors in the target node;

[0054] The second binding policy is used to indicate that the container resource p is exclusively bound to the first multi-core processor in the target node;

[0055] The third binding policy is used to indicate that the container resource p is bound to the second multi-core processor in the target node, and the second multi-core processor is not exclusively bound;

[0056] The fourth binding policy is used to indicate that the container resource p is bound to the target resource node of the target node. In one aspect, an embodiment of the present application provides a computer device comprising a processor, an input device, an output device, and a memory; the memory stores a computer program; when the computer program is executed by the processor, the node resource processing method described above is executed.

[0057] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned node resource processing method is executed.

[0058] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned node resource processing method is executed.

[0059] In an embodiment of the present application, when a resource scheduling request is received, a node resource topology information set is obtained, which includes the node resource topology information of N nodes. The node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component; based on the resource scheduling request and the node resource topology information set, resource analysis and processing are performed on the N nodes, and a target node is determined from the N nodes, which target node meets the resource requirements of the resource object requested to be scheduled by the resource scheduling request; based on the resource scheduling request, resources in the target node are pre-allocated, and resource annotation information of the resource object is generated based on the resource allocation result of the target node, and the resource annotation information is used to indicate: the agent component in the target node binds the allocated resource node with the resource object in the resource scheduling request in the target node; wherein, the bound resource node is scheduled as the resource object to perform data processing. It can be seen that, on the one hand, the agent component in each node in the present application can collect the node resource topology information of the corresponding node and store it in the scheduling component, so that the scheduling component can specifically perceive the topological structure of the node resources, so that when the scheduling component performs resource allocation, it can reasonably allocate resources according to the topological structure of the node resources; on the other hand, when the agent component is responsible for resource allocation, it can not only take into account the topological structure of the node resources, but also meet the resource scheduling requirements, further improve the rationality and accuracy of resource allocation, thereby improving the resource scheduling effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technical objects in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0061] Figure 1 This is a schematic diagram of the principle of a node resource processing solution provided in an embodiment of the present application;

[0062] Figure 2 This is a schematic diagram of the structure of a node provided in an embodiment of the present application;

[0063] Figure 3 This is a schematic diagram of the architecture of a node resource processing system provided in an embodiment of the present application;

[0064] Figure 4 This is a flowchart of a node resource processing method provided in an embodiment of the present application;

[0065] Figure 5 This is a complete interactive flow chart of node resource processing provided by an embodiment of the present application;

[0066] Figure 6a This is a configuration flow chart of an instance dimension provided in an embodiment of the present application;

[0067] Figure 6b This is a configuration flow chart of a database instance provided in an embodiment of the present application;

[0068] Figure 7 This is a schematic diagram of a scheduling process of a scheduling component provided in an embodiment of the present application;

[0069] Figure 8a This is a schematic diagram of an agent perspective of an agent component provided in an embodiment of the present application;

[0070] Figure 8b This is a schematic diagram of an agent perspective of another agent component provided in an embodiment of the present application;

[0071] Figure 9 This is a schematic diagram of a process for continuous tuning of an agent component provided in an embodiment of the present application;

[0072] Figure 10 This is a schematic diagram of the structure of a node resource processing device provided in an embodiment of the present application;

[0073] Figure 11 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0075] In this application, the term "module" or "unit" refers to a computer program or part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the functions of the module or unit.

[0076] This application provides a node resource processing solution for a cloud computing platform architecture (such as a NUMA architecture). The cloud computing platform architecture includes N nodes, one of which is deployed with an agent component, and one of which includes one or more resource nodes. The cloud computing platform architecture also includes a scheduling component, which is used to schedule resources for the N nodes, where N is a positive integer. Specifically, this application can allocate and schedule node resources according to the node resource topology and resource scheduling requests, so that node resources can be reasonably allocated and the scheduling results meet user needs, thereby improving resource scheduling performance. Please refer to Figure 1 , Figure 1 This is a schematic diagram of the principle of a node resource processing solution provided by the embodiment of the present application. Figure 1 The specific architecture and implementation principles of this application solution are briefly explained.

[0077] First, yes Figure 1 The cloud computing platform architecture shown is briefly described.

[0078] like Figure 1 The cloud computing platform architecture shown in FIG. 1 includes N nodes, one of which is an agent component deployed in each node; and a scheduler component is also deployed in the cloud computing platform architecture. The agent component and the scheduler component interact through the resource service component (ApiServer). Figure 1 Introduce and explain each component involved.

[0079] ①Agent component (agent component).

[0080] The Agent component is deployed in each node in the form of k8s daemonset.

[0081] The Agent component mainly includes two functions (Discovery and AllocateManager):

[0082] Discovery: Regularly collects node resource topology information into the Node Resource Topology Information Set (NRT);

[0083] AllocateManager: Responsible for allocating CPU / memory / device resources on the node based on the resource allocation results in the resource annotation information, thereby binding the allocated resource node to the resource object in the resource scheduling request.

[0084] ②Scheduler component (scheduling component).

[0085] The Scheduler component implements the NodeReosurceTopology (NRT) scheduling plug-in based on the scheduler-framework and runs in the form of k8s deployment.

[0086] Scheduler component has the following functions:

[0087] Combine the node resource topology information of all nodes in NRT, filter out nodes that do not meet the scheduling requirements, and select the best node (i.e., target node) from the nodes that meet the scheduling requirements;

[0088] The resource allocation result corresponding to the selected target node is recorded in the resource annotation information.

[0089] ③ApiServer component (resource service component).

[0090] The ApiServer component is a crucial component in k8s (Kubernetes), and it's key to implementing declarative APIs. The core functionality of the ApiServer in Kubernetes is to provide HTTP REST interfaces for adding, deleting, modifying, querying, and watching various resource objects (pods, services, etc.) in Kubernetes.

[0091] Then, the implementation principle of the node resource processing solution is explained.

[0092] ① Node resource topology information collection: Specifically, an agent component (Agent component) is deployed in a node. The Agent component is used to regularly summarize the node resource topology information in the node, and send the node resource topology information to the scheduling component (Scheduler component) by sending it to the ApiServer (resource service component).

[0093] ② Acquisition of the node resource topology information set: The Scheduler component has the function of watching pods. That is, whenever a new resource scheduling request is received, it can be considered that the Scheduler component has watched a new pod request. Then, the Scheduler component can obtain the node resource topology information set, which includes the node resource topology information (NRT) of N nodes.

[0094] ③ Determination of target node: Specifically, the Scheduler component performs resource analysis and processing on N nodes based on the resource scheduling request and the node resource topology information set, and determines the target node from the N nodes; the target node meets the resource requirements of the resource object (such as a container) requested to be scheduled by the resource scheduling request.

[0095] ④ Generation of resource annotation information: Specifically, the Scheduler component pre-allocates resources in the target node based on the resource scheduling request, and generates resource annotation information (pod annotation) of the resource object based on the resource allocation result of the target node.

[0096] ⑤ Resource Binding: The Agent component receives the pod annotation generated by the Scheduler component through the ApiServer (resource service component). Based on the resource allocation results recorded in the pod annotation, it binds the allocated resource node to the resource object of the resource scheduling request in the target node. The bound resource node is used to execute the data processing services provided by the resource object, such as data computing services, data analysis services, data storage services, and other services.

[0097] As can be seen from the above, on the one hand, the agent component in each node in the present application can collect the node resource topology information of the corresponding node and store it in the scheduling component, so that the scheduling component can specifically perceive the topological structure of the node resources, so that when the scheduling component performs resource allocation, it can reasonably allocate resources according to the topological structure of the node resources; on the other hand, when the agent component is responsible for resource allocation, it can not only take into account the topological structure of the node resources, but also meet the resource scheduling requirements, further improve the rationality and accuracy of resource allocation, thereby improving the resource scheduling effect.

[0098] The following is a detailed introduction to the background knowledge and key technical terms involved in this application.

[0099] 1. Introduction to the background knowledge of k8s (kubernetes).

[0100] The cloud computing platform architecture involved in the embodiments of this application is a cloud native system architecture of the k8s system. Specifically, the cloud computing platform architecture mentioned in this application is mainly the NUMA architecture under the k8s system. The relevant background knowledge involved in the NUMA architecture is described in detail below.

[0101] (1) Multi-core processor.

[0102] As the name suggests, a multi-core processor refers to a processor (Central Processing Unit, CPU) that includes multiple cores. Among them:

[0103] Socket / Package: This refers to the physically and mechanically separable CPU. Home PCs typically have a single socket, while servers typically support two sockets (often called "dual-socket"), though four or eight sockets are also possible. Sockets are typically connected via a high-speed bus (such as QPI Infinity Fabric).

[0104] Core: A complete, independently executed processing unit on a CPU. It's also the unit used by the operating system to schedule processes. Common home and server processors are multi-core processors. For example, an AMD EPYC 7763 processor has 64 cores, while an Intel Xeon Platinum 8380 has 40 cores.

[0105] (2) NUMA (Non-Uniform Memory Access).

[0106] Non-Uniform Memory Access (NUMA) is a cloud computing platform architecture that allows different CPUs to access different memory areas at different speeds. The node resource topology involved in this application mainly includes the relative positions of the CPU, memory, and PCI devices.

[0107] At present, the processors of computer devices usually adopt the NUMA architecture, that is, each socket is connected to the local memory through a memory controller, and accesses the remote memory (remote memory) belonging to other sockets through a high-speed bus between sockets. The directly connected CPU core and memory and other peripherals (such as network cards, GPUs) are called a NUMAdomain (or NUMA node). The memory access performance (including bandwidth and latency) in the same NUMA domain (intra-domain) is usually significantly higher than the performance across NUMA (inter-domain). This phenomenon is called the NUMA effect. Therefore, the embodiments of the present application are mainly aimed at the node resource processing solution provided under the NUMA architecture of the cloud native scenario, and combine the scheduling requirements under the NUMA architecture to jointly realize the reasonable allocation of node resources.

[0108] (3)Controller in k8s.

[0109] The Kubernetes system includes multiple types of controllers. A controller can be understood as the program itself that operates resources in Kubernetes. For example, controllers such as deployment and daemonset are built-in controllers in Kubernetes. In this application, both the Agent component and the Scheduler component can be considered a type of controller in Kubernetes. The Agent component can be considered a daemonset-type controller in Kubernetes, while the Scheduler component is a deployment-type controller in Kubernetes.

[0110] 2. Multi-level relationship of NUMA architecture.

[0111] The NUMA architecture mainly includes: nodes, resource nodes (NUMA nodes), and multi-core processors. Among them, a node includes one or more resource nodes, a resource node includes one or more processors (CPUs), and a processor includes multiple cores. In the embodiment of this application, for ease of explanation, an example is given in which a resource node corresponds to a multi-core processor (i.e., a resource node includes a processor, and the processor includes multiple cores). Please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a node provided in an embodiment of the present application. Figure 2 As shown, a node includes four resource nodes NUMA nodes: NUMA node0, NUMA node1, NUMA node2, and NUMA node3; and a resource node corresponds to a multi-core processor, such as NUMA node0 includes four cores: core0, core1, core2, and core3.

[0112] 3. Nodes and resource nodes.

[0113] A node is a physical machine or a virtual machine. A virtual machine is a complete computer system with complete hardware system functions, running in a completely isolated environment, simulated by software. Physical machines can be categorized by device type as: terminal devices or servers, such as CVMs (Cloud Virtual Machines). Based on business scenarios, physical machines can be categorized as: blockchain nodes, gaming devices, live streaming devices, and other computer devices. This application does not specifically limit the types of nodes.

[0114] A resource node (NUMA node), also known as a NUMA domain, primarily consists of directly connected CPU cores, memory, and other peripherals (such as network cards and GPUs). In practical applications, a node can be divided into multiple NUMA nodes based on the node's physical properties. For example, a node can be divided into four NUMA nodes (NUMA node0, NUMA node1, NUMA node2, and NUMA node3).

[0115] 4. Resource scheduling request, resource object (pod).

[0116] A resource scheduling request is an object request for scheduling a resource object to a node or resource node. Depending on affinity constraints, a resource scheduling request can be configured as a NUMA node affinity request, requesting that a resource object be scheduled to a specific NUMA node.

[0117] Resource object (pod): The so-called pod is the smallest deployable computing unit that can be created and managed in kubernetes. Among them, pod can usually be understood as a collection of containers, that is, a pod includes one or more container resources. Any container mentioned here usually has different service capabilities, such as database services, log services, management services, etc. Therefore, for CPU resources, the resource scheduling request is used to request that each container resource in the resource object pod be scheduled (bound) to the CPU core of the corresponding NUMAnode, such as scheduling container0 in the pod to Figure 2 On core0 in NUMA node0 shown, core0 bound to container0 can provide the data processing services of container0 for the current node, such as data computing services, data analysis services, data storage services, etc.

[0118] 5. Cloud technology.

[0119] The node resource processing solution proposed in this application involves a large number of data computing services and data storage services, so a large amount of computer operating costs are required. Then, cloud technology can be used to provide data computing services and data storage services for this solution, so that data processing can be better performed. Specifically, data computing services can be used to perform resource analysis and processing on N nodes, or data computing services can be used to pre-allocate resources in the target node; in addition, the node resource topology information collected by the proxy component can be stored in the NRT object of the scheduling component based on the data storage service, so as to facilitate subsequent analysis and scheduling of node resources based on the node resource topology information of each node in the NRT object. Among them, cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the application of cloud computing business model, which can form a resource pool, which is used on demand and is flexible and convenient. Among them, cloud technology can include cloud storage technology. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.

[0120] 6. Blockchain.

[0121] Blockchain is a new application model for computer technologies, including distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a series of data blocks linked using cryptographic methods. Each block contains information about a batch of network transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. The following explains concepts related to blockchain systems, blockchain nodes, and block structures.

[0122] In this application, the data processing process involves a lot of data, such as: node resource topology information set, resource allocation results, and resource annotation information, etc. Optionally, this application can send the above data to the blockchain for storage. Based on the blockchain's tamper-proof and traceable characteristics, the data can be prevented from being tampered with or leaked, thereby improving the data security and reliability of the node resource processing process.

[0123] It should be noted that many data are involved in the data processing process of this application, such as: node resource topology information set, resource allocation results, and resource annotation information, etc. When the above embodiments of this application are applied to specific products or technologies, user permission or consent must be obtained, and the relevant data collection, use and processing processes must comply with relevant local laws, regulations and standards, and comply with the principles of legality, legitimacy and necessity, and do not involve obtaining data types prohibited or restricted by laws and regulations. In some optional embodiments, the relevant data involved in the embodiments of this application are obtained after the object has been individually authorized. In addition, when obtaining the individual authorization of the object, the purpose of the relevant data involved must be indicated to the object.

[0124] The system architecture provided by this application is described in detail below with reference to the accompanying drawings.

[0125] See Figure 3 , Figure 3 This is a schematic diagram of the architecture of a node resource processing system provided by an embodiment of the present application. Figure 3 As shown, the architecture diagram of the node resource processing system may include at least: a terminal device cluster and a server 304. Among them, the terminal device cluster may include at least one terminal device, such as: terminal device 301, terminal device 302, terminal device 303, and so on. It should be understood that one terminal device corresponds to a node in this application, and each terminal device is deployed with an agent component for collecting the node resource topology information (NRT information) of the current terminal device, and a scheduling component is deployed in the server, that is, the server has the ability to perform resource scheduling for each terminal device; in addition, the embodiment of the present application does not specifically limit the number of terminal devices in the terminal device cluster, and the number of devices can be flexibly changed according to the different needs of business scenarios, such as the number of financial devices in the financial scenario is 2, and the number of game devices in the game scenario is 3. It should be noted that any terminal device can be directly or indirectly connected to the server 304 through wired or wireless communication.

[0126] Any computer device (terminal device, or server) in the node resource processing system provided in this application can be a mobile phone, tablet computer, laptop computer, PDA, mobile Internet device (MID), vehicle, vehicle-mounted equipment, roadside equipment, intelligent robot, aircraft, wearable device, such as smart watches, smart bracelets, pedometers and other smart devices, virtual reality equipment, etc.

[0127] Any computer device (terminal device, or server) in the node resource processing system provided in this application may also be a server. Specifically, a server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0128] It is understandable that the types of the various computer devices in the node resource processing system of the present application can be the same or different. For example, terminal device 301 can be a mobile phone, terminal device 302 can be a gaming device, and terminal device 303 can be a live broadcast device. For another example, terminal device 301, terminal device 302, and terminal device 303 can all be live broadcast devices. The present application does not limit the number and type of the various computer devices in the node resource processing system. The following briefly describes the specific process of the data processing scheme involved in the embodiment of the present application, taking terminal device 301 and server 304 as an example:

[0129] ① The agent component in any terminal device (such as terminal device 301) can regularly (such as once a week) collect the node resource topology information of the current device and send the node resource topology information to server 304.

[0130] ② When receiving the resource scheduling request, the server 304 may obtain a node resource topology information set, wherein the node resource topology information of each terminal device is stored in the node resource topology information set.

[0131] ③ Based on the resource scheduling request and the node resource topology information set, the server 304 performs resource analysis and processing on the N nodes, and determines the target node from the N nodes; the target node meets the resource requirements of the resource object requested for scheduling by the resource scheduling request.

[0132] ④ The server 304 pre-allocates resources in the target node based on the resource scheduling request, and generates resource annotation information of the resource object based on the resource allocation result of the target node.

[0133] ⑤ Server 304 sends the resource annotation information to terminal device 301. Based on the resource annotation information, the proxy component in terminal device 301 binds the allocated resource node to the resource object in the resource scheduling request. The bound resource node is used to execute the data processing services provided by the resource object, such as data computing services, data analysis services, and data storage services.

[0134] In one possible implementation, the node resource processing system provided in the present application can be deployed in a blockchain system, that is, the terminal device 301, the terminal device 302, the terminal device 303, and the server 304 can all be used as node devices in the blockchain system, and the relevant data involved in the above data processing process (such as node resource topology information, resource allocation results, and resource annotation information, etc.) are all stored on the blockchain, so that the specific processing flow of the node resource processing in the present application can be executed on the blockchain, which can not only ensure the fairness and justice of the node resource processing process, but also make the node resource processing process traceable, thereby improving the security and reliability of the node resource processing process.

[0135] The cloud computing platform architecture provided by the present application mainly includes N nodes, one of which is deployed with an agent component, and one of which includes one or more resource nodes; a scheduling component is also deployed in the cloud computing platform architecture, and the scheduling component is used to perform resource scheduling on the N nodes. Specifically, when a resource scheduling request is received, a node resource topology information set is obtained, and the node resource topology information set includes node resource topology information of N nodes, and the node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component; based on the resource scheduling request and the node resource topology information set, resource analysis and processing are performed on the N nodes, and a target node is determined from the N nodes, and the target node meets the resource requirements of the resource object requested for scheduling by the resource scheduling request; based on the resource scheduling request, resources in the target node are pre-allocated, and resource annotation information of the resource object is generated based on the resource allocation result of the target node, and the resource annotation information is used to indicate: the agent component in the target node binds the allocated resource node with the resource object in the resource scheduling request in the target node; wherein the bound resource node is scheduled to perform data processing as the resource object. It can be seen that, on the one hand, the agent component in each node in the present application can collect the node resource topology information of the corresponding node and store it in the scheduling component, so that the scheduling component can specifically perceive the topological structure of the node resources, so that when the scheduling component performs resource allocation, it can reasonably allocate resources according to the topological structure of the node resources; on the other hand, when the agent component is responsible for resource allocation, it can not only take into account the topological structure of the node resources, but also meet the resource scheduling requirements, further improve the rationality and accuracy of resource allocation, thereby improving the resource scheduling effect.

[0136] It can be understood that the node resource processing system described in the embodiment of the present application is for the purpose of more clearly illustrating the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. It is known to those skilled in the art that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems.

[0137] The following describes specific embodiments of the node resource processing solution with reference to the accompanying drawings.

[0138] See Figure 4 , Figure 4 This is a flow chart of a node resource processing method provided by an embodiment of the present application. Figure 3 The computer device (terminal device or server) of the system shown in FIG is executed, and the method is applied to a cloud computing platform architecture, which includes N nodes (such as Figure 3 In the terminal device), a node is deployed with an agent component, and a node includes one or more resource nodes; a scheduling component is also deployed in the cloud computing platform architecture, and the scheduling component is used to schedule resources for N nodes, where N is a positive integer. Figure 4 As shown, the node resource processing method mainly includes but is not limited to the following steps S401-S403:

[0139] S401: When a resource scheduling request is received, a node resource topology information set is obtained.

[0140] Among them, the node resource topology information set is stored in the scheduling component, and the node resource topology information set includes the node resource topology information (NodeReosurceTopology, NRT) of N nodes. The NRT information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component.

[0141] Among them, the scheduling component has the ability to watch pods. Whenever the scheduling component receives a newly created pod from an object, it can be considered that the scheduling component has received a resource scheduling request. The resource scheduling request is used to request the scheduling of the pod to the resource node in the cloud computing platform architecture. A pod is composed of one or more container resources (containers). The essence of the resource scheduling request is to schedule the various container resources in the pod to the NUMA node.

[0142] The following describes in detail the process of collecting node resource topology information.

[0143] In one possible implementation, the process of an agent component collecting node resource topology information for any node includes: determining the agent component's startup parameters (such as "discovery-period-seconds") and configuring the agent component's collection period (e.g., once a week or once a day) in the agent component's startup parameters; using the agent component to collect and process resource information from the node according to the configured collection period to obtain node collection information for the node; wherein the node collection information includes: the node's system resources and reserved resources; and calculating the difference between the node's system resources and reserved resources to obtain the node's node resource topology information. For example, if the system resources collected by the agent component are represented as capacity and the reserved resources are represented as reserved, then the node resource topology information refers to the node's available resources allocatable = capacity - reserved.

[0144] It should be noted that the node resource topology information collected by the agent component needs to be reported to the scheduling component for storage. The reporting structure is implemented through the custom k8s object NRT (NodeResourceTopology), which defines the amount of NUMA node resources distributed on a node.

[0145] Specifically, the reserved resources (reserved) of the agent component collection node mainly include the following two methods:

[0146] Method 1: Obtain startup parameters of the cloud computing platform architecture (kubelet), parse the startup parameters of the cloud computing platform architecture to obtain parsed parameters, and calculate the parsed parameters according to priority rules to obtain reserved resources of the node.

[0147] Method 2: Call the remote procedure call service to collect and process node information and obtain the node's reserved resources. This can be obtained through the kubelet's GRPC service (Google Remote Procedure Call) "GetAllocatableResources." This method is more efficient than the previous method and provides a more accurate representation of the reserved resources.

[0148] Based on this, the agent component in any node regularly collects the node resource topology information of the corresponding node according to the above process, and sends the node resource topology information to the node resource topology information set (NRT object) in the scheduling component for storage. It can be seen that the scheduling component can perceive the specific node resource topology structure based on the NRT object, so as to facilitate the subsequent reasonable allocation of node resources based on the NRT object.

[0149] It should be noted that in the process of collecting node resource topology information, when it is detected that the NUMA information on the node has changed, resulting in a change in the NRT object, the Agent component will send an update request to the Scheduler component through the ApiServer to trigger the Scheduler component to update the NUMA information of the corresponding node in the NRT object. That is, the update action performed here is an "edge-triggered" action.

[0150] S402: Based on the resource scheduling request and the node resource topology information set, perform resource analysis on N nodes and determine a target node from the N nodes.

[0151] In one possible implementation, the scheduling component performs resource analysis on N nodes based on a resource scheduling request and a node resource topology information set. The process of determining a target node from the N nodes includes: first, based on the resource scheduling request and the node resource topology information set, the N nodes are filtered according to a node resource topology management strategy to obtain K candidate nodes that meet affinity constraints; K is a positive integer and K≤N; then, the K candidate nodes are scored according to a scoring strategy to obtain scoring results for the K candidate nodes; the scoring node of any candidate node includes the score value of the corresponding candidate node; finally, the target node is determined based on the score values ​​of the K candidate nodes.

[0152] 1. Node filtering.

[0153] (1) The node resource topology management strategy is the "SingleNumaNode" strategy.

[0154] In one possible implementation, the resource objects requested for scheduling by the resource scheduling request include: X container resources, where X is a positive integer; when the node resource topology management policy is the first policy (e.g., the "SingleNumaNode" policy), the "SingleNumaNode" policy is used to indicate that a resource node in a node is to be used for resource scheduling; any one of the N nodes is represented as node i. Specifically, based on the resource scheduling request and the node resource topology information set, the scheduling component filters the N nodes according to the node resource topology management policy to obtain K candidate nodes that meet the affinity constraints, including:

[0155] ① Based on the node resource topology information set, analyze and obtain the available resources corresponding to each resource node in node i, and determine the maximum value of each available resource as the node's maximum available resource amount Y. For example, the scheduling component can determine the available resources of each resource node NUMAnode in any node i according to the node resource topology information of each node. For example, node i includes NUMA node0 and NUMAnode1, where the available resources of NUMA node0 include 6 cores; the available resources of NUMA node1 include 4 cores, then the maximum available resources of the current node i Y = 6 cores.

[0156] ② Among the N nodes, determine the nodes whose maximum available resources Y are less than the required resources X as the nodes to be filtered. For example, the N nodes include the first node and the second node. Assume that the required resources X = 5, where the maximum available resources Y1 of the first node is 6 and the maximum available resources Y2 of the second node is 3; then the second node can be selected as the node to be filtered.

[0157] ③ Filter one or more nodes to be filtered from the N nodes to obtain K candidate nodes that meet the affinity constraints.

[0158] In the above node filtering process, since the first strategy ("SingleNumaNode" strategy) is a strategy with NUMAnode affinity constraints, this strategy is used to indicate that the resources to be scheduled must be scheduled to the same NUMA node, so that each candidate node obtained by filtering meets the affinity constraints of the NUMA node, that is, in the embodiment of the present application, node resources can be scheduled according to the dimension of the NUMA node.

[0159] (2) The node resource topology management strategy is the "LeastNumaNode" strategy.

[0160] In one possible implementation, when the node resource topology management strategy is the second strategy (for example, the "LeastNumaNode" strategy), the "LeastNumaNode" strategy is used to indicate that one or more resource nodes in a node need to be used for resource scheduling. In a specific implementation, any candidate node among the K candidate nodes is represented as candidate node j. First, based on the node resource topology information of candidate node j, the amount of resources X to be scheduled indicated by the resource scheduling request can be allocated resources to obtain at least one resource node combination included in candidate node j, and a resource node combination includes one or more resource nodes; then, according to the preset selection conditions, the target resource node combination corresponding to candidate node j is selected from each resource node combination. The following is a detailed description of the process of determining the target resource node combination:

[0161] ① Enumerate all NUMA node combinations of each node. Specifically, any candidate node among the K candidate nodes is represented as candidate node j. Based on the node resource topology information of candidate node j, the amount of resources to be scheduled X indicated by the resource scheduling request can be allocated resources to obtain at least one resource node combination included in candidate node j. A resource node combination includes one or more resource nodes. For example, candidate node 1 includes NUMA node0, NUMA node1 and NUMAnode2, and the amount of available resources allocatable of NUMA node0 is 3c (i.e., 3 cores), the amount of available resources allocatable of NUMA node1 is 2c (i.e., 2 cores), and the amount of available resources allocatable of NUMA node2 is 2c (i.e., 2 cores); assuming that the amount of resources to be scheduled X is 5c (i.e., 5 containers), then the NUMA node combination (resource node combination) corresponding to candidate node 1 is shown in Table 1 below:

[0162] Table 1. NUMA node combination corresponding to candidate node 1

[0163]

[0164] As can be seen from Table 1 above, based on the amount of resources to be scheduled and the topological structure of node resources, the various NUMA node combinations included in the current node can be exhaustively enumerated.

[0165] ② Determine the optimal NUMA node combination (target resource node combination). Specifically, the target resource node combination corresponding to the candidate node j can be selected from each resource node combination according to the preset selection condition. The preset selection condition is used to indicate that the selection is made in a manner that minimizes the number of NUMA nodes across. For example, the "NUMA node0+NUMA node1" combination and the "NUMA node0+NUMA node2" combination in Table 1 above refer to the combination with the least number of NUMA nodes across. Alternatively, the preset selection condition can also be used to indicate that the distances between multiple NUMA node combinations are the shortest. For example, for the "NUMA node0+NUMA node1" combination and the "NUMA node0+NUMA node2" combination, the distances between the NUMA nodes in the "NUMAnode0+NUMAnode1" combination are the shortest.

[0166] 2. Node scoring.

[0167] (1) The node resource topology management strategy is the "SingleNumaNode" strategy.

[0168] When the node resource topology management policy is the "SingleNumaNode" policy, this policy is used to indicate whether NumaNode should be filled more fully or less sparsely. For example, if the resource scheduling request requires container resource X = 5c, the available resources of NumaNode1 in Node 1 are 8c, and the available resources of NumaNode2 in Node 2 are 5c, then when the "SingleNumaNode" policy is used to indicate whether NumaNode should be filled more fully, NumaNode2 in Node 2 can be used, thus filling NumaNode2 fully. Alternatively, when the "SingleNumaNode" policy is used to indicate whether NumaNode should be filled less sparsely, NumaNode1 in Node 1 can be used, thus filling NumaNode1 relatively more sparsely.

[0169] (2) The node resource topology management strategy is the "LeastNumaNode" strategy.

[0170] In one possible implementation, the scoring strategy includes a first scoring sub-strategy and a second scoring sub-strategy, and the scoring result of any candidate node includes a first scoring value and a second scoring value. The scheduling component scores the K candidate nodes according to the scoring strategy, and obtains the scoring results of the K candidate nodes, including:

[0171] ① Obtain K target resource node combinations for K candidate nodes; one candidate node corresponds to one target resource node combination. As shown above, under the "LeastNumaNode" strategy, each candidate node corresponds to one target resource node combination, which is the optimal NUMA node combination.

[0172] ② According to the first scoring sub-strategy, the number of resource nodes in the K target resource node combinations is calculated respectively to obtain the first scoring value of each candidate node. Among them, the first scoring sub-strategy is used to indicate that the candidate node with more resource nodes has a lower score. For example, the target NUMAnode combination of candidate node 1 includes: NUMA node0, NUMAnode1; the target NUMA node combination of candidate node 2 includes: NUMA node2; the target NUMA node combination of candidate node 3 includes: NUMA node3, NUMA node5; since candidate node 2 only contains one NUMA node, the first scoring value corresponding to candidate node 2 is the highest (such as 100). In addition, candidate node 1 and candidate node 3 both contain two NUMA nodes, so candidate node 1 and candidate node 3 have the same first scoring value (such as 80).

[0173] ③ Calculate the resource node distances in the K target resource node combinations according to the second scoring sub-strategy to obtain the second scoring value of each candidate node. For example, if the target NUMA node combination of candidate node 1 includes NUMAnode0 and NUMA node1, the target NUMA node combination of candidate node 2 includes NUMA node2, and the target NUMA node combination of candidate node 3 includes NUMA node3 and NUMA node5, then the second scoring value of candidate node 1 (e.g., 80) is higher than the second scoring value of candidate node 3 (e.g., 70).

[0174] 3. Determine the optimal node (target node).

[0175] In one possible implementation, the scheduling component determines one or more target candidate nodes with the largest first scores from the K candidate nodes based on the first scores of the K candidate nodes; if the number of target candidate nodes is 1, the target candidate node is determined as the target node; if the number of target candidate nodes is greater than 1, the target candidate node with the largest second score is determined as the target node based on the second scores of the respective target candidate nodes. For example, the scoring results of the respective candidate nodes determined according to the above steps are shown in Table 2 below:

[0176] Table 2. Scoring results of different nodes

[0177] node First scoring value Second scoring value Candidate node 1 100 100 Candidate node 2 100 80 Candidate node 3 80 70

[0178] As shown in Table 2 above, since the candidate nodes with the largest first score value (100) include candidate node 1 and candidate node 2, the second score values ​​of candidate node 1 and candidate node 2 can be further compared. Since the second score value of candidate node 1 (100) is greater than the second score value of candidate node 2 (80), candidate node 1 can be determined as the target node.

[0179] Based on this, during the resource analysis process, the scheduling component can analyze according to the specific topological structure of the node resources, and can also analyze according to the resource objects requested for scheduling by the resource scheduling request, so that the optimal node that meets the resource scheduling requirements can be determined from N nodes as the target node.

[0180] S403: Pre-allocate resources in the target node based on the resource scheduling request, and generate resource annotation information of the resource object based on the resource allocation result of the target node.

[0181] In one possible implementation, the node resource topology information of any node includes the available resource amount of each resource node in the corresponding node; the resource object requested for scheduling by the resource scheduling request includes: X container resources, any one of the X container resources is represented as container resource p; and the resource allocation result includes the allocated quantity of container resource p in one or more resource nodes.

[0182] In one possible implementation, when the scheduling component pre-allocates resources in the target node based on the resource scheduling request, it mainly includes the following steps: (1) determining the target resource node combination corresponding to the target node, the target resource node combination including one or more target resource nodes to be used. The target resource node combination corresponding to the target node is determined according to the above-mentioned "LeastNumaNode" strategy. The specific determination process can be referred to the content in the above-mentioned step S402 in detail, which will not be repeated here. For example, the NUMAnode combination corresponding to the target node includes: two target resource nodes NUMA node0 and NUMA node1. (2) According to the X container resources to be scheduled by the resource scheduling request and the available resources of each target resource node, determine the corresponding allocation quantity of the container resource p in each target resource node. Specifically, assuming that the amount of resources to be scheduled is X = 5c (i.e., 5 containers, and the resource demand for each container is 1c), and the available resources of NUMA node0 are allocatable = 3c, and the available resources of NUMA node1 are allocatable = 2c, then the allocation result of each container resource in NUMAnode is expressed as: NUMA node0 = 3c, NUMA node1 = 2c.

[0183] In one possible implementation, the resource annotation information includes multiple fields, including: a policy field and a resource allocation field. The proxy component generates resource annotation information (annotation) of the resource object (pod) based on the resource allocation result of the target node, mainly including the following steps: parsing the resource scheduling request, obtaining the resource binding policy of the container resource p, and writing the resource binding policy of the container resource p into the policy field of the resource annotation information; writing the corresponding allocation quantity of the container resource p in each target resource node into the resource allocation field of the resource annotation information. Among them, the resource scheduling request indicates the resource binding policy of each container resource, and the resource binding policy here includes: a first binding policy, a second binding policy, a third binding policy, and a fourth binding policy. Specifically:

[0184] ① The first binding strategy (none): used to indicate that the container resource p is bound to all multi-core processors in the target node, that is, the container resource p is bound to all shared CPUs;

[0185] ② The second binding strategy (exclusive): used to indicate that the container resource p is exclusively bound to the first multi-core processor (exclusive cpu) in the target node, that is, the current container resource p is exclusively bound.

[0186] ③ The third binding strategy (fixed): used to indicate that the container resource p is bound to the second multi-core processor in the target node. The second multi-core processor is not exclusively bound. That is, the current container resource p is bound to the specified shared CPU (but this core cannot be exclusive).

[0187] ④ The fourth binding strategy (numa): is used to indicate that the container resource p is bound to the target resource node of the target node, that is, the current container resource p is bound to all shared CPUs in the NUMA node, where it is guaranteed that at least the number of CPU cores that meet the resource requirements exists in the NUMA node).

[0188] For example, the annotation of the pod generated by this application can be shown in Table 3 below:

[0189] Table 3. Resource annotation information

[0190]

[0191] As can be seen from Table 3 above, a resource object (pod) includes multiple container resources (container1, container2, container3, container4), and each container resource corresponds to a resource binding strategy and allocation quantity. For example, for container1, its resource binding strategy is an exclusive binding (exclusive) strategy, and the container1 needs to be scheduled to the cpu core5 in NUMA node0; for container2, its resource binding strategy is a numa binding strategy, and the container2 needs to be scheduled to the share cpu in NUMA node0, that is, it can share all the cpu cores of NUMA node0. It can be seen that this application not only supports exclusive core binding mode and unrestricted mode, but also supports binding numa mode and core binding but not exclusive mode, that is, this application supports more diverse resource binding modes.

[0192] Based on this, resource annotations for the current pod can be generated based on the resource binding policies and allocation data for each container resource in the resource object pod requested by the resource scheduling request. This shows that during the resource allocation process, the scheduling component can specifically perceive the specific topological structure of node resources and select the optimal node that meets the scheduling requirements. It then rationally allocates the number of NUMA nodes based on resource requirements, improving the rationality and accuracy of resource allocation.

[0193] In an embodiment of the present application, when a resource scheduling request is received, a node resource topology information set is obtained, which includes the node resource topology information of N nodes. The node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component; based on the resource scheduling request and the node resource topology information set, resource analysis and processing are performed on the N nodes, and a target node is determined from the N nodes, which target node meets the resource requirements of the resource object requested to be scheduled by the resource scheduling request; based on the resource scheduling request, resources in the target node are pre-allocated, and resource annotation information of the resource object is generated based on the resource allocation result of the target node, and the resource annotation information is used to indicate: the agent component in the target node binds the allocated resource node with the resource object in the resource scheduling request in the target node; wherein, the bound resource node is scheduled as the resource object to perform data processing. It can be seen that, on the one hand, the agent component in each node in the present application can collect the node resource topology information of the corresponding node and store it in the scheduling component, so that the scheduling component can specifically perceive the topological structure of the node resources, so that when the scheduling component performs resource allocation, it can reasonably allocate resources according to the topological structure of the node resources; on the other hand, when the agent component is responsible for resource allocation, it can not only take into account the topological structure of the node resources, but also meet the resource scheduling requirements, further improve the rationality and accuracy of resource allocation, thereby improving the resource scheduling effect.

[0194] See Figure 5 , Figure 5 This is a complete interactive flow chart of node resource processing provided by an embodiment of the present application. Specifically, the method is applied to a cloud computing platform architecture, which includes N nodes, one node is deployed with an agent component, and one node includes one or more resource nodes; the cloud computing platform architecture also has a scheduling component deployed, which is used to schedule resources for the N nodes. Figure 5 As shown, the present application mainly involves the interaction process between the Agent component (agent component), the Scheduler component (scheduling component), and the object client during the node resource processing process. The method mainly includes the following steps S501-S507:

[0195] S501: The agent component collects node resource topology information of the node.

[0196] In one possible implementation, the process of collecting node resource topology information of any node by the agent component includes: (1) determining the startup parameters of the agent component and configuring the collection period of the agent component in the startup parameters of the agent component (for example, once a week or once a day); (2) using the agent component to collect resource information of the node according to the configured collection period to obtain the node collection information of the node; wherein the node collection information includes: the system resources and reserved resources of the node; calculating the difference between the system resources (capacity) and the reserved resources (reserved) of the node to obtain the node resource topology information of the node, for example, the node resource topology information means that the resources allocatable = capacity - reserved that can be used by the node.

[0197] S502: The client configures a resource scheduling request.

[0198] The following describes how to configure resource affinity requirements under NUMA in the instance dimension (S502).

[0199] (1) Grouping and sharding configuration under the instance dimension.

[0200] See Figure 6a , Figure 6a This is a configuration flow chart of an instance dimension provided by the embodiment of this application. For the abstract model of the existence state service, the existence state service here refers to the service of unique identification and state management. Typical state services include: database, middleware, public system, etc. Figure 6a As shown, ① first, for an instance, the pods it includes can be grouped. The grouping rule here is generally to group pods of the same type into one group, that is, a group of pods with the same service is grouped into one group. For example, this application can specifically divide an instance into group A and group B. ② Secondly, under any group, it can be further divided into multiple shards (grouping of stateful services can further perform shard division), for example, group A can be divided into shard-0 and shard-1; in particular, when the pods in the group have no sharding requirements, they can degenerate into one shard, that is, the group and the shard are equal, such as group B is divided into a shard-0. ③ Finally, the pods that provide services can be directly run under each shard.

[0201] (2) Example of database instance configuration.

[0202] See Figure 6b , Figure 6b This is a configuration flow chart of a database instance provided in the embodiment of this application. Figure 6bAs shown, taking the clickhouse_instance database instance as an example, there is a pod for the clickhosue database itself under this instance, and there is also a pod in the zookeeper database to provide synchronization services for clickhosue; therefore, in this case, the pods under this instance can be divided into two groups: clickhosue group and zookeeper group. Among them, for the clickhouse group, there are two shards (s0 and s1) under it, and there are two pods in each shard, that is, the s0 shard includes the two pods ck-0 and ck-1. For the zookeeper group, there is one shard (that is, shard s0) under it, and there are three pods (zk-0, zk-1, zk-2) under this shard s0. In this case, the shards and groups are equal.

[0203] (3) This application supports configuring different NUMA affinity strategies for the grouping dimension under the instance.

[0204] NodeResourceTopologyManagerPolicy (node ​​resource topology management policy), supports:

[0205] "None": indicates that the group does not require NUMA affinity constraints;

[0206] "LeastNumaNode": indicates that pods within a group are optimally allocated to NUMA based on available resources.

[0207] "SingleNumaNode": indicates that the pods in the group can only be allocated to one and only one NUMA node;

[0208] NodeResourceTopologyScope, which supports:

[0209] "container": indicates that multiple containers within a pod are independent of each other, and resource allocation under the NUMA architecture is performed based on the container dimension.

[0210] "pod": Indicates that all containers in a pod are treated as a whole to perform resource allocation.

[0211] CPUTopologyManagerPolicy (resource binding policy), supports:

[0212] "none": means no CPU resources need to be bound;

[0213] "exclusive": indicates that the CPU core is bound and exclusive (other pods cannot use this core);

[0214] "numa": means binding all shared cores under the specified NUMA node;

[0215] "fixed": means binding the CPU core but not exclusively (allowing other pods to share the core).

[0216] MemoryTopologyManagerPolicy, which supports:

[0217] "none": means no memory resources need to be bound;

[0218] "numa": indicates binding the memory under the specified NUMA node.

[0219] As can be seen from the above S502, the embodiment of the present application provides a configuration method for configuring NUMA affinity requirements from the instance dimension, and according to the above configuration method, NUMA topology affinity requirements can be configured in the instance dimension, so that during the node resource scheduling process, pods that do not require NUMA affinity can be effectively distinguished, thereby reducing resource waste.

[0220] S503: The scheduling component receives the resource scheduling request sent by the client.

[0221] S504: The scheduling component performs resource analysis on the N nodes based on the resource scheduling request and the node resource topology information set, and determines a target node from the N nodes.

[0222] S505: The scheduling component pre-allocates resources in the target node based on the resource scheduling request, and generates resource annotation information of the resource object based on the resource allocation result of the target node.

[0223] The resource allocation process performed by the scheduling component is described in detail below with reference to the accompanying drawings (S504-S505).

[0224] See Figure 7 , Figure 7 This is a schematic diagram of a scheduling process of a scheduling component provided in an embodiment of the present application. Figure 7 As shown in the figure, the core goal of the scheduling component during resource scheduling is to filter out nodes that do not meet the requirements based on the NRT resource objects and the resource requirements of the pod, and then calculate the NUMA node allocation on the nodes that meet the requirements. Finally, the optimal node is selected and the NUMA node allocation result is written to the pod annotation. The specific stages are as follows:

[0225] (1) Filter.

[0226] Filter out nodes that do not conform to SingleNumaNode;

[0227] For LeastNumaNode, the most appropriate NUMA node (target resource node combination) on the node is selected. The entire process is as follows: First, all NUMA node combinations are exhaustively enumerated (with pruning optimization); then, combined with cache information, the optimal NUMA node combination is selected according to the preset selection criteria. The preset selection criteria are used to indicate the selection method based on the minimum number of NUMA nodes (commonly known as the "narrowest" selection method); alternatively, the preset selection criteria can be used to indicate the selection method based on the shortest distance between multiple NUMA node combinations (commonly known as the "closest" selection method).

[0228] (2) Score.

[0229] None: all nodes have the same score;

[0230] SingleNumaNode: provides two strategies, tending to fill the NUMA node more fully or less sparsely;

[0231] LeastNumNode: Two-level scoring strategy, including the first-level scoring sub-strategy and the second-level scoring sub-strategy: First, give priority to high scores to the node with the least number of NUMA nodes (the best case is to have only one NUMA node); second, give priority to high scores to the node with the closest distance.

[0232] (3) Reserve.

[0233] Reserve the NUMA node resources requested by the pod in the cache shared by the scheduling plugin. For example, after determining the optimal NUMA node resources (such as NUMA node 0), you can reserve the resources in NUMA node 0 in the cache to facilitate subsequent resource binding.

[0234] (4) Release (UnReserve).

[0235] If scheduling fails, release the NUMA node resources reserved in the cache for the pod request.

[0236] (5) Preallocation (Prebind).

[0237] Write the determined NUMA node information to the pod's annotation.

[0238] As can be seen from the above S504-S505, the scheduling component can specifically perceive the NUMA architecture during resource scheduling, and allocate resources based on the specific NUMA architecture and scheduling requirements, thereby reasonably allocating resources and improving the rationality and accuracy of resource scheduling.

[0239] S506: The proxy component receives the resource annotation information sent by the scheduling component.

[0240] S507: The proxy component receives the scheduling component and performs resource binding processing in the target node according to the resource annotation information.

[0241] The specific process of pod resource binding performed by the proxy component is described in detail below (S507).

[0242] In one possible implementation, the process of the proxy component performing resource binding mainly includes: ① calling the proxy component in the target node to obtain the target node resource topology information of the target node; ② based on the allocated quantity of the container resource p indicated by the resource allocation field in the resource annotation information, and the available resources of each target resource node, verifying each target resource node, for example, the proxy component performs an admit verification: used to verify whether the actual topological resources on the target node meet the request of the container in the pod; ③ If the verification passes, the container runtime interface (Container Runtime Interface, CRI interface) is called to bind the container resource p in the operating system of the target node according to the resource binding policy and allocated quantity of the container resource p indicated by the policy field in the resource annotation information.

[0243] (1) Introduce several common keywords:

[0244] allocatable: Available resources on the node except for the system reserved resources (reserved);

[0245] Share cpus: public cpu pool, corresponding to exclusive cpu;

[0246] exclusive cpus: An exclusive cpu pool, each cpu inside is used by a unique container.

[0247] (2) When allocating CPU, the following four resource binding strategies are mainly supported:

[0248] none: bind all shared CPUs;

[0249] exclusive: bind exclusive cpu;

[0250] fixed: bind to the specified shared CPU (this core cannot be exclusive);

[0251] numa: Bind the shared CPU within numa (the NUMA node to be bound must contain at least the required number of CPU resources).

[0252] (3) Assume that four pods (pod-0, pod-1, pod-2, and pod-3) with a requirement of 1 cpu core are scheduled. Each pod uses the four strategies above. Pod-0, pod-1, and pod-2 are all scheduled to NUMA node 0, with cpu core 5 reserved by the system. Figure 8a , Figure 8a This is a schematic diagram of an agent perspective of an agent component provided in an embodiment of the present application. Figure 8a As shown in the figure, assume that the current node node0 includes: NUMA node 0 and NUMA node 1, and each NUMA node includes 6 CPU cores, that is, NUMA node 0 includes: cores 0-5, and NUMA node 1 includes: cores 6-11. Then, from the perspective of the Agent component, the cores used by the CPUs of different pods are as follows:

[0253] exclusive pod-0: 0 (exclusive core 0);

[0254] fixed pod-1:2 (fixed on core 2);

[0255] numa pod-2: 1-4 (shares cores 1-4, which are all on NUMA node 0);

[0256] none pod-3: 1-4, 6-11 (shares all cores on node 0, that is, in addition to cores 1-4 on NUMA node 0, there are cores 6-11 on NUMA node 1).

[0257] Furthermore, the scheduling perspective at this time is shown in Table 3 below:

[0258] Table 3. Scheduling perspective of the Scheduler component

[0259]

[0260] As shown in Table 3, the available resources allocatable in NUMA node 0 are 5 cores (i.e., 5 cores, because core 5 of the total 6 cores is reserved by the system). In addition, requested is 3 cores because only pod-0, pod-1, and pod-2 of the four requested pods are allocated to NUMA node 0. Furthermore, since the resources in NUMA node 1 are not reserved by the system, the available resources allocatable in NUMA node 1 are 6 cores.

[0261] (4) If an exclusive pod-4 is received at this time, this pod-4 uses core 1 and is also allocated to NUMA node 0. Figure 8b , Figure 8b This is a schematic diagram of another proxy perspective of a proxy component provided in an embodiment of the present application. Figure 8b As shown in the figure, the cores used by the CPUs of different pods from the perspective of the Agent component are as follows:

[0262] exclusive pod-0:0 (exclusive core 0, unchanged);

[0263] exclusive pod-4: 1 (newly added pod, exclusively occupies core 1);

[0264] fixed pod-1: 2 (fixed on core 2, unchanged);

[0265] numa pod-2: 2-4 (delete the exclusive core 1);

[0266] none pod-3: 2-4, 6-11 (delete the exclusive core 1).

[0267] (5) Furthermore, the Agent component also has the ability to continuously reconcile. The following is a detailed description of the continuous reconciliation process. Figure 9 , Figure 9 This is a flow chart of continuous tuning of a proxy component provided by an embodiment of the present application. Figure 9As shown, continuous tuning is performed to dynamically adjust resource allocation for all pods at the operating system level by reconciling the state (the state persisted to the file) and the cgroupstate (the in-memory state of the system cgroups). The goal of continuous tuning is to adjust the cgroupstate to be consistent with the state. Furthermore, persistence is required to ensure crash recovery, that is, to be able to recover the state in the file using the content recorded in the file after a system failure.

[0268] Therefore, the Agent component needs to continuously adjust Rreconcile during the resource binding process. First, the CPU core list of pods that do not bind cores may need to remove these cores due to exclusive CPU cores. Second, some binding results must be manually adjusted by modifying the pod annotations. However, the cpu_policy (resource binding policy) of these modified pods cannot be changed.

[0269] As can be seen from the above S507, the proxy component can perform specific allocation of CPU resources based on the resource annotation information of the resource object. Since the resource annotation information records the resource binding policy and allocation quantity of each container resource, during the resource binding process, the proxy component can allocate CPU resources according to the resource binding policy and allocation quantity specified by the object, thereby ensuring that after the pod is scheduled to a certain node, the NUMA resources on the node will definitely meet the pod requirements, which can improve the effectiveness and rationality of resource scheduling.

[0270] In summary, this application provides a comprehensive solution for resource decision-making and binding in the NUMA architecture for cloud-native scenarios. It can effectively address various deficiencies and problems in the current NUMA resource binding solution in Kubernetes, and has the following benefits:

[0271] 1) The scheduling component perceives the NUMA architecture and makes resource decisions: The scheduling component perceives the NUMA topology of the node and participates in resource allocation under the NUMA architecture. This effectively improves scheduling accuracy and helps find the target node that best meets the NUMA affinity requirements of the current pod. Furthermore, it ensures that after a pod is scheduled to a target node, the NUMA resources on that target node can meet the pod's resource requirements.

[0272] 2) Instance-level scope: NUMA topology affinity requirements can be configured at the instance level, effectively distinguishing pods that do not require NUMA affinity and reducing resource waste.

[0273] 3) Support for diverse CPU resource binding modes: For CPU resources, this application supports not only exclusive core binding mode and unlimited mode, but also numa binding mode and core binding but not exclusive mode;

[0274] 4) No limit on the number of NUMA nodes: This application does not limit the number of NUMA nodes and can adapt to all different NUMA architectures, with a wider range of applicable scenarios;

[0275] 5) Upgrade does not affect existing instances: Since the Agent component of this application does not involve k8s's management of the instance lifecycle, its upgrade will not affect existing instances.

[0276] The following describes the node resource processing solution provided in the embodiments of the present application.

[0277] See Figure 10 , Figure 10 This is a schematic diagram of the structure of a node resource processing device provided in an embodiment of the present application. Figure 10 As shown, the node resource processing device 1000 can be applied to the computer device mentioned in the aforementioned embodiment (such as an agent component or a scheduling component). Specifically, the node resource processing device 1000 can be a computer program (including program code) running in a computer device, for example, the node resource processing device 1000 is an application software; the node resource processing device 1000 can be used to execute the corresponding steps in the node resource processing method provided in the embodiment of the present application. In specific implementation, the node resource processing device 1000 is applied to a cloud computing platform architecture, which includes N nodes, an agent component deployed in one node, and one or more resource nodes in one node; a scheduling component is also deployed in the cloud computing platform architecture, and the scheduling component is used to perform resource scheduling on N nodes, where N is a positive integer; the node resource processing device 1000 includes:

[0278] The acquisition unit 1001 is configured to acquire a node resource topology information set upon receiving a resource scheduling request; wherein the node resource topology information set includes node resource topology information of N nodes, and the node resource topology information of any node is collected by an agent component deployed in the corresponding node and stored in the scheduling component;

[0279] The processing unit 1002 is configured to perform resource analysis on N nodes based on the resource scheduling request and the node resource topology information set, and determine a target node from the N nodes; the target node satisfies the resource requirements of the resource object requested for scheduling by the resource scheduling request;

[0280] The processing unit 1002 is further configured to pre-allocate resources in the target node based on the resource scheduling request, and generate resource annotation information of the resource object based on the resource allocation result of the target node; the resource annotation information is used to instruct the agent component in the target node to bind the allocated resource node to the resource object in the resource scheduling request in the target node;

[0281] The bound resource node is used to execute the data processing service provided by the resource object.

[0282] In a possible implementation, during the process of the agent component collecting node resource topology information of any node, the processing unit 1002 is configured to perform the following operations:

[0283] Determine the startup parameters of the agent component and configure the collection cycle of the agent component in the startup parameters of the agent component;

[0284] According to the configured collection cycle, the agent component is used to collect and process the resource information of the node to obtain the node collection information of the node; wherein the node collection information includes: the system resources and reserved resources of the node;

[0285] The difference between the system resources and the reserved resources of the node is calculated to obtain the node resource topology information of the node.

[0286] In a possible implementation, in a process of using an agent component to collect resource information of a node and obtain reserved resources of the node, the processing unit 1002 is configured to perform the following operations:

[0287] Obtaining startup parameters of the cloud computing platform architecture, parsing the startup parameters of the cloud computing platform architecture to obtain parsed parameters, and calculating the parsed parameters according to priority rules to obtain reserved resources of the node; or,

[0288] Call the remote procedure call service to collect and process information on the node and obtain the reserved resources of the node.

[0289] In one possible implementation, the processing unit 1002 performs resource analysis on N nodes based on the resource scheduling request and the node resource topology information set, and determines a target node from the N nodes to perform the following operations:

[0290] Based on the resource scheduling request and the node resource topology information set, N nodes are filtered according to the node resource topology management strategy to obtain K candidate nodes that meet the affinity constraints; K is a positive integer and K≤N;

[0291] Scoring the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes; the scoring node of any candidate node includes the score value of the corresponding candidate node;

[0292] Based on the scores of K candidate nodes, the target node is determined.

[0293] In one possible implementation, the resource objects requested for scheduling by the resource scheduling request include: X container resources, where X is a positive integer; when the node resource topology management policy is the first policy, the first policy is used to indicate that a resource node in a node needs to be used for resource scheduling; any one of the N nodes is represented as node i;

[0294] The processing unit 1002 filters the N nodes according to the node resource topology management policy based on the resource scheduling request and the node resource topology information set to obtain K candidate nodes that meet the affinity constraint, and performs the following operations:

[0295] Based on the node resource topology information set, analyze and obtain the available resource amounts corresponding to each resource node in the node i, and determine the maximum value of each available resource amount as the maximum available resource amount Y of the node;

[0296] Among the N nodes, the nodes whose maximum available resource amount Y is less than the required resource amount X are determined as nodes to be filtered;

[0297] One or more nodes to be filtered are determined from the N nodes and filtered to obtain K candidate nodes that meet the affinity constraint.

[0298] In one possible implementation, when the node resource topology management policy is the second policy, the second policy is used to indicate that one or more resource nodes in a node need to be used for resource scheduling; any candidate node among the K candidate nodes is represented as candidate node j; and the processing unit 1002 is further used to perform the following operations:

[0299] Based on the node resource topology information of the candidate node j, the resource amount X indicated by the resource scheduling request is allocated to obtain at least one resource node combination included in the candidate node j; a resource node combination includes one or more resource nodes;

[0300] According to the preset selection conditions, the target resource node combination corresponding to the candidate node j is selected from each resource node combination.

[0301] In one possible implementation, the scoring strategy includes a first scoring substratum and a second scoring substratum, and the scoring result of any candidate node includes a first scoring value and a second scoring value. The processing unit 1002 scores the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes, which are used to perform the following operations:

[0302] Get K target resource node combinations of K candidate nodes; one candidate node corresponds to one target resource node combination;

[0303] According to the first scoring strategy, the number of resource nodes in the K target resource node combinations is calculated to obtain the first scoring value of each candidate node;

[0304] According to the second scoring sub-strategy, the resource node distances in the K target resource node combinations are calculated respectively to obtain the second scoring value of each candidate node.

[0305] In one possible implementation, the processing unit 1002 determines a target node based on the scores of the K candidate nodes, and performs the following operations:

[0306] Based on the first scores of the K candidate nodes, determining one or more target candidate nodes with the largest first scores from the K candidate nodes;

[0307] If the number of target candidate nodes is 1, the target candidate node is determined as the target node;

[0308] If the number of target candidate nodes is greater than 1, based on the second scores of the respective target candidate nodes, a target candidate node with the largest second score is determined as the target node.

[0309] In one possible implementation, the node resource topology information of any node includes the available resource amount of each resource node in the node; the resource object requested for scheduling by the resource scheduling request includes: X container resources, any container resource among the X container resources is represented as container resource p; the resource allocation result includes the allocated quantity of container resource p in one or more resource nodes;

[0310] The processing unit 1002 pre-allocates resources in the target node based on the resource scheduling request, and is configured to perform the following operations:

[0311] Determine a target resource node combination corresponding to the target node, where the target resource node combination includes one or more target resource nodes to be used;

[0312] According to the X container resources to be scheduled in the resource scheduling request and the available resources of each target resource node, the corresponding allocation quantity of the container resource p in each target resource node is determined.

[0313] In one possible implementation, the resource annotation information includes multiple fields, including a policy field and a resource allocation field. The processing unit 1002 generates resource annotation information of the resource object based on the resource allocation result of the target node, and is used to perform the following operations:

[0314] Parse the resource scheduling request, obtain the resource binding policy of the container resource p, and write the resource binding policy of the container resource p into the policy field of the resource annotation information;

[0315] The allocated quantity of the container resource p in each target resource node is written into the resource allocation field of the resource annotation information.

[0316] In a possible implementation, the processing unit 1002 is further configured to perform the following operations:

[0317] Calling the proxy component in the target node to obtain the target node resource topology information of the target node;

[0318] Verify each target resource node based on the allocated quantity of the container resource p indicated by the resource allocation field in the resource annotation information and the available resources of each target resource node;

[0319] If the verification passes, the container resource p is bound in the operating system of the target node according to the resource binding policy and allocation quantity of the container resource p indicated by the policy field in the resource annotation information, and the container runtime interface is called.

[0320] In a possible implementation, the target node includes one or more target resource nodes, and any target resource node includes at least one multi-core processor; the resource binding strategy includes: a first binding strategy, a second binding strategy, a third binding strategy, and a fourth binding strategy; wherein,

[0321] The first binding policy is used to indicate that the container resource p is bound to all multi-core processors in the target node;

[0322] The second binding policy is used to indicate that the container resource p is exclusively bound to the first multi-core processor in the target node;

[0323] The third binding policy is used to indicate that the container resource p is bound to the second multi-core processor in the target node, and the second multi-core processor is not exclusively bound;

[0324] The fourth binding policy is used to indicate that the container resource p is bound to the target resource node of the target node.

[0325] In an embodiment of the present application, when a resource scheduling request is received, a node resource topology information set is obtained, which includes the node resource topology information of N nodes. The node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component; based on the resource scheduling request and the node resource topology information set, resource analysis and processing are performed on the N nodes, and a target node is determined from the N nodes, which target node meets the resource requirements of the resource object requested to be scheduled by the resource scheduling request; based on the resource scheduling request, resources in the target node are pre-allocated, and resource annotation information of the resource object is generated based on the resource allocation result of the target node, and the resource annotation information is used to indicate: the agent component in the target node binds the allocated resource node with the resource object in the resource scheduling request in the target node; wherein, the bound resource node is scheduled as the resource object to perform data processing. It can be seen that, on the one hand, the agent component in each node in the present application can collect the node resource topology information of the corresponding node and store it in the scheduling component, so that the scheduling component can specifically perceive the topological structure of the node resources, so that when the scheduling component performs resource allocation, it can reasonably allocate resources according to the topological structure of the node resources; on the other hand, when the agent component is responsible for resource allocation, it can not only take into account the topological structure of the node resources, but also meet the resource scheduling requirements, further improve the rationality and accuracy of resource allocation, thereby improving the resource scheduling effect.

[0326] See Figure 11 , Figure 11 1 is a structural diagram of a computer device provided in an embodiment of the present application. The computer device 1100 is used to execute the steps performed by the proxy component or the scheduling component in the aforementioned method embodiment. The computer device 1100 includes: one or more processors 1101; one or more input devices 1102, one or more output devices 1103 and a memory 1104. The above-mentioned processors 1101, input devices 1102, output devices 1103 and memory 1104 are connected via a bus 1105. Among them, the memory 1104 is used to store computer programs, and the computer programs include program instructions. Specifically, the processor 1101 is used to call the program instructions stored in the memory 1104 to perform the following operations:

[0327] When a resource scheduling request is received, a node resource topology information set is obtained; wherein the node resource topology information set includes node resource topology information of N nodes, and the node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component;

[0328] Based on the resource scheduling request and the node resource topology information set, perform resource analysis on N nodes and determine the target node from the N nodes; the target node meets the resource requirements of the resource object requested by the resource scheduling request;

[0329] Pre-allocating resources in the target node based on the resource scheduling request, and generating resource annotation information of the resource object based on the resource allocation result of the target node; the resource annotation information is used to instruct the agent component in the target node to bind the allocated resource node with the resource object in the resource scheduling request in the target node;

[0330] The bound resource node is used to execute the data processing service provided by the resource object.

[0331] In a possible implementation, during the process of the agent component collecting node resource topology information of any node, the processor 1101 is configured to perform the following operations:

[0332] Determine the startup parameters of the agent component and configure the collection cycle of the agent component in the startup parameters of the agent component;

[0333] According to the configured collection cycle, the agent component is used to collect and process the resource information of the node to obtain the node collection information of the node; wherein the node collection information includes: the system resources and reserved resources of the node;

[0334] The difference between the system resources and the reserved resources of the node is calculated to obtain the node resource topology information of the node.

[0335] In a possible implementation, in a process of using an agent component to collect resource information of a node to obtain reserved resources of the node, the processor 1101 is configured to perform the following operations:

[0336] Obtaining startup parameters of the cloud computing platform architecture, parsing the startup parameters of the cloud computing platform architecture to obtain parsed parameters, and calculating the parsed parameters according to priority rules to obtain reserved resources of the node; or,

[0337] Call the remote procedure call service to collect and process information on the node and obtain the reserved resources of the node.

[0338] In one possible implementation, the processor 1101 performs resource analysis on N nodes based on the resource scheduling request and the node resource topology information set, and determines a target node from the N nodes to perform the following operations:

[0339] Based on the resource scheduling request and the node resource topology information set, N nodes are filtered according to the node resource topology management strategy to obtain K candidate nodes that meet the affinity constraints; K is a positive integer and K≤N;

[0340] Scoring the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes; the scoring node of any candidate node includes the score value of the corresponding candidate node;

[0341] Based on the scores of K candidate nodes, the target node is determined.

[0342] In one possible implementation, the resource objects requested for scheduling by the resource scheduling request include: X container resources, where X is a positive integer; when the node resource topology management policy is the first policy, the first policy is used to indicate that a resource node in a node needs to be used for resource scheduling; any one of the N nodes is represented as node i;

[0343] Based on the resource scheduling request and the node resource topology information set, the processor 1101 filters the N nodes according to the node resource topology management policy to obtain K candidate nodes that meet the affinity constraint, and performs the following operations:

[0344] Based on the node resource topology information set, analyze and obtain the available resource amounts corresponding to each resource node in the node i, and determine the maximum value of each available resource amount as the maximum available resource amount Y of the node;

[0345] Among the N nodes, the nodes whose maximum available resource amount Y is less than the required resource amount X are determined as nodes to be filtered;

[0346] One or more nodes to be filtered are determined from the N nodes and filtered to obtain K candidate nodes that meet the affinity constraint.

[0347] In one possible implementation, when the node resource topology management policy is the second policy, the second policy is used to indicate that one or more resource nodes in a node need to be used for resource scheduling; any candidate node among the K candidate nodes is represented as candidate node j; and the processor 1101 is further used to perform the following operations:

[0348] Based on the node resource topology information of the candidate node j, the resource amount X indicated by the resource scheduling request is allocated to obtain at least one resource node combination included in the candidate node j; a resource node combination includes one or more resource nodes;

[0349] According to the preset selection conditions, the target resource node combination corresponding to the candidate node j is selected from each resource node combination.

[0350] In one possible implementation, the scoring strategy includes a first scoring substratum and a second scoring substratum, and the scoring result of any candidate node includes a first scoring value and a second scoring value. The processor 1101 scores the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes, which are used to perform the following operations:

[0351] Get K target resource node combinations of K candidate nodes; one candidate node corresponds to one target resource node combination;

[0352] According to the first scoring strategy, the number of resource nodes in the K target resource node combinations is calculated to obtain the first scoring value of each candidate node;

[0353] According to the second scoring sub-strategy, the resource node distances in the K target resource node combinations are calculated respectively to obtain the second scoring value of each candidate node.

[0354] In one possible implementation, the processor 1101 determines a target node based on the scores of the K candidate nodes, and performs the following operations:

[0355] Based on the first scores of the K candidate nodes, determining one or more target candidate nodes with the largest first scores from the K candidate nodes;

[0356] If the number of target candidate nodes is 1, the target candidate node is determined as the target node;

[0357] If the number of target candidate nodes is greater than 1, based on the second scores of the respective target candidate nodes, a target candidate node with the largest second score is determined as the target node.

[0358] In one possible implementation, the node resource topology information of any node includes the available resource amount of each resource node in the node; the resource object requested for scheduling by the resource scheduling request includes: X container resources, any container resource among the X container resources is represented as container resource p; the resource allocation result includes the allocated quantity of container resource p in one or more resource nodes;

[0359] The processor 1101 pre-allocates resources in the target node based on the resource scheduling request, and is configured to perform the following operations:

[0360] Determine a target resource node combination corresponding to the target node, where the target resource node combination includes one or more target resource nodes to be used;

[0361] According to the X container resources to be scheduled in the resource scheduling request and the available resources of each target resource node, the corresponding allocation quantity of the container resource p in each target resource node is determined.

[0362] In one possible implementation, the resource annotation information includes multiple fields, including a policy field and a resource allocation field. The processor 1101 generates resource annotation information of the resource object based on the resource allocation result of the target node, and is used to perform the following operations:

[0363] Parse the resource scheduling request, obtain the resource binding policy of the container resource p, and write the resource binding policy of the container resource p into the policy field of the resource annotation information;

[0364] The allocated quantity of the container resource p in each target resource node is written into the resource allocation field of the resource annotation information.

[0365] In a possible implementation, the processor 1101 is further configured to perform the following operations:

[0366] Calling the proxy component in the target node to obtain the target node resource topology information of the target node;

[0367] Verify each target resource node based on the allocated quantity of the container resource p indicated by the resource allocation field in the resource annotation information and the available resources of each target resource node;

[0368] If the verification passes, the container resource p is bound in the operating system of the target node according to the resource binding policy and allocation quantity of the container resource p indicated by the policy field in the resource annotation information, and the container runtime interface is called.

[0369] In a possible implementation, the target node includes one or more target resource nodes, and any target resource node includes at least one multi-core processor; the resource binding strategy includes: a first binding strategy, a second binding strategy, a third binding strategy, and a fourth binding strategy; wherein,

[0370] The first binding policy is used to indicate that the container resource p is bound to all multi-core processors in the target node;

[0371] The second binding policy is used to indicate that the container resource p is exclusively bound to the first multi-core processor in the target node;

[0372] The third binding policy is used to indicate that the container resource p is bound to the second multi-core processor in the target node, and the second multi-core processor is not exclusively bound;

[0373] The fourth binding policy is used to indicate that the container resource p is bound to the target resource node of the target node.

[0374] In an embodiment of the present application, when a resource scheduling request is received, a node resource topology information set is obtained, which includes the node resource topology information of N nodes. The node resource topology information of any node is collected by the agent component deployed in the corresponding node and stored in the scheduling component; based on the resource scheduling request and the node resource topology information set, resource analysis and processing are performed on the N nodes, and a target node is determined from the N nodes, which target node meets the resource requirements of the resource object requested to be scheduled by the resource scheduling request; based on the resource scheduling request, resources in the target node are pre-allocated, and resource annotation information of the resource object is generated based on the resource allocation result of the target node, and the resource annotation information is used to indicate: the agent component in the target node binds the allocated resource node with the resource object in the resource scheduling request in the target node; wherein, the bound resource node is scheduled as the resource object to perform data processing. It can be seen that, on the one hand, the agent component in each node in the present application can collect the node resource topology information of the corresponding node and store it in the scheduling component, so that the scheduling component can specifically perceive the topological structure of the node resources, so that when the scheduling component performs resource allocation, it can reasonably allocate resources according to the topological structure of the node resources; on the other hand, when the agent component is responsible for resource allocation, it can not only take into account the topological structure of the node resources, but also meet the resource scheduling requirements, further improve the rationality and accuracy of resource allocation, thereby improving the resource scheduling effect.

[0375] In addition, it should be noted here that: the embodiment of the present application also provides a computer storage medium, and a computer program is stored in the computer storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the method in the corresponding embodiment above, so it will not be described in detail here. For technical details not disclosed in the computer storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located in one place, or, executed on multiple computer devices distributed in multiple locations and interconnected by a communication network.

[0376] According to one aspect of the present application, embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, so that the computer device can perform the methods described in the corresponding embodiments above. Therefore, these methods will not be described in detail here.

[0377] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD) or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0378] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A node resource processing method, characterized in that: Applied to a cloud computing platform architecture, the cloud computing platform architecture includes N nodes, an agent component is deployed in each node, and each node includes one or more resource nodes; The cloud computing platform architecture further includes a scheduling component configured to schedule resources for the N nodes, where N is a positive integer. The method includes: When a resource scheduling request is received, a node resource topology information set is obtained; wherein the node resource topology information set includes node resource topology information of the N nodes, and the node resource topology information of any of the nodes is collected by an agent component deployed in the corresponding node and stored in the scheduling component; Based on the resource scheduling request and the node resource topology information set, performing resource analysis processing on the N nodes, and determining a target node from the N nodes; the target node meets the resource requirements of the resource object requested to be scheduled by the resource scheduling request; Pre-allocating resources in the target node based on the resource scheduling request, and generating resource annotation information of the resource object based on the resource allocation result of the target node; the resource annotation information is used to instruct: the agent component in the target node to bind the allocated resource node with the resource object in the resource scheduling request in the target node; The bound resource node is used to execute the data processing service provided by the resource object.

2. The method according to claim 1, wherein The process of the agent component collecting node resource topology information of any node includes: Determining startup parameters of the proxy component, and configuring a collection period of the proxy component in the startup parameters of the proxy component; According to the configured collection cycle, the agent component is used to collect resource information of the node to obtain node collection information of the node; wherein the node collection information includes: system resources and reserved resources of the node; The difference between the system resources and the reserved resources of the node is calculated to obtain node resource topology information of the node.

3. The method according to claim 2, wherein The process of using the proxy component to collect and process resource information of a node to obtain the reserved resources of the node includes: Obtaining startup parameters of the cloud computing platform architecture, parsing the startup parameters of the cloud computing platform architecture to obtain parsed parameters, and calculating the parsed parameters according to priority rules to obtain reserved resources of the node; or, The remote procedure call service is called to collect information about the node and obtain the reserved resources of the node.

4. The method according to claim 1, wherein The performing resource analysis and processing on the N nodes based on the resource scheduling request and the node resource topology information set, and determining a target node from the N nodes, includes: Based on the resource scheduling request and the node resource topology information set, filtering the N nodes according to the node resource topology management policy to obtain K candidate nodes that meet the affinity constraint; K is a positive integer and K≤N; Scoring the K candidate nodes according to the scoring strategy to obtain scoring results of the K candidate nodes; the scoring node of any candidate node includes the score value of the corresponding candidate node; A target node is determined based on the scores of the K candidate nodes.

5. The method according to claim 4, wherein The resource objects requested for scheduling by the resource scheduling request include: X container resources, where X is a positive integer; the node resource topology management strategy is a first strategy, where the first strategy is used to indicate that a resource node in a node needs to be used for resource scheduling; any one of the N nodes is represented as node i; The filtering process of the N nodes based on the resource scheduling request and the node resource topology information set according to the node resource topology management strategy to obtain K candidate nodes that meet the affinity constraint includes: Based on the node resource topology information set, analyzing and obtaining the available resource amounts corresponding to each resource node in the node i, and determining the maximum value of each available resource amount as the maximum available resource amount Y of the node; Among the N nodes, determine the nodes whose maximum available resource amount Y is less than the required resource amount X as nodes to be filtered; One or more nodes to be filtered are determined from the N nodes and filtered to obtain K candidate nodes that meet the affinity constraint.

6. The method according to claim 4, wherein The node resource topology management strategy is a second strategy, and the second strategy is used to indicate that one or more resource nodes in a node need to be used for resource scheduling; any candidate node among the K candidate nodes is represented as candidate node j; and the method further includes: Based on the node resource topology information of the candidate node j, resource allocation is performed on the amount of resources to be scheduled X indicated by the resource scheduling request to obtain at least one resource node combination included in the candidate node j; a resource node combination includes one or more resource nodes; According to the preset selection conditions, the target resource node combination corresponding to the candidate node j is selected from each resource node combination.

7. The method according to claim 6, wherein The scoring strategy includes a first scoring sub-strategy and a second scoring sub-strategy, and the scoring result of any candidate node includes a first scoring value and a second scoring value; the scoring process of the K candidate nodes according to the scoring strategy to obtain the scoring results of the K candidate nodes includes: Obtain K target resource node combinations of the K candidate nodes; one candidate node corresponds to one target resource node combination; Calculate the number of resource nodes in the K target resource node combinations according to the first scoring sub-strategy to obtain a first scoring value for each candidate node; The resource node distances in the K target resource node combinations are calculated respectively according to the second scoring sub-strategy to obtain a second scoring value for each candidate node.

8. The method according to claim 7, wherein Determining the target node based on the scores of the K candidate nodes includes: Based on the first scores of the K candidate nodes, determining one or more target candidate nodes with the largest first scores from the K candidate nodes; If the number of the target candidate nodes is 1, the target candidate node is determined as the target node; If the number of the target candidate nodes is greater than 1, based on the second scores of the respective target candidate nodes, a target candidate node with the largest second score is determined as the target node.

9. The method according to claim 1, wherein The node resource topology information of any node includes the available resource amount of each resource node in the node; the resource object requested for scheduling by the resource scheduling request includes: X container resources, any container resource of the X container resources is represented as container resource p; the resource allocation result includes the allocated quantity of the container resource p in one or more resource nodes; The pre-allocating resources in the target node based on the resource scheduling request includes: Determine a target resource node combination corresponding to the target node, where the target resource node combination includes one or more target resource nodes to be used; According to the X container resources to be scheduled in the resource scheduling request and the available resource amount of each target resource node, the corresponding allocation quantity of the container resource p in each target resource node is determined.

10. The method according to claim 9, wherein The resource annotation information includes multiple fields, including: a policy field and a resource allocation field; the resource annotation information of the resource object generated based on the resource allocation result of the target node includes: Parsing the resource scheduling request to obtain a resource binding policy for the container resource p, and writing the resource binding policy for the container resource p into the policy field of the resource annotation information; The corresponding allocation quantity of the container resource p in each target resource node is written into the resource allocation field of the resource annotation information.

11. The method according to claim 10, wherein The method further comprises: Calling the proxy component in the target node to obtain target node resource topology information of the target node; Verify each target resource node based on the allocated quantity of the container resource p indicated by the resource allocation field in the resource annotation information and the available resource quantity of each target resource node; If the verification passes, the container resource p is bound in the operating system of the target node according to the resource binding policy and allocation quantity of the container resource p indicated by the policy field in the resource annotation information, and the container runtime interface is called.

12. The method according to claim 11, wherein The target node includes one or more target resource nodes, any of which includes at least one multi-core processor; the resource binding strategy includes: a first binding strategy, a second binding strategy, a third binding strategy, and a fourth binding strategy; wherein, The first binding policy is used to indicate that the container resource p is bound to all multi-core processors in the target node; The second binding policy is used to indicate that the container resource p is exclusively bound to the first multi-core processor in the target node; The third binding policy is used to indicate that the container resource p is bound to the second multi-core processor in the target node, and the second multi-core processor is not exclusively bound; The fourth binding policy is used to instruct the container resource p to be bound to the target resource node of the target node.

13. A node resource processing device, characterized in that: Applied to a cloud computing platform architecture, the cloud computing platform architecture includes N nodes, an agent component is deployed in each node, and each node includes one or more resource nodes; The cloud computing platform architecture further includes a scheduling component configured to schedule resources for the N nodes, where N is a positive integer. The device includes: an acquiring unit, configured to acquire a node resource topology information set upon receiving a resource scheduling request; wherein the node resource topology information set includes node resource topology information of the N nodes, and the node resource topology information of any of the nodes is collected by an agent component deployed in the corresponding node and stored in the scheduling component; a processing unit configured to perform resource analysis processing on the N nodes based on the resource scheduling request and the node resource topology information set, and determine a target node from the N nodes; the target node satisfies the resource requirements of the resource object requested to be scheduled by the resource scheduling request; The processing unit is further configured to pre-allocate resources in the target node based on the resource scheduling request, and generate resource annotation information of the resource object based on the resource allocation result of the target node; the resource annotation information is used to instruct the proxy component in the target node to bind the allocated resource node with the resource object in the resource scheduling request in the target node; The bound resource node is used to execute the data processing service provided by the resource object.

14. A computer device, characterized in that: include: storage devices and processors; a memory storing one or more computer programs; A processor, configured to load the one or more computer programs to implement the node resource processing method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the node resource processing method according to any one of claims 1 to 12.

16. A computer program product, characterized in that The computer program product comprises a computer program, and the computer program is suitable for being loaded by a processor and executing the node resource processing method according to any one of claims 1 to 12.

Citation Information

Cited By

  • Distributed computing resource dynamic allocation method and system, electronic equipment and storage medium

    CN121092332A