Resource binding method and device based on dynamic isolation rule, equipment and medium

Through the resource binding method based on dynamic isolation rules, the problem that resource binding and isolation configuration in the existing technology cannot adapt to the dynamic requirements of tasks is solved, and the precise matching of resources and task requirements and the efficient operation of the system are achieved.

CN119988020APending Publication Date: 2025-05-13镁佳(北京)科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510085100.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing Linux resource binding and isolation configuration methods are based on static rules and cannot adapt to changes in dynamic tasks, resulting in inefficient resource utilization or failure in task scheduling.

Method used

The resource binding method based on dynamic isolation rules is adopted, and the resource isolation rules are constructed by obtaining the dynamic resource requirements matrix of task description information, and the resource status of candidate nodes is monitored in real time to optimize task characteristics and node characteristics, and determine the node binding path and target isolation configuration of the task.

Benefits of technology

It achieves close matching of resource allocation and task requirements, improves resource utilization efficiency, ensures the isolation of high-priority tasks and fair distribution of low-priority tasks, and improves the adaptability and operation efficiency of the system in dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988020A_ABST
    Figure CN119988020A_ABST
Patent Text Reader

Abstract

The invention relates to the field of resource scheduling, and discloses a resource binding method and device based on a dynamic isolation rule, equipment and a medium, and the method comprises the steps: obtaining a dynamic resource demand matrix corresponding to task description information in a current system; constructing a resource isolation rule corresponding to each task based on the dynamic resource demand matrix, and screening the initial nodes in the node resource pool by using the resource isolation rules to obtain a candidate node list; monitoring the resource state of each candidate node in the candidate node list, and constructing a corresponding dynamic resource state matrix according to the resource state; and analyzing the dynamic resource demand matrix and the dynamic resource state matrix to obtain optimized task features and node features, and determining a node binding path and target isolation configuration of each task according to the optimized task features and node features. The problems that system resource binding and isolation configuration cannot adapt to dynamic scenes, task dynamic requirements cannot be flexibly responded, and configuration cannot be optimized in real time are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of resource scheduling, and in particular to a resource binding method, device, equipment and medium based on dynamic isolation rules. Background Art

[0002] In the current booming fields of cloud computing, edge computing, and high-performance computing, the Linux operating system has been widely used due to its powerful performance and flexibility. In this context, the importance of its resource management and scheduling mechanism has become increasingly prominent, becoming one of the key technologies in this field.

[0003] However, most of the current Linux resource binding and isolation configuration methods are based on static rules, such as using control groups and namespaces to achieve resource division and restriction. This static method lacks the ability to flexibly respond to dynamic changes in task requirements. When resource requirements fluctuate greatly during task operation due to a sudden increase in load or an increase in concurrent tasks, the resource allocation strategy cannot be adjusted in time, which can easily lead to low resource utilization efficiency or task scheduling failure. Moreover, even if some methods combine monitoring tools to obtain system operation status data, they are limited to data monitoring and perform poorly in real-time optimization of resource binding and isolation configuration, especially in ensuring the isolation of high-priority tasks and achieving fair allocation of low-priority tasks. There are obvious shortcomings. Summary of the invention

[0004] In view of this, the embodiments of the present invention provide a resource binding method, device, equipment and medium based on dynamic isolation rules to solve the problem that the system resource binding and isolation configuration cannot adapt to dynamic scenarios, cannot flexibly respond to dynamic task requirements, and cannot optimize the configuration in real time.

[0005] In a first aspect, an embodiment of the present invention provides a resource binding method based on dynamic isolation rules, the method comprising:

[0006] Get the dynamic resource requirement matrix corresponding to the task description information in the current system;

[0007] Based on the dynamic resource demand matrix, the resource isolation rules corresponding to each task are constructed, and the initial nodes in the node resource pool are screened using the resource isolation rules to obtain a list of candidate nodes;

[0008] Monitor the resource status of each candidate node in the candidate node list, and build a corresponding dynamic resource status matrix based on the resource status;

[0009] The dynamic resource demand matrix and the dynamic resource status matrix are analyzed to obtain the optimized task characteristics and the optimized node characteristics, and the node binding path and target isolation configuration of each task are determined based on the optimized task characteristics and the optimized node characteristics.

[0010] Furthermore, the dynamic resource requirement matrix corresponding to the task description information in the current system is obtained, including:

[0011] Extracting resource requirements of each task from the task description information, wherein the resource requirements include a first resource type and a corresponding required quantity;

[0012] Constructing a resource requirement vector of the task according to the first resource type and the corresponding required quantity;

[0013] Obtain the priority and task type of the task, and construct a task description vector for each task using the priority, task type, and resource requirement vector;

[0014] Using a preset dynamic adjustment algorithm to adjust the resource requirement vector in each task description vector to obtain a dynamic resource requirement vector;

[0015] The dynamic resource requirement vectors corresponding to each task are combined into a dynamic resource requirement matrix.

[0016] Furthermore, a resource isolation rule corresponding to each task is constructed based on the dynamic resource requirement matrix, including:

[0017] Acquire the initial resource state of each initial node in the node resource pool, wherein the initial resource state includes the second resource type and the corresponding available quantity;

[0018] Calculate a dynamic weight corresponding to the constraint condition according to the second resource type and the corresponding available quantity;

[0019] Generate resource isolation rules based on constraints and corresponding dynamic weights.

[0020] Furthermore, the resource isolation rules are used to screen the initial nodes in the node resource pool to obtain a list of candidate nodes, including:

[0021] Get the initial resource status of the initial node in the node resource pool;

[0022] The resource isolation rule is used to match the initial resource state with the dynamic resource demand vector corresponding to each task in the dynamic resource demand matrix to obtain the resource matching degree between the task and the initial node;

[0023] For each initial node, calculate the ratio of the required quantity to the available quantity in each resource type, and calculate the load imbalance degree based on the ratio and the preset uniform ratio;

[0024] Calculate the first adaptation score between the task and the initial node according to the load imbalance and resource matching;

[0025] A candidate node list is screened out from the node resource pool based on the first adaptation score.

[0026] Furthermore, the dynamic resource demand matrix and the dynamic resource status matrix are analyzed to obtain optimized task characteristics and optimized node characteristics, including:

[0027] Analyze the dynamic resource demand matrix, obtain the first correlation between each task, and construct the task subgraph using the first correlation;

[0028] Analyze the dynamic resource state matrix to obtain the second association relationship between each candidate node, and use the second association relationship to construct a node subgraph;

[0029] Generate optimized task features and optimized node features based on the task subgraph and the node subgraph.

[0030] Furthermore, the node binding path and target isolation configuration of each task are determined according to the optimized task characteristics and the optimized node characteristics, including:

[0031] Calculate a second adaptation score between each task and the candidate node based on the optimized task characteristics and the optimized node characteristics;

[0032] Determine the binding nodes corresponding to each task in the candidate node list according to the second adaptation score, and obtain multiple task node groups;

[0033] Determine the initial binding path and target isolation configuration corresponding to each task based on the task node group.

[0034] Furthermore, the initial binding path and target isolation configuration corresponding to each task are determined according to the task node group, including:

[0035] Check whether there is a conflict in the initial binding path of the current task;

[0036] If there is a conflict, determine the replacement node corresponding to the task with the conflict in the candidate node list, and perform the node allocation operation in a loop until there is no conflict in the initial binding path. Use the initial binding path as the node binding path of the current task, and use the resource isolation rule corresponding to the current task as the target isolation configuration.

[0037] In a second aspect, an embodiment of the present invention provides a resource binding device based on dynamic isolation rules, the device comprising:

[0038] An acquisition module is used to obtain the dynamic resource demand matrix corresponding to the task description information in the current system;

[0039] A construction module is used to construct resource isolation rules corresponding to each task based on the dynamic resource demand matrix, and use the resource isolation rules to screen the initial nodes in the node resource pool to obtain a list of candidate nodes;

[0040] A monitoring module is used to monitor the resource status of each candidate node in the candidate node list and build a corresponding dynamic resource status matrix according to the resource status;

[0041] The analysis module is used to analyze the dynamic resource demand matrix and the dynamic resource status matrix to obtain the optimized task characteristics and the optimized node characteristics, and determine the node binding path and target isolation configuration of each task based on the optimized task characteristics and the optimized node characteristics.

[0042] In a third aspect, an embodiment of the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method of the first aspect or any corresponding embodiment thereof.

[0044] The method provided in the embodiment of the present application has the following beneficial effects:

[0045] The method provided in the embodiment of the present application can accurately grasp the actual demand of each task for various resources at the current moment by obtaining the dynamic resource demand matrix corresponding to the task description information. This provides the most basic data support for the subsequent reasonable allocation and scheduling of resources based on task requirements, so that resource allocation can closely fit the real-time dynamics of the task, avoid blind allocation of resources, and improve the matching degree of resources and task requirements, thereby improving the overall resource utilization efficiency. By constructing resource isolation rules based on the dynamic resource demand matrix, it is ensured that the formulation of rules is based on the actual needs of the task, rather than general fixed rules. These rules can effectively guarantee the isolation of resources between different tasks and prevent resource interference between tasks. Using these rules to screen the initial nodes, candidate nodes that can meet the task resource requirements and isolation requirements can be quickly located from a large number of nodes, narrowing the scope of subsequent processing, improving the efficiency and accuracy of resource binding, and providing guarantee for the stable operation of the task.

[0046] The method provided in the embodiment of the present application can obtain the real-time dynamic information of node resources in time by monitoring the resource status of candidate nodes in real time and constructing a dynamic resource status matrix. This makes it possible to not only consider the needs of tasks in the resource allocation decision-making process, but also to make more reasonable decisions in combination with the actual status of node resources. Avoid the situation where resources are over-allocated or irrationally allocated due to lack of understanding of the node resource status, and further improve the scientific nature of resource allocation and the stability of the system. By analyzing the two matrices, deep-level task and node features are excavated. The optimized features can more comprehensively and accurately reflect the actual situation of tasks and nodes. Based on these features, the node binding path and target isolation configuration are determined, and the optimal matching of tasks and node resources can be achieved, ensuring that the isolation of high-priority tasks is guaranteed, and low-priority tasks can also obtain fair resource allocation. The success rate of task scheduling is improved, and the overall resource allocation strategy of the system is optimized, which improves the adaptability and operating efficiency of the system in dynamic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 is a flow chart of a resource binding method based on dynamic isolation rules according to an embodiment of the present invention;

[0049] Figure 2 is a workflow diagram of a TND-Net network architecture according to an embodiment of the present invention;

[0050] Figure 3 is a schematic diagram of a process of binding nodes and resources according to an embodiment of the present invention;

[0051] Figure 4 is a structural block diagram of a resource binding device based on dynamic isolation rules according to an embodiment of the present invention;

[0052] Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0054] According to an embodiment of the present invention, a resource binding method, apparatus, device and medium based on dynamic isolation rules are provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0055] In this embodiment, a resource binding method based on dynamic isolation rules is provided. Figure 1 is a flow chart of a resource binding method based on dynamic isolation rules according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0056] Step S11, obtaining a dynamic resource requirement matrix corresponding to the task description information in the current system.

[0057] In the embodiment of the present application, step S11 includes the following steps A1-A5:

[0058] Step A1: extracting resource requirements of each task from task description information, wherein the resource requirements include a first resource type and a corresponding required quantity.

[0059] It should be noted that in the Linux operating system, task description information comes from multiple aspects. For example, when a user submits a task (such as running a specific application, executing a data processing script, etc.), the system will record various related information, which constitutes the task description information.

[0060] Specifically, the specific process of extracting resource requirements may include: first, it is necessary to determine the various resource types related to the task, that is, the first resource type. The first resource type includes but is not limited to: CPU resources, for example, the task requires a specific number of CPU cores, or has certain requirements on the CPU's computing power, such as requiring the CPU main frequency to reach a certain value, etc.; memory resources, the task requires a clear memory size, such as requiring 2GB of memory to run; disk resources, including disk space requirements, for example, the task needs to store a certain amount of data on the disk, requiring 100MB of disk space, and may also involve disk read and write speed requirements, etc.; network resources, such as the task requires a stable 10Mbps network bandwidth for data transmission during operation, or has specific requirements for network latency, etc.; GPU resources, if the task involves graphics processing, deep learning, etc., a specific model or performance GPU is required, as well as GPU video memory requirements, etc.

[0061] Secondly, for each resource type identified, further extract the corresponding required quantity from the task description information. This can be achieved by analyzing the task configuration file, command line parameters, system logs, etc. For example, for the image processing task, by analyzing its configuration file, it is determined that it requires the use of 4 CPU cores, 8GB of memory, 100MB of disk space for temporary storage of intermediate results, 10Mbps network bandwidth for uploading processed images, and a GPU with 2GB of video memory is required to accelerate the processing process. In this example, the resource requirements are: CPU resources, 4 cores; memory resources, 8GB; disk resources, 100MB; network resources, 10Mbps; GPU resources, video memory 2GB.

[0062] Step A2: construct a resource requirement vector of the task according to the first resource type and the corresponding required quantity.

[0063] It should be noted that the resource requirement vector is a method of expressing the requirements of a task for various resources in the form of a mathematical vector. It can present the specific requirements of a task in different resource dimensions, facilitate subsequent mathematical operations and processing, and provide a convenient data structure for resource allocation and management.

[0064] Specifically, the specific process of constructing the resource requirement vector includes: determining the dimension of the vector according to the extracted first resource type. For example, in the above-mentioned image processing task example, the first resource type includes CPU resources, memory resources, disk resources, network resources and GPU resources, so the resource requirement vector has five dimensions. According to the determined resource type order, the corresponding demand quantities are filled into the vector in sequence. Taking the image processing task as an example, the constructed resource requirement vector It can be expressed as:

[0065]

[0066] Among them, the first element of the vector represents the required quantity of CPU resources; the second element represents the required quantity of memory resources; the third element represents the required quantity of disk resources; the fourth element represents the required quantity of network resources; and the fifth element represents the required quantity of GPU resources.

[0067] In this way, the various resource requirements of the task are integrated into a vector to form the resource requirement vector of the task. This vector can intuitively reflect the specific requirements of the task for different resources, and provides a basic data representation for other steps in the subsequent resource binding and isolation configuration method based on dynamic isolation rules (such as building a task description vector, dynamically adjusting resource requirements, matching with node resources, etc.), so that the entire resource management process can be efficiently processed and optimized in a mathematical and algorithmic way.

[0068] Step A3, obtaining the priority and task type of the task, and constructing a task description vector for each task using the priority, task type and resource requirement vector.

[0069] It should be noted that in the Linux operating system, task priority can be determined in a variety of ways. For example, the priority of a task can be assigned based on factors such as its importance and urgency. Specific ways to obtain task priority may include viewing the attribute settings of the task, the relevant configuration in the system scheduling policy, etc. Task types can be divided according to the nature and characteristics of the task. Task types include computationally intensive (such as scientific computing, big data processing, and other tasks, which mainly consume CPU resources), I / O intensive (such as tasks such as database operations that frequently read and write disks, which have high requirements for disk I / O performance), and network intensive (such as tasks involving large amounts of network data transmission, which rely on network bandwidth and stability). The task type can be determined by analyzing the execution process of the task, the operations involved, and the usage pattern of system resources. For example, for image processing tasks, since they involve a large amount of pixel calculation and processing, they can be classified as computationally intensive tasks; while for file download tasks, they mainly rely on network transmission and can be classified as network intensive tasks.

[0070] Specifically, the process of constructing a task description vector may include: after obtaining the priority P of the task j 、Task Type C j and the resource demand vector constructed previously Finally, this information is combined to construct the task description vector T j :

[0071]

[0072] As an example, take the image processing task mentioned above as an example, assuming that the task priority is high (P j = high), the task type is computationally intensive (C j = compute-intensive), its resource requirement vector is Then the constructed task description vector T j It can be expressed as: j =[4, 8GB, 100MB, 10Mbps, 2GB, high, compute-intensive].

[0073] Step A4: Use a preset dynamic adjustment algorithm to adjust the resource requirement vector in each task description vector to obtain a dynamic resource requirement vector.

[0074] It should be noted that the components of the preset dynamic adjustment algorithm include scene influence factors and resource fluctuation prediction models Among them, the scenario impact factor is used to quantify the additional impact of different dynamic scenarios on resource requirements. For example, when a task runs in a high-load network environment or a high-concurrency environment, its resource requirements may change, and this factor can be obtained through actual scenario data statistics or empirical modeling. For example, for network-intensive tasks, in a high-concurrency network environment, the scenario impact factor of its network bandwidth may increase to reflect the increase in the task's demand for network bandwidth at this time. The resource fluctuation prediction model is implemented by a set of time series prediction networks, which are directly set as time series RNN (recurrent neural network) models in this application. This model is used to predict fluctuations in resource requirements and predict future trends in resource requirements based on historical data and the laws of time series. For example, by analyzing the resource usage of tasks in different time periods over the past period of time, the RNN model can predict how the task's demand for resources such as CPU and memory may change in the next period of time.

[0075] Specifically, the process of adjusting the resource requirement vector includes: for each resource requirement vector in the task description vector The resource demand vector after adjustment is It can be expressed as:

[0076] As an example, taking the image processing task as an example, assuming that the original resource requirement vector After analysis and calculation, it is determined that the current network is in a high-load state, and the scenario influencing factors of network bandwidth The part corresponding to the network bandwidth is 0.5 (indicating a 50% increase in network bandwidth demand). The resource fluctuation prediction model It is predicted that under the current situation, the number of CPU cores required will increase by 1, the memory requirement will increase by 1GB, the disk space requirement will increase by 50MB, the network bandwidth requirement will increase by 5Mbps, and the GPU memory requirement will remain unchanged. The adjusted resource requirement vector The CPU resources are: 4+1=5 cores; memory resources: 8GB+1GB=9GB; disk resources: 100MB+50MB=150MB; network resources: 10Mbps+5Mbps=15Mbps; GPU resources: 2GB (unchanged), that is

[0077] Step A5: Combine the dynamic resource requirement vectors corresponding to each task into a dynamic resource requirement matrix.

[0078] Specifically, assuming that there are m tasks in the system, the dynamic resource demand vector of each task after adjustment is These vectors can be combined into a matrix T dymamic , expressed as:

[0079]

[0080] Among them, t′ i,j It represents the specific requirements of the i-th task after dynamic adjustment in resource dimension j, such as the number of CPU cores, memory size, network bandwidth, etc.

[0081] The dynamic demand matrix provided in the embodiment of the present application provides the system with an overall view of the dynamic resource requirements of all tasks, clearly showing the requirements of each task for different resources, and is an important basis for resource binding and isolation configuration. The system can use this to combine the node resource pool status for reasonable allocation and scheduling. Moreover, because the matrix elements are dynamically adjusted, they can better adapt to changes in the system operating environment, respond to fluctuations in task requirements in a timely manner, and improve system flexibility and resource utilization efficiency. For example, when the task resource requirements change due to external factors, it can quickly reflect and provide real-time and accurate information to the resource management module to make corresponding adjustments and optimizations.

[0082] Step S12: construct resource isolation rules corresponding to each task based on the dynamic resource demand matrix, and use the resource isolation rules to screen the initial nodes in the node resource pool to obtain a candidate node list.

[0083] In the embodiment of the present application, constructing resource isolation rules corresponding to each task based on the dynamic resource demand matrix includes the following steps B1-B3:

[0084] Step B1, obtaining the initial resource status of each initial node in the node resource pool, wherein the initial resource status includes the second resource type and the corresponding available quantity.

[0085] It should be noted that in the Linux system, the node resource pool refers to the collection of all nodes available for allocation in the system (which can be understood as computing nodes or server nodes, etc.). These nodes provide various resources, such as CPU, memory, disk, network, etc., to support the operation of tasks.

[0086] Specifically, to obtain the initial resource status of the initial node, we must first identify the initial nodes in the node resource pool. These nodes can be physical servers or computing resource units in the form of virtual machines. Corresponding to the first resource type mentioned in the resource requirements of the task, the second resource type also includes common system resource types such as CPU, memory, disk, and network. For each second resource type, obtain its available quantity on the initial node. This can be obtained through the system's resource monitoring tool or related management interface. For example, through system commands or management software, it is queried that a certain node currently has 8 available CPU cores, 16GB of available memory, 500GB of available disk space, 100Mbps of available network bandwidth, etc.

[0087] Step B2: Calculate the dynamic weight corresponding to the constraint condition according to the second resource type and the corresponding available quantity.

[0088] Specifically, the formula and parameters for dynamic weight calculation are as follows:

[0089]

[0090] Among them, ω i is the resource type weight, indicating the importance or sensitivity of different resource types; j is the task priority, reflecting the importance and urgency of the task; β is the adjustment parameter of resource tension, which is used to control the rule convergence speed during task competition; R j,k is the demand of task j for resource i, which comes from the dynamic resource demand matrix constructed previously.

[0091] As an example, suppose that the CPU resource requirement of a task j is r CPU,j For 4 cores, the available CPU resources on node k are R CPU,k For 8 cores, the CPU resource type weight ω CPU Set to 2 (indicating that CPU resources are relatively important), the priority of task j is p j =3 (assuming a higher priority), and the resource intensity adjustment parameter β is set to 0.5. Substituting these values ​​into the formula, calculate the dynamic weight W corresponding to the CPU resource CPU ≈5.29. Similarly, the dynamic weight of the task on other resources (such as memory, disk, network, etc.) on the node can be calculated according to the above method.

[0092] Step B3, generating resource isolation rules according to the constraints and the corresponding dynamic weights.

[0093] Specifically, the constraints include exclusive resource constraints and shared resource constraints, which are:

[0094] Exclusive resource constraint: r i,j ≤R i,k , represents the demand r of task j for resource i i,j Cannot exceed the available amount of resources on the node R j,k This is to ensure that when tasks are assigned to nodes, there is no over-allocation of resources, that is, there is sufficient available capacity on the node to satisfy the exclusive resources requested by the task.

[0095] Shared resource constraints: Where N is the number of shared tasks. This means that for shared resources, the demand for resources by a task cannot exceed the available amount of resources on the node divided by the number of tasks sharing the resource.

[0096] On the basis of satisfying the above constraints, the resource isolation rules are generated using the calculated dynamic weight W. The formula is as follows:

[0097]

[0098] Maximize means to maximize the sum of the product of the weights of all resources and the degree of resource matching; match(r i,j , R i,k ) can be understood as the degree of match between the demand of task j for resource i and the available amount of resource i on node k. If the demand of a task for a resource is completely met by the available amount on the node, the match value may be 1 (or a higher value determined by the specific matching algorithm); if the demand is much greater than the available amount, the match value may be smaller (even 0).

[0099] Assume that for a task j, there are three resource requirements: CPU, memory, and disk. After step B2, their dynamic weights on a node k are calculated to be W CPU =3,W MEM =2,W DISK =1, and a matching algorithm is used to calculate the matching degrees of these three resources on the node. CPU =0.8, match MEM =0.6, match MEM= 0.4. According to the formula, the sum of the product of the resource matching degree and weight of the task on this node is 4. The system will perform such calculations on each node in the node resource pool to obtain the sum of the product of the resource matching degree and weight of each node for the task.

[0100] The resource isolation rule finally generated is to select the node with the largest sum as one of the candidate nodes for the task (if there are multiple tasks, other factors and rules need to be considered to determine the final resource allocation and isolation scheme). At the same time, the specific resource allocation and isolation method of the task on the node is determined according to the exclusive resource constraints and shared resource constraints, such as which resources are exclusively allocated to the task, which resources are shared with other tasks, and the proportion or quantity of sharing, etc.

[0101] In the embodiment of the present application, the initial nodes in the node resource pool are screened using the resource isolation rule to obtain a candidate node list, including the following steps C1-C5:

[0102] Step C1, obtaining the initial resource state of the initial node in the node resource pool.

[0103] Specifically, the initial resource status of the initial node in the node resource pool includes resource type and available quantity. Obtaining the initial resource status of the initial node is the basis of the entire candidate node screening process. Only by clearly knowing the availability of various resources on each initial node can we compare and match them with the dynamic resource requirements of the task in the subsequent steps, and then determine whether the node is suitable for running a specific task. This information provides key data support for the subsequent resource matching degree calculation, load imbalance degree calculation, and final candidate node screening. It enables the system to make scientific and reasonable node selection decisions based on the actual resource status and task requirements, so as to ensure that the task can run on a node with sufficient resources, improve the success rate of task execution and the overall performance of the system. For example, if a task requires a large amount of memory resources, and a certain initial node has little available memory, then in the subsequent screening process, the node is unlikely to be selected as a candidate node.

[0104] Step C2, using resource isolation rules to match the initial resource state with the dynamic resource demand vector corresponding to each task in the dynamic resource demand matrix, to obtain the resource matching degree between the task and the initial node.

[0105] Specifically, the resource isolation rule consists of two parts: exclusivity and sharing. Dynamic weight calculation is used to comprehensively consider factors such as resource type weight, task priority, and resource intensity to determine the priority and constraints of resource allocation. i,j , R i,k), calculate the resource matching degree between the initial resource state and the dynamic resource demand vector corresponding to each task in the dynamic resource demand matrix. The resource matching function is used to represent the demand r of task i for resource i i,j and the remaining amount R of resource i on node k i,k The matching degree is calculated as follows:

[0106]

[0107] For example, for task j's CPU resource requirement r CPU,j For 4 cores, the remaining amount of CPU resources on node k is R CPU,k For 6 cores, Indicates that the CPU resources on node k fully meet the needs of task i; if the remaining amount of CPU resources on node k is R CPU,k For 3 cores, It means that the CPU resources on node k can only partially meet the needs of task j.

[0108] Step C3: for each initial node, calculate the ratio of the required quantity to the available quantity in each resource type, and calculate the load imbalance degree based on the ratio and a preset uniform ratio.

[0109] Specifically, first, for each initial node k, for each resource type i, the ratio of the required quantity of the resource type to the available quantity is calculated.

[0110] Assume that there are m types of resources on node k (such as CPU, memory, disk, network, etc.), and the required quantity for resource type i is r i,j (dynamic resource demand vector from task j), the available quantity is R i,k (from the initial resource state of node k), the proportion of this resource type is:

[0111]

[0112] in, Represents the sum of available quantities of all resource types on node k.

[0113] Then, the load imbalance (R k ), the calculation formula is as follows:

[0114]

[0115] The meaning of this formula is: first calculate the difference between the proportion of each resource in the node and the ideal uniform distribution proportion, then sum the squares of these differences, and finally take the square root. If all resources on node k are evenly distributed (that is, the proportion of each resource is close), the imbalance is close to 0; on the contrary, if a certain type of resource is overused or some resources are left over, resulting in uneven resource distribution, the imbalance will increase significantly.

[0116] Step C4, calculating a first adaptation score between the task and the initial node according to the load imbalance degree and the resource matching degree.

[0117] Specifically, the first adaptation score S between task j and initial node k j,k The calculation formula is as follows:

[0118]

[0119] Among them, R i,k represents the first adaptation score between task j and node i; m is the number of resource types; α i is the weight of resource type i, indicating the importance weight of this resource type; match(r i,j , R i,k ) is the resource matching function; λ is the load balancing adjustment coefficient, which is used to control the importance of load balancing; imbalance(R k ) is the resource load imbalance degree of node k.

[0120] As an example, assume that task j has two resource requirements: CPU and memory. The CPU requirement is 4 cores and the memory requirement is 8GB. The number of CPU cores available on node k is 6, and the available memory is 12GB. The CPU resource type weight α CPU is 0.6, and the memory resource type weight is α MEM is 0.4. The calculation shows that the CPU resource matching degree is 1 and the memory resource matching degree is 1. Assume that the load imbalance (R k )=0.2, then the first adaptation score S j,k S can be calculated j,k ≈0.91.

[0121] Step C5: Filter out a candidate node list from the node resource pool based on the first adaptation score.

[0122] Specifically, after completing step C4, the first adaptation score S of each task and each initial node in the node resource pool has been obtained. j,kFor example, for task j, the first adaptation score with node k1 may be calculated to be 0.8, the first adaptation score with node k2 may be 0.6, and the first adaptation score with node k3 may be 0.9. Such calculations are performed for all nodes in the node resource pool.

[0123] The method for selecting candidate nodes may include: sorting all nodes from high to low according to the first adaptation score with the task; selecting several nodes with the highest order as candidate nodes, for example, selecting the first n (n is the preset number of candidate nodes) nodes with higher first adaptation scores. In addition, in order to further reduce the allocation failure rate, a soft constraint function g(S j,k ), the calculation formula is as follows:

[0124]

[0125] Among them, S min is the adaptation threshold, μ is the soft constraint attenuation coefficient. j,k Although it does not meet the previously selected criteria, it can also be allowed to enter the candidate node pool N if it meets the conditions after calculation by the soft constraint function. candidates (T j ).

[0126] Step S13, monitoring the resource status of each candidate node in the candidate node list, and constructing a corresponding dynamic resource status matrix according to the resource status.

[0127] It should be noted that the monitoring object is each node in the screened candidate node list, and the monitoring content covers the utilization and availability of resources such as CPU, memory, GPU, bandwidth and storage I / O, such as the CPU usage percentage and the number of idle cores, the used amount of memory and the remaining available amount, the inflow and outflow rate of network bandwidth and the available bandwidth, etc. The monitoring method is to deploy a lightweight monitoring agent on each candidate node to collect resource status information in real time, and the agent has little impact on system performance.

[0128] Specifically, for each candidate node k, the node resource state vector R is generated according to the available amount of various resources monitored in real time. k =[r k,1 , r k,2 , ..., r k,n ] T , where r k,i represents the real-time available amount of resource type i for node k. For example, the real-time available amount of CPU for node k is r k,CPU , the real-time available amount is r k,MEM wait.

[0129] Sampling frequency of node resource status Δtk Dynamically adjusted according to system load, the sampling interval is determined by the following formula:

[0130]

[0131] Among them, τ is the basic sampling period, ρ k is the load ratio of node k, and γ is the load sensitivity parameter. k When the sampling interval is larger, k It will be shortened, thereby increasing the sampling frequency, ensuring that the latest information on resource status can be obtained in a timely manner when the node load changes greatly, ensuring the high real-time performance of resource status.

[0132] In order to reduce the communication overhead in a distributed environment, an incremental update synchronization mechanism is adopted. Each node only uploads the change in resource status ΔR k , the calculation formula is:

[0133] ΔR k =R k (t)-R k (t-Δt)

[0134] Among them, R k (t) is the resource status of node k at the current moment, R k (t-Δt) is the resource status at the last sampling. If ||ΔR k If ||1 is less than the threshold ∈, the synchronization of the node is skipped to avoid unnecessary communication overhead.

[0135] The resource state vectors of all candidate nodes are summarized into a dynamic resource state matrix R, which is defined as:

[0136]

[0137] Where K represents the total number of candidate nodes.

[0138] To ensure the real-time and consistency of resource status data, an update method based on weighted timestamp is adopted. The status update weight of each node is ω t Determined by its latest update time, the weight calculation formula is:

[0139]

[0140] Where T is the current global time, t k is the last update time of node k, and λ is the time decay factor. Each resource value R′ in the resource state matrix R i,k Updated to:

[0141]

[0142] In this way, the data update process is smoothed while retaining the reference role of historical data for current decision-making. In this step, a real-time node resource monitoring mechanism is designed to complete the collection and dynamic update of node resource status in the cluster. Each node collects the utilization and availability of multi-dimensional resources such as CPU, memory, GPU, bandwidth, and storage I / O through a lightweight monitoring agent to generate a node resource status vector R k The resource collection frequency is dynamically adjusted in combination with the node load to ensure the real-time performance of resource data in high-load scenarios. To reduce the communication overhead in large-scale clusters, this step proposes a synchronization mechanism based on incremental updates, which only uploads the change in resource status ΔR k ,avoiding redundant data transmission.,The collected resource data are summarized into a dynamic resource status matrix R,,and the real-time and consistency of the matrix are ensured through a timestamp-based weighted update method,while taking into account the smoothness of historical data.

[0143] In addition, this step introduces an anomaly detection mechanism based on resource status. By calculating the abnormal value η of the node resource k , can quickly identify abnormal changes in resource usage and trigger dynamic adjustment mechanisms to ensure the stability and robustness of the system. The output of this step is: dynamic resource state matrix R, a comprehensive matrix used to describe the real-time resource state of each node, providing key input for subsequent task allocation and resource binding. Incremental resource change log ΔR k , used to record the changes in node resources and provide support for abnormal monitoring and optimization.

[0144] Step S14, analyze the dynamic resource demand matrix and the dynamic resource status matrix to obtain optimized task characteristics and optimized node characteristics, and determine the node binding path and target isolation configuration of each task based on the optimized task characteristics and optimized node characteristics.

[0145] It should be noted that this method proposes a task-node adaptive graph-level decoupling network (TND-Net) to meet the needs of dynamic resource binding and isolation configuration scenarios. Its overall idea includes: decoupling tasks and nodes into task subgraphs and node subgraphs, modeling the relationship between the two respectively and fusing features through interaction modules; learning the complex relationship between task dynamic requirements and node resource status layer by layer through multi-scale interaction design to achieve global optimization; designing a dynamic weight adjustment module with dynamic edge weights to adapt to real-time changes and solve problems such as resource competition; setting an overall loss function with multiple objectives such as comprehensive resource adaptability, load balancing, task priority and isolation conflict, and improving task binding path and isolation configuration through multi-objective optimization.

[0146] The overall architecture of the task-node adaptive graph-level decoupling network (TND-Net) consists of a task subgraph module, a node subgraph module, and a task-node interaction module.

[0147] The task subgraph module models tasks as nodes and conflict relationships as edges, and updates task features with the help of message passing, covering dynamic priorities and conflict characteristics. The node subgraph module also models nodes as nodes and resource sharing relationships as edges, optimizes node features through message passing, and displays load balancing and resource allocation status. The task-node interaction module is based on the first two modules, models dynamic adaptation relationships, calculates adaptation scores, and optimizes binding paths through two-way message passing. The final generated path can take into account the fitness, task priority, and node load balancing goals.

[0148] The specific process of this architecture, such as Figure 2 As shown in the figure, it includes: the task subgraph module first models the conflict and priority relationship between tasks, and then optimizes and generates dynamically adjusted task features; the node subgraph module then models the resource sharing and load balancing between nodes, thereby optimizing and generating load-balanced node features; then the task-node interaction module models the dynamic adaptation relationship between tasks and nodes; finally, the task-node binding path is generated and optimized based on the adaptation score. This architecture achieves efficient adaptation and binding of tasks and nodes through the collaborative work of various modules.

[0149] In the embodiment of the present application, the dynamic resource demand matrix and the dynamic resource status matrix are analyzed to obtain optimized task characteristics and optimized node characteristics, including the following steps D1-D3:

[0150] Step D1, analyzing the dynamic resource demand matrix, obtaining the first association relationship between the tasks, and constructing a task subgraph using the first association relationship.

[0151] Specifically, the input of this step is the dynamic resource demand matrix T dynamic , which contains the dynamic resource demand information of each task, such as the demand of tasks in different resource dimensions (such as the number of CPU cores, memory size, number of GPU cores, GPU memory, etc.); and task conflict information Conflict (T j , T k ), used to represent task T j and Task T k The intensity of the conflict over resources and the priority of tasks P j Among them, the quantitative formula of task conflict information is as follows:

[0152]

[0153] Among them, R is the set of all resources that need to be considered; T j,r Represents task T j Demand on resource r; Min(T j,r , T k,r) is the smaller value of the task and the resource requirement; Max(T j,r , T k,r ) is the larger value of the resource requirements of the two. This formula quantifies the intensity of the conflict between tasks by calculating the sum of the proportions of the requirements on each resource. For example, if two tasks need GPU cores or memory, the larger the intersection, the stronger the conflict.

[0154] In the task subgraph, the task is modeled as a node V of the graph structure. T , the task characteristics of each node Among them, T′ j Contains information such as the dynamic requirements of the task, P j is the task priority. The conflict relationship between tasks is modeled as edge E T , the quantized value of the edge is Conflict(T j , T k ), that is, task T j and T k The intensity of the conflict between them.

[0155] The calculation formula of edge weight is:

[0156]

[0157] Among them, SharedResource(T j , T k ) is task T j and T k The intersection number of required resources, TotalResource(T j ) is task T j of all resource requirements. The higher the value, the stronger the conflict between tasks. For example, if two tasks require more of the same GPU cores or memory, the edge weight between them will be larger.

[0158] Feature aggregation function f T The specific explanation is:

[0159]

[0160] in, Represents the feature connection operation (dimensional concatenation of vectors), and Combined into a new feature vector, W T is the weight matrix of the task subgraph, which is used to learn the projection or mapping after feature connection. It is task T j and T k The edge weight of , represents the conflict intensity or other relationship weight between the two. T Calculate Tj The task receives the k The amount of information is used to update T j characteristics.

[0161] The features of each task are gradually updated through the message passing mechanism, where the calculation formula of the message passing mechanism is as follows:

[0162]

[0163] in, It is task T j The features at the kth layer represent the dynamic requirements, priorities, and other information of the task. j ) is task T j The neighbor task set of T j For tasks connected by edges, Aggregate is a feature aggregation function, where summation is performed to integrate all neighbors’ messages and update task features.

[0164] After the above process, this module outputs the dynamic priority adjustment of each task and the conflict optimization characteristics between tasks. These outputs provide an important basis for the subsequent analysis of the adaptation relationship between tasks and nodes and the generation of task-node binding paths, so that the characteristics of tasks can better reflect their actual situation in a dynamic resource environment, including the dynamic change of priority and the optimization of conflict relationships with other tasks, which helps to allocate resources and schedule tasks more accurately.

[0165] Step D2, analyzing the dynamic resource state matrix, obtaining the second association relationship between each candidate node, and constructing a node subgraph using the second association relationship.

[0166] Specifically, the input of this step is the node resource status matrix R, which contains the resource status information of each candidate node, such as the real-time available resources such as CPU, memory, GPU, bandwidth and storage I / O on the node; as well as the node-to-node shared resource ratio Shared (N i , N j ) and the current node load Load(N i ).

[0167] In the node subgraph, the nodes are modeled as an undirected graph G N =(V N , E N ) in the node V N , the characteristics of each node Where R i Indicates the resource status of the node (such as the available amount of various resources, etc.), Load (N i ) is node N iThe shared relationship between nodes is modeled as edge E N , the quantized value of the edge is the ratio of shared resources between nodes Shared(N i , N j ), that is, node N i and N j The amount of resources shared.

[0168] Message passing mechanism: edge weights The calculation formula is:

[0169]

[0170] Among them, Shared(N i , N j ) is node N i and N j The amount of shared resources, Load(N i ) and Load(N j ) are nodes N i and N j The current load. The higher the value, the higher the ratio of shared resources between the two nodes and the lower the load pressure. For example, if two nodes share a large amount of GPU memory and their own load is relatively low, the edge weight between them will be larger.

[0171] Aggregate function f N It is expressed as:

[0172]

[0173] in, Represents the feature connection operation. and Combined into a new feature vector, W N is the weight matrix of the node subgraph, which is used to learn the relationship between node features. is node N i and N j The edge weights of f represent the degree of resource sharing between them. N Compute Node N i Received from node N j The amount of information.

[0174] The characteristics of each node are gradually updated through the message passing mechanism, where the calculation formula of the message passing mechanism is as follows:

[0175]

[0176] in, is node N jThe features at the kth layer represent the resource status, load pressure and other information of the node, N(N j ) is node N j The neighbor node set of N j There are nodes connected by edges. Aggregate is a feature aggregation function. The summation is performed here to integrate all neighbors' messages and update the node features.

[0177] Step D3, generating optimized task features and optimized node features based on the task subgraph and the node subgraph.

[0178] Specifically, limit task T j The message is only in the candidate node list N candidates (T j ) propagates within. Here N candidates (T j ) is task T j The preliminary screening result of task T j The neighbor node set of is changed to the list of candidate nodes that are initially screened. For example, for a task, after the previous screening, several candidate nodes that may be suitable for running the task are determined, then the message transmission of the task is only carried out between these candidate nodes, narrowing the calculation scope.

[0179] In the original task-node fitness S j,k Based on this, the importance weight of the candidate nodes is added Edge Weight The calculation formula becomes:

[0180]

[0181] in, is the priority weight of the candidate node, which is used to measure the node N i The ranking importance in the candidate list is calculated as:

[0182]

[0183] Among them, Rank(N i ) is the ranking of the node in the candidate list; N candidates (T j ) is the total number of candidate nodes. The higher the ranking, the greater the weight.

[0184] In the embodiment of the present application, the process of task-to-node message transmission and feature update includes:

[0185] Compute task to node messages The formula is:

[0186]

[0187] in:

[0188]

[0189] in, It is task T j Feature vector at the kth layer; is node N i Feature vector at the kth layer; It is task T j and node N i The edge weight of W indicates the degree of fit between the two. T-N is a linear transformation matrix used to project the concatenated features of tasks and nodes into the message passing space, which is a feature concatenation operation; Connect tasks and node features.

[0190] The task feature update formula is:

[0191]

[0192] Among them, σ is the activation function; W T Aggregate is an aggregation function (such as summation) used to aggregate messages from nodes to tasks; N(T j ) is task T j In this way, the task features are updated according to the messages from the nodes, incorporating the node information to obtain the optimized task features.

[0193] In the embodiment of the present application, the process of message transmission and feature update from node to task includes:

[0194] Node to Task Messages The calculation formula is:

[0195]

[0196] in:

[0197]

[0198] is node N i The feature vector at the kth layer, It is task T j The feature vector at the kth layer, is node N i and Task T j The edge weight of W represents the strength of node adaptation to the task. N→Tis a linear transformation matrix used to project the concatenated features of node tasks into the message passing space.

[0199] The node feature update formula is:

[0200]

[0201] Among them, σ is the activation function, W N is the weight matrix of node feature update, Aggregate is the aggregation function, N(N i ) is node N i The node features are updated according to the messages from the tasks, and the information of the tasks is integrated to obtain the optimized node features.

[0202] Through the two-way message transmission and feature update mechanism between tasks and nodes, optimized task features and optimized node features are finally generated based on task subgraphs and node subgraphs. These optimized features can better reflect the adaptation relationship and mutual influence between tasks and nodes, provide a more accurate and rich information basis for the subsequent generation of task-node binding paths, etc., and help achieve better resource allocation and task scheduling.

[0203] In the embodiment of the present application, determining the node binding path and target isolation configuration of each task according to the optimized task characteristics and the optimized node characteristics includes the following steps E1-E3:

[0204] Step E1, calculating a second adaptation score between each task and a candidate node based on the optimized task features and the optimized node features.

[0205] Specifically, task T j With node N i The second adaptation score S(T j ,N i ) consists of two parts, namely:

[0206]

[0207] in, It is task T j and node N i The edge weight reflects the static adaptability between the two. This edge weight may be calculated in the previous step (such as when constructing the task subgraph and node subgraph) based on some factors (such as resource sharing degree, conflict intensity, etc.). It reflects the adaptability of the task and the node in some inherent attributes to a certain extent. For example, if the demand of the task for a certain resource is related to the abundance of the resource on the node, the edge weight will reflect this relationship. is a dynamic adaptation score calculated based on task and node characteristics, where and are the feature vectors of the task and the node respectively, which are the result of message passing and feature update. This part of the dynamic adaptation score includes information such as the degree of match between the task demand and the node resource status. For example, the degree of match can be measured by calculating the cosine similarity between the task resource demand vector and the node resource status vector. For example, if the task requires a large amount of memory resources and there is sufficient available memory on the node, then this part of the score will be high, reflecting a good match between the task and the node in terms of resource demand and supply.

[0208] The method provided in the embodiment of the present application combines static adaptability and dynamic adaptability to obtain a second adaptability score that can more comprehensively and accurately evaluate the degree of adaptability between each task and the candidate node. It not only takes into account some relatively fixed relationships and attributes between tasks and nodes, but also takes into account the dynamic matching of task requirements and the real-time resource status of nodes under the current system state, thus avoiding the problem of inaccurate adaptability evaluation caused by considering only a single factor.

[0209] Step E2, determining the binding nodes corresponding to each task in the candidate node list according to the second adaptation score, and obtaining a plurality of task node groups.

[0210] Specifically, for each task T j , from the candidate node set N candidates (T j ) selects the node with the highest score as the binding node. Specifically, it is to find the node that meets the following conditions:

[0211]

[0212] Among them, S(T j , N i ) is the second adaptation score between the task and the node calculated in step E1. For example, if task T1 has three candidate nodes N1, N2, and N3, and their second adaptation scores with task T1 are calculated to be 0.7, 0.8, and 0.6 respectively, then task T1 will select node N2 as the binding node.

[0213] If multiple nodes have the same score, the node with the most remaining resources will be given priority. This is to further consider the resource abundance of the node when the adaptability is the same, so as to better meet the resource requirements of the task and the resource utilization efficiency of the system. For example, the candidate nodes N4 and N5 of task T2 have the second adaptation score of 0.9 with task T2, but the remaining resources of node N4 (such as the number of CPU cores, memory, etc.) are more than those of node N5, then node N4 will be selected as the binding node of task T2.

[0214] Through the above method, a corresponding binding node is determined for each task, so that multiple task node groups are obtained, such as (T1, N2), (T2, N4), etc. These task node groups represent the corresponding relationship between the tasks and the binding nodes.

[0215] Step E3, determining the initial binding path and target isolation configuration corresponding to each task according to the task node group.

[0216] Specifically, the initial binding path B contains the binding nodes and allocated resources of each task, namely:

[0217] B={(T j ,N(T j ), Allocated Resources)|T j ∈T}

[0218] Among them, T j is the task, N(T j ) is the binding node of the task, and Allocated Resources is the resources allocated to the task. The resource allocation calculation method is:

[0219]

[0220] Among them, T′ j It is task T j The resource demand vector of is the node N(T j )’s remaining available resources.

[0221] After determining the task and node binding and resource allocation, it is necessary to determine the isolation configuration of each task target based on the task characteristics (such as exclusivity, resource sharing requirements, etc.) and node resource isolation rules (such as exclusivity and sharing constraints). For example, for exclusive GPU tasks, the GPU resources of the bound node are exclusive; for sharable memory tasks, the memory sharing method and restrictions are determined according to the sharing constraints to achieve reasonable isolation and utilization of resources, and ensure smooth execution of tasks and system stability.

[0222] In the embodiment of the present application, step E3 includes the following steps E31-E32:

[0223] Step E31, detecting whether there is a conflict in the initial binding path of the current task.

[0224] Specifically, for each node N i , calculate the cumulative usage of resources:

[0225]

[0226] Among them, T(N i) is currently bound to node N i Task collection; Allocated Resources (T j , N i ) is for task T j At node N i resources allocated on it.

[0227] Check if any resource exceeds the node limit, that is:

[0228]

[0229] Among them, R i,available,r Is the available amount of resources on the node.

[0230] Step E32, if there is a conflict, determine the replacement node corresponding to the task with the conflict in the candidate node list, and execute the node allocation operation in a loop until there is no conflict in the initial binding path, use the initial binding path as the node binding path of the current task, and use the resource isolation rule corresponding to the current task as the target isolation configuration.

[0231] Specifically, if a resource conflict is detected on a node, a low-priority task recovery measure is taken. i The tasks are sorted by priority P j Sort from low to high, and then recycle (unbind) the lowest priority task until the conflict is resolved. For example, if tasks T3 (priority 3), T4 (priority 2), and T5 (priority 1) are bound to node N2 and there is a resource conflict, unbind task T5 first and check whether the conflict is resolved. If not, unbind task T4 until the conflict is resolved.

[0232] Task T to be recycled j Redistribute to other candidate nodes, that is:

[0233]

[0234] Among them, N candidates (T j )\N i Represents task T j The candidate node set is the node N with the current conflict removed i ; is the candidate node N i available resources; The resources used by the node.

[0235] For tasks that require exclusive resources, check whether there are other tasks using the resource on the bound node, that is:

[0236]

[0237] If an exclusive conflict is detected, reselect the candidate node:

[0238]

[0239] Among them, S(T j , N i ) is task T j With node N i The adaptation score.

[0240] According to the task priority P(T j ) and node load, select the isolation mode. Isolation modes include Exclusive, Shared, and Flexible.

[0241] The specific selection logic is:

[0242]

[0243] Then, specific resource constraints are determined based on different resources, such as CPU binding mode (binding to a specific CPU core or setting a maximum usage ratio), GPU core locking (exclusive or shared and setting time slices), memory limits (allocating a fixed amount of memory or setting a maximum upper limit), network isolation (allocating a separate network namespace), etc.

[0244] Finally, the isolation rules for each task are output:

[0245] C={(T j , Isolation Mode, Resource Constraints)|T j ∈T}

[0246] The initial binding path is used as the node binding path of the current task, and the resource isolation rule corresponding to the current task is used as the target isolation configuration.

[0247] In summary, the task-node binding path B and the initial isolation configuration C are completed. Among them, the task-node binding path B contains the specific node bound to each task and the allocated resource information, and the output format is: B = {(T j ,N(T j ), Allocated Resources)|T j ∈T}; The initial isolation configuration C covers the resource isolation rules of each task, including isolation mode and resource restrictions, and the output format is: C = {(T j , Isolation Mode, Resource Constraints)|Tj ∈T}. The overall process includes selecting the optimal node based on the adaptation score and dynamically handling conflicts and exclusive requirements to generate binding paths, as well as allocating resources to tasks and defining isolation rules to ensure the fairness and dynamic adjustment capability of resource allocation.

[0248] In the embodiment of the present application, the resource and node binding process is as follows: Figure 3 As shown in the figure, it mainly includes: first extracting task resource requirements, then generating a resource requirement matrix and dynamically adjusting it, then generating dynamic resource isolation rules, and then screening initial candidate nodes. After that, it collects and updates node resource status in real time, builds a task-node adaptive graph-level decoupling network, generates task-node binding paths, detects and dynamically adjusts resource conflicts, specifies task isolation modes, and finally outputs binding paths and isolation rules. The entire process is closely linked to achieve efficient and reasonable task resource allocation and scheduling.

[0249] In this embodiment, a resource binding device based on dynamic isolation rules is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0250] This embodiment provides a resource binding device based on dynamic isolation rules, such as Figure 4 As shown, including:

[0251] An acquisition module 41 is used to acquire a dynamic resource requirement matrix corresponding to the task description information in the current system;

[0252] A construction module 42 is used to construct a resource isolation rule corresponding to each task based on the dynamic resource demand matrix, and use the resource isolation rule to screen the initial nodes in the node resource pool to obtain a candidate node list;

[0253] A monitoring module 43 is used to monitor the resource status of each candidate node in the candidate node list and construct a corresponding dynamic resource status matrix according to the resource status;

[0254] The analysis module 44 is used to analyze the dynamic resource demand matrix and the dynamic resource status matrix to obtain optimized task characteristics and optimized node characteristics, and determine the node binding path and target isolation configuration of each task based on the optimized task characteristics and optimized node characteristics.

[0255] Furthermore, an acquisition module 41 is used to extract the resource requirements of each task from the task description information, wherein the resource requirements include a first resource type and a corresponding requirement quantity; construct a resource requirement vector for the task based on the first resource type and the corresponding requirement quantity; obtain the priority and task type of the task, and use the priority, task type and resource requirement vector to construct a task description vector for each task; use a preset dynamic adjustment algorithm to adjust the resource requirement vector in each task description vector to obtain a dynamic resource requirement vector; and combine the dynamic resource requirement vectors corresponding to each task into a dynamic resource requirement matrix.

[0256] Further, the construction module 42 includes a generation submodule and a screening submodule;

[0257] A generation submodule is used to obtain the initial resource status of each initial node in the node resource pool, wherein the initial resource status includes the second resource type and the corresponding available quantity; calculate the dynamic weight corresponding to the constraint condition based on the second resource type and the corresponding available quantity; and generate resource isolation rules based on the constraint condition and the corresponding dynamic weight.

[0258] The screening submodule is used to obtain the initial resource state of the initial node in the node resource pool; use the resource isolation rule to match the initial resource state with the dynamic resource demand vector corresponding to each task in the dynamic resource demand matrix to obtain the resource matching degree between the task and the initial node; for each initial node, calculate the ratio of the required quantity in the available quantity of each resource type, and calculate the load imbalance based on the ratio and the preset uniform ratio; calculate the first adaptation score between the task and the initial node according to the load imbalance and the resource matching degree; and screen out a list of candidate nodes from the node resource pool based on the first adaptation score.

[0259] Further, the analysis module 44 includes an optimization submodule and a detection submodule;

[0260] The optimization submodule is used to analyze the dynamic resource demand matrix, obtain the first relationship between each task, and use the first relationship to build a task subgraph; analyze the dynamic resource status matrix, obtain the second relationship between each candidate node, and use the second relationship to build a node subgraph; generate optimized task features and optimized node features based on the task subgraph and the node subgraph.

[0261] The detection submodule is used to detect whether there is a conflict in the initial binding path of the current task; if there is a conflict, the replacement node corresponding to the task with the conflict is determined in the candidate node list, and the node allocation operation is performed cyclically until there is no conflict in the initial binding path, and the initial binding path is used as the node binding path of the current task, and the resource isolation rule corresponding to the current task is used as the target isolation configuration.

[0262] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system).

[0263] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0264] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0265] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0266] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0267] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0268] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0269] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A resource binding method based on dynamic isolation rules, characterized in that: The method comprises: Get the dynamic resource requirement matrix corresponding to the task description information in the current system; Constructing resource isolation rules corresponding to each task based on the dynamic resource demand matrix, and using the resource isolation rules to screen the initial nodes in the node resource pool to obtain a candidate node list; Monitoring the resource status of each candidate node in the candidate node list, and constructing a corresponding dynamic resource status matrix according to the resource status; The dynamic resource demand matrix and the dynamic resource status matrix are analyzed to obtain optimized task characteristics and optimized node characteristics, and the node binding path and target isolation configuration of each task are determined according to the optimized task characteristics and the optimized node characteristics.

2. The method according to claim 1, characterized in that The step of obtaining a dynamic resource requirement matrix corresponding to the task description information in the current system includes: Extracting resource requirements of each task from the task description information, wherein the resource requirements include a first resource type and a corresponding required quantity; Constructing a resource requirement vector for the task according to the first resource type and the corresponding required quantity; Obtaining the priority and task type of the task, and constructing a task description vector for each task using the priority, the task type and the resource requirement vector; Using a preset dynamic adjustment algorithm to adjust the resource requirement vector in each of the task description vectors to obtain a dynamic resource requirement vector; The dynamic resource requirement vectors corresponding to each task are combined into a dynamic resource requirement matrix.

3. The method according to claim 1, characterized in that The constructing of resource isolation rules corresponding to each task based on the dynamic resource requirement matrix includes: Acquire an initial resource state of each initial node in the node resource pool, wherein the initial resource state includes a second resource type and a corresponding available quantity; Calculate a dynamic weight corresponding to the constraint condition according to the second resource type and the corresponding available quantity; Generate resource isolation rules based on the constraints and corresponding dynamic weights.

4. The method according to any one of claims 2 to 3, characterized in that: The method of screening the initial nodes in the node resource pool by using the resource isolation rule to obtain a candidate node list includes: Obtaining the initial resource state of the initial node in the node resource pool; Using the resource isolation rule, the initial resource state is matched with the dynamic resource demand vector corresponding to each task in the dynamic resource demand matrix to obtain the resource matching degree between the task and the initial node; For each of the initial nodes, calculate the ratio of the required quantity to the available quantity in each resource type, and calculate the load imbalance degree based on the ratio and a preset uniform ratio; Calculate a first adaptation score between the task and the initial node according to the load imbalance degree and the resource matching degree; A candidate node list is screened out from the node resource pool based on the first adaptation score.

5. The method according to claim 1, characterized in that The analyzing the dynamic resource demand matrix and the dynamic resource status matrix to obtain optimized task characteristics and optimized node characteristics includes: Analyze the dynamic resource demand matrix to obtain a first association relationship between each task, and construct a task subgraph using the first association relationship; Analyze the dynamic resource state matrix to obtain a second association relationship between each candidate node, and construct a node subgraph using the second association relationship; An optimized task feature and an optimized node feature are generated based on the task subgraph and the node subgraph.

6. The method according to claim 1, characterized in that The determining of the node binding path and the target isolation configuration of each task according to the optimized task characteristics and the optimized node characteristics includes: Calculate a second adaptation score between each task and the candidate node based on the optimized task characteristics and the optimized node characteristics; Determine the binding nodes corresponding to each of the tasks in the candidate node list according to the second adaptation score, and obtain multiple task node groups; An initial binding path and a target isolation configuration corresponding to each task are determined according to the task node group.

7. The method according to claim 6, characterized in that Determining the initial binding path and target isolation configuration corresponding to each task according to the task node group includes: Check whether there is a conflict in the initial binding path of the current task; If there is a conflict, the replacement node corresponding to the task with the conflict is determined in the candidate node list, and the node allocation operation is performed in a loop until there is no conflict in the initial binding path, and the initial binding path is used as the node binding path of the current task, and the resource isolation rule corresponding to the current task is used as the target isolation configuration.

8. A resource binding device based on dynamic isolation rules, characterized in that: The device comprises: An acquisition module is used to obtain the dynamic resource demand matrix corresponding to the task description information in the current system; A construction module, used to construct a resource isolation rule corresponding to each task based on the dynamic resource demand matrix, and use the resource isolation rule to screen the initial nodes in the node resource pool to obtain a candidate node list; A monitoring module, used to monitor the resource status of each candidate node in the candidate node list, and construct a corresponding dynamic resource status matrix according to the resource status; The analysis module is used to analyze the dynamic resource demand matrix and the dynamic resource status matrix to obtain optimized task characteristics and optimized node characteristics, and determine the node binding path and target isolation configuration of each task based on the optimized task characteristics and the optimized node characteristics.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Reliable configuration method for regularly executing tasks

    CN120216209A