Edge Node Scheduling System, Edge Node Scheduling Method and Edge Node
By building a node resource scheduling model in the edge node scheduling system and using reinforcement learning technology for intelligent scheduling, the shortcomings of edge computing technology in computing power coordination and intelligent scheduling are solved, efficient resource utilization and adaptability are achieved, and the needs of complex application scenarios are met.
Patent Information
- Application Number
- CN202411920485.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing edge computing technology has shortcomings in computing power coordination and intelligent scheduling, resulting in insufficient utilization of edge node computing resources, low GPU utilization rate, and difficult to meet computing power requirements. In addition, traditional fixed accuracy models and simple scheduling strategies cannot take into account the requested delay, accuracy and service level goals (SLO), and lack support for a variety of application scenarios and adaptive learning capabilities.
It provides an edge node scheduling system, including data request module, management backend module, performance collector, scheduler, predictor and controller. By building a node resource scheduling model, using reinforcement learning technology, building reward functions based on throughput and service delay targets, iterative training is performed, and the trained node resource scheduling model is obtained, which is used to allocate tasks in a dynamic computing resource environment and realize intelligent scheduling.
It has achieved the improvement of the system's throughput and resource utilization without damaging the service level goal (SLO), adapting to diversified scenarios, reducing the interference caused by resource competition, and improving the stability and response speed of the system.
Smart Images

Figure CN119363753B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and in particular to an edge node scheduling system, an edge node scheduling method, and an edge node. Background Art
[0002] Edge computing, as a distributed computing architecture, compared with the traditional cloud computing mode, migrates application programs from cloud central nodes to edge nodes for processing, getting closer to user terminals, thus significantly improving the data processing and transmission speed and reducing latency. The edge computing model shows significant advantages in aspects such as real-time data processing, security, privacy protection, scalability, location awareness, and traffic optimization. However, with the rapid development of the industrial Internet and the continuous expansion of application scenarios, the application complexity at the edge is increasing day by day, the number of edge nodes is growing rapidly, the access devices are becoming more complex, and the amount of generated data is also expanding sharply, which brings unprecedented challenges to the management of edge devices and applications.
[0003] Specifically, the complex and changeable network topology structure, the unstable network access environment, and the highly customized and rapidly iterated update requirements all pose higher requirements for edge computing. In addition, the heterogeneous hardware characteristics of edge gateways also bring additional difficulties to resource management. Against this background, the existing edge computing technologies have obvious deficiencies in computing power coordination and intelligent scheduling. For example, edge nodes do not make full use of computing resources, the GPU utilization rate is low, and it is difficult to meet the computing power requirements; the traditional fixed-precision model and simple scheduling strategy cannot take into account the latency, accuracy, and service level objective (SLO) of requests; the lack of support for multiple application scenarios and adaptive learning capabilities limits the practical value and application prospects of the technology; at the same time, when the existing concurrent model framework processes model parallelism, it often ignores the problems of resource competition and interference with model inference, and fails to make full use of the advantages of parallel models. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides an edge node scheduling system, an edge node scheduling method, and an edge node. The edge node scheduling system includes:
[0005] A data request module, configured to receive multiple request data from end-side devices. Each request data includes a service delay target metric value, and place the multiple request data into different dynamic task request queues according to their data types;
[0006] The management backend module is used to retrieve request data from the dynamic task request queue, and batch-process the request data by invoking a concurrent inference model according to the computing resource situation and task requirements of the current node to obtain network performance metric values; wherein, the network performance metric values include concurrent inference model parameters, queue backlog data, accuracy, throughput, and service latency target metric values;
[0007] The performance collector is used to monitor and collect the network performance metric values and computing resource utilization in real time and at regular intervals;
[0008] The scheduler is used to process the network performance metric values and the computing resource utilization to obtain a node resource scheduling strategy for task allocation in a dynamic computing resource environment;
[0009] The predictor is used to process the network performance metric values and the computing resource utilization to obtain an evaluation result of the current node resource scheduling strategy output by each iteration calculation of the scheduler, and feedback the evaluation result to the scheduler to change the probability of the scheduler selecting the current node resource scheduling strategy;
[0010] And the controller is used to switch the concurrent inference model for data processing and reallocate network computing resources according to the node resource scheduling strategy;
[0011] Wherein, the scheduler is used to process the network performance metric values and the computing resource utilization to obtain a node resource scheduling strategy for task allocation in a dynamic computing resource environment, specifically including:
[0012] Construct a node resource scheduling model, input the network performance metric values and the computing resource utilization as system operation state data into the node resource scheduling model, construct a reward function based on throughput and service latency targets, and perform iterative training with the goal of maximizing the value of the reward function to obtain a trained node resource scheduling model;
[0013] Based on the trained node resource scheduling model, input the current system operation state data to obtain the action with the maximum reward value, that is, the node resource scheduling strategy.
[0014] In an embodiment of the present invention, obtaining the trained node resource scheduling model specifically includes:
[0015] Construct a state space an action space and a policy space wherein represents the current time interval The constraint objectives of the concurrent inference model accuracy, the number of concurrent inference model instances, the batch size of each concurrent inference model, the amount of unprocessed requests in the task queue, the accuracy of the end of the concurrent inference model output, the average throughput, the service latency, the CPU usage rate, the memory occupancy rate, the GPU usage rate, and the GPU video memory usage rate; Indicates the current time interval The accuracy of the concurrent inference model to be run, the number of concurrent inference model instances, and the batch size specified for model operation in it;
[0016] The said policy space Initialize the parameters of the current Q-network, and use the current Q-network to calculate the current state The Q-values of all possible actions below , and select an action according to the ε-greedy policy a , execute the action in the environment a , calculate the reward r through the said reward function and select the next state , according to the reward r and the next state , use the target Q-network to calculate the selected action a The target Q-value of ;
[0017] Use the target Q-value And the Q-value predicted by the current Q-network Calculate the loss value, and update the parameters of the current Q-network by backpropagation with minimizing the said loss value until the network converges or reaches the maximum number of iteration rounds, and obtain the trained node resource scheduling model;
[0018] Among them, the current Q-network is a neural network structure with multiple convolutional layers and multiple fully connected layers, and the structure of the target Q-network is the same as that of the current Q-network.
[0019] In an embodiment of the present invention, select an action according to the ε-greedy policy a , specifically including:
[0020] Set the initial value ε of the exploration probability, generate a random number n between 0 and 1, and judge the size of n and ε:
[0021] If , randomly select an action from the action space A a ;
[0022] If , select the predicted action with the maximum Q-value a .
[0023] In an embodiment of the present invention, the said reward function For weighing the throughput and service latency objectives of the system, its expression is as follows:
[0024] ,
[0025] wherein, 、 、 are all hyperparameters; represents the accuracy of throughput; represents the accuracy of the service latency objective; represents the throughput at the i-th time period; represents the service latency objective at the i-th time period; represents the GPU utilization rate at the i-th time period.
[0026] In an embodiment of the present invention, the expression of the target Q value y is as follows:
[0027] ,
[0028] wherein, is the discount factor, represents the updated state, is the action for executing the maximum Q value according to the state , represents the target Q network.
[0029] In an embodiment of the present invention, the predictor is used to perform data processing on the network performance metric value and the computing resource utilization rate to obtain an evaluation result of the current node resource scheduling policy output by the scheduler for each iterative calculation, specifically including:
[0030] Obtain the action a selected by the scheduler in the previous state s and the previous policy , where the action a is the current node resource scheduling policy. Based on the current node resource scheduling policy, the predictor uses a pre-trained two-layer deep neural network to obtain the predicted values of the system service latency objective and throughput;
[0031] Calculate the percentage difference between the predicted values of the service latency objective and throughput and the actual values of the service latency objective and throughput , and obtain the evaluation result of the current node resource scheduling policy according to the comparison result between the percentage difference and a set threshold; meanwhile, update the parameters of the two-layer deep neural network through the percentage difference .
[0032] Based on the same inventive concept, the present invention also provides an edge node scheduling method, which uses the above-mentioned edge node scheduling system to implement node resource allocation. The method includes the following steps:
[0033] S1: Receive multiple request data from the end-side device. Each request data contains a service delay target metric value, and put the multiple request data into different dynamic task request queues according to their data types;
[0034] S2: Take out the request data from the dynamic task request queue, and call the concurrent inference model to batch process the request data according to the computing resource situation and task requirements of the current node to obtain network performance metric values; wherein, the network performance metric values include concurrent inference model parameters, queue backlog data, accuracy, throughput, and service delay target metric values;
[0035] S3: Real-time monitor and regularly collect the network performance metric values and computing resource utilization rates;
[0036] S4: Process the network performance metric values and the computing resource utilization rates to obtain a node resource scheduling strategy for task allocation in a dynamic computing resource environment;
[0037] S5: According to the node resource scheduling strategy, switch the concurrent inference model for data processing and re-allocate network computing resources;
[0038] S6: Process the network performance metric values and the computing resource utilization rates to obtain an evaluation result of the current node resource scheduling strategy output by each iterative calculation, and feedback the evaluation result to step S4 to change the probability of selecting the current node resource scheduling strategy.
[0039] In an embodiment of the present invention, in step S4, the method for obtaining a node resource scheduling strategy for task allocation in a dynamic computing resource environment includes:
[0040] S41: Build a node resource scheduling model, input the network performance metric values and the computing resource utilization rates as system operation state data into the node resource scheduling model, build a reward function based on throughput and service delay targets, and perform iterative training with the goal of maximizing the value of the reward function to obtain a trained node resource scheduling model;
[0041] S42: Based on the trained node resource scheduling model, input the current system operation state data to obtain the action with the maximum reward value, that is, the node resource scheduling strategy.
[0042] The present invention also provides an edge node, which includes a processor, a memory, and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the above-mentioned edge node scheduling method.
[0043] The above technical solution of the present invention has the following advantages compared with the prior art:
[0044] 1. Meet service level objectives: Ensure that each request segment can achieve the service level objective (SLO) while taking into account throughput and latency, thereby significantly improving the reliability of the system.
[0045] 2. Fully improve resource utilization: By using different models to call the various hardware resources of the GPU, it can quickly respond under real-time or near-real-time conditions, effectively reduce the cost of upward transmission, and thus improve the overall user experience.
[0046] 3. Achieve intelligent scheduling: The constructed node resource scheduling model learns each scheduling operation as an experience, so that it can intelligently select the model accuracy, batch_size of each model, and the number of concurrent instances in subsequent scheduling, further improving the accuracy of scheduling.
[0047] 4. Adapt to diverse scenarios: After each scenario change, the node resource scheduling model can autonomously learn and automatically select an appropriate model combination according to the current load pressure without manual re-selection of the model, thereby reducing the workload of model selection and enhancing the adaptability and robustness of the system in the face of unknown or changing environments.
[0048] 5. Reduce interference caused by resource contention: During each scheduling process, the current action space is predicted to effectively prevent interference caused by resource competition. This measure can avoid performance degradation caused by resource conflicts, thereby enhancing the stability and response speed of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the drawings, where
[0050] Figure 1 is a schematic diagram of the working principle of the edge node scheduling system provided in Embodiment 1 of the present invention;
[0051] Figure 2 is a flowchart of the edge node scheduling method provided in Embodiment 2 of the present invention;
[0052] Description of the reference numerals in the drawings: 10, data request module; 20, management backend module; 30, performance collector; 40, scheduler; 50, predictor; 60, controller. Detailed implementation manners
[0053] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.
[0054] Embodiment 1
[0055] Refer to Figure 1 As shown, the present invention provides an edge node scheduling system, and the edge node scheduling system includes:
[0056] A data request module 10, configured to receive a plurality of request data from end-side devices (such as smart phones, sensors, etc.), each request data includes a service latency objective (SLO) metric value, that is, the performance requirement for this inference (such as response time, accuracy, etc.), and put the plurality of request data into different dynamic task request queues according to their data types;
[0057] A management backend module 20, configured to take out request data from the dynamic task request queue, and call a concurrent inference model to batch-process the request data according to the computing resource situation and task requirements of the current node, so as to obtain network performance metric values; wherein, the network performance metric values include concurrent inference model parameters, queue backlog data, accuracy, throughput, and service latency objective metric values;
[0058] A performance collector 30, configured to monitor and collect the network performance metric values and computing resource utilization in real time and regularly;
[0059] A scheduler 40, configured to process the network performance metric values and the computing resource utilization to obtain a node resource scheduling strategy for allocating tasks in a dynamic computing resource environment;
[0060] A predictor 50, configured to process the network performance metric values and the computing resource utilization to obtain an evaluation result of the current node resource scheduling strategy output by each iterative calculation of the scheduler 40, and feedback the evaluation result to change the probability of the scheduler selecting the current node resource scheduling strategy;
[0061] And a controller 60, configured to switch the concurrent inference model to process data and reallocate network computing resources according to the node resource scheduling strategy.
[0062] As can be seen from the above technical solutions, the present invention sends the request data of the terminal device to the edge node nearby through the data request module 10, avoiding the latency caused by long-distance data transmission to the cloud, reducing network congestion and bandwidth occupancy; in view of the fact that different request data may need to be processed by different inference models, the data in the request is classified according to the model requirements through the data request module 10, and the classified data is stored in different dynamic queues. These queues are highly flexible, and their size and quantity can be dynamically adjusted according to the inference requirements and data volume in the actual application scenario, thus ensuring the efficiency and flexibility of data processing.
[0063] Furthermore, the AI inference task is executed on the management backend module 20 of the edge node. According to the real-time computing resource status of the edge node (such as CPU, memory, GPU, etc.) and the task requirements, the AI inference task is intelligently scheduled, avoiding waste and overload of computing resources, meeting the latency and SLO accuracy requirements for different requests, and improving throughput, thus improving resource utilization.
[0064] The performance collector 30 can monitor and regularly collect network performance metric values and computing resource utilization in real time, providing comprehensive performance data support for the system. These performance data not only help the system administrator to timely understand the running status of the system, but also can provide accurate data input for the scheduler 40, so that it can formulate a more reasonable node resource scheduling strategy.
[0065] The predictor 50 can predict the resource interference situation that the scheduler 40 may receive during the data processing process and send the prediction result to the scheduler 40 for decision-making. This function not only improves the anti-interference ability of the system, but also ensures that the system can quickly respond when facing resource tension or emergencies, and assists the scheduler 40 to formulate the optimal scheduling strategy.
[0066] The controller 60 can flexibly switch the concurrent inference model for data processing according to the node resource scheduling strategy and reallocate network computing resources. This mechanism ensures that the system can quickly adjust the resource allocation strategy in the face of different tasks and performance requirements to meet various complex business needs.
[0067] Furthermore, the management backend module 20 includes functions such as loading models, AI inference, and controlling models, aiming to use deep learning frameworks such as TensorRT, libtorch, onnx, paddle, and tvm to call the underlying hardware for inference.
[0068] Furthermore, the scheduler 40 is used to process the network performance metric values and the computing resource utilization to obtain a node resource scheduling strategy for task allocation in a dynamic computing resource environment, specifically including:
[0069] Build a node resource scheduling model based on reinforcement learning, input the network performance metric values and the computing resource utilization rate as system operation status data into the node resource scheduling model, construct a reward function based on the throughput and service latency objectives, and perform iterative training with the goal of maximizing the value of the reward function to obtain a trained node resource scheduling model;
[0070] Based on the trained node resource scheduling model, input the current system operation status data to obtain the action with the maximum reward value, that is, the node resource scheduling strategy;
[0071] The current Q-network is a neural network structure with multiple convolutional layers and multiple fully connected layers, and the structure of the target Q-network is the same as that of the current Q-network.
[0072] Among them, obtaining the trained node resource scheduling model specifically includes:
[0073] Construct a state space an action space and a policy space , where represents the accuracy of the concurrent inference model, the number of concurrent inference model instances, the batch size of each concurrent inference model, the number of requests not yet processed in the task queue, the accuracy of the output end of the concurrent inference model, the average throughput, the constraint objectives of the service latency, the CPU utilization rate, the memory occupancy rate, the GPU utilization rate, and the GPU video memory utilization rate in the current time interval i; represents the accuracy of the concurrent inference model to be run, the number of concurrent inference model instances, and the batch size specified for model operation in the current time interval i;
[0074] The policy space Initialize the parameters of the current Q-network, use the current Q-network to calculate the Q-values of all possible actions in the current state s , and select an action according to the ε-greedy policy a , specifically including:
[0075] Set the initial value ε of the exploration probability, generate a random number n between 0 and 1, and judge the size of n and ε:
[0076] If , randomly select an action from the action space A a ;
[0077] If , select the predicted action with the maximum Q-value a , that is ;
[0078] Execute an action in the environment a , calculate the reward r and select the next state through the reward function , according to the reward r and the next state , use the target Q-network to calculate the selected action a The target Q-value y of:
[0079] ,
[0080] wherein, is the discount factor, represents the updated state, is the action that executes the maximum Q-value according to the state ; represents the target Q-network;
[0081] Use the target Q-value y and the Q-value predicted by the current Q-network Calculate the loss value, and update the parameters of the current Q-network by backpropagation through minimizing the loss value until the network converges or reaches the maximum number of iterations, obtaining the trained node resource scheduling model.
[0082] Specifically, the reward function is used to balance the system throughput and service latency objectives, and its expression is as follows:
[0083] ,
[0084] wherein, 、 、 are all hyperparameters; represents the accuracy of the throughput; represents the accuracy of the service latency objective; represents the throughput at the i-th time period; represents the service latency objective at the i-th time period; represents the GPU utilization rate at the i-th time period.
[0085] In this embodiment, the predictor 50 can accurately predict the balance state between the system throughput and the service latency objective. This prediction result can intuitively show whether the scheduler 40 has adopted the optimal policy solution. In addition, by comparing the predicted value with the true value, the predictor 50 can obtain in advance the potential interference that these changes may have on the inference model when the accuracy of the concurrent inference model, the batch size (batch_size) of each inference model, and the number of concurrency carried in the management backend module 20 all change.
[0086] Further, the predictor 50 is used to process the network performance metric value and the computing resource utilization rate to obtain an evaluation result of the current node resource scheduling policy output by the scheduler 40 for each iterative calculation, specifically including:
[0087] Obtain the action a selected by the scheduler 40 in the previous state s and the previous policy The action a is the current node resource scheduling policy. Based on the current node resource scheduling policy, the predictor 50 calculates the predicted values of the system service delay target and throughput using a pre-trained two-layer deep neural network (DNN), and calculates the percentage difference between the predicted values of the service delay target and throughput and the actual values of the service delay target and throughput , according to the percentage difference and the comparison result with the set threshold, obtain the evaluation result of the node resource scheduling policy of the scheduler 40; if is greater than the set threshold, it indicates that the current node resource scheduling policy will cause resource interference, and the predictor 50 will notify the scheduler 40 to reduce the probability of selecting this node resource scheduling policy; at the same time, through the percentage difference to update the parameters of the two-layer deep neural network. This process ensures that the predictor 50 can continuously learn and adapt to the changing environment, thereby continuously enhancing the node resource scheduling ability of the scheduler 40.
[0088] In this embodiment, the controller 60 realizes node resource allocation according to the node resource scheduling policy calculated by the scheduler 40, that is:
[0089] In the face of different business scenarios and performance requirements, it can flexibly select a higher-precision model (such as the FP32 model) to ensure data accuracy, or select a lower-precision but higher-processing-efficiency model (such as INT8, FP16 model) to accelerate the data processing process;
[0090] Dynamically adjust the operating parameters of the model, and appropriately increase or decrease the batch size (batch_size) according to the current system load and data processing requirements, and adjust the number of concurrent model instances;
[0091] After performing the model configuration switch, the controller is responsible for starting or stopping the model instance. This function ensures that the system can quickly and accurately adjust resource allocation when switching the model configuration, so as to meet new business needs and performance requirements.
[0092] Taking a specific example, assume that the original running configuration of the system is the Tensor RT fp16 model with 2 parallel instances and a batch size of 16; and the Tensor RT int8 model with 1 parallel instance and the same batch size of 16. According to the calculation result of the scheduler 40, the controller 60 may switch these configurations to the Tensor RT int8 model with the number of parallel instances increased to 2 and the batch size remaining unchanged at 16; meanwhile, a new Tensor RT int4 model is added as another option with 1 parallel instance but the batch size adjusted to 4. Such a switch not only optimizes the model's accuracy and efficiency but also further improves the overall performance of the system by adjusting the number of concurrent instances and the batch size.
[0093] Embodiment 2
[0094] Based on the same inventive concept as Embodiment 1, the present invention also provides an edge node scheduling method, which uses the system described in Embodiment 1 for node resource allocation. Refer to Figure 2 as shown, the method includes the following steps:
[0095] S1: Receive multiple request data from the end-side device, each request data containing a service latency target metric value, and place the multiple request data into different dynamic task request queues according to their data types;
[0096] S2: Take out the request data from the dynamic task request queue, and call the concurrent inference model to batch process the request data to obtain network performance metric values; wherein, the network performance metric values include concurrent inference model parameters, queue backlog data, accuracy, throughput, and service latency target metric values;
[0097] S3: Monitor in real time and collect the network performance metric values and computing resource utilization regularly;
[0098] S4: Process the network performance metric values and the computing resource utilization to obtain a node resource scheduling strategy for task allocation in a dynamic computing resource environment;
[0099] S5: According to the node resource scheduling strategy, switch the concurrent inference model for data processing and reallocate network computing resources;
[0100] S6: Process the network performance metric values and the computing resource utilization to obtain an evaluation result of the current node resource scheduling strategy output by each iterative calculation, and feedback the evaluation result to step S4 to change the probability of selecting the current node resource scheduling strategy.
[0101] Further, in step S4, the method for obtaining a node resource scheduling policy for task allocation in a dynamic computing resource environment includes:
[0102] S41: Construct a node resource scheduling model based on reinforcement learning. Input the network performance metric values and the computing resource utilization rate as system operation state data into the node resource scheduling model. Construct a reward function based on the throughput and service latency targets, and perform iterative training with the goal of maximizing the value of the reward function to obtain a trained node resource scheduling model.
[0103] S42: Based on the trained node resource scheduling model, input the current system operation state data to obtain the action with the maximum reward value, that is, the node resource scheduling policy.
[0104] In this embodiment, the method for processing the network performance metric values and the computing resource utilization rate to obtain the evaluation result of the current node resource scheduling policy output by each iterative calculation is the same as the method in Embodiment 1 for processing the network performance metric values and the computing resource utilization rate by a predictor to obtain the evaluation result of the current node resource scheduling policy output by each iterative calculation of the scheduler, and will not be elaborated here.
[0105] Embodiment 3
[0106] The present invention also provides an edge node, which includes a processor, a memory, and a bus system. The processor and the memory are connected through the bus system. The memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the steps of the edge node scheduling method described in Embodiment 2 above.
[0107] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0108] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.
[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in a block or multiple blocks.
[0111] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. An edge node scheduling system, characterized in that: Includes the following modules: A data request module is used to receive multiple request data from a terminal device, each request data includes a service delay target indicator value, and put the multiple request data into different dynamic task request queues according to their data types; A management backend module is used to take out request data from the dynamic task request queue, call the concurrent reasoning model to batch process the request data according to the computing resource situation of the current node and the task requirements, and obtain a network performance index value; wherein the network performance index value includes concurrent reasoning model parameters, queue backlog data, accuracy, throughput and service delay target index values; A performance collector is used to monitor in real time and regularly collect network performance indicator values and computing resource utilization; A scheduler, used for performing data processing on the network performance indicator value and the computing resource utilization rate to obtain a node resource scheduling strategy for allocating tasks in a dynamic computing resource environment, including: Constructing a node resource scheduling model, inputting the network performance indicator value and the computing resource utilization rate as system operation status data into the node resource scheduling model, constructing a reward function based on throughput and service delay targets, and performing iterative training with the goal of maximizing the value of the reward function to obtain a trained node resource scheduling model; Based on the trained node resource scheduling model, the current system operation status data is input to obtain the action with the maximum reward value, that is, the node resource scheduling strategy; wherein the reward function R(s,a) is used to weigh the system throughput and service delay targets, and its expression is as follows: Among them, α, β, and γ are all hyperparameters; acc i Indicates the accuracy of throughput; acc slo represents the accuracy of the service delay target; T i represents the throughput of the i-th period; slo i represents the service delay target for the i-th period; Indicates the GPU utilization rate during the i-th period; s indicates the current system operation status data, including the accuracy of the concurrent reasoning model currently running, the number of concurrent reasoning model instances, the batch size of each concurrent reasoning model, the number of unprocessed requests in the task queue, the accuracy of the concurrent reasoning model output results, the average throughput, the constraint target of the service delay, the CPU utilization rate, the memory occupancy rate, the GPU utilization rate, and the GPU video memory utilization rate; a indicates the selected node resource scheduling strategy, including the accuracy of the concurrent reasoning model to be run, the number of concurrent reasoning model instances, and the batch size specified for the model run; A predictor is used to perform data processing on the network performance indicator value and the computing resource utilization rate to obtain an evaluation result of the current node resource scheduling strategy output by the scheduler for each iteration calculation, and feed the evaluation result back to the scheduler to change the probability of the scheduler selecting the current node resource scheduling strategy; wherein obtaining the evaluation result of the current node resource scheduling strategy output by the scheduler for each iteration calculation includes: Obtaining a current node resource scheduling strategy of the scheduler, based on which the predictor uses a pre-trained two-layer deep neural network to obtain predicted values of system service delay targets and throughput; Calculate the difference percentage Δx between the predicted value of the service delay target and the throughput and the actual value of the service delay target and the throughput, and obtain an evaluation result of the current node resource scheduling strategy according to a comparison result of the difference percentage Δx and a set threshold; at the same time, update the parameters of the two-layer deep neural network by the difference percentage Δx; and a controller, which is used to switch the concurrent reasoning model to perform data processing and reallocate network computing resources according to the node resource scheduling strategy.
2. The edge node scheduling system according to claim 1, characterized in that: The trained node resource scheduling model is obtained, which specifically includes: Construct state space S, action space A and strategy space π(s|a), where s i ∈S represents the accuracy of the concurrent reasoning model in the current time interval i, the number of concurrent reasoning model instances, the batch size of each concurrent reasoning model, the number of unprocessed requests in the task queue, the accuracy of the concurrent reasoning model output results, the average throughput, the constraint target of service delay, CPU usage, memory occupancy, GPU usage, and GPU video memory usage; a i ∈A represents the concurrent reasoning model accuracy, number of concurrent reasoning model instances, and batch size specified for model running in the current time interval i; The strategy space π(s|a) initializes the parameters of the current Q network, uses the current Q network to calculate the Q values Q(s,a) of all possible actions in the current state s, selects action a according to the ε-greedy strategy, executes action a in the environment, calculates the reward r through the reward function, and selects the next state s', and uses the target Q network to calculate the target Q value y of the selected action a according to the reward r and the next state s'; The loss value is calculated using the target Q value y and the Q value Q(s,a) predicted by the current Q network, and the parameters of the current Q network are updated by backpropagation by minimizing the loss value until the network converges or reaches the maximum iteration round, thereby obtaining a trained node resource scheduling model; The current Q network is a neural network structure of multiple convolutional layers and multiple fully connected layers, and the structure of the target Q network is the same as that of the current Q network.
3. The edge node scheduling system according to claim 2, characterized in that: Select action a according to the ε-greedy strategy, which includes: Set the initial value ε of the exploration probability, generate a random number n between 0 and 1, and determine the size of n and ε: If n<ε, randomly select an action a from the action space A; If n ≥ ε, select the predicted action a with the largest Q value.
4. The edge node scheduling system according to claim 2, characterized in that: The expression of the target Q value y is as follows: Among them, γ∈[0,1] is the discount factor, s′ represents the updated state, a′ is the action with the maximum Q value according to the state s′, Q target represents the target Q network.
5. An edge node scheduling method, characterized in that: Node resource allocation is implemented using the edge node scheduling system as described in any one of claims 1 to 4, the method comprising the following steps: S1: receiving multiple request data from a terminal device, each request data including a service delay target indicator value, and placing the multiple request data into different dynamic task request queues according to their data types; S2: taking out request data from the dynamic task request queue, calling the concurrent reasoning model to batch process the request data according to the computing resource situation of the current node and the task requirements, and obtaining a network performance index value; wherein the network performance index value includes concurrent reasoning model parameters, queue backlog data, accuracy, throughput and service delay target index values; S3: real-time monitoring and regular collection of the network performance index value and computing resource utilization; S4: performing data processing on the network performance indicator value and the computing resource utilization rate to obtain a node resource scheduling strategy for allocating tasks in a dynamic computing resource environment; S5: switching the concurrent reasoning model to perform data processing and reallocating network computing resources according to the node resource scheduling strategy; S6: Perform data processing on the network performance indicator value and the computing resource utilization to obtain an evaluation result of the current node resource scheduling strategy output by each iterative calculation, and feed the evaluation result back to step S4 to change the probability of selecting the current node resource scheduling strategy.
6. The edge node scheduling method according to claim 5, characterized in that: In step S4, the method for obtaining a node resource scheduling strategy for allocating tasks in a dynamic computing resource environment includes: S41: construct a node resource scheduling model, input the network performance indicator value and the computing resource utilization rate as system operation status data into the node resource scheduling model, construct a reward function based on throughput and service delay targets, perform iterative training with the goal of maximizing the value of the reward function, and obtain a trained node resource scheduling model; S42: Based on the trained node resource scheduling model, the current system operation status data is input to obtain the action with the maximum reward value, that is, the node resource scheduling strategy.
7. An edge node, characterized in that: The edge node includes a processor, a memory and a bus system, the processor and the memory are connected through the bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory to implement the edge node scheduling method described in claim 5 or claim 6.
Citation Information
Patent Citations
Self-adaptive batch processing and parallel scheduling system for edge device deep learning model reasoning
CN115454585A
Resource scheduling strategy optimization method based on AI training task indexes
CN118733274A