Server-unaware computing reinforcement learning resource configuration and training method and system
Patent Information
- Application Number
- CN202510383169.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]尽管所述方案尝试把强化学习的演员进程分配到高性能CPU服务器上,将学习者进程分配到高性能GPU服务器上进行分布式训练,但是这些工作仍然采用以服务器为中心的模式,演员进程的数目N和学习者进程的数目M是事先固定的,缺乏弹性扩展演员进程的能力,且训练成本较高
[0059]1、本发明结合服务器无感知计算进行强化学习训练,相较于纯服务器训练可以根据用户成本需求和强化学习算法的需求,弹性调整演员的数目。
Smart Images

Figure CN122840155A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically relating to a method and system for configuring and training reinforcement learning resources in server-insensitive computing. More specifically, it relates to a method and related apparatus for reinforcement learning training in a server-insensitive computing scenario. Background Technology
[0002] In recent years, artificial intelligence (AI) technologies have sparked a research boom. Reinforcement learning, as an important branch of AI, has been widely applied in robotics, automatic control, and large language models. Reinforcement learning focuses on maximizing the cumulative reward of an agent through continuous interaction with its environment. The agent collects information about the state of the environment, makes decisions to change the environment based on its current strategy, and receives reward feedback from the environment to iteratively update its strategy.
[0003] Currently, reinforcement learning model training can be summarized as the following iterative steps: (1) The actor process, executed on the CPU, interacts with the environment and samples the trajectory; (2) The actor process sends the trajectory to the learner process, executed on the GPU; (3) The learner process trains based on the trajectory data and updates the model parameters; (4) The learner process synchronizes the latest model parameters to the actor process and repeats step (1). In the same-track strategy, i.e., on-policy reinforcement learning algorithms, there is a data dependency between steps (1) and (3). That is, when (1) is not completed, the GPU used in step (3) is idle, and when (3) is not completed, the CPU used in (1) is also idle. In the off-track strategy, i.e., off-policy reinforcement learning algorithms, the time of steps (1) and (3) may overlap. In reinforcement learning model training, developers often focus on meeting a deadline, i.e., training the model before the deadline, and minimizing the training cost as much as possible.
[0004] From a training environment perspective, existing reinforcement learning training is server-centric, lacking flexibility and unable to flexibly adjust resources to reduce training costs while meeting deadlines. For example, some studies rent large numbers of servers equipped with high-performance CPUs and GPUs for reinforcement learning training to meet training deadlines, resulting in high training costs.
[0005] Patent document CN117172093A discloses a strategy optimization method and apparatus for Linux system kernel configuration based on machine learning. The scheme includes: collecting load data based on the current parameter configuration of the Linux system; identifying the load type of the current load based on the load data; recommending the latest parameter configuration for the current load based on the historical best recommendation of the current load; generating a parameter space of the latest parameter configuration; and using a parameter recommendation optimization model to recommend parameters of the Linux system within the parameter space to obtain parameter recommendation results.
[0006] Patent document CN114861826A discloses a large-scale reinforcement learning training framework system based on distributed design. This scheme includes generating training data on a large scale and with high concurrency by running multiple actor modules in a distributed manner and utilizing the large amount of CPU computing resources of the cluster. This breaks through the physical limitations of single-machine training and greatly improves the data generation efficiency in reinforcement learning.
[0007] Although the proposed scheme attempts to distribute reinforcement learning actor processes to high-performance CPU servers and learner processes to high-performance GPU servers for distributed training, these efforts still adopt a server-centric model. The number of actor processes N and learner processes M are fixed in advance, lacking the ability to flexibly expand actor processes, and the training cost is high.
[0008] Therefore, how to provide a reinforcement learning training framework that can flexibly adjust training resources, and minimize the cost of reinforcement learning training while ensuring training performance, is an urgent problem to be solved. Summary of the Invention
[0009] To address the shortcomings of existing technologies, the purpose of this invention is to provide a server-aware computing-based reinforcement learning resource allocation and training method and system.
[0010] A server-aware computation-based reinforcement learning resource allocation method according to the present invention includes:
[0011] Step S1: Establish performance and cost models based on distributed reinforcement learning algorithms;
[0012] Step S2: Based on the performance model and cost model, sample reinforcement learning performance data on a hardware platform equipped with CPU and GPU;
[0013] Step S3: Fit the model hyperparameters based on the reinforcement learning performance data;
[0014] Step S4: Based on the model hyperparameters, calculate the optimal training cost on a hardware platform equipped with CPU and GPU, and output the corresponding resource configuration; the resource configuration includes: the number of CPUs and the number of GPUs.
[0015] Preferably, the distributed reinforcement learning algorithm refers to the distributed actor-learner reinforcement learning algorithm, i.e., the Actor-Learner reinforcement learning algorithm;
[0016] In step S1, a performance model and a cost model containing hyperparameters are established based on the distributed actor-learner reinforcement learning algorithm.
[0017] The performance model includes: a performance model for the same-track strategy and a performance model for the different-track strategy; the performance model for the same-track strategy includes the PPO algorithm; the performance model for the different-track strategy includes the IMPALA algorithm.
[0018] The mathematical expression for the total time of the performance model of the same-track strategy is:
[0019] T sum =T a +T b +T c +T d
[0020] Among them, T a Indicates the sampling time; T b Indicates the trajectory transmission time; T c Indicates training time; T d Indicates the parameter synchronization time;
[0021] For the performance model of the aforementioned off-track strategy, the p-norm is used to model the coverage relationship between sampling time and training time;
[0022] The coverage relationship, i.e. the coverage ratio, is adjusted by the hyperparameter;
[0023] The mathematical expression for the covering relationship is:
[0024]
[0025] in, This indicates the coverage relationship, and λ represents the hyperparameter.
[0026] The mathematical expression for the cost model is:
[0027] C sum =T a (d,k1)*α*k1+T c (d,k2)*β*k2+T sum *γ
[0028] Among them, C sum α represents the total cost; β represents the price per second of a server-independent compute instance; β represents the price per second of a GPU server; γ represents the average price per second of a VPC network; k1 represents the number of CPU processes, k2 represents the number of GPU processes; d represents the amount of data represented by the linear model; T a (d,k1) represents the sampling time for k1 CPU processes; T c (d,k2) represents the training time with k2 GPU processes.
[0029] Preferably, in step S2, n rounds of simulated training are performed to collect reinforcement learning performance data; the reinforcement learning performance data includes: the execution time of the sampling process in each round of reinforcement learning, the transmission time of the sampling process sending the sampled data to the training process, the execution time of the training process, the time of the training process updating the updated model weight parameters back to the sampling process, and the total time of each round of reinforcement learning; wherein, n is greater than or equal to 5;
[0030] In step S3, based on the reinforcement learning performance data, the model hyperparameters are fitted and obtained; the fitting method is polynomial fitting or linear regression fitting.
[0031] The model hyperparameters include: random variable a1, random variable a2, random variable b1, random variable b2 and hyperparameter λ;
[0032] In step S3, the R-squared test is used to verify the accuracy of the fit. It is determined whether the correlation coefficient of the variables, i.e., the R-squared value, is greater than 0.6. If the result is yes, then step S4 is executed; if the result is no, then the process ends.
[0033] Preferably, in step S4, the solution time for the optimal training cost is less than the maximum deadline; the mathematical expression for the optimal training cost is:
[0034] argmin (k1,k2) C sum ,stT sum ≤T ddl
[0035] Where, argmin (k1,k2) This indicates that C sum That is, the total cost is minimized at (k1,k2); st represents the value at time T. sum ≤T ddl Under the constraints, T ddl This indicates the maximum deadline.
[0036] The reinforcement learning training method for server-aware computing provided by the present invention employs the reinforcement learning resource allocation method for server-aware computing as described in claim 4, comprising:
[0037] Step A1: Allocate N server-insensitive computing instances and run actor processes and learner processes; N is a constant;
[0038] Step A2: Release the server-insensitive computing instance;
[0039] Step A3: Employ reinforcement learning resource allocation methods to train the learner process and update the model's gradient parameters;
[0040] Step A4: Based on the gradient parameters of the model, elastically allocate N' new server-insensitive computing instances, and re-execute step A1 until the model converges; N' is a constant;
[0041] In step A4, the model converges when the model's loss function is less than a preset threshold or reaches a preset maximum number of training rounds.
[0042] Preferably, in step A1, the number of actor processes is expressed mathematically as follows:
[0043]
[0044] Where, N i N represents the number of computational instances in the i-th round; max This represents the maximum number of server-insensitive computing instances supported by the system; r represents the total number of training rounds.
[0045] Preferably, in step A4, the range of the elastic allocation is constrained by the maximum number of processes supported by the server-insensitive computing platform and the current training batch size, and the range is [1, 128].
[0046] Preferably, in step A4, the default memory allocated to the new server-invisible computing instance is 1GB; it is determined whether the default memory can run the actor process. If the result is yes, no processing is performed and the actor process is run; if the result is no, the default memory is multiplied by multiples of 2 until the actor process is run.
[0047] A server-aware computational reinforcement learning resource allocation subsystem according to the present invention includes:
[0048] Modeler module: Builds performance and cost models based on distributed reinforcement learning algorithms;
[0049] Sampler module: Based on the performance model and cost model, sample reinforcement learning performance data on a hardware platform equipped with CPU and GPU;
[0050] Fitter module: Fits model hyperparameters based on the reinforcement learning performance data;
[0051] Solver module: Based on the model hyperparameters, calculates the optimal training cost on a hardware platform equipped with CPU and GPU, and outputs the corresponding resource configuration; the resource configuration includes: the number of CPUs and the number of GPUs.
[0052] According to the present invention, a server-aware computing-based reinforcement learning training system is capable of triggering the operation of a server-aware computing-based reinforcement learning resource allocation subsystem, including:
[0053] Module A1: Allocate N server-insensitive computing instances to run actor processes and learner processes; N is a constant;
[0054] Module A2: Release the server-side non-transparent computing instance;
[0055] Module A3: Triggers the operation of the reinforcement learning resource allocation subsystem, initiates the reinforcement learning training process for the learners, and updates the gradient parameters of the model;
[0056] Module A4: Based on the gradient parameters of the model, N' new server-insensitive computing instances are elastically allocated, and Module A1 is reactivated until the model converges; N' is a constant.
[0057] In module A4, the model converges when the model's loss function is less than a preset threshold or reaches a preset maximum number of training rounds.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] 1. This invention combines server-insensitive computing for reinforcement learning training, which, compared to pure server training, allows for flexible adjustment of the number of actors based on user cost requirements and the needs of the reinforcement learning algorithm.
[0060] 2. This invention is a reinforcement learning resource allocation system that balances cost and performance, and can achieve the optimal resource allocation for training cost while ensuring training performance.
[0061] 3. When allocating server-insensitive computing instances, this invention attempts to allocate memory by doubling from the minimum available memory. This doubling memory expansion method makes this invention suitable for reinforcement learning environments with high memory consumption, using the minimum amount of memory and reducing the training cost of the model.
[0062] 4. The present invention adopts a training method using hybrid servers and server-insensitive computing instances. Compared with training using only servers, it saves the idle time of the server CPU and can significantly reduce costs. In other words, the training method using hybrid servers and server-insensitive computing instances can flexibly adjust the resource allocation during the training process compared with traditional methods, greatly reduce the execution time, and achieve the theoretically optimal training cost while ensuring training performance.
[0063] 5. This invention targets asynchronous reinforcement learning training algorithms, namely IMPALA and Gorila, and uses p-norm to model the coverage relationship of training time. Compared with the prior art, it expands the scope of supported algorithms, making this invention applicable not only to training systems using synchronous reinforcement learning training algorithms, but also to training systems using asynchronous reinforcement learning training algorithms. Attached Figure Description
[0064] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0065] Figure 1 This is a schematic diagram of the software flow provided by the present invention;
[0066] Figure 2 This is a schematic diagram of the software architecture provided by the present invention;
[0067] Figure 3 This is a schematic diagram of the hardware device structure provided by the present invention. Detailed Implementation
[0068] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0069] This invention is a reinforcement learning resource allocation system that balances cost and performance for server-insensitive computing scenarios, comprising: a reinforcement learning training method combining server-insensitive computing, and a reinforcement learning resource allocation system that balances cost and performance.
[0070] The modeler of this invention establishes performance and cost models for classic reinforcement learning algorithms. A sampler performs several rounds of simulated training to collect reinforcement learning performance data. A fitter then fits the model hyperparameters based on the performance data. Finally, a solver is run to solve for the resource allocation parameters corresponding to the theoretical minimum training cost that meets the training deadline. This invention employs a training method using hybrid servers and server-insensitive computing instances. Compared to traditional methods, this allows for flexible adjustment of resource allocation during training, significantly reducing execution time, and achieving the theoretically optimal training cost while ensuring training performance.
[0071] A server-aware computation-based reinforcement learning resource allocation method according to the present invention includes:
[0072] Step S1: Establish performance and cost models based on distributed reinforcement learning algorithms;
[0073] Step S2: Based on the performance model and cost model, sample reinforcement learning performance data on a hardware platform equipped with server-insensitive computing instances, i.e. CPUs and computing instances, i.e. GPUs.
[0074] Step S3: Fit the model hyperparameters based on the reinforcement learning performance data;
[0075] Step S4: Based on the model hyperparameters, calculate the optimal training cost on a hardware platform equipped with server-agnostic computing instances, i.e., CPUs and computing instances, i.e. GPUs, and output the corresponding resource configuration; the resource configuration includes: the number of CPUs and the number of GPUs.
[0076] Specifically, in step S3, based on the reinforcement learning performance data, the model hyperparameters are fitted and obtained; the fitting method is polynomial fitting or linear regression fitting.
[0077] The model hyperparameters include: random variable a1, random variable a2, random variable b1, random variable b2 and hyperparameter λ;
[0078] In other words, the model hyperparameters include: sampling time: T a (d,k)=a1+b1*d / k, training time: T c (d,k)=a2+b2*d / k contains a1,b1,a2,b2. The hyperparameter λ in the equation.
[0079] A reinforcement learning training method combining server-insensitive computation, provided by the present invention, includes:
[0080] Step 1.1: Allocate N server-side non-aware computing instances, run the actor process, and obtain the training trajectory.
[0081] Step 1.2: The actor process sends the training trajectory to the M learner processes running on the high-performance GPU server.
[0082] Step 1.3: Release the server-invisible computing instance.
[0083] Step 1.4: The learner process performs reinforcement learning training and updates the model gradient parameters.
[0084] Step 1.5: Based on the latest gradient parameters, elastically allocate N' new server-insensitive computing instances, and repeat Step 1.1 until the model converges.
[0085] Specifically, the actor process and the learner process, namely 'Actor-Learner', is a reinforcement learning training architecture;
[0086] In this process, the Actor process is responsible for sampling data, and the learner process is responsible for training the model. Sources: https: / / arxiv.org / abs / 1906.04585 and https: / / arxiv.org / abs / 1802.01561.
[0087] The training of the reinforcement learning model, namely 'Actor-Learner', includes: (1) the actor process executed on the CPU interacts with the environment and samples the trajectory; (2) the actor process sends the trajectory to the learner process executed on the GPU; (3) the learner process trains based on the trajectory data and updates the model parameters; (4) the learner process synchronizes the latest model parameters to the actor process and repeats step (1).
[0088] Specifically, in step 1.5, the range of elasticity is constrained by the maximum number of processes supported by the server-insensitive computing platform and the current training batch size, and in typical tests, the range is [1, 128].
[0089] A reinforcement learning resource allocation system that balances cost and performance, provided by the present invention, includes a modeler module, a sampler module, a fitter module, and a solver module.
[0090] The fitter module uses polynomial fitting or linear regression fitting.
[0091] The solver module uses the SLSQP gradient descent algorithm or the L-BFGS-B gradient descent algorithm;
[0092] Step 2.1: The modeler builds performance and cost models for distributed reinforcement learning algorithms.
[0093] Step 2.2: The sampler undergoes several rounds of simulated training to sample reinforcement learning performance data.
[0094] Step 2.3: The fitter fits the model hyperparameters based on the performance data.
[0095] Step 2.4: The solver calculates the theoretically optimal training cost while ensuring that the training time does not exceed the maximum deadline, and outputs the corresponding resource allocation.
[0096] Specifically, in step 2.4, the solver uses argmin (k1,k2) C sum ,stT sum ≤T ddl This is to ensure that the maximum deadline is not exceeded, which is T. ddl The solver solves the problem under these constraints.
[0097] Specifically, the actor process in step 1.1 corresponds to a virtual CPU of the server-invisible computing instance provided by the cloud platform, and the learner process in step 1.2 corresponds to a GPU on the server. The server-invisible computing instance and the server cluster are connected via a high-speed network.
[0098] Specifically, the number of actor processes in step 1.1 is not fixed. It can be flexibly adjusted as the model training process progresses.
[0099] Assume that the maximum number of server-independent computing instances supported by the system is N. max If the total number of training rounds is r, then the mathematical expression for calculating the number of instances in the i-th round is:
[0100]
[0101] Where * represents product.
[0102] Specifically, the relationship between sample data and the adjustment of the number of actor processes; the focus of this invention is not on the specific adjustment method, but rather on providing a "capability" to dynamically adjust during training after deploying actor processes on server-side computational instances—a capability not available in training systems deployed on servers. This capability is orthogonal to the algorithm layer's adjustment of the number of actor processes; dynamically increasing the number of actor processes during training can improve model training performance, and the system described in this invention can support this algorithm.
[0103] In steps 1.1 and 1.5, the default memory of the new server-invisible computing instance is allocated as 1GB. When the default memory is insufficient to run the actor process, the default memory is increased in multiples of 2: 2GB, 4GB, 8GB, etc., until the actor process can be run.
[0104] Specifically, step 2.1 includes:
[0105] Step 2.1.1: The modeler builds a performance model with hyperparameters for the distributed actor-learner reinforcement learning algorithm implemented based on the Ray framework. The performance model uses a linear model containing random variables a1, a2, b1, and b2 to characterize the relationship between the data volume d, the number of CPU processes k1, the number of GPU processes k2, and sampling and training times. Sampling time: T a (d,k)=a1+b1*d / k, training time: T c (d,k)=a2+b2*d / k.
[0106] Specifically, the performance model characterizes the number of resources, i.e., the number of CPUs and GPUs, and the training time, i.e., T. sum The relationship.
[0107] Specifically, the performance model "distributed actor-learner reinforcement learning algorithm" is divided into on-policy and off-policy types. The modeler must be able to model both types of algorithms. The following T... sum =T a +T b +T c +T d This is the performance model under the on-policy condition.
[0108] For on-policy reinforcement learning training algorithms, such as the proximal policy optimization algorithm and the PPO algorithm, the total time T of the performance model is... sum It is the sampling time T a Trajectory transmission time T b Training time T c Parameter synchronization time T d The sum of the four stages, i.e., T sum =T a +T b +T c +T d Among them, T b With T d It can be obtained through sampling.
[0109] For off-policy reinforcement learning training algorithms, such as the importance-sampling-based distributed reinforcement learning training algorithm IMPALA, the sampling time T... a and training time T c There may be overlap, and the total training time of the model cannot be obtained by simply adding them together.
[0110] Specifically, using the p-norm This model uses the overlap between sampling time and training time to model the relationship between them. The overlap ratio is adjusted by regulating the hyperparameter λ. When λ = 1, it degenerates into T. sum =T a +T b +T c +T d This corresponds to the case where there is no coverage of sampling time and training time. When λ approaches infinity, T... sum =max(T) a +T b ,T c +T d This corresponds to the situation where the sampling time and training time are completely covered.
[0111] Specifically, a highly adaptable model should be able to broadly represent different reinforcement learning algorithms. Reinforcement learning algorithms are divided into two types: on-policy and off-policy. For on-policy reinforcement learning algorithms, "sampling" and "training" are events that occur sequentially on the same timeline. Therefore, modeling the performance model, i.e., the relationship between total training time and resources, simply involves adding up the times of each stage.
[0112] For off-policy reinforcement learning training algorithms, the sampling time T mentioned above... a and training time T c These are not events that occur sequentially on the same timeline, but rather events that occur in parallel. Therefore, the direct summation of events in the on-policy case to obtain T_sum cannot be applied to off-policy reinforcement learning algorithms and cannot accurately represent the total training time of off-policy algorithms.
[0113] Step 2.1.2: The modeler builds a cost model with hyperparameters for the distributed actor-learner reinforcement learning algorithm implemented based on the Ray framework. The total cost C of the cost model. sum It is the sum of the CPU and GPU usage costs of the elastic computing instance, and the CPU and memory usage costs of the server-independent computing application platform. The mathematical expression is:
[0114] C sum =T a (d,k1)*α*k1+T c (d,k2)*β*k2+T sum *γ
[0115] Where α, β, and γ represent the price per second of server-insensitive computing instances, the price per second of GPU servers, and the average price per second of VPC networks, respectively.
[0116] Specifically, the sampler in step 2.2 runs reinforcement learning simulation training five times in a server-to-server non-aware computing environment.
[0117] Specifically, step 2.3 includes:
[0118] Step 2.3.1: The fitter reads the simulated training data from Step 2.2 and performs regression analysis.
[0119] Step 2.3.2: The fitter determines the values of the hyperparameters α1, a2, b1, b2 of the performance model and cost model described in Step 2.1. For off-policy reinforcement learning training scenarios, the fitter determines the hyperparameter λ in the p-norm model based on the ratio of the actual sampling time to the training time.
[0120] The hyperparameter λ is used to characterize the "coverage ratio"; the larger the hyperparameter λ is, the larger the coverage ratio is. When the hyperparameter λ exceeds 100, it can be approximately considered that the coverage is complete.
[0121] A smaller hyperparameter λ indicates a lower coverage ratio; a hyperparameter λ of 1 indicates no coverage. Once this hyperparameter λ is determined, the modeling of the off-policy scenario is complete. The performance model can be used to estimate the total runtime T under other resource configurations (GPU, CPU) quantities. sum .
[0122] After the sampler runs reinforcement learning W times, we can obtain the execution time (T) for W groups. a1 ,T b1 ,T c1 ,T d1 ,T sum1 ),(T a2 ,T b2 ,T c2 ,T d2 ,T sum2 ),……(T aW ,T bW ,T cw ,T dW ,T sumW These W sets of data can be used to fit the function. The value of the parameter λ. The fitting method used is BFGS, short for Broyden-Fletcher-Goldfarb-Shanno.
[0123] Step 2.3.3: The fitter uses the R-squared test to verify the accuracy of the fit. When the correlation coefficient of the variables, i.e. the R-squared value is greater than 0.6, the hyperparameter is considered reliable.
[0124] Specifically, step 2.4 includes:
[0125] Step 2.4.1: The solver uses a quasi-Newton method, namely L-BFGS-B, to perform gradient descent to find the feasible region of optimal resource allocation, including the range E = {(k1,k2)|k1,k2∈N} for both CPU and GPU. The solver's optimization objective is to minimize the training cost while ensuring training performance, mathematically expressed as: argmin (k1,k2) C sum ,stT sum ≤T ddl .
[0126] Step 2.4.2: The solver enumerates in the above feasible region E to find the resource configurations k1 and k2 that minimize training cost while ensuring training performance.
[0127] Specifically, the server is an Elastic Compute Instance (ECS) equipped with several vCPUs and GPUs; the server-aware computing instance is a computing instance of the server-aware computing application platform, namely, a computing instance of SAE; the connection between the server and the server-aware computing instance is through a Virtual Private Cloud (VPC).
[0128] Example 1, as Figure 1 As shown in Figure a, this invention combines server-insensitive computing with reinforcement learning training methods. The following section uses a specific training process as an example. Figure 1 a. A detailed description of the training process for reinforcement learning:
[0129] In step 101, the server-side computation instance runs the actor process to obtain the training trajectory and then executes step 102.
[0130] In step 102, the actor process sends the training trajectory to the learner process running on the high-performance GPU server, proceeding to step 103.
[0131] In step 103, the server-unaware computing instance is released, and step 104 is performed.
[0132] In step 104, the learner process performs reinforcement learning training, updates gradient parameters, and proceeds to step 105.
[0133] In step 105, it is determined whether the model has converged. If it has converged, the process ends. Otherwise, step 106 is executed. Specifically, the convergence condition of the model is that the model's loss function is less than a preset threshold or reaches a preset maximum number of training epochs.
[0134] In step 106, a new server-agnostic computing instance is allocated based on the latest gradient parameters, and step 101 is executed.
[0135] like Figure 1 As shown in b, this is the specific process of the resource allocation optimization system of the present invention. The following section uses a specific optimization process as an example, combined with… Figure 1 b provides a detailed description of resource allocation optimization:
[0136] In step 201, the modeler establishes a performance model and a cost model with hyperparameters for the distributed actor-learner proximal policy optimization algorithm (PPO) implemented based on the Ray framework. The performance model uses a linear model with random variables to characterize the relationship between data volume and sampling time and training time. The total time of the performance model is the sum of the four stages: sampling time, trajectory transmission time, training time, and parameter synchronization time, i.e., T. sum The mathematical expression is:
[0137] T sum =T a +T b +T c +T d
[0138] The total cost of the cost model is the sum of the CPU and GPU usage costs of the elastic computing instance, the CPU and memory usage costs of the server-insensitive computing application platform, and then step 202 is executed.
[0139] In step 202, the sampler runs a reinforcement learning training session in a server-to-server seamless computing environment and records the training performance data. Then, step 203 is executed.
[0140] In step 203, it is determined whether the number of sampling rounds has reached 5; if the result is yes, step 204 is executed; if the result is no, step 202 is executed.
[0141] In step 204, the fitter uses regression analysis based on the training performance data to derive the hyperparameters of the performance model and cost model, thus obtaining the training time T. sum Training cost C sum To determine the relationship between the data volume d and the number of CPUs, i.e., GPUs (k1, k2), proceed to step 205.
[0142] The training cost C sum The mathematical expression is:
[0143] C sum =T a (d,k1)*0.01*k1+T c (d,k2)*0.02*k2+T sum *0.005
[0144] T a(d,k1)=0.183+0.2*d / k1
[0145] T c (d,k2)=0.192+0.3*d / k2
[0146] In step 205, the R-squared test is used to verify that the correlation coefficient of the fit is 0.83, which is greater than 0.6, and then step 206 is executed.
[0147] In step 206, the solver solves the optimization problem by combining the L-BFGS-B gradient descent method and the enumeration method based on the performance model and the cost model; thus obtaining the training time T for each round. sum Minimize the training cost C under the constraint of less than 1 second. sum The CPU configuration k1 is 48, and the GPU configuration k2 is 3.
[0148] The mathematical expression for the optimization problem is:
[0149]
[0150] like Figure 2 As shown, the reinforcement learning training system of this invention, which combines server-side seamless computing, is divided into a control layer and an execution layer. The user sends commands to the command parser in the control layer, which then calls the modules in the execution layer to complete the training. First, the sampler performs five simulations based on the model parameters, and the sampled data is sent to the fitter. The fitter performs regression analysis to determine the hyperparameters of the performance and cost models, and sends them, along with additional parameters such as time constraints, to the solver. The solver, under the conditions of satisfying time constraints, solves for the parameter configuration of the theoretically optimal cost and outputs it to the user.
[0151] like Figure 3 As shown, this invention combines an elastic computing instance, namely ECS, and a server-insensitive computing application instance, namely SAE, through a virtual private network, namely VPC, into a reinforcement learning training system.
[0152] This invention executes the actor process of reinforcement learning in SAE, releasing it immediately after each round to save training costs. This invention executes the learner process of reinforcement learning in ECS. ECS and SAE are connected via a VPC network. Sampled data from the actor process is transmitted to the learner through the VPC network, and the learner's latest parameters are also transmitted to the new actor through the VPC network. Users interact with the console via Python commands to set training tasks and obtain training results.
[0153] The present invention also provides a server-aware computing reinforcement learning resource allocation subsystem, which can be implemented by executing the process steps of the server-aware computing reinforcement learning resource allocation method. That is, those skilled in the art can understand the server-aware computing reinforcement learning resource allocation method as a preferred embodiment of the server-aware computing reinforcement learning resource allocation subsystem.
[0154] The present invention also provides a server-agnostic computing reinforcement learning training system, which can be implemented by executing the process steps of the server-agnostic computing reinforcement learning training method. That is, those skilled in the art can understand the server-agnostic computing reinforcement learning training method as a preferred embodiment of the server-agnostic computing reinforcement learning training system.
[0155] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0156] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A server-insensitive computing method for reinforcement learning resource allocation, characterized in that, include: Step S1: Establish performance and cost models based on distributed reinforcement learning algorithms; Step S2: Based on the performance model and cost model, sample reinforcement learning performance data on a hardware platform equipped with CPU and GPU; Step S3: Fit the model hyperparameters based on the reinforcement learning performance data; Step S4: Based on the model hyperparameters, calculate the optimal training cost on a hardware platform equipped with CPU and GPU, and output the corresponding resource configuration; The resource configuration includes the number of CPUs and the number of GPUs.
2. The server-insensitive computing reinforcement learning resource allocation method according to claim 1, characterized in that, The distributed reinforcement learning algorithm refers to the distributed actor-learner reinforcement learning algorithm, i.e., the Actor-Learner reinforcement learning algorithm. In step S1, a performance model and a cost model containing hyperparameters are established based on the distributed actor-learner reinforcement learning algorithm. The performance model includes: a performance model for the same-track strategy and a performance model for the different-track strategy; the performance model for the same-track strategy includes the PPO algorithm; the performance model for the different-track strategy includes the IMPALA algorithm. The mathematical expression for the total time of the performance model of the same-track strategy is: T sum =T a +T b +T c +T d Among them, T a Indicates the sampling time; T b Indicates the trajectory transmission time; T c Indicates training time; T d Indicates the parameter synchronization time; For the performance model of the aforementioned off-track strategy, the p-norm is used to model the coverage relationship between sampling time and training time; The coverage relationship, i.e. the coverage ratio, is adjusted by the hyperparameter; The mathematical expression for the covering relationship is: in, This indicates the coverage relationship, and λ represents the hyperparameter. The mathematical expression for the cost model is: C sum =T a (d,k1)*α*k1+T c (d,k2)*β*k2+T sum *c Among them, C sum α represents the total cost; β represents the price per second of a server-independent compute instance; β represents the price per second of a GPU server; γ represents the average price per second of a VPC network; k1 represents the number of CPU processes, k2 represents the number of GPU processes; d represents the amount of data represented by the linear model; T a (d,k1) represents the sampling time for k1 CPU processes; T c (d,k2) represents the training time with k2 GPU processes.
3. The server-insensitive computing reinforcement learning resource allocation method according to claim 2, characterized in that, In step S2, n rounds of simulated training are performed to collect reinforcement learning performance data. The reinforcement learning performance data includes: the execution time of the sampling process in each round of reinforcement learning, the transmission time of the sampling process sending the sampled data to the training process, the execution time of the training process, the time of the training process updating the updated model weight parameters back to the sampling process, and the total time of each round of reinforcement learning; where n is greater than or equal to 5; In step S3, based on the reinforcement learning performance data, the model hyperparameters are fitted and obtained; the fitting method is polynomial fitting or linear regression fitting. The model hyperparameters include: random variable a1, random variable a2, random variable b1, random variable b2 and hyperparameter λ; In step S3, the R-squared test is used to verify the accuracy of the fit. It is determined whether the correlation coefficient of the variables, i.e., the R-squared value, is greater than 0.
6. If the result is yes, then step S4 is executed; if the result is no, then the process ends.
4. The server-insensitive computing reinforcement learning resource allocation method according to claim 3, characterized in that, In step S4, the solution time for the optimal training cost is less than the maximum deadline; the mathematical expression for the optimal training cost is: argmin (k1,k2) C sum ,s.t.T sum ≤T ddl Where, argmin (k1,k2) This indicates that C sum That is, the total cost is minimized at (k1,k2); st represents the value at time T. sum ≤T ddl Under the constraints, T ddl This indicates the maximum deadline.
5. A server-insensitive computing reinforcement learning training method, characterized in that, The reinforcement learning resource allocation method using server-insensitive computing as described in claim 4 includes: Step A1: Allocate N server-insensitive computing instances and run actor processes and learner processes; N is a constant; Step A2: Release the server-insensitive computing instance; Step A3: Employ reinforcement learning resource allocation methods to train the learner process and update the model's gradient parameters; Step A4: Based on the gradient parameters of the model, elastically allocate N' new server-insensitive computing instances, and re-execute step A1 until the model converges; N' is a constant; In step A4, the model converges when the model's loss function is less than a preset threshold or reaches a preset maximum number of training rounds.
6. The server-insensitive computing reinforcement learning training method according to claim 5, characterized in that, In step A1, the number of actor processes is expressed mathematically as follows: Where, N i N represents the number of computational instances in the i-th round; max This represents the maximum number of server-insensitive computing instances supported by the system; r represents the total number of training rounds.
7. The server-insensitive computing reinforcement learning training method according to claim 6, characterized in that, In step A4, the range of the elastic allocation is constrained by the maximum number of processes supported by the server-insensitive computing platform and the current training batch size, and the range is [1, 128].
8. The server-insensitive computing reinforcement learning training method according to claim 7, characterized in that, In step A4, the default memory allocated to the new server-invisible computing instance is 1GB; it is determined whether the default memory is sufficient to run the actor process. If the result is yes, no further processing is performed, and the actor process is run. If the result is negative, the default memory is increased by multiples of 2 until the actor process is run.
9. A server-insensitive computing-based reinforcement learning resource allocation subsystem, characterized in that, include: Modeler module: Builds performance and cost models based on distributed reinforcement learning algorithms; Sampler module: Based on the performance model and cost model, sample reinforcement learning performance data on a hardware platform equipped with CPU and GPU; Fitter module: Fits model hyperparameters based on the reinforcement learning performance data; Solver module: Based on the model hyperparameters, calculates the optimal training cost on a hardware platform equipped with CPU and GPU, and outputs the corresponding resource configuration; the resource configuration includes: the number of CPUs and the number of GPUs.
10. A server-insensitive computation reinforcement learning training system, characterized in that, Triggering the operation of the server-insensitive computing reinforcement learning resource allocation subsystem as described in claim 9 includes: Module A1: Allocate N server-insensitive computing instances to run actor processes and learner processes; N is a constant; Module A2: Release the server-side non-transparent computing instance; Module A3: Triggers the operation of the reinforcement learning resource allocation subsystem, initiates the reinforcement learning training process for the learners, and updates the gradient parameters of the model; Module A4: Based on the gradient parameters of the model, N' new server-insensitive computing instances are elastically allocated, and Module A1 is reactivated until the model converges; N' is a constant. In module A4, the model converges when the model's loss function is less than a preset threshold or reaches a preset maximum number of training rounds.
Citation Information
Patent Citations
Large-scale reinforcement learning training framework system based on distributed design
CN114861826A
Strategy optimization method and device for Linux system kernel configuration based on machine learning
CN117172093A