Non-Uniform Memory Access Resource Allocation Method Based on Parallel Adaptive Auction Algorithm
Through parallel adaptive auction algorithm and DQN model, combined with heterogeneous multi-threading technology, the parallelization problem of resource allocation in NUMA architecture is solved, efficient and stable resource allocation is achieved, and system performance and computing efficiency are improved.
Patent Information
- Application Number
- CN202410987081.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-23
AI Technical Summary
在非统一内存访问(NUMA)架构中,现有算法难以有效并行化处理带有位置约束的资源分配问题,导致计算效率低下和资源分配不均衡。
Using a parallel adaptive auction algorithm, combined with deep Q network (DQN) model and heterogeneous multithreading technology, the resource allocation is optimized through the auction mechanism in economics to achieve independent competition and efficient allocation of tasks.
It significantly improves computing efficiency, improves the parallelism and stability of resource allocation, optimizes system performance, especially in large-scale task processing, which improves computing performance by 20 times and resource allocation efficiency by 19%.
Smart Images

Figure CN118981372B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer resource scheduling, and particularly relates to a non-uniform memory access resource allocation method based on a parallel adaptive auction algorithm. Background Art
[0002] The non-uniform memory access (NUMA) architecture is a computer memory hardware design architecture for multi-processor systems. In this architecture, processors and memory are organized into multiple nodes, each node having its own memory and multiple processors. The NUMA architecture allows processors to directly access the memory within the local node, thereby reducing memory access latency and improving system performance. In the field of computer underlying operating systems, the NUMA architecture is widely applied to scenarios such as memory management and task scheduling. The knapsack problem is an important combinatorial optimization problem and is widely applied to fields such as production scheduling, cloud computing, network communication, and asset management. In the field of computer underlying operating systems, different types of knapsack problems can be applied to memory management and task scheduling. In this case, task scheduling can be conceptualized as a dynamic and real-time online market, in which various tasks (or "buyers") compete for limited resources (or "goods") with the aim of optimizing and allocating the limited resources according to requirements and priorities.
[0003] To better optimize the resource allocation problem in the NUMA architecture, researchers have proposed various extended versions of the knapsack problem, including the multidimensional knapsack problem, the unbounded knapsack problem, the bounded knapsack problem, and the multiple knapsack problem, etc. These extended versions can more accurately simulate complex scenarios in the real world. However, in production scheduling and hardware resource allocation, the relative position constraints between items must be considered. This extended knapsack problem with position constraints increases the complexity of the problem, and not only the value and weight of the items need to be considered, but also specific constraints need to be imposed on the positions of the items in the solution. The non-uniform memory access (NUMA) architecture is an important application of the extended knapsack problem and is a computer memory hardware design architecture for multi-processor systems. In the NUMA architecture, processors and memory are organized into multiple nodes, each node having its own memory and multiple processors (as Figure 1 shown); this design allows processors to directly access the memory within its local node without having to cross other nodes, which helps to reduce memory access latency and improve system performance. Compared with the traditional system architecture, the NUMA architecture has more advantages in terms of scalability and reducing memory latency.
[0004] In a Non-Uniform Memory Access (NUMA) architecture, optimizing memory access latency is crucial for improving the overall system performance. Therefore, solving the knapsack problem with location constraints is of great significance for the system's performance and energy efficiency. However, such complex large-scale optimization problems require allocating resources to limited nodes, and traditional optimization methods include dynamic programming algorithms and greedy algorithms. These methods have the following drawbacks in a parallel computing environment:
[0005] For the dynamic programming algorithm, it has the disadvantages of reference interdependence, storage conflicts, and uneven load: Since the current solution usually depends on the solutions of previously computed sub-problems, the interdependence between sub-problems makes it difficult to independently partition sub-problems into subtasks for parallel computing; at the same time, the dynamic programming algorithm usually needs to frequently access and update storage structures (such as tables or arrays), and in a parallel environment, multiple processors or threads accessing and modifying the same storage location simultaneously may lead to conflicts and inconsistencies; also, different stages of dynamic programming may require different amounts of computation, which makes it difficult to evenly distribute the workload among processors.
[0006] For the greedy algorithm, it has the disadvantages of local optimality, non-online sequential decision-making, and lack of a global perspective: It is an algorithm that takes the current optimal solution in each step of the selection. However, since it makes locally optimal decisions in each step, these decisions usually need to depend on the current state; and when involving a sequential decision-making process, each decision is based on the previous one, and this continuity makes parallelization complex; also, since it usually does not have a global perspective and only focuses on the optimal solution of the current step, it may lead to inconsistent or inefficient decisions in a parallel environment. In addition, in a complex and uncertain environment, making robust decisions for different states is crucial; to solve these two problems, researchers have proposed some methods, such as GPU-accelerated dynamic programming algorithms and distributed computing optimization schemes for sub-problem decomposition. However, these methods cannot fully decouple subtasks. Summary of the Invention
[0007] The main objective of the present invention is to overcome the drawbacks and deficiencies of the prior art, and provide a non-uniform memory access resource allocation method based on a parallel adaptive auction algorithm, which combines economic theory and computer science to achieve decentralized and efficient resource allocation.
[0008] To achieve the above objective, the present invention adopts the following technical solutions:
[0009] On the one hand, a non-uniform memory access resource allocation method based on a parallel adaptive auction algorithm is provided, including the following steps:
[0010] When the operating system starts, obtain the node information and task information in the non-uniform memory access architecture; the node information includes the number of nodes, the starting position of the nodes, and the ending position of the nodes; the task information includes the number of tasks and the number of nodes occupied by the tasks.
[0011] Regard the resource problem of the non-uniform memory access architecture as a knapsack problem with position constraints, and use the parallel adaptive auction algorithm for optimization to obtain the allocation result.
[0012] As a preferred technical solution, the resource allocation problem of the non-uniform memory access architecture is regarded as an extended knapsack problem with position constraints, which is expressed as:
[0013]
[0014] W i =(N - i) 0.5 , 0 ≤ i < N,
[0015]
[0016] where (position i < position j + width j ) and (position j < position i + width i ),
[0017] Among them, Score() is the objective function; W i is the weight of the i-th task; width i is the amount of resources occupied by the i-th task; x i is the status of the i-th task. When x i = 1, it means the i-th task is selected. When x i = 0, it means the i-th task is not selected; N is the number of tasks; C is the knapsack capacity, that is, the maximum amount of resources that all selected tasks can occupy cannot exceed; position i is the position information of the i-th task, l i is the lower limit of the position of the i-th task, u i is the upper limit of the position of the i-th task; x j is the status of the j-th task; position j is the position information of the j-th task.
[0018] As a preferred technical solution, the process of using the parallel adaptive auction algorithm for optimization is as follows:
[0019] Use the trained DQN model to select the best α value for each task that suits the current environment and state; the best α value is used to adjust the bidding price of each task.
[0020] Adopt heterogeneous multi-threading to calculate the bidding prices of the starting positions of each task in parallel for bidding, and store them in the task list; the calculation formula for the bidding price is: bid i =W i *α i , where bid i is the bidding price of the i-th task, and α i is the best α value of the i-th task.
[0021] Allocate tasks based on the task list to obtain the final starting positions and the list of the best α values of all tasks.
[0022] As a preferred technical solution, the DQN model includes a current Q network and a target Q network with the same architecture but different parameters, and an experience replay buffer.
[0023] The current Q network is used to predict the Q values of each possible bidding action of each task in the current task state, and select the bidding action with the highest Q value as the optimal bidding action in the current task state.
[0024] The target Q network is used to calculate the target Q value in the next task state, and the parameters of the target Q network are regularly copied from the current Q network.
[0025] The experience replay buffer is used to store the experiences in the training process of the DQN model and update the parameters of the current Q network.
[0026] As a preferred technical solution, the use of the trained DQN model to select the best α value for each task that suits the current environment and state is specifically as follows:
[0027] Randomly extract a small batch of experience data from the experience replay buffer; each piece of experience data includes the task state, bidding action, reward of the task at the current time step, and the task state, bidding action at the next time step.
[0028] Use the target Q network in the trained DQN model to calculate the target Q value of the task, and select the bidding action with the highest target Q value of the task as the best α value in the current task state; the calculation formula for the target Q value is:
[0029] Q(state t ,action t )=reward t +γmax(action t+1 )*Q(state t+1, action t+1 ; θ - ),
[0030] where Q(state t , action t ) is the target Q-value of the task at the current time step t, state t is the task state of the task at the current time step t, action t is the bidding action of the task at the current time step t, reward t is the reward of the task at the current time step t; γ is the discount factor, used to balance the proportion of the reward at the current time step and the reward at the next time step; Q(state t+1 , action t+1 ; θ - ) is the target Q-value of the task at the next time step t+1 with parameter θ - ; max(action t+1 ) is the maximum target Q-value under the next task state predicted by the target Q-network; θ - is the parameter of the target Q-network at the next time step t+1; state t+1 is the task state of the task at the next time step t+1; action t+1 is the bidding action of the task at the next time step t+1.
[0031] As a preferred technical solution, the DQN model uses the mean squared error function as the loss function, expressed as:
[0032]
[0033] δ i = Q(state t , action t ) - Q(state t , action t ; θ),
[0034] where δ i is the TD error of the i-th experience, representing the difference between the predicted Q-value of the current Q-network and the target Q-value of the target Q-network; θ is the parameter of the current Q-network; Num ER is the number of experiences in the experience replay buffer; Q(state t , action t ) is the target Q-value output by the target Q-network, Q(state t , action t ; θ) is the predicted Q-value output by the current Q-network under parameter θ; state tis the task status of the task at the current time step t, action t is the bidding action of the task at the current time step t.
[0035] As a preferred technical solution, heterogeneous multi-threading is used to parallelly calculate the bidding prices of the starting positions of each task for bidding, specifically:
[0036] Initialize a task list tasks, a thread pool thread_pool, and a shared result structure shared_results; the task list tasks contains all the subtasks that need to be executed, that is, the bidding tasks after each task calculates the bidding price of the starting position; the thread pool thread_pool is used to allocate threads to execute the subtasks; the shared result structure shared_results is used to store the execution results of each subtask;
[0037] Allocate a heterogeneous thread async_thread for each subtask from the thread pool and execute the subtask, store the execution result of the subtask into shared_results until all threads have finished executing;
[0038] Process the execution results of all subtasks in shared_results to generate the final result final_result.
[0039] As a preferred technical solution, the tasks are allocated based on the task list, specifically:
[0040] When the task list is not empty, take out the task pattern with the highest bidding price and the task starting position pos from the task list;
[0041] Judge whether the position range from the task starting position pos to the task ending position pos + grid[pattern][0] is covered, where grid[pattern][0] represents the amount of resources required for the task pattern;
[0042] If it is not covered, update the final starting position final_positions[pattern] and the coverage status covered of the task;
[0043] If it is already covered, update the task list, remove the tasks conflicting with the current position, and re-stack the task list;
[0044] Repeat the above steps until all tasks are allocated, and return the final starting positions final_positions of all tasks.
[0045] On the other hand, a non-uniform memory access resource allocation system based on a parallel adaptive auction algorithm is provided, including an information acquisition module and a solution allocation module;
[0046] The information acquisition module is used to obtain node information and task information in the non-uniform memory access architecture when the operating system starts; the node information includes the number of nodes, the starting position of the nodes, and the ending position of the nodes; the task information includes the number of tasks and the number of nodes occupied by the tasks;
[0047] The solution allocation module is used to regard the resource problem of the non-uniform memory access architecture as a knapsack problem with position constraints, and use a parallel adaptive auction algorithm for optimal solution to obtain the allocation result.
[0048] On the other hand, the present invention also provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned non-uniform memory access resource allocation method based on a parallel adaptive auction algorithm.
[0049] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0050] 1. Heterogeneous multi-threaded parallel acceleration computing:
[0051] When dealing with the extended knapsack problem with position constraints, it is difficult for traditional methods to effectively decompose the main problem into independent parallel computing subtasks, resulting in low efficiency. However, through the idea of the auction mechanism, the present invention makes the bidding strategies of each task independent of each other, and can decompose the main problem into independent parallel computing subtasks, significantly improving the computing efficiency.
[0052] 2. Efficient parallel adaptive auction algorithm:
[0053] Traditional algorithms have limitations in parallel processing and resource allocation, especially when dealing with a large amount of data and complex constraints; while the present invention breaks through the parallelization difficulties of traditional algorithms for the first time, and based on the auction mechanism in economics, realizes more effective and efficient allocation of memory resources; this way of combining economic theory and computer science provides a new perspective for parallel processing.
[0054] 3. Adaptive price bidding strategy based on the DQN model:
[0055] Traditional auction mechanisms are based on static optimal solution strategies and lack research on dynamic problems; and the stability of traditional DQN models is limited when implementing the dynamic optimal bidding strategy of tasks in the auction mechanism. The present invention combines the DQN model with the solution of the advertiser strategy under the auction mechanism, and optimizes the stability of the traditional DQN model, realizing the autonomous decision-making of tasks based on local information, thus achieving decentralized and efficient resource allocation. Description of the Drawings
[0056] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0057] Figure 1 It is a schematic diagram of a non-uniform memory access (NUMA) architecture in the prior art.
[0058] Figure 2 It is a flowchart of a non-uniform memory access resource allocation method based on a parallel adaptive auction algorithm in the embodiment.
[0059] Figure 3 It is a framework diagram of the parallel adaptive auction algorithm in the embodiment of the present invention.
[0060] Figure 4 It is a schematic diagram of parallel and concurrent computing in the embodiment of the present invention.
[0061] Figure 5 It is a comparison diagram of the average computing time between the parallel adaptive auction algorithm in the embodiment of the present invention and the existing algorithm.
[0062] Figure 6 It is a comparison diagram of the acceleration between the parallel adaptive auction algorithm in the embodiment of the present invention and the existing algorithm. Detailed implementation manners
[0063] In order to enable those skilled in the art of this technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0064] Referring to "embodiments" in the present application means that the specific features, structures or characteristics described in conjunction with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art understand explicitly and implicitly that the embodiments described in the present application can be combined with other embodiments.
[0065] The following will first explain some technical terms in the present application. Others not elaborated should be explained according to the conventional technical terms in the art, unless otherwise specified.
[0066] Non-Uniform Memory Access (NUMA): It can be abstracted as a knapsack problem with location constraints. It is a special scenario of the knapsack problem and is difficult to solve by other algorithms. Non-Uniform Memory Access (NUMA) is a computer memory design for multi-processor systems, where each processor or core has its own local memory. In a NUMA architecture, a processor can access its local memory faster than accessing the memory of other processors. This design aims to improve the performance and scalability of multi-processor systems.
[0067] Auction algorithm / mechanism: It is an important mechanism design in game theory and has won the Nobel Prize in Economics many times. Although it is a method of mechanism design in itself, this application borrows and optimizes it as a method of resource allocation rather than as a mechanism design.
[0068] Parallel auction mechanism: Optimize the auction mechanism based on the two characteristics of parallelism and adaptability. It is an algorithm improvement for the knapsack problem with location constraints. Traditional dynamic programming and greedy algorithms are usually difficult to parallelize, mainly due to their algorithm structures and processing methods. However, this invention no longer relies on traditional methods such as dynamic programming, but for the first time performs algorithm parallelization under the auction mechanism, significantly improving the calculation efficiency.
[0069] Parallel computing: A computing architecture in which multiple processors execute computing tasks simultaneously to improve computing speed and efficiency. This method is usually used to handle large-scale and complex computing problems, such as scientific simulations, big data analysis, and complex graphics processing. Data parallelism: In data parallelism, the data set is divided into smaller parts and then processed in parallel; each processor performs the same operation on different subsets of data. In parallel computing, a large problem is usually decomposed into many small parts, which can be independently calculated on different processors simultaneously.
[0070] DQN (Deep Q-Network) model: It combines deep learning and reinforcement learning methods and is mainly used to solve decision-making problems with high-dimensional input spaces. It is an important breakthrough in the field of reinforcement learning. The DQN model uses a deep neural network to approximate the Q-value function, which enables the algorithm to handle high-dimensional input spaces and learn from complex environments.
[0071] Experience replay: DQN breaks the correlation between data by storing the experiences (states, actions, rewards, etc.) of tasks in a memory bank and randomly sampling from this bank during training, which helps improve learning stability and efficiency. Fixed Q-target: To further improve learning stability, the DQN model uses two networks: the current Q-network for current value estimation and the target Q-network for target value estimation; the parameters of the target Q-network are periodically updated from the current Q-network, which helps reduce the correlation between the target and the estimated value.
[0072] In a Non-Uniform Memory Access (NUMA) architecture, optimizing memory access latency is crucial for improving the overall system performance. Therefore, the solution to the knapsack problem with location constraints plays a vital role in the system's performance and energy efficiency. The goal of this application is to allocate appropriate node resources to handle different applications, thereby improving system performance. Based on this, this application proposes a resource allocation method that can not only maximize system performance but also take into account location constraints (i.e., reduce memory latency and ensure that the CPU and memory performance within the node can support task execution), ensuring that processes with different performance requirements can be allocated to appropriate nodes, thus optimizing system performance. As Figure 2 shown, the non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm in this embodiment includes the following steps:
[0073] When the operating system starts, obtain the node information and task information in the non-uniform memory access architecture; where the node information includes the number of nodes, the starting position of the node, and the ending position of the node; the task information includes the number of tasks and the number of nodes occupied by the tasks (i.e., the resource occupancy of the tasks);
[0074] Regard the node allocation problem of the non-uniform memory access architecture as a knapsack problem with location constraints, and use the parallel adaptive auction algorithm to solve it to obtain the allocation result.
[0075] Furthermore, compared with Uniform Memory Access (UMA), the Non-Uniform Memory Access (NUMA) architecture has lower memory access latency because the processor can access the local memory node faster than accessing a remote node, thereby improving system performance. To make full use of this advantage, the present invention adopts an extended knapsack problem with location constraints. When allocating processes on nodes with different CPU performance and local memory capacities, consider the characteristics of the local memory to ensure that processes with different performance requirements can be allocated to appropriate nodes, thus optimizing system performance. Therefore, based on the local memory advantage of the above hardware architecture and the different performance requirements of processes, the goal of the present invention is to maximize the utilization rate of system resources in the NUMA architecture under the premise of meeting location constraints. Therefore, the node allocation problem of the NUMA architecture is regarded as an extended knapsack problem with location constraints, expressed as:
[0076]
[0077] W i =(N - i) 0.5 , 0 ≤ i < N,
[0078]
[0079] where (position i < position j + width j ) and (position j < position i + width i ),
[0080] where Score() is the objective function, which needs to maximize the score of the task, and a high score indicates a higher efficiency of resource allocation; W i is the weight of the i-th task; width i is the amount of resources occupied by the i-th task (in the context of the knapsack problem, width usually refers to the resources required by the task, such as time, space, bandwidth, etc., so in the constraint conditions, the sum of the widths of all selected tasks cannot exceed the knapsack capacity C); x i is the state of the i-th task. When x i = 1, it means the i-th task is selected, and when x i = 0, it means the i-th task is not selected; N is the number of tasks; C is the knapsack capacity, that is, the maximum amount of resources that the resources occupied by all selected tasks cannot exceed; position i is the position information of the i-th task, l i is the lower limit of the position of the i-th task, u i is the upper limit of the position of the i-th task; x j is the state of the j-th task; position j is the position information of the j-th task.
[0081] As Figure 3 shown, in order to prevent tasks with a large width in the previous process from being blocked due to the inability to obtain continuous nodes, the present invention assigns priorities in sequence. W i represents this weight; different processes have different weights to compete for nodes; of course, the memory and CPU performance required by different tasks in the process are also different, and the starting node is used for position constraints; the nodes required by these processes are also different, Figure 3 which are represented by continuous blocks in, thereby reducing the memory latency between continuous nodes. Optimizing the memory access latency in the NUMA architecture is crucial for maximizing the overall system performance. For example, Task 1( Figure 3Agent1) requires two nodes to complete. In the NUMA architecture, only nodes with starting positions 0 and 4 can achieve a match; once a match is successful, the node will be occupied and unable to complete other tasks; the utilization rate of system resources in the NUMA architecture is measured by the total score Score(), which seems boring but is the basic process.
[0082] Furthermore, for the constructed extended knapsack problem with position constraints, the present application uses the parallel adaptive auction (PAA-EKP) algorithm for optimal solution. The core idea of this algorithm is to regard task allocation as a market competition process, and allocate each resource (i.e., a node in the NUMA architecture) to the bidder with the highest bid (i.e., a task in the NUMA architecture). This method realizes the independent competition between tasks and promotes parallel processing; and the bidders evaluate resources according to their own needs and submit bid prices calculated in parallel using local information; then each node simultaneously allocates the resource to the bidder with the highest bid, thus achieving high parallelism. The steps are as follows:
[0083] Use the trained DQN model to select the best α value for each task that is suitable for the current environment and state, and use this to bid for the final starting position of the task; the best α value is used to adjust the bid price of each task in order to obtain a more favorable position in resource allocation;
[0084] Adopt heterogeneous multi-threading to calculate the bid prices for the starting positions of each task in parallel for bidding and store them in the task list; the bid price calculation formula is: bid i =W i *α i , where bid i is the bid price of the i-th task, and α i is the best α value of the i-th task;
[0085] Allocate tasks based on the task list to obtain the final starting positions of all tasks and the list of best α values.
[0086] Furthermore, the Deep Q-Network (DQN) is superior to heuristic and dynamic programming methods in complex environments and reduces computational complexity. By using a neural network for continuous state space approximation, the DQN model can effectively handle discrete and continuous state spaces, making it more scalable in large-scale problems. To improve stability and reduce spatio-temporal correlations between data, the DQN model in this application includes a current Q-network and a target Q-network with the same architecture but different parameters, as well as a replay buffer; among them, the current Q-network is used to predict the Q-values of each possible bidding action for each task in the current task state, and select the bidding action with the highest Q-value as the optimal bidding action in the current task state, responsible for selecting the optimal bidding action and calculating the loss at each time step; while the target Q-network is used to calculate the target Q-value in the next task state, which helps to improve the stability of learning; at the same time, the replay buffer is used to store the experiences during the training process of the DQN model and update the parameters of the Q-network.
[0087] Furthermore, the trained DQN model is used to select the best α value for each task that adapts to the current environment and state, specifically:
[0088] The replay buffer stores the experiences of the task during the training process in a data structure called the experience replay buffer; first, a small batch (mini-batch) of experience data is randomly sampled from the replay buffer, and each piece of experience data contains information such as the task state, bidding action, reward, and the task state and bidding action at the next time step of the task at the current time step; by this method, the temporal correlation between data can be broken, thereby improving the stability and efficiency of learning.
[0089] The action value function output by the target Q-network, that is, the Q-value, is calculated for each task; this value represents the expected utility or return that the task can obtain by taking a specific bidding action in a given state; during the auction process, the task can use the target Q-network in the trained DQN model to calculate the target Q-value of the task, and then select the bidding action with the highest target Q-value of the task as the best α value in the current task state, find the optimal bidding strategy in the auction, thereby improving the resource allocation and task allocation results; among them, the calculation formula of the target Q-value is:
[0090] Q(state t ,action t )=reward t +γmax(action t+1 )*Q(state t+1 ,action t+1 ;θ - ),
[0091] where, Q(statet , action t ) is the target Q-value of the task at the current time step t, state t is the task state of the task at the current time step t, action t is the bidding action of the task at the current time step t, reward t is the reward of the task at the current time step t; γ is the discount factor, used to balance the proportion of the reward of the task at the current time step and the reward of the next time step; Q(state t+1 , action t+1 ; θ - ) is the target Q-value of the task at the next time step t+1, parameter θ - ; max(action t+1 ) is the maximum target Q-value in the next task state predicted by the target Q network; θ - is the parameter of the target Q network at the next time step t+1; state t+1 is the task state of the task at the next time step t+1; action t+1 is the bidding action of the task at the next time step t+1.
[0092] During the training process, the parameters of the current Q network are continuously updated, while the parameters of the target Q network are periodically copied from the current Q network; in this embodiment, the mean squared error (MSE) function is used as the loss function to minimize the loss between the predicted Q-value and the target Q-value, expressed as:
[0093]
[0094] δ i = Q(state t , action t ) - Q(state t , action t ; θ),
[0095] where, δ i is the TD (Temporal Difference Error) error of the i-th experience, representing the difference between the predicted Q-value of the current Q network and the target Q-value of the target Q network; θ is the parameter of the current Q network; Num ER is the number of experiences in the experience replay buffer; Q(state t , action t ) is the target Q-value output by the target Q network, Q(state t , action t ; θ) is the predicted Q-value output by the current Q network under the parameter θ; statet is the task status for the task at the current time step t, action t is the bidding action for the task at the current time step t.
[0096] In the traditional Q-Learning algorithm, the Q-value update process involves the Q-values of the current and next states, which may lead to oscillations. The DQN model of this application decouples the Q-value calculation from the network parameter update process by implementing a target Q-network, thereby improving the stability of learning.
[0097] Furthermore, this application also uses heterogeneous multi-threading to calculate the bidding prices for the starting positions of each task in parallel for bidding, while concurrent computing cannot; as Figure 4 shown, in concurrent computing, multiple tasks alternate within the same time period, but in fact only one task is being executed at a certain moment, and the other tasks are waiting. This method simulates multi-tasking on a single-core processor but does not significantly improve the computing performance. Parallel computing greatly improves the computing speed and performance by utilizing available resources. Heterogeneous multi-threading can overcome the limitations of parallel computing and improve the performance of the optimization algorithm, which is crucial for solving large-scale complex problems where time and resources are of utmost importance. The steps for heterogeneous multi-threading parallel acceleration computing are as follows:
[0098] Initialize a task list tasks, a thread pool thread_pool, and a shared result structure shared_results; among them, the task list tasks contains all the subtasks that need to be executed, that is, the bidding tasks after each task calculates the bidding price for the starting position; the thread pool thread_pool is used to allocate threads to execute the subtasks; the shared result structure shared_results is used to store the execution results of each subtask;
[0099] Allocate a heterogeneous thread async_thread for each subtask from the thread pool and execute the subtask, and store the execution result of the subtask into shared_results until all threads have finished execution;
[0100] Process the execution results of all subtasks in shared_results to generate the final result final_result.
[0101] Furthermore, after obtaining the task list, the tasks are allocated, specifically as follows:
[0102] First, determine whether the task list is empty. When the task list is not empty, take out the task pattern with the highest bidding price and the task starting position pos from the task list;
[0103] Determine whether the position range from the start position pos of the task to the end position pos + grid[pattern][0] is covered, where grid[pattern][0] represents the amount of resources required for task pattern;
[0104] If it is not covered, update the final start position final_positions[pattern] and the coverage status covered of the task;
[0105] If it is already covered, update the task list, remove the tasks conflicting with the current position, and re-stack the task list;
[0106] Repeat the above steps until all tasks are allocated, and return the final start positions final_positions of all tasks.
[0107] In summary, the present invention can solve the extended knapsack problem with position constraints, which is crucial for various applications, including operating systems, production scheduling, especially online markets, where efficient resource allocation is essential for economic optimization. The present invention proposes an innovative and efficient parallel adaptive auction (PAA-EKP) algorithm to solve this problem. By combining the non-uniform memory access (NUMA) architecture concept and task scheduling for analysis and experiments, aiming at the problem that traditional optimization methods involve sequential dependencies, making it difficult to decompose the main problem into independent parallel computing subtasks, the present invention strategically optimizes the bid price through a robust DQN model; and the bid prices can be calculated and submitted in parallel, accelerating the allocation process of the auction. The results show that in large-scale tasks involving a large number of tasks, the present invention improves the computing performance by 20 times compared with dynamic programming and improves the resource allocation by 19%. As Figure 5 、 Figure 6As shown in the figure, in this embodiment, the PAA-EKP algorithm proposed in this application is used to conduct a comparative experiment with existing algorithms, and the average calculation time and acceleration are used to verify the performance of each method. It can be seen that the PAA-EKP algorithm proposed in the present invention is significantly superior to other algorithms. Among them, the multi-threaded and single-threaded situations of the PAA-EKP algorithm are experimentally compared. In addition, it is compared with the commonly used greedy algorithm, first fit algorithm, and dynamic programming algorithm in terms of average computational time and speedup factor, reflecting the performance improvement of different algorithms under the same conditions, and verifying the significant advantages of the PAA-EKP algorithm proposed in the present invention in terms of performance. It can be seen that the multi-threaded PAA-EKP algorithm has obvious computational efficiency and performance improvement when processing large-scale tasks.
[0108] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0109] Based on the same idea as the non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm in the above embodiment, the present invention also provides a non-uniform memory access resource allocation system based on the parallel adaptive auction algorithm, which can be used to execute the non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm. For the sake of convenience of description, in the structural schematic diagram of the non-uniform memory access resource allocation system embodiment based on the parallel adaptive auction algorithm, only the parts related to the embodiment of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than those shown, or combine some components, or arrange different components.
[0110] Another embodiment of the present invention provides a non-uniform memory access resource allocation system based on the parallel adaptive auction algorithm, including an information acquisition module and a solution allocation module;
[0111] The information acquisition module is used to obtain node information and task information in the non-uniform memory access architecture when the operating system starts; wherein, the node information includes the number of nodes, the starting position of the nodes, and the ending position of the nodes; the task information includes the number of tasks and the number of nodes occupied by the tasks;
[0112] The allocation solution module is used to regard the resource problem of the non-uniform memory access architecture as a knapsack problem with location constraints, and uses a parallel adaptive auction algorithm to optimize and solve it to obtain the allocation result.
[0113] It should be noted that the technical features and beneficial effects described in the above-mentioned embodiment of the non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm are applicable to the embodiment of the non-uniform memory access resource allocation system based on the parallel adaptive auction algorithm. For specific contents, please refer to the description in the embodiment of the method of the present invention, which will not be repeated here. It is hereby declared. In addition, in the implementation of the non-uniform memory access resource allocation system based on the parallel adaptive auction algorithm in the above-mentioned embodiment, the logical division of each program module is only an example. In actual application, according to needs, for example, for the configuration requirements of the corresponding hardware or the convenience of software implementation, the above-mentioned function allocation can be completed by different program modules, that is, the internal structure of the non-uniform memory access resource allocation system based on the parallel adaptive auction algorithm is divided into different program modules to complete all or part of the functions described above.
[0114] In one embodiment, a computer-readable storage medium is provided, which stores a program in a memory. When the program is executed by a processor, a non-uniform memory access resource allocation method based on a parallel adaptive auction algorithm is implemented, specifically:
[0115] When the operating system starts, node information and task information in the non-uniform memory access architecture are obtained; wherein the node information includes the number of nodes, the starting position of the nodes, and the ending position of the nodes; the task information includes the number of tasks and the number of nodes occupied by the tasks;
[0116] The resource problem of non-uniform memory access architecture is regarded as a knapsack problem with location constraints, and a parallel adaptive auction algorithm is used to optimize and solve it to obtain the allocation result.
[0117] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0118] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0119] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for non-uniform memory access resource allocation based on a parallel adaptive auction algorithm, characterized in that Including the following steps: When the operating system starts, obtain the node information and task information in the non-uniform memory access architecture; the node information includes the number of nodes, the starting position of the nodes, and the ending position of the nodes; the task information includes the number of tasks and the number of nodes occupied by the tasks; Regard the resource problem of the non-uniform memory access architecture as a knapsack problem with position constraints, and use a parallel adaptive auction algorithm for optimal solution to obtain the allocation result; The resource allocation problem of the non-uniform memory access architecture is regarded as an extended knapsack problem with position constraints, which is expressed as: W i = (N - i) 0.5 , 0 ≤ i < N, where (position i <position j + width j ) and (position j <position i + width i ), Among them, Score() is the objective function; W i is the weight of the i-th task; width i is the amount of resources occupied by the i-th task; x i is the status of the i-th task. When x i = 1, it means the i-th task is selected. When x i = 0, it means the i-th task is not selected; N is the number of tasks; C is the backpack capacity, that is, the maximum amount of resources that the resources occupied by all selected tasks cannot exceed; position i is the position information of the i-th task, l i is the lower limit of the position of the i-th task, u i is the upper limit of the position of the i-th task; x j is the status of the j-th task; position j is the position information of the j-th task; The process of using the parallel adaptive auction algorithm for optimal solution is as follows: Use the trained DQN model to select the best α value suitable for the current environment and state for each task; the best α value is used to adjust the bidding price of each task; Use heterogeneous multi-threading to calculate the bid price of the starting position of each task in parallel for bidding, and store it in the task list; the bid price calculation formula is: bid i = W i * α i , where bid i is the bid price of the i-th task, and α i is the optimal α value of the i-th task; Allocate tasks based on the task list to obtain the final starting position and the list of the best α values of all tasks; The DQN model includes a current Q network and a target Q network with the same architecture but different parameters, and a replay buffer; The current Q network is used to predict the Q values of each possible bidding action of each task in the current task state, and select the bidding action with the highest Q value as the optimal bidding action in the current task state; The target Q network is used to calculate the target Q value in the next task state, and the parameters of the target Q network are periodically copied from the current Q network; The replay buffer is used to store the experience in the training process of the DQN model and update the parameters of the current Q network; The process of using the trained DQN model to select the best α value suitable for the current environment and state for each task is specifically as follows: Randomly extract a small batch of experience data from the replay buffer; each piece of experience data includes the task state, bidding action, reward of the task at the current time step, and the task state, bidding action at the next time step; Use the target Q network in the trained DQN model to calculate the target Q value of the task, and select the bidding action with the highest target Q value of the task as the best α value in the current task state; the calculation formula of the target Q value is: Q(state t , action t ) = reward t + γ max(action t+1 ) * Q(state t+1 , action t+1 ; θ - ), Among them, Q(state t , action t ) is the target Q-value of the task at the current time step t, state t is the task state of the task at the current time step t, action t is the bidding action of the task at the current time step t, reward t is the reward of the task at the current time step t; γ is the discount factor, which is used to balance the proportion of the reward of the task at the current time step and the reward of the next time step; Q(state t+1 , action t+1 ; θ - ) is the target Q-value of the task at the next time step t + 1 with parameter θ - ; max(action t+1 ) is the maximum target Q-value under the next task state predicted by the target Q network; θ - is the parameter of the target Q network at the next time step t + 1; state t+1 is the task state of the task at the next time step t + 1; action t+1 is the bidding action of the task at the next time step t + 1.
2. The non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm according to claim 1, characterized in that, The DQN model uses the mean squared error function as the loss function, which is expressed as: δ i = Q(state t , action t ) - Q(state t , action t ; θ), Among them, δ i is the TD error of the i-th experience, representing the difference between the predicted Q value of the current Q network and the target Q value of the target Q network; θ is the parameter of the current Q network; Num ER is the number of experiences in the experience replay buffer; Q(state t , action t ) is the target Q value output by the target Q network, Q(state t , action t ; θ) is the predicted Q value output by the current Q network under the parameter θ; state t is the task state of the task at the current time step t, and action t is the bidding action of the task at the current time step t.
3. The non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm according to claim 1, characterized in that Use heterogeneous multi-threading to parallelly calculate the bidding price of the starting position of each task for bidding, specifically as follows: Initialize a task list tasks, a thread pool thread_pool, and a shared result structure shared_results; the task list tasks contains all the subtasks that need to be executed, that is, the bidding tasks after each task calculates the bidding price of the starting position; the thread pool thread_pool is used to allocate threads to execute the subtasks; the shared result structure shared_results is used to store the execution results of each subtask; Allocate a heterogeneous thread async_thread for each subtask from the thread pool and execute the subtask, and store the execution result of the subtask into shared_results until all threads are executed; Process the execution results of all subtasks in shared_results to generate the final result final_result.
4. The non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm according to claim 1, characterized in that The task allocation based on the task list is specifically as follows: When the task list is not empty, take out the task pattern with the highest bid price and the task start position pos from the task list; Determine whether the position range from the task start position pos to the task end position pos + grid[pattern][0] is covered, where grid[pattern][0] represents the amount of resources required for task pattern; If it is not covered, update the final start position final_positions[pattern] and the coverage status covered of the task; If it is already covered, update the task list, remove the tasks conflicting with the current position, and re-stack the task list; Repeat the above steps until all tasks are allocated, and return the final start positions final_positions of all tasks.
5. A non-uniform memory access resource allocation system based on a parallel adaptive auction algorithm, characterized in that, Applied to the non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm according to any one of claims 1-4, the system includes an information acquisition module and a solution allocation module; The information acquisition module is used to obtain node information and task information in the non-uniform memory access architecture when the operating system starts; the node information includes the number of nodes, the start position of the nodes, and the end position of the nodes; the task information includes the number of tasks and the number of nodes occupied by the tasks; The solution allocation module is used to regard the resource problem of the non-uniform memory access architecture as a knapsack problem with position constraints, and use the parallel adaptive auction algorithm for optimization to obtain the allocation result; Regarding the resource allocation problem of the non-uniform memory access architecture as an extended knapsack problem with position constraints, it is expressed as: W i = (N - i) 0.5 , 0 ≤ i < N, where (position i < position j + width j ) and (position j < position i + width i ), Among them, Score() is the objective function; W i is the weight of the i-th task; width i is the amount of resources occupied by the i-th task; x i is the status of the i-th task. When x i = 1, it means the i-th task is selected. When x i = 0, it means the i-th task is not selected; N is the number of tasks; C is the backpack capacity, that is, the maximum amount of resources that the resources occupied by all selected tasks cannot exceed; position i is the position information of the i-th task, l i is the lower limit of the position of the i-th task, u i is the upper limit of the position of the i-th task; x j is the status of the j-th task; position j is the position information of the j-th task; The process of using the parallel adaptive auction algorithm for optimization is as follows: Use the trained DQN model to select the best α value suitable for the current environment and state for each task; the best α value is used to adjust the bid price of each task; Use heterogeneous multi-threading to calculate the bid price of the starting position of each task in parallel for bidding, and store it in the task list; the bid price calculation formula is: bid i = W i *α i , where bid i is the bid price of the i-th task, and α i is the optimal α value of the i-th task; Allocate tasks based on the task list to obtain the final start positions of all tasks and the list of the best α values; The DQN model includes a current Q network and a target Q network with the same architecture but different parameters, and a replay buffer; The current Q network is used to predict the Q values of each possible bidding action of each task in the current task state, and select the bidding action with the highest Q value as the optimal bidding action in the current task state; The target Q network is used to calculate the target Q value in the next task state, and the parameters of the target Q network are periodically copied from the current Q network; The replay buffer is used to store the experience in the training process of the DQN model and update the parameters of the current Q network; The process of using the trained DQN model to select the best α value suitable for the current environment and state for each task is specifically as follows: Randomly sample a mini-batch of experience data from the experience replay buffer; each piece of experience data contains the task state, bidding action, reward at the current time step of the task, and the task state and bidding action at the next time step. Use the target Q-network in the trained DQN model to calculate the target Q-value of the task, and select the bidding action with the highest target Q-value of the task as the best α value in the current task state; the calculation formula of the target Q-value is: Q(state t , action t ) = reward t + γ max(action t+1 ) * Q(state t+1 , action t+1 ; θ - ), Among them, Q(state t , action t ) is the target Q-value of the task at the current time step t, state t is the task state of the task at the current time step t, action t is the bidding action of the task at the current time step t, reward t is the reward of the task at the current time step t; γ is the discount factor, which is used to balance the proportion of the reward of the task at the current time step and the reward of the next time step; Q(state t+1 , action t+1 ; θ - ) is the target Q-value of the task at the next time step t + 1 with parameter θ - ; max(action t+1 ) is the maximum target Q-value under the next task state predicted by the target Q network; θ - is the parameter of the target Q network at the next time step t + 1; state t+1 is the task state of the task at the next time step t + 1; action t+1 is the bidding action of the task at the next time step t + 1.
6. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements the non-uniform memory access resource allocation method based on the parallel adaptive auction algorithm described in any one of claims 1-4.