Resource allocation method, resource allocation device and storage medium
By using reinforcement learning to obtain resource allocation actions in heterogeneous multi-core processor systems, the problem of high energy delay product caused by resource matching difficulties is solved, and the reasonable allocation of resources and performance improvement is achieved.
Patent Information
- Application Number
- CN202311776191.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
During the operation of the operating system, heterogeneous multi-core processors are difficult to effectively match complex software and hardware resources, resulting in poor energy delay product (EDP).
By obtaining the current operating status of the processor system and the pre-stored action status information obtained through reinforcement learning, finding the corresponding maximum expected reward value and target resource allocation action, the expected reward value is the reciprocal of EDP, and the processor system is controlled to perform target resource allocation action to reassign resources.
The rational allocation of resources is achieved, the energy delay product (EDP) is reduced, and the operational performance of the processor system is improved.
Smart Images

Figure CN120196423A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a resource allocation method, a resource allocation device, and a storage medium. Background Art
[0002] The problem of resource allocation for heterogeneous multi-core processors is a relatively complex problem. During the operation of an operating system, it is necessary to effectively match complex software and hardware resources to obtain the desired operating performance. Any incorrect matching between the resource requirements and allocations of an application during runtime will result in a suboptimal Energy-Delay-Product (EDP).
[0003] Therefore, how to reasonably allocate resources and reduce EDP has become an urgent problem to be solved currently. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a resource allocation method, a resource allocation device, and a storage medium to achieve reasonable resource allocation and reduce EDP.
[0005] To achieve the above purpose, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, the embodiments of this application provide a resource allocation method, including: obtaining the current operating state of a processor system, where the processor system includes multiple processors; obtaining pre-stored action state information, where the action state information includes expected reward values corresponding to different resource allocation actions selected under different operating states, the action state information is obtained through reinforcement learning of the action state information to be learned, and the expected reward value is the reciprocal of the Energy-Delay-Product (EDP); in the action state information, finding the maximum expected reward value corresponding to the current operating state and the target resource allocation action corresponding to the maximum expected reward value; controlling the processor system to execute the target resource allocation action to reallocate resources.
[0007] In a second aspect, the embodiments of this application provide a resource allocation device, including: a processor, a memory, and a program or instruction stored on the memory and executable on the processor, where the program or instruction, when executed by the processor, implements the steps of the resource allocation method described in the first aspect embodiments of this application.
[0008] In a third aspect, the embodiments of this application provide a readable storage medium, where a program or instruction is stored on the readable storage medium, and the program or instruction, when executed by a processor, implements the steps of the resource allocation method described in the first aspect embodiments of this application.
[0009] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects:
[0010] In the embodiments of the present application, according to the current operating state of the processor system, in the action-state information pre-stored and obtained through reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions selected under different operating states, the corresponding maximum expected reward value and the corresponding target resource allocation action are searched. The expected reward value is the reciprocal of the EDP. The processor system is controlled to execute the target resource allocation action to re-allocate resources. In the embodiments of the present application, according to the current operating state of the processor system, the corresponding maximum expected reward value and the corresponding target resource allocation action are searched in the action-state information, that is, the target resource allocation action corresponding to the smallest EDP is searched, and resource re-allocation is achieved based on the target resource allocation action. Since the target resource allocation action is the resource allocation action corresponding to the smallest EDP in the current operating state, reasonable allocation of resources is achieved and the EDP is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0012] Figure 1 is a schematic flowchart of a resource allocation method provided by an embodiment of the present application;
[0013] Figure 2 is a schematic flowchart of a resource allocation method provided by another embodiment of the present application;
[0014] Figure 3 is a schematic flowchart of a resource allocation method provided by another embodiment of the present application;
[0015] Figure 4 is a schematic diagram of a reinforcement learning framework provided by an embodiment of the present application;
[0016] Figure 5 is a schematic diagram of a hybrid cache architecture provided by an embodiment of the present application;
[0017] Figure 6 is a schematic structural diagram of a resource allocation device provided by an embodiment of the present application;
[0018] Figure 7 is a schematic structural diagram of a resource allocation device provided by another embodiment of the present application;
[0019] Figure 8 is a schematic structural diagram of a resource allocation device provided by an embodiment of the present application. Detailed implementation manners
[0020] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0021] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, the "and / or" in the present application means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after. It should be noted that the data involved in the present application are all obtained on the premise of obtaining user authorization.
[0022] A multi-core processor refers to two or more processors integrated on a single chip, also known as a multi-core processor on a chip. Multi-core processors have characteristics such as high performance, high energy efficiency, and low power consumption, and have gradually become the mainstream of the market. With the emergence of research on the specialization of some cores, fault tolerance processing during program operation, power management, etc., heterogeneous multi-core processors have emerged. Different-sized processor cores are placed on heterogeneous multi-core processors, and there are significant differences in the structures, performance, power consumption, etc. of these cores. Compared with homogeneous multi-core processors, heterogeneous multi-core systems pose greater challenges to operating system programming, but have greater advantages in improving multi-thread throughput, single-thread performance, and reducing power consumption.
[0023] The resource allocation problem of heterogeneous multi-core processors is a relatively complex problem. During the operation of the operating system, it is necessary to effectively match complex software and hardware resources to obtain the desired operating performance. Any incorrect matching between the resource requirements and allocations of the application programs during operation will result in sub-optimal EDP. Therefore, how to reasonably allocate resources and reduce EDP has become an urgent problem to be solved. For this reason, the present application proposes a resource allocation method, device, resource allocation device, and storage medium to achieve reasonable resource allocation and reduce EDP.
[0024] The following will detail the technical solutions provided by the embodiments of the present application in conjunction with the drawings.
[0025] Figure 1 It is a schematic flowchart of a resource allocation method provided for an embodiment of the present application. As Figure 1As shown in the figure, the resource allocation method according to the embodiment of the present application may specifically include the following steps:
[0026] S101. Obtain the current operating state of the processor system, where the processor system includes multiple processors.
[0027] In the embodiment of the present application, the execution subject of the resource allocation method according to the embodiment of the present application is a resource allocation device, which may be set in a semiconductor device, specifically in a resource allocation controller of the semiconductor device, such as a Last Level Cache (LLC) resource allocation controller.
[0028] The processor system in the embodiment of the present application may include multiple processors, that is, a multi-core processor system. The sizes, structures, performance power consumptions, etc. of each processor core are quite different, that is, it may be a heterogeneous multi-core processor system.
[0029] Obtain the current operating state s of the processor system, and the operating state s corresponds to a set of operating parameters, such as TPI (Time-Per-Instruction), EDP, etc.
[0030] As a metric of the current operating state, TPI is sensitive to the frequency change of the processor core per clock cycle and the capture of global effects (such as contention latency).
[0031] EDP is the product of the energy Eavg and the average delay time Td, that is, EDP = Eavg × Td = PDP × Td. Among them, PDP (Power-Delay-Product) is the product of the power consumption Pavg and the average delay time Td, that is, PDP = Pavg × Td.
[0032] S102. Obtain the pre-stored action state information, where the action state information includes the expected reward values corresponding to different resource allocation actions in different operating states. The action state information is obtained by performing reinforcement learning on the action state information to be learned, and the expected reward value is the reciprocal of the energy delay product EDP.
[0033] In the embodiment of the present application, reinforcement learning, as one of the techniques of machine learning, is applicable to sequential decision-making problems and is suitable for optimizing the search for long-term cumulative rewards. Reinforcement learning is a computational model for reward learning by interacting with the system. When the samples required by the system cannot be determined, it interacts with the system and uses the observed values evaluated as good or bad.
[0034] The action state information (Q-table) obtained by reinforcement learning in advance is stored in the processor system. The action state information includes different operating states s iSelect different resource allocation actions a below i The corresponding expected reward value Q(s i , a i ). The expected reward value Q is the reciprocal of the EDP, i.e., Q = 1 / EDP.
[0035] The action state information is stored in any one of the forms such as a table, a list, and a matrix.
[0036] When the action state information is stored in the form of a table, this table is denoted as the action state table. The action state table is shown in Table 1. The first column includes multiple different operating states s i , and the first row includes multiple different resource allocation actions a i . The remaining multiple expected reward values Q(s i , a i ) form the action state matrix.
[0037] Table 1
[0038] Status / Action a1 a2 a3 … a140 a141 s1 Q(s1,a1) Q(s1,a2) Q(s1,a3) … Q(s1,a140) Q(s1,a141) s2 Q(s2,a1) Q(s2,a2) Q(s2,a3) … Q(s2,a140) Q(s2,a141) …
[0039] Here, it should be noted that the number of resource allocation actions in the action state information is determined according to the number of processors in the processor system. Specifically, a vector is used to represent the request allocation status of each processor core, i.e., allocation / unchanged / release (1 / 0 / -1). C processor cores share cache resources, and their request allocation status can be represented by a vector of C elements. The resource allocation action represents the change to the current resource allocation policy for each core, corresponding to the request allocation status of each core. For example, in a 6-core (cores C0 to C5) processor system, [1, 0, 0, -1, 0, 0] represents a request allocation status, where core C0 obtains additional cache resources after core C3 releases cache resources. In this fixed shared cache technology, when one core is allocated an additional cache resource (or cache path), another core must release a cache resource. The set of resource allocation actions (i.e., the action space) Action = <[1, 0, 0, -1, 0, 0], [0, 1, 0, -1, 0, 0], …, [1, 1, 1, -1, -1, -1], [0, 0, 0, 0, 0, 0]>. For a 6-core processor system, there are 141 resource allocation actions (a1 to a141).
[0040] In the action state information, the number of operating states is determined according to the number of processors and the number of resources. For example, if each processor core has S operating states, then for a 6-core processor system, there are S 6 operating states.
[0041] The action state information is a state information in the reinforcement learning process. During the reinforcement learning process, the action state information will be continuously updated, and the update depends on the running state before the update. For the specific process, please refer to Figure 2 the relevant descriptions in it, which will not be elaborated here.
[0042] S103. In the action state information, search for the maximum expected reward value corresponding to the current running state and the target resource allocation action corresponding to the maximum expected reward value.
[0043] In the embodiment of the present application, the maximum expected reward value corresponding to the current running state is the maximum one among the multiple expected reward values corresponding to the current running state. The target resource allocation action is the resource allocation action corresponding to the maximum expected reward value, that is, the expected reward value when the target resource allocation action is executed in the current running state is the maximum expected reward value.
[0044] For example, if the current running state is s1, then among the 141 expected reward values Q(s1,a1), Q(s1,a2), …, Q(s1,a141) corresponding to s1, the maximum expected reward value, such as Q(s1,a2), is determined as the maximum expected reward value, and the resource allocation action a2 corresponding to the maximum expected reward value Q(s1,a2) is determined as the target resource allocation action.
[0045] S104. Control the processor system to execute the target resource allocation action to reallocate resources.
[0046] In the embodiment of the present application, control the processor system to execute the target resource allocation action determined in step S104. For example, if the target resource allocation action is [1,0,0,-1,0,0], the processor system releases one resource of processor core C3 and allocates one resource to core C0 to achieve the reallocation of resources.
[0047] In summary, for the resource allocation method in the embodiment of the present application, according to the current running state of the processor system, in the pre-stored action state information obtained through reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions in different running states, search for the corresponding maximum expected reward value and the corresponding target resource allocation action. The expected reward value is the reciprocal of EDP. Control the processor system to execute the target resource allocation action to reallocate resources. In the embodiment of the present application, according to the current running state of the processor system, search for the corresponding maximum expected reward value and the corresponding target resource allocation action in the action state information, that is, search for the target resource allocation action corresponding to the minimum EDP, and based on the target resource allocation action, achieve resource reallocation. Since the target resource allocation action is the resource allocation action corresponding to the minimum EDP in the current running state, the reasonable allocation of resources is realized and the EDP is reduced.
[0048] Figure 2 A flowchart of a resource allocation method provided for another embodiment of the present application. As Figure 2 shown, based on the embodiment shown in Figure 1 , the resource allocation method of the embodiment of the present application may specifically include the following steps:
[0049] S201, perform reinforcement learning on the action state information to be learned, obtain the action state information after reinforcement learning, and store the action state information after reinforcement learning.
[0050] In the embodiment of the present application, the above Figure 1 shown embodiment describes the use process or execution process of the action state information after reinforcement learning. This step S201 describes the reinforcement learning process of the action state information.
[0051] S202, obtain the current operating state of the processor system, where the processor system includes multiple processors.
[0052] S203, obtain the pre-stored action state information, where the action state information includes the expected reward values corresponding to different resource allocation actions in different operating states, the action state information is obtained by performing reinforcement learning on the action state information to be learned, and the expected reward value is the reciprocal of the energy-delay product EDP.
[0053] S204, in the action state information, search for the maximum expected reward value corresponding to the current operating state and the target resource allocation action corresponding to the maximum expected reward value.
[0054] S205, control the processor system to execute the target resource allocation action to reallocate resources.
[0055] In the embodiment of the present application, steps S202 - S205 are the same as steps S101 - S104 in the above embodiment, and will not be elaborated here.
[0056] Furthermore, as Figure 3 shown, the above step S201 "perform reinforcement learning on the action state information to be learned, obtain the action state information after reinforcement learning, and store the action state information after reinforcement learning" may specifically include the following steps:
[0057] S301, initialize the action state information to be learned, where the action state information to be learned includes the expected reward values corresponding to different resource allocation actions in different operating states.
[0058] In the embodiments of the present application, the action state information to be learned is initialized. For example, the action state table shown in Table 1 is initialized, that is, each Q value in the action state table shown in Table 1 is initialized to a preset value. Since the selection of the initial value does not affect the reinforcement result at the end of the reinforcement learning (i.e., the action state information after reinforcement learning), these preset values can be set randomly. For example, each Q value can be initialized to 0.
[0059] S302. Obtain the to-be-updated operating state of the processor system. The initial value of the to-be-updated operating state is a randomly selected operating state.
[0060] In the embodiments of the present application, during the first round of learning, a random operating state is selected as the to-be-updated operating state. During the learning process other than the first round, the next operating state determined in the previous round of learning is determined as the to-be-updated operating state.
[0061] S303. Obtain the to-be-updated resource allocation action, the actual reward value, and the next operating state corresponding to the to-be-updated operating state. The next operating state is the operating state of the processor system after executing the to-be-updated resource allocation action in the to-be-updated operating state.
[0062] In the embodiments of the present application, the to-be-updated resource allocation action corresponding to the to-be-updated operating state can be determined by any one of the following two methods. Each time it is determined, different methods can be used, or the same method can be used:
[0063] As a first feasible implementation method, a random resource allocation action can be selected as the to-be-updated resource allocation action.
[0064] As a second feasible implementation method, the maximum reward resource allocation action corresponding to the to-be-updated operating state can be used as the to-be-updated resource allocation action. The maximum reward resource allocation action satisfies the following condition: the actual reward value corresponding to the operating state after the to-be-updated operating state executes the maximum reward resource allocation action is the largest. The calculation process of the actual reward value can be referred to the relevant description in the following embodiments and will not be elaborated here.
[0065] In some embodiments, it can be determined which of the above methods is used to determine the to-be-updated resource allocation action corresponding to the to-be-updated operating state by judging a random number. Specifically: obtain the random number generated by the processor system. The random number is a randomly generated number. If the random number is less than the preset greedy coefficient ε, a random resource allocation action is selected as the to-be-updated resource allocation action. If the random number is equal to or greater than the greedy coefficient ε, the maximum reward resource allocation action corresponding to the to-be-updated operating state is used as the to-be-updated resource allocation action.
[0066] Those skilled in the art can understand that the greedy coefficient is a mechanism in reinforcement learning. The greedy coefficient has no direct connection with the formula of reinforcement learning and can be set by the user before the reinforcement learning process. The greedy coefficient generally needs to be set according to experience, which is obtained through a large number of reinforcement learning, iterations, and observations. Through the greedy coefficient, a resource allocation action can be randomly selected as the resource allocation action to be updated, which can improve the convergence speed of the reinforcement learning process and the efficiency of reinforcement learning.
[0067] It should be noted here that the maximum reward resource allocation action can be determined through the test process before the reinforcement learning process, and the maximum reward resource allocation action corresponding to each state is stored. Then, during the reinforcement learning process, the maximum reward resource allocation action corresponding to the running state to be updated is obtained from the stored data. The maximum reward resource allocation action can also be determined in real-time during the reinforcement learning process.
[0068] In some embodiments, the actual reward value corresponding to the running state to be updated can be determined through the test process before the reinforcement learning process, and the actual reward value corresponding to each state is stored. Then, during the reinforcement learning process, the actual reward value corresponding to the running state to be updated is obtained from the stored data. That is, obtaining the actual reward value corresponding to the running state to be updated includes: obtaining the actual reward value corresponding to the running state to be updated stored in advance.
[0069] In some embodiments, the actual reward value corresponding to the running state to be updated can also be determined in real-time during the reinforcement learning process.
[0070] The actual reward value corresponding to the running state can be calculated through the reward function. Taking the actual reward value corresponding to the running state to be updated as an example, before obtaining the actual reward value corresponding to the running state to be updated stored in advance, the following steps for calculating the actual reward value are also included: obtaining the time per instruction TPI of the processor system in the running state to be updated; calculating the actual reward value corresponding to the running state to be updated according to TPI and the reward function.
[0071] Among them, the above step "calculating the actual reward value corresponding to the running state to be updated according to TPI and the reward function" can specifically include the following steps: calculating the EDP of the processor system in the running state to be updated according to TPI; calculating the actual reward value corresponding to the running state to be updated according to EDP and the reward function.
[0072] Specifically, the following formula (1) can be used to calculate EDP:
[0073] EDP = (E / I) * (t / I) = (E / I) * TPI (1)
[0074] Where E represents energy consumption, I represents the number of executed instructions, and t represents the time to execute I instructions.
[0075] Then the actual reward value R is calculated using the following formula (2) (i.e., reward function):
[0076] R=1 / EDP+f (2)
[0077] Among them, f represents the parameters for modeling and calculating the CPU information, memory and other parameters occupied by the running state. It is information biased towards the specific application level and is used to assist EDP as a reward function. The specific calculation method of f is defined by the specific application. For example, f can be at least one of the following parameters: the number of CPUs occupied by the process, the size of the CPU cache value, and the calculated value of multiple system information.
[0078] S304 , searching the action state information to be learned for the expected reward value to be updated corresponding to the running state to be updated and the resource allocation action to be updated, and searching for the maximum expected reward value corresponding to the next running state.
[0079] In the embodiment of the present application, assuming that the running state to be updated is s and the resource allocation action to be updated is a, the expected reward value Q(s, a) corresponding to s and a is searched in the action state information as the expected reward value to be updated. Assuming that the running state (i.e., the next running state) after the running state s to be updated executes the resource allocation action a to be updated is s', the largest one among the multiple, for example, 141, expected reward values Q corresponding to s' is searched in the action state information, i.e., the maximum expected reward value, which can be expressed as max Q(s', a').
[0080] S305, updating the expected reward value to be updated in the action state information to be learned by using reinforcement learning according to the expected reward value to be updated, the actual reward value and the maximum expected reward value corresponding to the next state.
[0081] In the embodiment of the present application, the following formula (3) (i.e., the reinforcement learning formula) can be used to update the expected reward value to be updated:
[0082] New Q(s, a)=(1-α)Q(s, a)+α[r+γmax Q(s', a')] (3)
[0083] Among them, New Q(s, a) is the updated expected reward value to be updated, α is the parameter used to control the proportion of past and future rewards in the learning process, γ represents the attenuation value of future rewards, and r is the actual reward value R.
[0084] S306, updating the state to be updated to the next running state.
[0085] In an embodiment of the present application, the next operating state s' is used as the new state to be updated, and the process returns to step S302 to re-obtain the operating state to be updated of the processor system, and perform the next round of reinforcement learning. After multiple iterations like this, until a preset loop learning end condition is reached, for example, the action state table no longer updates, then the next round of reinforcement learning is no longer performed, and the current action state information to be learned is stored as the action state information after reinforcement learning for subsequent use or execution processes.
[0086] To clearly illustrate the above reinforcement learning process, the following description is made in conjunction with Figure 4 the framework of the reinforcement learning shown. As Figure 4 shown, a resource allocation controller (as an agent) obtains the current operating state of the processor system as the operating state s to be updated, obtains the resource allocation action a to be updated corresponding to the operating state s to be updated. For example, a random selection is made from multiple resource allocation actions in the action state table as the resource allocation action a to be updated, obtains the TPI in the operating state s to be updated, calculates the actual reward value r corresponding to the operating state s to be updated in combination with the reward function, controls the processor system to execute the resource allocation action a to be updated, obtains the next operating state s', looks up the corresponding maximum expected reward value max Q(s', a') in the action state table according to s', and looks up the expected reward value Q(s, a) corresponding to s and a in the action state table, and updates Q(s, a) according to Q(s, a), r, and max Q(s', a').
[0087] It should be noted here that in the embodiment of the present application, the resource is the specified allocation cache resource E that needs to be allocated to each processor after removing the fixed shared cache resource F from the shared cache resource W corresponding to the processor system. The fixed shared cache resource F and the specified allocation cache resource E are allocated and controlled by different resource allocation controllers.
[0088] Specifically, Figure 5 is a six-core processor, and the shared cache resource is dynamically partitioned by allocating caches to the cores. The cache resources allocated to some cores may be more than needed, and the partitioning of the shared storage may cause some shared cache fragmentation (when the linear cache is unreasonably allocated, some caches are not used by the current core and cannot be used by other cores, and such idle caches are called cache fragments), which is relatively common for application programs with a low cache occupancy rate. To avoid fragmentation, a hybrid cache architecture method needs to be constructed. In the hybrid method, the fixed shared cache resource in the shared cache resource is always shared and cannot be specifically allocated to any one core to solve the uneven relative cache demand of the cores, and still allows them to occupy the fixed shared cache resource when needed. The resources other than the fixed shared cache resource in the shared cache resource are specified and allocated to each core.
[0089] Suppose there are 27 cache resources shared among 6 cores, with 7 resources being fixed shared cache resources. For the 20 specified allocated cache resources, there are a total of 6 20 possible cache allocation methods. Since there are a large number of possibilities in this allocation method and the cache reallocation strategy cannot be directly determined, therefore, in the embodiments of the present application, the shared cache state is modeled, and the established model needs to capture the cache reallocation changes on all shared cores and form a feasible state space (i.e., the set of running states).
[0090] The fixed shared cache resources and the specified allocated cache resources are allocated and controlled by different resource allocation controllers, i.e., different agents. Multi-agent is also known as multi-agent system. In a multi-agent system, each agent has independence and autonomy, can solve the given sub-problems, reason and plan independently and select appropriate strategies, and affect the environment in a specific way; the multi-agent system supports distributed applications, so it has good modularity, easy scalability and simple and flexible design, overcomes the management and expansion difficulties caused by building a large system, and can effectively reduce the total cost of the system; in the implementation process of the multi-agent system, instead of pursuing a single large and complex system, it constitutes multiple levels and diversified agents according to the object-oriented method, and the result is to reduce the complexity of the system and the complexity of problem-solving for each agent; the multi-agent system is a system that emphasizes coordination, and each agent solves large-scale complex problems through mutual coordination; the multi-agent system is also an integrated system, which uses information integration technology to integrate the information of each subsystem to complete the integration of complex systems. In a multi-agent system, each agent communicates with each other, coordinates with each other, and solves problems in parallel, so it can effectively improve the problem-solving ability; the multi-agent technology breaks the limitation of only using one expert system in the field of artificial intelligence. In the MAS (Mobile Agent Serve) environment, different experts in different fields may collaborate to solve a problem that a single expert cannot solve well, improving the problem-solving ability of the system; the agents are heterogeneous and distributed. They can be different individuals and organizations, developed using different design methods and computer languages, and thus may be completely heterogeneous and distributed. The processing is asynchronous. Since each agent is autonomous, each agent has its own process and runs asynchronously according to its own running mode.
[0091] The coordinated management of multiple resources requires a system to explore a vast resource allocation decision space, and it is almost infeasible for a single reinforcement learning agent to solve this problem. Therefore, the embodiments of this application adopt distributed multi-agents to solve this problem and stimulate the needs of multi-agents. An efficient way to model a multi-resource management system is to model the agents corresponding to each resource, so that each reinforcement learning agent only focuses on optimizing its corresponding resource. The multi-agents (which can be homogeneous or heterogeneous) simultaneously execute their respective optimal resource management.
[0092] In summary, the resource allocation method of the embodiments of this application, according to the current operating state of the processor system, searches in the pre-stored action state information obtained through reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions selected under different operating states, for the corresponding maximum expected reward value and the corresponding target resource allocation action. The expected reward value is the reciprocal of the energy delay product (EDP). It controls the processor system to execute the target resource allocation action to reallocate resources. The embodiments of this application search for the corresponding maximum expected reward value and the corresponding target resource allocation action in the action state information according to the current operating state of the processor system, that is, search for the target resource allocation action corresponding to the minimum EDP, and reallocate resources based on this target resource allocation action. Since the target resource allocation action is the resource allocation action corresponding to the minimum EDP in the current operating state, the reasonable allocation of resources is achieved and the EDP is reduced.
[0093] Figure 6 It is a schematic structural diagram of a resource allocation device provided by an embodiment of this application. As Figure 6 shown, the resource allocation device 600 of the embodiments of this application may specifically include: a first acquisition module 601, a second acquisition module 602, a search module 603, and an allocation module 604. Among them:
[0094] The first acquisition module 601 is used to acquire the current operating state of the processor system, and the processor system includes multiple processors.
[0095] The second acquisition module 602 is used to acquire the pre-stored action state information, which includes the expected reward values corresponding to different resource allocation actions selected under different operating states. The action state information is obtained through reinforcement learning of the action state information to be learned, and the expected reward value is the reciprocal of the energy delay product (EDP).
[0096] The search module 603 is used to search in the action state information for the maximum expected reward value corresponding to the current operating state and the target resource allocation action corresponding to the maximum expected reward value.
[0097] An allocation module 604 for controlling the processor system to perform a target resource allocation action to reallocate resources.
[0098] In the embodiments of the present application, for the specific processes of each module and unit in the resource allocation device of the embodiments of the present application to implement their functions, reference may be made to the relevant descriptions in the above embodiments of the resource allocation method, which will not be elaborated here.
[0099] In summary, the resource allocation device of the embodiments of the present application, according to the current operating state of the processor system, looks up the corresponding maximum expected reward value and the corresponding target resource allocation action in the pre-stored action state information obtained through reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions selected under different operating states. The expected reward value is the reciprocal of the EDP. It controls the processor system to execute the target resource allocation action to reallocate resources. The embodiments of the present application look up the corresponding maximum expected reward value and the corresponding target resource allocation action in the action state information according to the current operating state of the processor system, that is, look up the target resource allocation action corresponding to the minimum EDP, and realize resource reallocation based on the target resource allocation action. Since the target resource allocation action is the resource allocation action corresponding to the minimum EDP in the current operating state, reasonable allocation of resources is achieved and the EDP is reduced.
[0100] Figure 7 It is a schematic structural diagram of a resource allocation device provided in another embodiment of the present application. As Figure 7 shown, on the basis of the embodiment shown in Figure 6 In the resource allocation device 600 of the embodiments of the present application, the second acquisition module 602 specifically includes:
[0101] An initialization unit 701 for initializing the action state information to be learned, and the action state information to be learned includes the expected reward values corresponding to different resource allocation actions selected under different operating states.
[0102] A first acquisition unit 702 for acquiring the to-be-updated operating state of the processor system, and the initial value of the to-be-updated operating state is a randomly selected operating state.
[0103] A second acquisition unit 703 for acquiring the to-be-updated resource allocation action, the actual reward value, and the next operating state corresponding to the to-be-updated operating state, and the next operating state is the operating state of the processor system after executing the to-be-updated resource allocation action in the to-be-updated operating state.
[0104] A search unit 704 for searching in the action state information to be learned for the to-be-updated expected reward value corresponding to the to-be-updated operating state and the to-be-updated action, and for searching for the maximum expected reward value corresponding to the next state.
[0105] The first update unit 705 is configured to update the to-be-updated expected reward value in the to-be-learned action state information by means of reinforcement learning according to the to-be-updated expected reward value, the actual reward value, and the maximum expected reward value corresponding to the next state.
[0106] The second update unit 706 is configured to update the to-be-updated state to the next running state, trigger the first acquisition unit 702 to re-execute the step of acquiring the to-be-updated running state of the processor system until a preset loop learning end condition is reached, and store the current to-be-learned action state information.
[0107] In a feasible implementation manner of the embodiment of the present application, the number of resource allocation actions is determined according to the number of processors in the processor system; and / or, the number of running states is determined according to the number of processors and the number of resources in the processor system.
[0108] In a feasible implementation manner of the embodiment of the present application, the second acquisition unit 703 is further configured to: acquire a random number generated by the processor system; if the random number is less than a preset greedy coefficient, randomly select a resource allocation action as the to-be-updated resource allocation action; if the random number is equal to or greater than the greedy coefficient, use the maximum reward resource allocation action corresponding to the to-be-updated running state as the to-be-updated resource allocation action, where the maximum reward resource allocation action satisfies the following condition: the actual reward value corresponding to the running state after the to-be-updated running state executes the maximum reward resource allocation action is the largest.
[0109] In a feasible implementation manner of the embodiment of the present application, the second acquisition unit 703 is further configured to: acquire the actual reward value corresponding to the to-be-updated running state stored in advance, and before acquiring the actual reward value corresponding to the to-be-updated running state stored in advance, acquire the time per instruction TPI of each instruction of the processor system in the to-be-updated running state; calculate the actual reward value corresponding to the to-be-updated running state according to TPI and the reward function.
[0110] In a feasible implementation manner of the embodiment of the present application, the second acquisition unit 703 is further configured to: calculate the corresponding EDP according to TPI; calculate the actual reward value corresponding to the to-be-updated running state according to EDP and the reward function.
[0111] In a feasible implementation manner of the embodiment of the present application, the resource is the specified allocation cache resource that needs to be allocated to each processor after removing the fixed shared cache resource from the shared cache resource corresponding to the processor system, and the fixed shared cache resource and the specified allocation cache resource are allocated and controlled by different resource allocation controllers.
[0112] In the embodiments of the present application, for the specific processes of each module and unit in the resource allocation device of the embodiments of the present application to implement their functions, reference may be made to the relevant descriptions in the above embodiments of the resource allocation method, which will not be elaborated herein.
[0113] In summary, the resource allocation device of the embodiments of the present application, according to the current operating state of the processor system, searches for the corresponding maximum expected reward value and the corresponding target resource allocation action in the action state information pre-stored and obtained through reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions selected under different operating states. The expected reward value is the reciprocal of the EDP. It controls the processor system to execute the target resource allocation action to re-allocate resources. The embodiments of the present application search for the corresponding maximum expected reward value and the corresponding target resource allocation action in the action state information according to the current operating state of the processor system, that is, search for the target resource allocation action corresponding to the minimum EDP, and realize resource re-allocation based on the target resource allocation action. Since the target resource allocation action is the resource allocation action corresponding to the minimum EDP in the current operating state, reasonable allocation of resources is achieved and the EDP is reduced.
[0114] The embodiments of the present application also provide a resource allocation device. As Figure 8 shown, the resource allocation device 600 may vary greatly due to configuration or performance differences. It may include one or more processors 801 and a memory 802. One or more application programs or data may be stored in the memory 802. Among them, the memory 802 may be transient storage or persistent storage. The application programs stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the resource allocation device 600. Further, the processor 801 may be set to communicate with the memory 802 and execute a series of computer-executable instructions in the memory 802 on the resource allocation device 600. The resource allocation device 600 may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, and one or more keyboards 806.
[0115] Specifically in this embodiment, the resource allocation device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the resource allocation device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions for:
[0116] Obtain the current operating state of the processor system, where the processor system includes multiple processors;
[0117] Obtain the pre-stored action state information, where the action state information includes the expected reward values corresponding to different resource allocation actions selected under different operating states. The action state information is obtained by performing reinforcement learning on the action state information to be learned, and the expected reward value is the reciprocal of the energy-delay product EDP;
[0118] In the action state information, search for the maximum expected reward value corresponding to the current operating state and the target resource allocation action corresponding to the maximum expected reward value;
[0119] Control the processor system to execute the target resource allocation action to reallocate resources.
[0120] The resource allocation device according to the embodiment of the present application, based on the current operating state of the processor system, searches for the corresponding maximum expected reward value and the corresponding target resource allocation action in the pre-stored action state information obtained by reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions selected under different operating states. The expected reward value is the reciprocal of EDP, and controls the processor system to execute the target resource allocation action to reallocate resources. According to the embodiment of the present application, based on the current operating state of the processor system, searches for the corresponding maximum expected reward value and the corresponding target resource allocation action in the action state information, that is, searches for the target resource allocation action corresponding to the minimum EDP, and realizes resource reallocation based on the target resource allocation action. Since the target resource allocation action is the resource allocation action corresponding to the minimum EDP in the current operating state, reasonable allocation of resources is achieved and EDP is reduced.
[0121] The embodiment of the present application also proposes a readable storage medium, on which one or more computer programs are stored. The one or more computer programs include instructions. When the program or instructions are executed by a processor in a resource allocation device including multiple application programs, the processor in the resource allocation device can execute each process of the above resource allocation method embodiment, and is specifically used to execute:
[0122] Obtain the current operating state of the processor system, where the processor system includes multiple processors;
[0123] Obtain the pre-stored action state information, where the action state information includes the expected reward values corresponding to different resource allocation actions selected under different operating states. The action state information is obtained by performing reinforcement learning on the action state information to be learned, and the expected reward value is the reciprocal of the energy-delay product EDP;
[0124] In the action state information, search for the maximum expected reward value corresponding to the current running state and the target resource allocation action corresponding to the maximum expected reward value;
[0125] Control the processor system to execute the target resource allocation action to reallocate resources.
[0126] The readable storage medium of the embodiment of the present application, according to the current running state of the processor system, searches in the pre-stored action state information obtained by reinforcement learning, which includes the expected reward values corresponding to different resource allocation actions in different running states, for the corresponding maximum expected reward value and the corresponding target resource allocation action. The expected reward value is the reciprocal of the EDP. Control the processor system to execute the target resource allocation action to reallocate resources. According to the current running state of the processor system, the embodiment of the present application searches in the action state information for the corresponding maximum expected reward value and the corresponding target resource allocation action, that is, searches for the target resource allocation action corresponding to the minimum EDP, and realizes resource reallocation based on the target resource allocation action. Since the target resource allocation action is the resource allocation action corresponding to the minimum EDP in the current running state, the reasonable allocation of resources is realized and the EDP is reduced.
[0127] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0128] For the convenience of description, the above devices are described by dividing them into various units according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0129] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or a device for implementing the functions specified in one block or more blocks.
[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or a device for implementing the functions specified in one block or more blocks.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or a device for implementing the functions specified in one block or more blocks.
[0133] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0134] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0135] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0136] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0137] This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0138] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, they are described relatively simply, and the relevant parts can refer to the description of the method embodiments.
[0139] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A resource allocation method, characterized in that, Including: Obtain the current operating state of the processor system, where the processor system includes multiple processors; Obtain the pre-stored action state information, where the action state information includes the expected reward values corresponding to different resource allocation actions selected under different operating states, the action state information is obtained through reinforcement learning of the action state information to be learned, and the expected reward value is the reciprocal of the energy-delay product EDP; In the action state information, find the maximum expected reward value corresponding to the current operating state and the target resource allocation action corresponding to the maximum expected reward value; Control the processor system to execute the target resource allocation action to reallocate resources.
2. The method according to claim 1, wherein The number of the resource allocation actions is determined according to the number of the processors in the processor system; and / or, the number of the operating states is determined according to the number of the processors and the number of the resources.
3. The method according to claim 1, wherein Before obtaining the pre-stored action state information, it further includes: Initialize the action state information to be learned, where the action state information to be learned includes the expected reward values corresponding to different resource allocation actions selected under different operating states; Obtain the operating state to be updated of the processor system, and the initial value of the operating state to be updated is randomly selected from the operating states; Obtain the resource allocation action to be updated, the actual reward value, and the next operating state corresponding to the operating state to be updated, where the next operating state is the operating state of the processor system after executing the resource allocation action to be updated in the operating state to be updated; In the action state information to be learned, find the expected reward value to be updated corresponding to the operating state to be updated and the action to be updated, and find the maximum expected reward value corresponding to the next state; According to the expected reward value to be updated, the actual reward value, and the maximum expected reward value corresponding to the next state, update the expected reward value to be updated in the action state information to be learned by using the method of reinforcement learning; Update the operating state to be updated to the next operating state, return to the step of obtaining the operating state to be updated of the processor system, and repeat until a preset loop learning end condition is reached, and store the current action state information to be learned.
4. The method according to claim 3, wherein The obtaining of the resource allocation action to be updated corresponding to the operating state to be updated includes: Obtain the random number generated by the processor system; If the random number is less than a preset greedy coefficient, randomly select one of the resource allocation actions as the resource allocation action to be updated; If the random number is equal to or greater than the greedy coefficient, use the resource allocation action with the maximum reward corresponding to the operating state to be updated as the resource allocation action to be updated, and the resource allocation action with the maximum reward satisfies the following condition: the actual reward value corresponding to the operating state after the operating state to be updated executes the resource allocation action with the maximum reward is the largest.
5. The method according to claim 3, characterized in that, The obtaining of the actual reward value corresponding to the operating state to be updated includes: Obtain the actual reward value corresponding to the operating state to be updated pre-stored; Before obtaining the actual reward value corresponding to the to-be-updated operation state stored in advance, it further includes: Obtaining the time TPI spent by each instruction of the processor system in the to-be-updated operation state; Calculating the actual reward value corresponding to the to-be-updated operation state according to the TPI and the reward function.
6. The method according to claim 5, characterized in that, The calculating the actual reward value corresponding to the to-be-updated operation state according to the TPI and the reward function includes: Calculating the EDP of the processor system in the to-be-updated operation state according to the TPI; Calculating the actual reward value corresponding to the to-be-updated operation state according to the EDP and the reward function.
7. The method according to claim 1, characterized in that The resource is the specified allocation cache resource that needs to be allocated to each processor after removing the fixed shared cache resource from the shared cache resource corresponding to the processor system. The fixed shared cache resource and the specified allocation cache resource are allocated and controlled by different resource allocation controllers.
8. The method according to claim 1, wherein The action state information is stored in any one of the forms of a table, a list, and a matrix.
9. A resource allocation device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.
10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.