Depth reinforcement learning calculation system and calculation method thereof

By performing experience collection and network update processes in parallel in deep reinforcement learning accelerator, the problems of low hardware utilization and long latency in the prior art are solved, and more efficient hardware utilization and shorter latency are achieved.

CN120020822APending Publication Date: 2025-05-20IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410094627.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-01-23
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

Existing deep reinforcement learning accelerators can only handle training or inferences separately, resulting in low hardware utilization and long latency.

Method used

By designing a deep reinforcement learning calculation system, the processor uses the processor to execute the experience collection process and network update process in parallel, and achieve simultaneous experience collection and network update.

Benefits of technology

It effectively improves hardware utilization, reduces latency, and improves the overall execution efficiency of deep reinforcement learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020822A_ABST
    Figure CN120020822A_ABST
Patent Text Reader

Abstract

The invention discloses a deep reinforcement learning calculation system and a calculation method thereof. The calculation method comprises the following steps: initializing an environment and a model; executing the experience collection process and the network updating process in parallel, and judging whether the experience collection process and the network updating process reach termination conditions or not; continuously executing the experience collection process and the network updating process in parallel in response to the situation that neither the experience collection process nor the network updating process reaches the termination condition; and in response to one of the experience collection process and the network updating process reaching a termination condition, ending execution of the experience collection process and the network updating process. The experience collection process comprises obtaining the current state of the environment; according to a current strategy of the model, calculating based on the current observed value to determine a current action; and returning the current action to the environment. The network updating process comprises obtaining a previous state of the environment and a previous strategy of the model; deciding current data based on the previous state operation; and updating a previous policy of the model to the current policy based on the current data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a computing technology, and particularly to a computing system and a computing method for deep reinforcement learning. Background Art

[0002] Currently, the deep reinforcement learning algorithm of artificial intelligence is widely applied in machine vision and industrial robotic arms. Reinforcement learning is a machine learning technology, which is characterized by the ability to autonomously explore in an unknown environment and train an agent model that can solve multi-step decision-making problems. The agent makes actions based on the observation of the current environmental state and obtains feedback of rewards. The agent gradually updates the strategy of selecting actions according to the obtained reward information to maximize the rewards obtained in the environment.

[0003] Among them, reinforcement learning involves two different steps: experience collection and network update. In the experience collection step, it is necessary to evaluate the input information of the environment according to the current strategy to determine the action to be executed in the next step in the environment; while in the network update step, operations are performed based on the data collected in the environment in the past to update the current model's strategy. Among them, the experience collection step requires inference operations, and the network update step requires training operations.

[0004] Currently, existing deep reinforcement learning accelerators usually adopt a method of using the same set of computing resources to support training and inference, but only one of training or inference can be processed at the same time, which makes the two parts of deep reinforcement learning have to be carried out alternately. Among them, the inference stage faces problems of low hardware utilization and long latency due to the small number of batches and the need to wait for environmental responses, resulting in a longer execution time of the overall reinforcement learning algorithm in the existing architecture. Therefore, how to solve the problem that the existing deep reinforcement learning accelerators can only process training or inference separately with a long latency will be a topic that needs to be broken through. Summary of the Invention

[0005] The present disclosure provides a calculus system for deep reinforcement learning, including: a memory, an input / output interface, and a processor. The memory is used to store the previous state of the environment, the previous policy of the model, an inference program, and a training program. The processor is coupled to the memory and the input / output interface, and is used to perform initializing the environment and the model through the input / output interface; reading the inference program and the training program from the memory, where the inference program corresponds to an experience collection process and the training program corresponds to a network update process; executing the experience collection process and the network update process in parallel, and determining whether the experience collection process and the network update process reach termination conditions; in response to neither the experience collection process nor the network update process reaching the termination conditions, continuously executing the experience collection process and the network update process in parallel; and in response to one of the experience collection process and the network update process reaching the termination condition, ending the execution of the experience collection process and the network update process. The experience collection process includes obtaining the current state of the environment through the input / output interface, where the current state includes a current reward value and a current observation value; determining a current action based on the current observation value according to the current policy of the model; and transmitting the current action back to the environment through the input / output interface. The network update process includes obtaining the previous state of the environment and the previous policy of the model from the memory, where the previous state includes a previous action, a previous reward value, and a previous observation value; determining current data based on the previous state; and updating the previous policy of the model to the current policy based on the current data.

[0006] In one embodiment, the processor further includes an inference processing module and a training processing module. The inference processing module is used to read the inference program from the memory and execute the experience collection process; the training processing module is used to read the training program from the memory and execute the network update process.

[0007] In one embodiment, when the processor executes the experience collection process, it is further used to perform: determining whether the number of executions of the experience collection process reaches an execution count threshold; and in response to the execution count reaching the execution count threshold, determining that the experience collection process reaches the termination condition.

[0008] In one embodiment, when the processor executes the network update process, it is further used to perform: determining whether the number of executions of the network update process reaches an execution count threshold; and in response to the execution count reaching the execution count threshold, determining that the network update process reaches the termination condition.

[0009] In one embodiment, when the processor executes the experience collection process, it is further used to perform: when the environment receives the current action through the input / output interface, determining whether the success rate corresponding to the current state of the environment reaches a success rate threshold; and in response to the success rate reaching the success rate threshold, determining that the experience collection process reaches the termination condition.

[0010] In one embodiment, when the processor executes the network update process, it is further used to: determine the current action based on the previous observation values according to the current policy of the model; when the environment receives the current action through the input / output interface, determine whether the success rate corresponding to the current state of the environment reaches the success rate threshold; and in response to the success rate reaching the success rate threshold, determine that the experience collection process reaches the termination condition.

[0011] The present disclosure provides a calculus method for deep reinforcement learning, including initializing the environment and the model; executing the experience collection process and the network update process in parallel, and determining whether the experience collection process and the network update process reach the termination condition; in response to neither the experience collection process nor the network update process reaching the termination condition, continuously execute the experience collection process and the network update process in parallel; and in response to one of the experience collection process and the network update process reaching the termination condition, end the execution of the experience collection process and the network update process. The experience collection process includes obtaining the current state of the environment, where the current state includes the current reward value and the current observation value; determining the current action based on the current observation value according to the current policy of the model; and transmitting the current action back to the environment. The network update process includes obtaining the previous state of the environment and the previous policy of the model, where the previous state includes the previous action, the previous reward value, and the previous observation value; determining the current data based on the previous state; and updating the previous policy of the model to the current policy based on the current data.

[0012] Based on the above, the calculus system and the calculus method for deep reinforcement learning according to the present invention provide a solution that integrates training and inference. By simultaneously performing experience collection and network update in a parallel processing manner, the hardware utilization rate is effectively improved and the latency time is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are included to provide a further understanding of the present invention, and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.

[0014] Figure 1 is an architecture diagram of a calculus system for deep reinforcement learning illustrated according to an embodiment of the present disclosure;

[0015] Figure 2 is a flowchart of a calculus method for deep reinforcement learning illustrated according to an embodiment of the present invention.

[0016] REFERENCE NUMERAL DESCRIPTION OF THE DRAWINGS

[0017] 1: Calculus system;

[0018] 11: Memory;

[0019] 111: Model;

[0020] 12: Input / Output Interface;

[0021] 13: Processor;

[0022] 131: Inference Processing Module;

[0023] 132: Training Processing Module;

[0024] 14: Data Transmission Interface;

[0025] 2: Calculation Method;

[0026] S21, S23, S25, S27, S231~S234, S251~S254: Steps;

[0027] 9: Environment. Detailed Implementation Manner

[0028] Figure 1 is an architecture diagram of a calculation system 1 for deep reinforcement learning illustrated according to an embodiment of the present disclosure. Please refer to Figure 1 , the calculation system 1 for deep reinforcement learning includes a memory 11, an input / output interface 12, a processor 13, and a data transmission interface 14. In practice, the calculation system 1 for deep reinforcement learning can be implemented by a computer device, such as a desktop computer, a notebook computer, a tablet computer, a workstation, etc., which are computer devices with computing functions, display functions, and networking functions, and the present disclosure is not limited thereto.

[0029] The memory 11 is used to store the previous state of the environment 9, the previous policy of the model 111, the inference program, and the training program. In practice, the memory 11 is, for example, a Static Random-Access Memory (SRAM), a Dynamic Random Access Memory (DRAM), or other memories, and the present disclosure is not limited thereto.

[0030] The processor 13 is coupled to the memory 11 and the input / output interface 12 through the data transmission interface 14. In practice, the processor 13 can be a Central Processing Unit (CPU), a micro-processor, or an embedded controller, and the present disclosure is not limited thereto.

[0031] The processor 13 is used to execute the calculation method of deep reinforcement learning. Figure 2 is a flowchart of a calculation method 2 for deep reinforcement learning illustrated according to an embodiment of the present invention. Figure 2 The calculation method 2 can be achieved by Figure 1The processor 13 of the computing system 1 executes its process. The computing method 2 of deep reinforcement learning includes steps S21, S23, S25, S27, S231 to S234, and S251 to S254. Next, please refer to Figure 1 and Figure 2 simultaneously, and the computing method 2 of deep reinforcement learning will be described.

[0032] After the processor 13 starts to execute the computing method 2, in step S21, the processor 13 initializes the environment 9 and the model 111 through the input / output interface 12.

[0033] After the processor 13 initializes the environment 9 and the model 111, the processor 13 reads the inference program and the training program from the memory 11. The inference program corresponds to the experience collection process, and the training program corresponds to the network update process. In the known deep reinforcement learning computing method, the experience collection process is to evaluate the current state of the environment (including the reward value and the observation value) according to the policy of the current network model to determine the action to be performed in the environment next, and the network update process is to perform operations based on the data of the previous states (including actions, reward values, and observation values) collected in the environment in the past to update the policy of the current network model.

[0034] After the processor 13 obtains the inference program and the training program from the memory 11, the processor 13 executes the experience collection process of step S23 and the network update process of step S25 in parallel. After the experience collection process of step S23 is executed once, the processor 13 will judge whether the experience collection process reaches the termination condition in step S234. Similarly, after the network update process of step S25 is executed once, the processor 13 will also judge whether the network update process reaches the termination condition in step S254.

[0035] In the known deep reinforcement learning computing method, only the inference program or the training program can be executed at the same time, and they cannot be executed simultaneously. However, the deep reinforcement learning computing system 1 and its computing method 2 disclosed in the present invention can execute the inference program and the training program simultaneously by the processor 13 through parallel processing to parallel process the experience collection and the network update. As long as a processor capable of parallel computing for different programs simultaneously can be used to implement the processor 13 of the deep reinforcement learning computing system 1 disclosed in the present invention to execute the inference program and the training program in parallel.

[0036] In an embodiment, the processor 13 of the deep reinforcement learning computing system 1 disclosed in the present invention can be a multi-task processor, which is used to execute the inference program and the training program simultaneously to parallel process the experience collection and the network update.

[0037] In another embodiment, the processor 13 of the calculus system 1 for deep reinforcement learning disclosed by the present invention includes an inference processing module 131 and a training processing module 132. The inference processing module 131 is used to read an inference program from the memory and execute an experience collection process. The training processing module 132 is used to read a training program from the memory and execute a network update process.

[0038] In response to that neither the experience collection process in step S23 nor the network update process in step S25 reaches the termination condition, the processor 13 continuously and concurrently executes the experience collection process in step S23 (steps S231 to S233) and the network update process in step S25 (steps S251 to S253). In response to that one of the experience collection process in step S23 and the network update process in step S25 reaches the termination condition, in step S27, the processor 13 ends the execution of the experience collection process in step S23 and the network update process in step S25.

[0039] Next, the experience collection process in step S23 and the network update process in step S25 will be described respectively. Although steps S23 and S25 are described separately, the experience collection process in step S23 and the network update process in step S25 are executed concurrently by the processor 13.

[0040] First, the experience collection process in step S23 will be described. In step S231, the processor 13 obtains the current state of the environment 9 through the input / output interface 12. The current state includes the current reward value and the current observation value. After the processor 13 obtains the current state of the environment 9, in step S232, according to the current policy of the model 111, it calculates based on the current observation value to determine the current action. Once the current action is determined, in step S233, the processor 13 transmits the current action back to the environment 9 through the input / output interface 12.

[0041] Then, in step S234, the processor 13 determines whether the experience collection process in step S23 reaches the termination condition. In one embodiment, the processor 13 determines whether the number of executions of the experience collection process in step S23 reaches an execution number threshold (for example: 10,000 times). If the number of executions of the experience collection process in step S23 by the processor 13 is less than 10,000 times, the experience collection process in step S23 is iteratively executed. In response to the number of executions reaching the execution number threshold, the processor 13 determines that the experience collection process in step S23 reaches the termination condition. Once the experience collection process in step S23 reaches the termination condition, even if the network update process in step S25 has not reached the termination condition, the processor 13 will directly execute step S27 to end the execution of the calculus method 2.

[0042] In another embodiment, when the environment 9 receives the current action determined in step S232 of the experience collection process through the input / output interface 12, the environment 9 generates a new current state (i.e., the current reward value and the current observation value). The processor 13 determines whether the success rate corresponding to the current state of the environment 9 reaches the success rate threshold. In response to the success rate reaching the success rate threshold, the processor 13 determines that the experience collection process of step S23 reaches the termination condition. Once the experience collection process of step S23 reaches the termination condition, even if the network update process of step S25 has not reached the termination condition, the processor 13 directly executes step S27 to end the execution of the calculation method 2.

[0043] Next, the network update process of step S25 will be described. In step S251, the processor 13 obtains the previous state of the environment 9 and the previous policy of the model 111 from the memory 11. The previous state includes the previous action, the previous reward value, and the previous observation value. It should be specifically noted that "previous" here is earlier than "current" in time. In other words, the network update process of step S25 does not use the current state of the environment 9 and the finally determined current action obtained in the experience collection process of step S23, but uses the previous state that already existed in the environment 9 and the previous policy that already existed in the model 111 before the processor 13 started to execute the experience collection process of step S23 and the network update process of step S25 in parallel.

[0044] In step S252, the processor 13 calculates based on the previous state to determine the current data. In step S253, the processor 13 updates the previous policy of the model 111 to the current policy based on the current data.

[0045] Next, in step S254, the processor 13 determines whether the network update process of step S25 reaches the termination condition. In one embodiment, the processor 13 determines whether the number of executions of the network update process of step S25 reaches the execution number threshold (e.g., 10,000 times). If the number of executions of the network update process of step S25 by the processor 13 is less than 10,000 times, the network update process of step S25 is iteratively executed. In response to the number of executions reaching the execution number threshold, the processor 13 determines that the network update process of step S25 reaches the termination condition. Once the network update process of step S25 reaches the termination condition, even if the experience collection process of step S23 has not reached the termination condition, the processor 13 directly executes step S27 to end the execution of the calculation method 2.

[0046] In another embodiment, the processor 13 determines the current action of the environment 9 based on the previous observed numerical operations according to the current policy of the model 111. After the environment 9 receives the determined current action through the input / output interface 12, the environment 9 generates a new current state (i.e., the current reward value and the current observed value), and the processor 13 determines whether the success rate corresponding to the current state of the environment 9 reaches the success rate threshold. In response to the success rate reaching the success rate threshold, the processor 13 determines that the network update process in step S25 reaches the termination condition. Once the network update process in step S25 reaches the termination condition, even if the experience collection process in step S23 has not reached the termination condition, the processor 13 directly executes step S27 to end the execution of the calculus method 2.

[0047] Based on the above, the calculus system and the calculus method of the deep reinforcement learning according to the present invention provide a solution integrating training and inference. The processor concurrently executes the experience collection process and the network update process to simultaneously perform experience collection and network update, effectively improving the hardware utilization rate and reducing the latency time.

Claims

1. A deep reinforcement learning algorithm system, comprising: Memory, which is used to store the previous state of the environment, the previous policy of the model, the inference program, and the training program; Input / output interface; as well as A processor is coupled to the memory and the input / output interface, and is configured to execute: Initializing the environment and the model through the input / output interface; Reading the inference program and the training program from the memory, wherein the inference program corresponds to an experience collection process, and the training program corresponds to a network update process; Executing the experience collection process and the network update process in parallel, and determining whether the experience collection process and the network update process have reached a termination condition; In response to the experience collection process and the network update process both not reaching the termination condition, continuing to execute the experience collection process and the network update process in parallel; as well as In response to one of the experience collection process and the network update process reaching the termination condition, terminating the experience collection process and the network update process; The experience collection process includes: Obtain the current state of the environment through the input / output interface, where the current state includes the current reward value and the current observation value; According to the current policy of the model, determine the current action based on the current observation value calculation; and Returning the current action to the environment via the input / output interface; The network update process includes: Obtain the previous state of the environment and the previous policy of the model from the memory, the previous state including a previous action, a previous reward value, and a previous observation value; Determine current data based on the previous state operation; and The previous policy of the model is updated to the current policy based on the current data.

2. The deep reinforcement learning algorithm system according to claim 1, wherein the processor further comprises: An inference processing module, used for reading the inference program from the memory and executing the experience collection process; as well as The training processing module is used to read the training program from the memory and execute the network update process.

3. The deep reinforcement learning algorithm system according to claim 1, wherein when the processor executes the experience collection process, it is also used to execute: Determine whether the number of executions of the experience collection process reaches the execution number threshold; and In response to the execution number reaching the execution number threshold, it is determined that the experience collection process reaches the termination condition.

4. The deep reinforcement learning algorithm system according to claim 1, wherein when the processor executes the network update process, it is also used to execute: Determine whether the number of executions of the network update process reaches the execution number threshold; and In response to the execution number reaching the execution number threshold, it is determined that the network update process reaches the termination condition.

5. The deep reinforcement learning algorithm system according to claim 1, wherein when the processor executes the experience collection process, it is also used to execute: After the environment receives the current action through the input / output interface, determining whether the success rate corresponding to the current state of the environment reaches a success rate threshold; and In response to the success rate reaching the success rate threshold, it is determined that the experience collection process reaches the termination condition.

6. The deep reinforcement learning algorithm system according to claim 1, wherein when the processor executes the network update process, it is also used to: According to the current strategy of the model, the current action is determined based on the previous observation value calculation; After the environment receives the current action through the input / output interface, determining whether the success rate corresponding to the current state of the environment reaches a success rate threshold; and In response to the success rate reaching the success rate threshold, it is determined that the experience collection process reaches the termination condition.

7. A method for deep reinforcement learning, comprising: Initialize the environment and model; Execute the experience collection process and the network update process in parallel, and determine whether the experience collection process and the network update process have reached the termination condition. In response to the experience collection process and the network update process both not reaching the termination condition, continuing to execute the experience collection process and the network update process in parallel; as well as In response to one of the experience collection process and the network update process reaching the termination condition, terminating the experience collection process and the network update process; The experience collection process includes: Get the current state of the environment, which includes the current reward value and the current observation value; According to the current policy of the model, determine the current action based on the current observation value calculation; and Return the current action to the environment; The network update process includes: Get the previous state of the environment and the previous strategy of the model, where the previous state includes the previous action, the previous reward value, and the previous observation value; Determine current data based on the previous state operation; and The previous policy of the model is updated to the current policy based on the current data.

8. The deep reinforcement learning algorithm according to claim 7, wherein the step of determining whether the experience collection process reaches the termination condition further comprises: Determine whether the number of executions of the experience collection process reaches the execution number threshold; as well as In response to the execution number reaching the execution number threshold, it is determined that the experience collection process reaches the termination condition.

9. The deep reinforcement learning algorithm according to claim 7, wherein the step of determining whether the experience collection process reaches the termination condition further comprises: After the environment receives the current action, it determines whether the success rate corresponding to the current state of the environment reaches a success rate threshold; In response to the success rate reaching a success rate threshold, it is determined that the experience collection process reaches the termination condition.

10. The deep reinforcement learning algorithm according to claim 7, wherein the step of determining whether the network update process reaches the termination condition further comprises: Determine whether the number of executions of the network update process reaches the execution number threshold; as well as In response to the execution number reaching the execution number threshold, it is determined that the network update process reaches the termination condition.

11. The deep reinforcement learning algorithm according to claim 7, wherein the step of determining whether the network update process reaches the termination condition further comprises: According to the current strategy of the model, the current action is determined based on the previous observation value calculation; as well as After the environment receives the current action, it determines whether the success rate corresponding to the current state of the environment reaches a success rate threshold; In response to the success rate reaching a success rate threshold, it is determined that the experience collection process reaches the termination condition.