The present invention discloses a
reinforcement learning parallel processing accelerator, an acceleration method and an electronic device, which relate to the technical field of accelerators. The controller writes the current view features, at least two groups of batch historical view features and instruction sequences into corresponding memories respectively. The instruction loading and distribution component reads and
decodes the instruction sequences and distributes parameters and start calculation instructions to the calculation components. The data loading control component selects the required
feature data from the memory according to the parameters of each calculation layer and loads it into the corresponding feature cache. After receiving the start calculation instruction, the calculation components read
multiple view feature data simultaneously and perform
parallel processing in combination with the parameters, which can greatly improve the
data processing efficiency. Through the mutual cooperation of the above components, the characteristics of the
reinforcement learning model can be deeply analyzed, and according to different
reinforcement learning tasks and data characteristics, the parameters and
processing flow can be flexibly adjusted, so as to improve the
processing efficiency of the reinforcement learning model, reduce the
resource utilization rate of the accelerator at the same time, and have a fast response speed.