A Reinforcement Learning Parallel Processing Accelerator, an Acceleration Method and an Electronic Device

By designing a reinforcement learning parallel processing accelerator, using the combination of controller, instruction load distribution components, data load control components and calculation components, the problem of poor acceleration effect of existing accelerators when processing reinforcement learning models is solved, and efficient data processing and acceleration effects are achieved.

CN119718696BActive Publication Date: 2025-06-17LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510245683.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-17
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

When handling reinforcement learning models, existing accelerators fail to fully utilize the parallel processing characteristics of the model, resulting in poor acceleration results.

Method used

A reinforcement learning parallel processing accelerator is designed, including a controller, an instruction loading distribution component, a data loading control component and a computing component. By writing the current view features and historical view features to memory respectively, quickly reading and decoding the instruction sequence, loading feature data into the feature cache, and performing parallel processing after receiving the enabled calculation instruction.

Benefits of technology

By processing multiple view feature data in parallel, making full use of hardware resources, it significantly improves data processing efficiency, reduces processing time, accelerates the operation of reinforcement learning algorithms, and is suitable for rapid decision-making in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718696B_ABST
    Figure CN119718696B_ABST
Patent Text Reader

Abstract

The present invention discloses a reinforcement learning parallel processing accelerator, an acceleration method and an electronic device, which relate to the technical field of accelerators. The controller writes the current view features, at least two groups of batch historical view features and instruction sequences into corresponding memories respectively. The instruction loading and distribution component reads and decodes the instruction sequences and distributes parameters and start calculation instructions to the calculation components. The data loading control component selects the required feature data from the memory according to the parameters of each calculation layer and loads it into the corresponding feature cache. After receiving the start calculation instruction, the calculation components read multiple view feature data simultaneously and perform parallel processing in combination with the parameters, which can greatly improve the data processing efficiency. Through the mutual cooperation of the above components, the characteristics of the reinforcement learning model can be deeply analyzed, and according to different reinforcement learning tasks and data characteristics, the parameters and processing flow can be flexibly adjusted, so as to improve the processing efficiency of the reinforcement learning model, reduce the resource utilization rate of the accelerator at the same time, and have a fast response speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of accelerators, and particularly to a reinforcement learning parallel processing accelerator, an acceleration method, and an electronic device. Background Art

[0002] The reinforcement learning model is a composite network in architecture, usually composed of a convolutional network, a Generative Pretrained Transformer (GPT), and a Feed Forward Neural Network (FFN). During the operation of the model, a large number of various computing operations are used to process various types of input data, but there are differences in the quantity scale of these data. Traditionally, such reinforcement learning models are computed using a Graphics Processing Unit (GPU) or some traditional accelerators. However, in the related technical solutions, the accelerator does not combine the structural characteristics of such models, and the acceleration effect is not good. Summary of the Invention

[0003] The purpose of the present invention is to provide a reinforcement learning parallel processing accelerator, an acceleration method, and an electronic device, which can deeply analyze the characteristics of the reinforcement learning model, improve the processing efficiency of the model, reduce the resource usage rate of the accelerator, and enhance the acceleration efficiency.

[0004] To solve the above technical problems, the present invention provides a reinforcement learning parallel processing accelerator, including:

[0005] A controller, configured to write the current view feature, at least two groups of batch historical view features, and an instruction sequence into corresponding memories respectively;

[0006] An instruction loading and distributing component, connected to the memory, configured to read the instruction sequence in the memory, decode the read instruction sequence, and distribute the obtained memory access parameters, calculation parameters, and start calculation instructions to a calculation component;

[0007] A data loading and control component, connected to the memory, the instruction loading and distributing component, and each calculation component respectively, configured to select required feature data from the memory according to the parameters of each layer and load it into a corresponding feature cache;

[0008] The calculation component, connected to the instruction loading and distributing component, configured to, after receiving the start calculation instruction, simultaneously read the current view feature data and at least two groups of batch historical view feature data in the feature cache, combine the corresponding memory access parameters and calculation parameters, perform parallel processing, obtain intermediate layer result data and final result data, and return them to the controller.

[0009] To solve the above technical problems, the present invention also provides an acceleration method for a reinforcement learning parallel processing accelerator, where the reinforcement learning parallel processing accelerator is the above-mentioned reinforcement learning parallel processing accelerator provided by the present invention, and the acceleration method includes:

[0010] The controller writes the current view features, at least two sets of batch historical view features, and the instruction sequence into the corresponding memories respectively;

[0011] The instruction loading and distribution component reads the instruction sequence in the memory, decodes the read instruction sequence, and distributes the obtained memory access parameters, calculation parameters, and start calculation instructions to the calculation components;

[0012] The data loading control component selects the required feature data from the memory according to the parameters of each layer and loads it into the corresponding feature cache;

[0013] After receiving the start calculation instruction, the calculation components simultaneously read the current view feature data and at least two sets of batch historical view feature data in the feature cache, and perform parallel processing in combination with the corresponding memory access parameters and calculation parameters to obtain the intermediate layer result data and the final result data, and return them to the controller.

[0014] As can be seen from the above technical solutions, an electronic device provided by the present invention includes the above-mentioned reinforcement learning parallel processing accelerator provided by the present invention.

[0015] The beneficial effects of the present invention are as follows: for the above-mentioned reinforcement learning parallel processing accelerator provided by the present invention, the controller writes the current view features, at least two sets of batch historical view features, and the instruction sequence into the corresponding memories respectively to prepare for subsequent processing; the instruction loading and distribution component can quickly read, decode the instruction sequence, and distribute the parameters and start calculation instructions to the calculation components; the data loading control component selects the required feature data from the memory according to the parameters of each calculation layer and loads it into the corresponding feature cache; after receiving the start calculation instruction, the calculation components simultaneously read the current view feature data and at least two sets of batch historical view feature data in the feature cache, and perform parallel processing in combination with the corresponding memory access parameters and calculation parameters. In this way, multiple view feature data are read simultaneously and parallel processing is performed in combination with the parameters, making full use of the hardware resources, greatly improving the data processing efficiency, reducing the processing time, accelerating the operation of the reinforcement learning algorithm, and being applicable to fast decision-making in complex scenarios. The above components have clear division of labor and cooperate with each other, can deeply analyze the characteristics of the reinforcement learning model, flexibly adjust the parameters and processing flow according to different reinforcement learning tasks and data characteristics, adapt to a variety of application scenarios, make the data reading, instruction processing, and calculation execution proceed in an orderly manner, improve the processing efficiency of the reinforcement learning model, and at the same time reduce the resource utilization rate of the accelerator, achieving a better acceleration ratio and fast response speed.

[0016] In addition, the present invention also provides corresponding acceleration methods and electronic devices for the reinforcement learning parallel processing accelerator, which have the same or corresponding technical features as the above-mentioned reinforcement learning parallel processing accelerator, and the effects are the same as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic structural diagram of the reinforcement learning parallel processing accelerator provided by the embodiment of the present invention;

[0019] Figure 2 Schematic topological structure diagram of the reinforcement learning parallel processing accelerator provided by the embodiment of the present invention;

[0020] Figure 3 Schematic structural diagram of the flattening and splicing component provided by the embodiment of the present invention;

[0021] Figure 4 Schematic structural diagram of the stacked parallel component provided by the embodiment of the present invention;

[0022] Figure 5 Schematic data processing diagram of the stacked parallel component provided by the embodiment of the present invention;

[0023] Figure 6 Schematic structural diagram of the query key-value splitting component provided by the embodiment of the present invention;

[0024] Figure 7 Schematic data processing diagram of the query key-value splitting component provided by the embodiment of the present invention;

[0025] Figure 8 Schematic structural diagram of the splitting and splicing component provided by the embodiment of the present invention;

[0026] Figure 9 Schematic data processing diagram of the splitting and splicing component provided by the embodiment of the present invention;

[0027] Figure 10 Schematic calculation flow diagram of the reinforcement learning parallel processing accelerator provided by the embodiment of the present invention (one);

[0028] Figure 11 Schematic calculation flow diagram of the reinforcement learning parallel processing accelerator provided by the embodiment of the present invention (two);

[0029] Figure 12Flowchart of the acceleration method for the reinforcement learning parallel processing accelerator provided by the embodiments of the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. Figure 1 Schematic diagram of the structure of the reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, as Figure 1 shown, the reinforcement learning parallel processing accelerator includes:

[0032] A controller, configured to write the current view feature, at least two sets of batch historical view features, and the instruction sequence into corresponding memories respectively;

[0033] An instruction loading and distributing component, connected to the memory, configured to read the instruction sequence in the memory, decode the read instruction sequence, and distribute the obtained memory access parameters, calculation parameters, and start calculation instructions to the calculation components;

[0034] A data loading and control component, respectively connected to the memory, the instruction loading and distributing component, and each calculation component, configured to select the required feature data from the memory according to the parameters of each layer and load it into the corresponding feature cache;

[0035] A calculation component, connected to the instruction loading and distributing component, configured to, after receiving the start calculation instruction, simultaneously read the current view feature data and at least two sets of batch historical view feature data in the feature cache, and perform parallel processing in combination with the corresponding memory access parameters and calculation parameters to obtain intermediate layer result data and final result data and return them to the controller.

[0036] In the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the controller writes the current view features, at least two sets of batch historical view features, and the instruction sequence into the corresponding memories respectively to prepare for subsequent processing; the instruction loading and distribution component can quickly read, decode the instruction sequence, and distribute the parameters and start the calculation instructions to the calculation component; the data loading control component selects the required feature data from the memory according to the parameters of each calculation layer and loads it into the corresponding feature cache; after receiving the start calculation instruction, the calculation component reads the current view feature data and at least two sets of batch historical view feature data in the feature cache at the same time, and combines the corresponding memory access parameters and calculation parameters for parallel processing. In this way, multiple view feature data are read at the same time and combined with the parameters for parallel processing, which makes full use of the hardware resources, greatly improves the data processing efficiency, reduces the processing time, accelerates the operation of the reinforcement learning algorithm, and is applicable to fast decision-making in complex scenarios. The above components have clear division of labor and cooperate with each other, can deeply analyze the characteristics of the reinforcement learning model, flexibly adjust the parameters and processing flow according to different reinforcement learning tasks and data characteristics, adapt to a variety of application scenarios, make the data reading, instruction processing, and calculation execution proceed in an orderly manner, improve the processing efficiency of the reinforcement learning model, and at the same time reduce the resource utilization rate of the accelerator, achieve a better acceleration ratio, and have a fast response speed.

[0037] It should be noted that the controller can select a Peripheral Component Interconnect Express (PCIE) controller, which is an interface for high-speed data transmission between the entire system and the outside. The memory can select a High-Bandwidth Memory (HBM), which is a high-performance storage chip. The present invention can send a data transmission instruction to the PCIE controller through a host computer. For example, after receiving the data transmission instruction from the host computer, the PCIE controller can pre-load the key parameters of the reinforcement learning network into the corresponding HBM memory. In this way, before the actual reinforcement learning calculation process starts, the relevant data and instructions are loaded into a specific HBM memory in advance, so that they can be quickly obtained for subsequent use, reducing the waiting time and improving the system operation efficiency.

[0038] The controller can write the current image features and historical image features into the corresponding feature input caches respectively through Direct Memory Access (DMA). Direct Memory Access allows external devices to directly transfer data with the computer memory without the continuous participation of the processor.

[0039] The above image can be a Bird's-Eye View (BEV), which refers to a scene view looking down from above the vehicle. BEV is applied in fields such as autonomous driving and computer vision. It can convert image data obtained from multiple perspectives (such as cameras at different angles around the vehicle) into a perspective similar to looking down from the air through algorithms to present environmental information, forming a top-view data representation. This data representation method can more intuitively display information such as the positions, shapes, and relative relationships of objects around the vehicle, which is helpful for subsequent tasks such as object detection and path planning. The autonomous driving perception model can use BEV features as input.

[0040] Figure 2 This is a schematic diagram of the topology structure of the reinforcement learning parallel processing accelerator provided by the embodiments of the present invention. As Figure 2 shown, the controller (such as a PCIE controller) can write the current view features (such as C-BEV), at least two sets of batch history view features (such as the first set of batch history view features H0-BEV and the second set of batch history view features H1-BEV), and the instruction sequence into the corresponding memories (such as HBM) through DMA transfer. Specifically, the controller can access memories 0-7 and memories 10-11 through H2D (Host to Device), that is, transfer the data of the host to these memories, and can also receive the data returned by memories 8-9 through D2H (Device to Host). The bit width of the data output by the memory can be 512 bits (bit).

[0041] It should be noted that the calculation of the present invention runs in an offline mode. During the process, the parameter distribution component can parse the instructions of each calculation layer in the set order to obtain the memory simulation parameters, calculation parameters, start instructions, etc. of each calculation layer, and transmit them to the corresponding calculation components. After receiving the start instructions, each calculation component reads the current BEV feature data and historical BEV feature data in the input cache at the same time, and reads the corresponding memory simulation parameters and calculation parameters for corresponding parallel calculation processing.

[0042] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the controller can also be used to write the state vector, action sequence, reward sequence, weight data, bias data, and embedding data into the corresponding memories respectively. Correspondingly, the data loading control component can also be used to select the required state vector, action sequence, and reward sequence from the memories according to the parameters of each layer and load them into the corresponding register groups, and write the weight data, bias data, and embedding data in the memories into the corresponding caches in sequence.

[0043] In implementation, the state vector represents state information, and the reward sequence stores the reward values obtained for the corresponding action sequence. The weight data is the strength of the connection between neurons, which determines the influence degree of the input data on the neuron output. The bias data is a constant term added in the neuron calculation, which can adjust the activation function of the neuron at different positions and ranges, increasing the flexibility and expressive power of the network. The embedding data can map high-dimensional discrete data (such as words, categories in text, etc.) to a low-dimensional continuous vector space. Through the controller, the present invention can write the state vector, action sequence, reward sequence, weight data, bias data, and embedding data into the corresponding memory (such as HBM) respectively, making the data storage regular and orderly, facilitating subsequent use, and also being conducive to the secure storage and management of data, reducing data chaos and errors. The data loading control component selects the required state vector, action sequence, and reward sequence from the memory according to the parameters of each layer and loads them into the corresponding register bank, which can quickly provide the necessary data for the operation and shorten the data search time. The weight data, bias data, and embedding data in the memory are sequentially written into the corresponding cache, enabling the processor to quickly access these frequently used data, reducing the waiting time, and improving the overall operation speed.

[0044] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiment of the present invention, the computing component can be used to perform parallel processing based on the state vector, action sequence, and reward sequence in the register bank, the weight data, bias data, and embedding data in the cache, and the read corresponding feature data, memory access parameters, and calculation parameters, to obtain the intermediate layer result data and the final result data; the intermediate layer result data is transferred and stored between each group of caches and the register bank.

[0045] In implementation, each computing component can perform parallel processing based on the state vector, action sequence, and reward sequence in the register bank, the weight data, bias data, and embedding data in the cache, and the read corresponding feature data, memory access parameters, and calculation parameters, to obtain the intermediate layer result data and the final result data. This can make full use of the hardware resources, execute multiple computing tasks simultaneously, greatly shorten the overall computing time, and is suitable for processing large-scale and complex computing tasks. The intermediate layer result data can be transferred and stored between the cache and the register bank, facilitating flexible invocation in different computing steps, being able to adapt to the complex computing processes of different algorithms and models, and not needing to reread the original data every time, improving the flexibility and convenience of data processing.

[0046] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: a First In First Out (FIFO) queue, which is used to convert the bus width of the intermediate layer result data and the final result data, and write the obtained data after conversion into the memory.

[0047] In implementation, as Figure 2 shown, the First In First Out queue is a rule or strategy for data storage and processing. Just like queuing, the data or elements that enter first will be processed or taken out first, and those that enter later will be processed or taken out later, to ensure that the data is processed in the order of its arrival, guaranteeing the certainty and fairness of the processing order. The present invention uses the First In First Out queue to convert the bus width of the intermediate layer result data and the final result data, and write the obtained data after conversion into the memory. In this way, the First In First Out queue can be used as a buffer area to temporarily store the intermediate layer and final result data, enabling the smooth transmission of data between modules with different speeds, and avoiding data loss or data conflicts. Figure 2 The data width written into the First In First Out queue in

[0048] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: a cache cross-selection component, which is respectively connected to each computing component, and may include a feature cache for respectively storing caches of weight data, bias data, and embedding data, and a register group for respectively storing state vectors, action sequences, and reward sequences; the cache cross-selection component can be used to select and schedule among the data sources of the feature cache, the cache for respectively storing weight data, bias data, and embedding data, and the register group.

[0049] In implementation, as Figure 2 shown, the cache cross-selection component may include a feature cache for respectively storing caches of weight data, bias data, and embedding data, and a register group for respectively storing state vectors, action sequences, and reward sequences, such as cache 1-3 corresponding to C, etc.; cache 1-3 corresponding to H0, etc.; cache 1-3 corresponding to H1, etc.; register groups corresponding to C, H0, and H1; and separate register groups 2-H0, 2-H1, 3-H0, 3-H1, 4-H0, 4-H1, register groups 5-8, etc. The cache cross-selection component can store different types of data in the feature cache, dedicated cache, and register group respectively, and perform unified selection and scheduling, making the storage and search of data more orderly. When a computing component needs specific data, the cache cross-selection component can quickly locate and provide it, reducing the data retrieval time and improving the data acquisition efficiency.

[0050] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the computing component may include: a General Matrix Multiply (GEMM) computing component respectively connected to the instruction loading and distribution component and the cache cross-selection component; the GEMM computing component includes a first matrix multiplier component for performing matrix multiplication calculation on the current view features, a second matrix multiplier component for performing matrix multiplication calculation on the first set of batch historical view features, and a third matrix multiplier component for performing matrix multiplication calculation on the second set of batch historical view features.

[0051] In implementation, as Figure 2 shown, the matrix multiplication computing component in the computing component may include multiple matrix multiplier components. The first matrix multiplier component, the second matrix multiplier component, and the third matrix multiplier component can perform matrix multiplication calculations on the current view features, the first set of batch historical view features, and the second set of batch historical view features respectively at the same time, realizing the parallelization of the calculation, greatly improving the calculation speed, being able to process a large amount of data in a shorter time, especially when processing tasks containing multi-view features such as complex image or video data, the calculation time can be significantly reduced. And different matrix multiplier components are responsible for different tasks, avoiding the bottleneck problem that may occur when a single component processes all calculation tasks. Each component focuses on a specific type of feature calculation, can more efficiently utilize the hardware resources, and improve the overall calculation efficiency.

[0052] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the first matrix multiplier component, the second matrix multiplier component, and the third matrix multiplier component may all include: one or a combination of a Scale part, a Mask part, an Atten_mask part, a Bias part, a Residual part, an Embedding part, and an Active part; the Scale part is used to perform a scaling operation on the data under the control of a scaling enable signal; the Mask part is used to perform a masking calculation on the data under the control of a masking enable signal; the Atten_mask part is used to perform an attention masking calculation on the data under the control of an attention masking enable signal; the Bias part is used to perform a bias calculation on the data under the control of a bias enable signal; the Residual part is used to perform a residual calculation on the data under the control of a residual enable signal; the Embedding part is used to perform an embedding calculation on the data under the control of an embedding enable signal; the Active part is used to perform an activation calculation on the data under the control of an activation enable signal.

[0053] In implementation, the scaling unit can perform flexible scaling operations on data according to the scaling enable signal, which can be used to adjust the dynamic range of the data, making the data more suitable for subsequent calculations and processing, and helping to improve the accuracy and stability of the calculations. The masking unit performs masking calculations on the data through the masking enable signal, and can selectively mask or retain certain features in the data, helping to highlight key information. The attention masking unit performs attention masking calculations under the control of the attention masking enable signal, enabling the model to pay more attention to important parts of the data, automatically allocate computing resources, and thus better capture long-sequence dependencies and key features in the data, enhancing the model's ability to understand and process complex data. The bias unit performs bias calculations on the data using the bias enable signal, and can add a fixed offset to the data, helping to adjust the distribution of the data and enabling the model to better fit the data. The residual unit performs residual calculations through the residual enable signal, which can fuse the original data with the processed data, effectively retaining the information in the original data, avoiding information loss or vanishing gradients in multi-layer calculations, helping to train deeper models, and improving the expressive power of the model. The embedding unit performs embedding calculations under the control of the embedding enable signal, which can transform the data into a low-dimensional vector space more suitable for model processing, represent discrete or high-dimensional data as continuous low-dimensional vectors, reduce the curse of dimensionality of the data, and at the same time extract the latent features of the data, facilitating model learning and utilization. The activation unit performs activation calculations on the data under the control of the activation enable signal, introducing non-linear factors into the model, enabling the model to learn complex non-linear relationships in the data, enhancing the expressive power of the model, enabling the model to fit various complex functions, and improving the accuracy and generalization ability of the model. The existence of these components enables the model to more effectively adjust parameters during the training process. Through various preprocessing and calculation operations on the data, the model converges more easily, reducing the consumption of training time and computing resources.

[0054] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the computing component may further include: a Flatten_Concat component respectively connected to the instruction loading and distribution component and the cache cross-selection component; the Flatten_Concat component includes a first Flatten_Concat sub-component for performing matrix multiplication calculations on the current view features, a second Flatten_Concat sub-component for performing matrix multiplication calculations on the first set of batch historical view features, and a third Flatten_Concat sub-component for performing matrix multiplication calculations on the second set of batch historical view features.

[0055] In implementation, as Figure 2As shown, the flattening and splicing component can include multiple flattening and splicing sub-components. Since different view features and batch history view features usually have different dimensions and shapes, the flattening and splicing sub-components can perform flattening operations on these features, converting multi-dimensional data into one-dimensional data for subsequent processing and calculation, enabling data from different sources to be integrated and analyzed in a unified format, and improving the efficiency and accuracy of data processing. Through the splicing operation, the current view feature, the first set of batch history view features, and the second set of batch history view features are combined, which can fuse information from different time dimensions and perspectives, enrich the feature representation of the data, help the model capture more comprehensive and complex patterns and relationships, and enhance the model's ability to understand and analyze the data.

[0056] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the first flattening and splicing sub-component, the second flattening and splicing sub-component, and the third flattening and splicing sub-component all include a first input port, a second input port, a channel control input section, a first data selection controller, a first multiplexer, and a first output port; the first input port is connected to the channel control input section and is used to sequentially read the feature data of each channel group through the channel control input section; the first data selection controller is connected to the first multiplexer and is used to rearrange the feature data of each channel group and transmit the rearranged feature data to the first multiplexer; the second input port is connected to the first multiplexer and is used to read the state vector and transmit it to the first multiplexer; the first output port is used to output the feature data and the state vector selected by the first multiplexer.

[0057] Figure 3 It is a schematic structural diagram of the flattening and splicing component provided by the embodiments of the present invention. As Figure 3As shown, the first input port can sequentially read the feature data of each channel group through the channel control input section. This enables the sub-component to flexibly process information from different channels, is suitable for processing data with multi-channel characteristics, can read and process the features of each channel targeted according to specific requirements, and improves the processing ability for complex data. The first data selection controller can rearrange the feature data of each channel group. This rearrangement operation can adjust the order and structure of the data according to different task and algorithm requirements, optimize the processing efficiency of the data in subsequent calculations, facilitate the model to better extract features, and enhance the configurability and adaptability of data processing. The second input port reads the state vector and transmits it to the first multiplexer, where it is fused with the rearranged feature data. This fusion method combines different types of information (feature data and state vector), can provide richer and more comprehensive information for subsequent calculations and analyses, and helps the model better understand the context and dynamic changes of the data. The first multiplexer can select the appropriate combination of feature data and state vector according to needs and output it through the first output port. This enables the sub-component to flexibly select and output the most useful data according to different conditions and task requirements, improves the efficiency and pertinence of data processing, and avoids unnecessary data transmission and calculations. The above structures of the first flattening and splicing sub-component, the second flattening and splicing sub-component, and the third flattening and splicing sub-component can improve the scalability and compatibility of the entire system.

[0058] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the computing component may further include: a stacked parallel (Stack) component respectively connected to the instruction loading and distribution component and the cache cross-selection component; the stacked parallel component includes a third input port, a fourth input port, a fifth input port, a second data selection controller, and a second output port; the third input port is used to read the state vectors corresponding to at least two groups of historical view features; the fourth input port is used to read the action sequences corresponding to at least two groups of historical view features; the fifth input port is used to read the reward sequences corresponding to at least two groups of historical view features; the second data selection controller is used to rearrange the data read from the third input port, the fourth input port, and the fifth input port, and output the rearranged data from the second output port.

[0059] Figure 4 It is a schematic structural diagram of the stacked parallel component provided by the embodiments of the present invention. As Figure 4As shown, the third input port (i.e., input ports 1-H0, 1-H1), the fourth input port (i.e., input ports 2-H0, 2-H1), and the fifth input port (i.e., input ports 3-H0, 3-H1) are respectively responsible for reading the state vectors, action sequences, and reward sequences corresponding to at least two sets of historical view features. This parallel reading mechanism can simultaneously obtain various types of data, greatly improving the data reading speed and processing efficiency. Compared with the sequential data reading method, it can collect various information required by the model in a shorter time, laying a foundation for subsequent rapid processing and analysis. Figure 5 It is a schematic diagram of data processing of the stacked parallel component provided by an embodiment of the present invention, as Figure 5 shown, the second data selection controller can rearrange the data from the three input ports. This can reduce the time overhead of data access, improve the data flow efficiency in the whole system, and thus enhance the overall data processing performance.

[0060] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by an embodiment of the present invention, the computing component may further include: a query key-value splitting (QKV_Split) component respectively connected to the instruction loading and distribution component and the cache cross-selection component; the query key-value splitting component includes a sixth input port, a shunt controller, a third output port, a fourth output port, and a fifth output port; the sixth input port is used to read the hidden layer features to be subjected to query key-value splitting; the shunt controller is used to sequentially divide the feature data read by the sixth input port into query cache feature data, key cache feature data, and value cache feature data; the third output port is used to output the query cache feature data; the fourth output port is used to output the key cache feature data; the fifth output port is used to output the value cache feature data.

[0061] Figure 6 It is a schematic structural diagram of the query key-value splitting component provided by an embodiment of the present invention. As Figure 6 shown, in the attention calculation, the query key-value splitting component inputs the hidden layer features to be subjected to query key-value splitting, and through the shunt controller, the read hidden layer features can be accurately divided into three groups of data, namely query cache feature data, key cache feature data, and value cache feature data. This clear classification method enables different types of data to be accurately identified and separated, providing a basis for subsequent specific processing of different types of data, avoiding the possibility of data confusion and incorrect processing, and improving the accuracy and reliability of data processing. Figure 7 It is a schematic diagram of data processing of the query key-value splitting component provided by an embodiment of the present invention, as Figure 7As shown, the third output port outputs query cache feature data; the fourth output port outputs key cache feature data; and the fifth output port outputs value cache feature data. Query cache feature data can be used for fast retrieval and matching operations, key cache feature data can be used for indexing and positioning, and value cache feature data can be used to store and obtain specific information content. This targeted processing can give full play to the role of each type of data and improve the processing efficiency and effect of data in various application scenarios.

[0062] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by an embodiment of the present invention, the computing component may also include: a split-concatenation (Split_Concat) component respectively connected to the instruction loading and distribution component and the cache cross selection component; the split-concatenation component includes a seventh input port, an eighth input port, a third data selection controller, a second multiplexer, and a sixth output port; the seventh input port is connected to the second multiplexer, and is used to read the state sequence corresponding to the current view feature, and transmit the state sequence corresponding to the current view feature to the second multiplexer; the eighth input port is connected to the third data selection controller, and is used to group the input hidden layer features into a reward sequence, a state sequence and an action sequence through the third data selection controller, and transmit them to the second multiplexer; the sixth output port is used to output the data selected by the second multiplexer.

[0063] Figure 8 The schematic diagram of the structure of the splitting and splicing assembly provided by the embodiment of the present invention. Figure 8 As shown, the seventh input port is responsible for reading the state sequence corresponding to the current view feature, and the eighth input port can group the input hidden layer features into reward sequences, state sequences, and action sequences. This design enables the component to process data from different sources and types at the same time, integrate multi-source data, and provide richer and more comprehensive information for subsequent analysis and calculation. The third data selection controller can group the hidden layer features into reward sequences, state sequences, and action sequences as needed, and the second multiplexer can select from these data according to specific needs. This flexible data grouping and selection mechanism allows the component to process and output data in a targeted manner according to different tasks and scenarios, thereby improving the efficiency and accuracy of data processing. Figure 9 A schematic diagram of data processing of the splitting and splicing components provided in an embodiment of the present invention, such as Figure 9As shown, the splitting and splicing component centralizes operations such as data splitting, grouping, and selection, and finally outputs the selected data through an output port. This can accurately group and flexibly combine these important sequences, providing strong support for the training of the reinforcement learning model. By reasonably selecting and outputting these sequence data, the model can better learn the relationship between the environmental state, actions, and rewards, thereby optimizing the strategy, improving the accuracy and efficiency of decision-making, reducing the number of data transmissions and complexity between different components, reducing the data transmission latency, and improving the overall efficiency of data processing.

[0064] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the computing component may further include: a Softmax component respectively connected to the instruction loading and distribution component and the cache cross-selection component; the Softmax component is used to perform Softmax operation on the read data.

[0065] In implementation, as Figure 2 shown, the Softmax component can uniformly map the data to a specific range or scale, making different data comparable, and the normalized data can make the model easier to converge, helping to improve the generalization ability and accuracy of the model.

[0066] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the computing component may further include: a Layer Normalization (LN) component respectively connected to the instruction loading and distribution component and the cache cross-selection component; the LN component is used to receive weight data and bias data and perform normalization processing in combination with the read data.

[0067] In implementation, as Figure 2 shown, the LN component can make the model more adaptable to different input data by performing normalization processing on each layer of data, reducing the model's dependence on a specific data distribution, thereby improving the generalization ability of the model. The weight data (LN_Weight) and bias data (LN_Bias) in the LN component are two important parameters in the LN operation, which are used to scale and translate the input data respectively, improving the adaptability and processing ability of the computing component to different data, and enhancing the reliability and stability of the model in various practical application scenarios.

[0068] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, two groups of memories are used as ping-pong caches to interactively store the calculation results and provide the feature inputs for the next layer.

[0069] In implementation, the ping-pong buffer mechanism enables the storage and reading of data to alternate between two sets of memories. When one set of memories provides data to the computing component as the input of the next-layer features, the other set can simultaneously store the calculation results, thus achieving seamless data connection, avoiding waiting time in the data transmission and processing process, improving the overall computing efficiency, and ensuring the continuity of the calculation.

[0070] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: a first selector, connected to the computing component, the cache cross-selection module, and the first-in-first-out queue.

[0071] In implementation, the first selector can flexibly select to transfer data from the computing component to the cache cross-selection module or the first-in-first-out queue, ensuring that data flows along a suitable path between different components. That is to say, the existence of the first selector helps to achieve parallel transmission and processing of data among the computing component, the cache cross-selection module, and the first-in-first-out queue, and speeds up the execution speed of the overall computing task.

[0072] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: a second selector, connected to the instruction loading and distribution component, the matrix multiplication calculation component, the cache cross-selection component, and the parameter cache for storing weight data, bias data, and embedding data.

[0073] In implementation, the second selector can accurately select and transfer instructions to the corresponding components, ensuring that components such as the matrix multiplication calculation component can obtain the required instructions in a timely manner and execute operations in sequence, improving the accuracy and efficiency of instruction execution. As Figure 2 shown, the weight cache, bias cache, and embedding cache respectively receive data from the controller (such as data with a bit width of 512bit). After cache processing, the output weight data, bias data, and embedding data (such as data with a bit width of 2048bit) are written into the parameter cache to provide data support for the matrix multiplication calculation component. The cache cross-selection component can provide attention key-value (KV) matrix data, etc. for the matrix multiplication calculation component. The attention KV matrix data and the weight data can be input into the matrix multiplication calculation component after being processed by the second selector.

[0074] Furthermore, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the controller can also be used to read the data obtained after the conversion of the first-in-first-out queue from the memory for data comparison.

[0075] In implementation, the controller compares the data after conversion by reading the data in the first-in-first-out queue. The controller can verify whether the data remains consistent during the conversion process, ensure that there are no errors or losses in the storage, transmission, and processing of the data, guarantee the accuracy and integrity of the data in the system, and provide a reliable data basis for subsequent operations and decisions.

[0076] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: an instruction control component, connected to the controller and the instruction loading and distribution component, and used to send a start command to the instruction loading and distribution component.

[0077] In implementation, as Figure 2 shown, the controller can write relevant instruction data to the instruction control component through register control, and the bit width of the register control channel can be 32bit. The relevant instruction data here may include debug instructions, reset instructions, start instructions, etc., and is used to control the operating state of the entire system. The present invention can send debug instructions, reset instructions, start calculation instructions, etc. to the instruction loading and distribution component to start and execute the calculation tasks of the entire network; the instruction loading and distribution component can return a completion signal, the number of running layers, and the calculation time to the instruction control component.

[0078] Further, in specific implementation, in the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the instruction control component is further used to notify the host computer in an interrupt manner after all layers of calculations are completed; or, it is further used to feedback the operation completion status to the host computer when receiving a query instruction from the host computer after all layers of calculations are completed.

[0079] In implementation, the instruction control component has the function of notifying the host computer in an interrupt manner or feedbacking the operation completion status according to the host computer query instruction after all layers of calculations are completed. Notifying the host computer in an interrupt manner enables the host computer to obtain the message at the first time after the calculation is completed, without the host computer continuously polling the underlying calculation status, which greatly saves the time and resources of the host computer. The instruction control component receives the host computer query instruction and feedbacks the operation completion status, providing a flexible way for the host computer to obtain information. The present invention adopts interrupt notification or query feedback, both of which can enable the host computer to clearly know the calculation task status, avoid misoperations or repeated operations caused by unclear information, and enhance the stability and reliability of the system.

[0080] In practical applications, the present invention analyzes the reinforcement learning network structure. Table 1-5 shows the calculation order of the network. Here, C is the number of blocks in each row of the block matrix, R is the number of rows of the block matrix, H is the number of attention heads. Let X be the feature matrix, W be the weight matrix, E be the residual matrix, b be the bias, D be the embedding, S be the scaling, M be the mask, and A be the attention mask. Conv represents convolution, the i in RAMi represents input, RAM represents the memory, ReLU represents the rectified linear, Softplus represents an activation function, RegGrp is the register group, and Att represents the attention score.

[0081] Table 1 Example 1 of Feature Cache Location

[0082]

[0083] Table 2 Example 2 of Feature Cache Location

[0084]

[0085] Table 3 Example 3 of Feature Cache Location

[0086]

[0087] Table 4 Example 4 of Feature Cache Location

[0088]

[0089] Table 5 Example 5 of Feature Cache Location

[0090]

[0091] Examples in Table 1-5 illustrate the switching order of the feature cache storage locations for each step of calculation in the network. Among them, according to the data cache requirements, the K matrix and the V matrix are stored in the cache -H0.

[0092] The bias, weight, and embedding use independent caches. The feature cache includes input caches of C-BEV, H0-BEV, and H1-BEV, -C-- -C, -H0-- -H0 and -H1-- -H1 feature calculation transfer cache and dedicated register groups RegGrp1, RegGrp2-H0—RegGrp4-H0, RegGrp2-H1—RegGrp4-H1, RegGrp5—RegGrp8. Among them, register groups 1 and 2 are the input state vectors, register groups 3 and 4 are the input action sequences and reward sequences, register group 5 is the state features obtained by C-BEV calculation, registers 6 and 7 are the distribution mean and distribution variance of the calculated policy network, and register group 8 is the value scalar of the calculated value network. In addition, when performing the split operation numbered 19 in Table 1, since three groups of matrices need to be generated, to reduce the total cache usage and because the K matrix and the V matrix will not participate in the calculation simultaneously, the K matrix and the V matrix can be placed on the same group of cache and stored in two regions. This avoids the waste of space caused by allocating independent caches for the K matrix and the V matrix respectively, significantly reducing the overall cache usage. In the case of limited memory resources, this optimization can free up more cache space for other data or operations, improving the memory usage efficiency of the entire system, reducing the overall memory access scheduling overhead of the system, and making the system structure more compact. Compared with fixedly allocating a large block of cache space for each matrix, this method can make more full use of each byte of the cache, making the cache space more effectively utilized, reducing the problem of partial space idleness that may be caused by fixed allocation, and enabling subsequent data access to hit the cache faster, reducing the number of cache misses and improving the calculation performance.

[0093] It should be noted that the present invention adopts an offline scheduling method. The host computer preloads the current view, at least two groups of batch history views, weight data, etc. into the memory, and then sends a start calculation command to the instruction control component. After the instruction control component sends an enable instruction to the instruction loading and distribution component, the instruction loading and distribution component controls each calculation component to complete the corresponding calculation process in sequence, and each component completes data loading, calculation, and result writing back according to the calculation parameters.

[0094] Figure 10 It is one of the schematic diagrams of the calculation process of the reinforcement learning parallel processing accelerator provided by the embodiment of the present invention. Figure 11 It is the second schematic diagram of the calculation process of the reinforcement learning parallel processing accelerator provided by the embodiment of the present invention. The calculation process may include: as Figure 10 shown, simultaneously perform convolutional feature extraction on the current image and at least two groups of batch history images. Assuming that convolutional feature extraction is performed on the C-BEV image, H0-BEV image, and H1-BEV image, C-BEV features, H0-BEV features, and H1-BEV features are obtained; Figure 10In this, N represents the number of cycles. Based on the C-BEV feature, H0-BEV feature, and H1-BEV feature, the C-state vector, H0-state vector, and H1-state vector are obtained. Then, matrix multiplication calculations are performed on the C-state vector, H0-state vector, and H1-state vector to obtain the C-state feature, H0-state feature, and H1-state feature. The C-state vector and the C-BEV feature are flattened and concatenated to obtain the C-concatenated feature, and matrix multiplication calculations are performed on it to obtain the C-state sequence. At the same time, the H0-state vector and the H0-BEV feature are flattened and concatenated to obtain the H0-concatenated feature, and matrix multiplication calculations are performed on it to obtain the H0-state sequence; the H1-state vector and the H1-BEV feature are flattened and concatenated to obtain the H1-concatenated feature, and matrix multiplication calculations are performed on it to obtain the H1-state sequence. Next, as Figure 11 shown, embedding processes are respectively performed according to the H0-state sequence and the H1-state sequence to obtain the H0-state feature embedding and the H1-state feature embedding, and then the H0-action sequence and the H1-action sequence are obtained. Embedding processes are performed on the H0-action sequence and the H1-action sequence to obtain the H0-action feature embedding and the H1-action feature embedding. Based on the action feature embedding, the reward sequence is obtained. Embedding processes are performed on the reward sequence to obtain the H0-embedded reward feature embedding and the H1-embedded reward feature embedding. Then, the H0-state feature embedding, the H0-action feature embedding, the H0-reward feature embedding, the H1-state feature embedding, the H1-action feature embedding, and the H1-reward feature embedding are stacked and calculated. After layer normalization processing and one-dimensional convolution operations are performed on the data after the stacking calculation, query-key-value splitting is performed to obtain the corresponding query, key, and value, attention scores are obtained under the attention mask, and then calculation results are obtained after operations such as the normalization exponential function. If all data blocks are calculated, layer normalization processing continues, splitting and concatenation processing is combined with the C-state sequence to obtain the state feature, and then a series of calculation steps are performed, including multiple fully-connected layers (FC), rectified linear units (ReLU), normalization exponential functions, activation functions (Softplus), etc., to obtain the distribution mean and distribution variance of the policy network, and finally the value scalar of the value network is obtained to complete the calculation process.

[0095] After the calculation process is completed, the host computer can be notified of the interruption or the host computer can query the completion status through the instruction control module, read the results, and then load the next set of feature data and issue a start calculation command. This cycle continues until all feature data is calculated. The calculation process is controlled by the instruction parameter distribution module according to the parameters obtained by decoding the instructions loaded in advance, without the participation of the host computer, which improves the processing efficiency and avoids interaction delays.

[0096] In the above embodiments, the reinforcement learning parallel processing accelerator is described in detail. Based on the same inventive concept, embodiments of the present invention also provide an acceleration method for the reinforcement learning parallel processing accelerator and corresponding embodiments of an electronic device.

[0097] Figure 12 Flowchart of the acceleration method for the reinforcement learning parallel processing accelerator provided by the embodiments of the invention. The acceleration method for the reinforcement learning parallel processing accelerator provided in this embodiment is as Figure 12 shown and includes the following steps:

[0098] S1201. The controller writes the current view feature, at least two sets of batch historical view features, and the instruction sequence into corresponding memories respectively.

[0099] S1202. The instruction loading and distribution component reads the instruction sequence in the memory, decodes the read instruction sequence, and distributes the obtained memory access parameters, calculation parameters, and start calculation instructions to the calculation components.

[0100] S1203. The data loading control component selects the required feature data from the memory according to the parameters of each layer and loads it into the corresponding feature cache.

[0101] S1204. After receiving the start calculation instruction, the calculation components simultaneously read the current view feature data and at least two sets of batch historical view feature data in the feature cache, and perform parallel processing in combination with the corresponding memory access parameters and calculation parameters to obtain intermediate layer result data and final result data and return them to the controller.

[0102] In the acceleration method of the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, the controller writes the current view features, at least two sets of batch historical view features, and the instruction sequence into the corresponding memories respectively to prepare for subsequent processing; the instruction loading and distribution component can quickly read, decode the instruction sequence, and distribute parameters and start calculation instructions to the calculation components; the data loading control component selects the required feature data from the memory according to the parameters of each calculation layer and loads it into the corresponding feature cache; after receiving the start calculation instruction, the calculation components simultaneously read the current view feature data and at least two sets of batch historical view feature data in the feature cache, and combine the corresponding memory access parameters and calculation parameters for parallel processing. In this way, multiple view feature data are read simultaneously and combined with parameters for parallel processing, making full use of hardware resources, greatly improving data processing efficiency, reducing processing time, accelerating the operation of the reinforcement learning algorithm, and being applicable to fast decision-making in complex scenarios. Through the mutual cooperation of the above components, the characteristics of the reinforcement learning model can be deeply analyzed, and according to different reinforcement learning tasks and data characteristics, the parameters and processing flow can be flexibly adjusted to adapt to a variety of application scenarios, enabling the data reading, instruction processing, and calculation execution to proceed in an orderly manner, improving the processing efficiency of the reinforcement learning model, while reducing the resource utilization rate of the accelerator, achieving a better acceleration ratio, and having a fast response speed.

[0103] Since the embodiments of the acceleration method part correspond to the embodiments of the reinforcement learning parallel processing accelerator part, please refer to the description of the embodiments of the reinforcement learning parallel processing accelerator part for the embodiments of the acceleration method part, and will not be elaborated here. And it has the same beneficial effects as the above-mentioned reinforcement learning parallel processing accelerator.

[0104] Further, in specific implementation, in the acceleration method of the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, while executing step S1201, it may further include: the controller writes the state vector, action sequence, reward sequence, weight data, bias data, and embedding data into the corresponding memories respectively. When executing step S1203, the data loading control component selects the required state vector, action sequence, and reward sequence from the memory according to the parameters of each layer and loads them into the corresponding register group, and writes the weight data, bias data, and embedding data in the memory into the corresponding cache in sequence. When executing step S1204, the calculation components perform parallel processing according to the state vector, action sequence, and reward sequence in the register group, the weight data, bias data, and embedding data in the cache, as well as the read corresponding feature data, memory access parameters, and calculation parameters, to obtain the intermediate layer result data and the final result data; the intermediate layer result data is transferred between the caches and the register group.

[0105] Further, in specific implementation, in the acceleration method of the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, after performing step S1204, it may further include: using a first-in-first-out queue to convert the intermediate layer result data and the final result data to the bus width, and writing the converted data into the memory.

[0106] Further, in specific implementation, in the acceleration method of the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: the cache cross-selection component selects and schedules among the data sources of the feature cache, cache, and register file. The cache cross-selection component is respectively connected to each computing component, and may include a feature cache for respectively storing weight data, bias data, and embedding data, and a register file for respectively storing state vectors, action sequences, and reward sequences.

[0107] Further, in specific implementation, in the acceleration method of the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: two groups of memories are used as ping-pong caches to alternately store the calculation results and provide the feature inputs for the next layer.

[0108] Further, in specific implementation, in the acceleration method of the above-mentioned reinforcement learning parallel processing accelerator provided by the embodiments of the present invention, it may further include: the controller reads the data obtained after the conversion of the first-in-first-out queue from the memory for data comparison.

[0109] For the more specific working processes of the above steps, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0110] Based on the same inventive concept, the embodiments of the present invention further provide an electronic device including the above-mentioned reinforcement learning parallel processing accelerator. Since the principle of the electronic device for solving problems is similar to that of the foregoing reinforcement learning parallel processing accelerator, the implementation of the electronic device may refer to the implementation of the reinforcement learning parallel processing accelerator, and the repeated parts will not be elaborated.

[0111] The various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments may be referred to each other.

[0112] Finally, it should also be noted that unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the terms "including", "comprising" and "having" and any other variants thereof in the present invention are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0113] For the foregoing embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and components involved are not necessarily essential to the present invention.

[0114] The above has introduced in detail the reinforcement learning parallel processing accelerator, acceleration method and electronic device provided by the present invention. The various embodiments in the specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that those of ordinary skill in the art in this technical field can make several improvements and modifications to the present invention without departing from the principle of the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A reinforcement learning parallel processing accelerator, characterized in that: include: A controller, used to write the current view features, at least two groups of batch history view features and the instruction sequence into corresponding memories respectively; An instruction loading and distributing component is connected to the memory, and is used to read the instruction sequence in the memory, decode the read instruction sequence, and distribute the obtained memory access parameters, calculation parameters and start calculation instructions to the calculation component; A data loading control component, connected to the memory, the instruction loading distribution component and each of the computing components, respectively, for selecting required feature data from the memory according to the parameters of each layer and loading it into a corresponding feature cache; The calculation component is connected to the instruction loading and distributing component, and is used to read the current view feature data and at least two groups of batch historical view feature data in the feature cache at the same time after receiving the start calculation instruction, and perform parallel processing in combination with corresponding memory access parameters and calculation parameters to obtain intermediate layer result data and final result data and return them to the controller; A cache cross selection component is connected to each of the computing components, including the feature cache, a cache for storing weight data, bias data and embedded data, and a register group for storing state vectors, action sequences and reward sequences; the cache cross selection component is used to select and schedule data sources among the feature cache, the cache for storing weight data, bias data and embedded data, and the register group; The calculation unit includes: a matrix multiplication calculation component connected to the instruction load distribution unit and the cache cross selection unit respectively; the matrix multiplication calculation component includes a first matrix multiplier component for performing matrix multiplication calculation for current view features, a second matrix multiplier component for performing matrix multiplication calculation for a first group of batch history view features, and a third matrix multiplier component for performing matrix multiplication calculation for a second group of batch history view features.

2. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The controller is further used to write the state vector, action sequence, reward sequence, weight data, bias data and embedded data into the corresponding memory respectively; The data loading control component is also used to select the required state vector from the memory according to the parameters of each layer, load the action sequence and reward sequence into the corresponding register group, and write the weight data, bias data and embedded data in the memory into the corresponding cache in sequence.

3. The reinforcement learning parallel processing accelerator according to claim 2, characterized in that: The computing component is used to perform parallel processing according to the state vector, action sequence and reward sequence in the register group, the weight data, bias data and embedded data in the cache, and the read corresponding feature data, memory access parameters and calculation parameters to obtain intermediate layer result data and final result data; The intermediate layer result data is transferred between each set of cache and register groups.

4. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: Also includes: The first-in-first-out queue is used to convert the bus bit width of the intermediate layer result data and the final result data, and write the data obtained after the conversion into the memory.

5. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The first matrix multiplier component, the second matrix multiplier component and the third matrix multiplier component each include: one or a combination of a scaling part, a mask part, an attention mask part, a bias part, a residual part, an embedding part and an activation part; The scaling unit is used to perform a scaling operation on the data under the control of a scaling enable signal; The mask part is used to perform mask calculation on the data under the control of the mask enable signal; The attention mask unit is used to perform attention mask calculation on the data under the control of the attention mask enable signal; The biasing unit is used to perform bias calculation on the data under the control of the bias enabling signal; The residual part is used to perform residual calculation on the data under the control of the residual enable signal; The embedding unit is used to perform embedding calculation on the data under the control of the embedding enable signal; The activation unit is used to perform activation calculation on the data under the control of the activation enable signal.

6. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The computing component further includes: a flattening and splicing component connected to the instruction load distribution component and the cache cross selection component respectively; The flattening and splicing component includes a first flattening and splicing subcomponent for performing matrix multiplication calculation on current view features, a second flattening and splicing subcomponent for performing matrix multiplication calculation on a first group of batch historical view features, and a third flattening and splicing subcomponent for performing matrix multiplication calculation on a second group of batch historical view features.

7. The reinforcement learning parallel processing accelerator according to claim 6, characterized in that: The first flattening and splicing subassembly, the second flattening and splicing subassembly and the third flattening and splicing subassembly each include a first input port, a second input port, a channel control input portion, a first data selection controller, a first multiplexer and a first output port; The first input port is connected to the channel control input unit and is used to read the characteristic data of each channel group one by one through the channel control input unit; The first data selection controller is connected to the first multiplexer and is used to rearrange the characteristic data of each channel group and transmit the rearranged characteristic data to the first multiplexer; The second input port is connected to the first multiplexer and is used to read the state vector and transmit it to the first multiplexer; The first output port is used to output the feature data and state vector selected by the first multiplexer.

8. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The computing component further includes: a stack parallel component connected to the instruction load distribution component and the cache cross selection component respectively; The stacked parallel component includes a third input port, a fourth input port, a fifth input port, a second data selection controller, and a second output port; The third input port is used to read the state vectors corresponding to at least two groups of historical view features; The fourth input port is used to read the action sequences corresponding to at least two groups of historical view features; The fifth input port is used to read reward sequences corresponding to at least two groups of historical view features; The second data selection controller is used to rearrange the data read from the third input port, the fourth input port, and the fifth input port, and output the rearranged data from the second output port.

9. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The computing component further includes: a query key value splitting component connected to the instruction load distribution component and the cache cross selection component respectively; The query key value splitting component includes a sixth input port, a branching controller, a third output port, a fourth output port and a fifth output port; The sixth input port is used to read the hidden layer features to be split into query keys; The branch controller is used to sequentially divide the feature data read by the sixth input port into query cache feature data, key cache feature data and value cache feature data; The third output port is used to output the query cache feature data; The fourth output port is used to output the key cache feature data; The fifth output port is used to output the value cache characteristic data.

10. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The computing component further includes: a splitting and splicing component connected to the instruction loading and distributing component and the cache cross selection component respectively; The splitting and splicing component includes a seventh input port, an eighth input port, a third data selection controller, a second multiplexer, and a sixth output port; The seventh input port is connected to the second multiplexer, and is used to read the state sequence corresponding to the current view feature, and transmit the state sequence corresponding to the current view feature to the second multiplexer; The eighth input port is connected to the third data selection controller, and is used to group the input hidden layer features into reward sequences, state sequences and action sequences through the third data selection controller, and transmit them to the second multiplexer; The sixth output port is used to output the data selected by the second multiplexer.

11. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The calculation component further includes: a normalized index component connected to the instruction load distribution component and the cache cross selection component respectively; The normalized index component is used to perform a normalized index operation on the read data.

12. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The computing component further includes: a layer normalization component connected to the instruction load distribution component and the cache cross selection component respectively; The layer normalization component is used to receive weight data and bias data, and perform normalization processing in combination with the read data.

13. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: The two groups of memories are used as ping-pong buffers to interactively store calculation results and provide feature input for the next layer.

14. The reinforcement learning parallel processing accelerator according to claim 4, characterized in that: Also includes: A first selector is connected to the calculation unit, the cache cross selection unit and the first-in-first-out queue.

15. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: Also includes: The second selector is connected to the instruction load distribution component, the matrix multiplication calculation component, the cache cross selection component, and a parameter cache for storing weight data, bias data, and embedding data.

16. The reinforcement learning parallel processing accelerator according to claim 4, characterized in that: The controller is also used to read the data obtained after the FIFO queue conversion from the memory for data comparison.

17. The reinforcement learning parallel processing accelerator according to claim 1, characterized in that: Also includes: An instruction control component is connected to the controller and the instruction loading and distributing component, and is used to send a start command to the instruction loading and distributing component.

18. The reinforcement learning parallel processing accelerator according to claim 17, characterized in that: The instruction control component is also used to notify the host computer in an interrupt manner after all layers of calculation are completed; or, after all layers of calculation are completed, when receiving a query instruction from the host computer, the operation completion status is fed back to the host computer.

19. A method for accelerating a reinforcement learning parallel processing accelerator, characterized in that: The reinforcement learning parallel processing accelerator is the reinforcement learning parallel processing accelerator according to any one of claims 1 to 18, and the acceleration method comprises: The controller writes the current view features, at least two sets of batch history view features and the instruction sequence into corresponding memories respectively; The instruction loading and distributing component reads the instruction sequence in the memory, decodes the read instruction sequence, and distributes the obtained memory access parameters, calculation parameters and start calculation instructions to the calculation component; The data loading control component selects required feature data from the memory according to the parameters of each layer and loads it into the corresponding feature cache; After receiving the start calculation instruction, the calculation component simultaneously reads the current view feature data and at least two groups of batch historical view feature data in the feature cache, combines the corresponding memory access parameters and calculation parameters, performs parallel processing, obtains the intermediate layer result data and the final result data, and returns them to the controller; The cache cross selection component selects and schedules data sources between feature caches, caches for storing weight data, bias data and embedded data, and register groups; The calculation unit includes: a matrix multiplication calculation component connected to the instruction load distribution unit and the cache cross selection unit respectively; the matrix multiplication calculation component includes a first matrix multiplier component for performing matrix multiplication calculation for current view features, a second matrix multiplier component for performing matrix multiplication calculation for a first group of batch historical view features, and a third matrix multiplier component for performing matrix multiplication calculation for a second group of batch historical view features.

20. An electronic device, characterized in that: Comprising the reinforcement learning parallel processing accelerator as described in any one of claims 1 to 18.