A reinforcement learning accelerator, acceleration method and electronic device
By designing a reinforcement learning accelerator that includes a controller, instruction control component, parameter distribution component and calculation component, the shortcomings of traditional accelerators in resource optimization and acceleration effects are solved, and efficient computing support and resource savings for reinforcement learning models are achieved.
Patent Information
- Application Number
- CN202510245688.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Traditional accelerators have poor resource optimization and acceleration effects when supporting the computing of each layer of the network of reinforcement learning models.
A reinforcement learning accelerator is designed, including a controller, instruction control component, parameter distribution component and calculation component. Through clear data and instruction transmission paths, it ensures efficient transmission of data and instructions, and improves computing efficiency through precise parameter distribution and calculation processing.
It realizes efficient computing support for reinforcement learning models, improves processing efficiency, reduces resource utilization, and enhances the universality and flexibility of the accelerator.
Smart Images

Figure CN119721153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of accelerators, and in particular to a reinforcement learning accelerator, an acceleration method and an electronic device. Background Art
[0002] Reinforcement learning model is a network model structure used for specific tasks (such as autonomous driving). Traditionally, such models are calculated using graphics processing units (GPUs) or some traditional accelerators. However, in related technical solutions, the accelerators do not combine the characteristics of the network structure, and do not achieve good resource optimization and acceleration effects. Summary of the invention
[0003] The purpose of the present invention is to provide a reinforcement learning accelerator, acceleration method and electronic device, which can fully, efficiently and compactly support the calculation of each layer of the network, improve the processing efficiency of such models, and reduce the resource utilization rate of the accelerator.
[0004] In order to solve the above technical problems, the present invention provides a reinforcement learning accelerator, comprising:
[0005] A controller, used to write the current image features and the historical image features into the corresponding input buffers, and write the instruction sequence into the instruction control component;
[0006] The instruction control component is connected to the controller and is used to send the instruction sequence to the parameter distribution component;
[0007] The parameter distribution component is connected to the controller and the instruction control component respectively, and is used to parse the instruction sequence and distribute the obtained memory access parameters, computing parameters and startup instructions of each computing layer to the computing component;
[0008] The computing component is connected to the parameter distribution component, and is used to read the current image feature data and historical image feature data in the input cache, the corresponding memory access parameters and calculation parameters in sequence after receiving the start instruction, perform corresponding calculation processing, obtain result data and return it to the controller.
[0009] In order to solve the above technical problems, the present invention further provides an acceleration method of a reinforcement learning accelerator, wherein the reinforcement learning accelerator is the above reinforcement learning accelerator provided by the present invention, and the acceleration method comprises:
[0010] The controller writes the current image features and the historical image features into the corresponding input buffers respectively, and writes the instruction sequence into the instruction control component; the instruction control component is connected to the controller;
[0011] The instruction control component sends the instruction sequence to the parameter distribution component; the parameter distribution component is connected to the controller and the instruction control component respectively;
[0012] The parameter distribution component parses the instruction sequence and distributes the obtained memory access parameters, computing parameters and startup instructions of each computing layer to the computing component; the computing component is connected to the parameter distribution component;
[0013] After receiving the start instruction, the computing component sequentially reads the current image feature data and the historical image feature data in the input cache, the corresponding memory access parameters and the computing parameters, performs corresponding computing processing, obtains the result data and returns it to the controller.
[0014] It can be seen from the above technical solution that an electronic device provided by the present invention includes the above reinforcement learning accelerator provided by the present invention.
[0015] The beneficial effect of the present invention is that, in the above-mentioned reinforcement learning accelerator provided by the present invention, the controller writes the current image and historical image features into the corresponding input cache respectively, so that the feature data is stored in order, which is convenient for subsequent computing components to read. At the same time, the instruction sequence is written into the instruction control component. This clear data and instruction transmission path ensures the efficient transmission of data and instructions, reduces the confusion and delay of data transmission, and improves the overall data processing efficiency of the system. The instruction control component receives the instruction sequence and sends it to the parameter distribution component. After the parameter distribution component parses the instruction sequence, it accurately distributes the computing layer access parameters, calculation parameters and startup instructions to the computing component. This process ensures that the computing component can accurately obtain the required parameters and instructions, avoids calculation errors or resource waste caused by improper instruction processing, and improves the accuracy and reliability of instruction execution. After receiving the startup instruction, the computing component reads the feature data in the input cache in turn, and performs calculation processing on the access parameters and calculation parameters, so that the computing component can flexibly perform various computing tasks according to different parameters, which can adapt to the diverse computing needs in reinforcement learning, enhance the versatility and flexibility of the accelerator, and help improve the operating efficiency and performance of the reinforcement learning algorithm. The entire accelerator deeply analyzes the characteristics of the reinforcement learning model, and can fully, efficiently and compactly support the calculations of each layer of the network, improve the processing efficiency of such models, do not require the participation of the host computer, avoid interaction delays, and reduce the resource utilization of the accelerator.
[0016] In addition, the present invention also provides a corresponding acceleration method and electronic device for a reinforcement learning accelerator, which have the same or corresponding technical features as the reinforcement learning accelerator mentioned above, and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A schematic diagram of the structure of a reinforcement learning accelerator provided in an embodiment of the present invention;
[0019] Figure 2 A schematic diagram of the topological structure of a reinforcement learning accelerator provided in an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of the structure of a matrix multiplication calculation component provided in an embodiment of the present invention;
[0021] Figure 4 A schematic diagram of the structure of a flattened splicing assembly provided in an embodiment of the present invention;
[0022] Figure 5 A schematic diagram of data processing of a flattened splicing component provided by an embodiment of the present invention;
[0023] Figure 6 A schematic diagram of a calculation process of a reinforcement learning accelerator provided by an embodiment of the present invention;
[0024] Figure 7 A second schematic diagram of the calculation process of the reinforcement learning accelerator provided by an embodiment of the present invention;
[0025] Figure 8 A flow chart of an acceleration method of a reinforcement learning accelerator provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Figure 1 A schematic diagram of the structure of a reinforcement learning accelerator provided in an embodiment of the present invention is shown in FIG. Figure 1 As shown, the reinforcement learning accelerator includes:
[0028] A controller, used to write the current image features and the historical image features into the corresponding input buffers, and write the instruction sequence into the instruction control component;
[0029] An instruction control component, connected to the controller, for issuing instruction sequences to the parameter distribution component;
[0030] The parameter distribution component is connected to the controller and the instruction control component respectively, and is used to parse the instruction sequence and distribute the obtained memory access parameters, calculation parameters and startup instructions of each computing layer to the computing component;
[0031] The computing component is connected to the parameter distribution component and is used to read the current image feature data and historical image feature data in the input cache, the corresponding memory access parameters and calculation parameters, perform corresponding calculation processing, obtain the result data and return it to the controller after receiving the start instruction.
[0032] In the above reinforcement learning accelerator provided by the embodiment of the present invention, the controller writes the current image and historical image features into the corresponding input cache respectively, so that the feature data is stored in order, which is convenient for subsequent computing components to read. At the same time, the instruction sequence is written into the instruction control component. This clear data and instruction transmission path ensures the efficient transmission of data and instructions, reduces the confusion and delay of data transmission, and improves the overall data processing efficiency of the system. The instruction control component receives the instruction sequence and sends it to the parameter distribution component. After the parameter distribution component parses the instruction sequence, it accurately distributes the computing layer access parameters, calculation parameters and startup instructions to the computing component. This process ensures that the computing component can accurately obtain the required parameters and instructions, avoids calculation errors or resource waste caused by improper instruction processing, and improves the accuracy and reliability of instruction execution. After receiving the startup instruction, the computing component reads the feature data in the input cache in turn, and performs calculation processing on the access parameters and calculation parameters, so that the computing component can flexibly perform various computing tasks according to different parameters, which can adapt to the diverse computing needs in reinforcement learning, enhance the versatility and flexibility of the accelerator, and help improve the operating efficiency and performance of the reinforcement learning algorithm. The entire accelerator deeply analyzes the characteristics of the reinforcement learning model, and can fully, efficiently and compactly support the calculations of each layer of the network, improve the processing efficiency of such models, do not require the participation of the host computer, avoid interaction delays, and reduce the resource utilization of the accelerator.
[0033] It should be noted that the controller can use a peripheral component interconnect express (PCIE) controller, which is an interface for high-speed data transmission between the entire system and the outside. The present invention can send a data transmission instruction to the PCIE controller through the host computer. For example, after receiving the data transmission instruction from the host computer, the PCIE controller can preload the reinforcement learning network calculation parameter instruction to the parameter distribution component or preload other key parameters to the corresponding cache, etc. In this way, before the actual reinforcement learning calculation process begins, the relevant data and instructions are loaded into a specific component or cache in advance, so that they can be quickly obtained during subsequent use, reducing waiting time and improving system operation efficiency.
[0034] The controller can write the current image features and historical image features into the corresponding feature input buffers through direct memory access (DMA). Direct memory access allows external devices to transfer data directly to computer memory without the need for continuous processor involvement.
[0035] The above image can be a bird's-eye view (BEV). BEV is used in fields such as autonomous driving and computer vision. It can convert image data obtained from multiple perspectives (such as cameras at different angles around the vehicle) through algorithms to present environmental information in a perspective similar to looking down from the air, forming a top view data expression. This data representation method can more intuitively display information such as the position, shape, and relative relationship of objects around the vehicle, which is helpful for subsequent tasks such as target detection and path planning. Autonomous driving perception models can use BEV features as input.
[0036] Figure 2 A schematic diagram of the topological structure of the reinforcement learning accelerator provided in an embodiment of the present invention. Figure 2 As shown, the controller exchanges data with other components through the sending channel and the receiving channel, and the bit width of the sending channel and the receiving channel can be 512 bits (bit). The controller can specifically write the current image features (such as the current BEV) features and the historical image features (such as the historical BEV) features into the corresponding feature input cache through DMA transmission, and write the instruction sequence into the instruction control component through register control. The bit width of the register control channel can be 32 bits. The instruction sequence here may include debug instructions, reset instructions, startup instructions, etc., which are used to control the operating status of the entire system. The present invention can send debug instructions, reset instructions, startup instructions, etc. to the parameter distribution component to start executing the computing tasks of the entire network.
[0037] It should be noted that the calculation of the present invention is run in an offline mode. During the process, the parameter distribution component can parse the instructions of each computing layer in a set order, obtain the simulation parameters, calculation parameters and startup instructions of each computing layer, and transmit them to the corresponding computing components. After receiving the startup instruction, each computing component sequentially reads the current BEV feature data and historical BEV feature data in the input cache, and reads the corresponding simulation parameters and calculation parameters to perform corresponding calculation processing.
[0038] Furthermore, in the specific implementation, in the above-mentioned reinforcement learning accelerator provided by the embodiment of the present invention, the controller can also be used to write the state vector, action sequence and reward sequence into the corresponding register group, and write the weight data, bias data and embedded data into the corresponding cache in sequence. Correspondingly, the computing component can be used to perform corresponding computing processing according to the state vector, action sequence and reward sequence in the register group, the weight data, bias data and embedded data in the cache, combined with the read feature data, memory access parameters and computing parameters, to obtain the computing results of each layer; the computing results of each layer are transferred between each group of caches and register groups.
[0039] In the implementation, the state vector represents the state information, and the reward sequence stores the reward value obtained by the corresponding action sequence. The present invention writes the state vector, action sequence and reward sequence into the corresponding register group through the controller, so as to facilitate the subsequent rapid reading and processing of these key data. Weight data is the strength of the connection between neurons, which determines the degree of influence of input data on neuron output. Bias data is a constant term added in neuron calculation, which can adjust the activation function of neurons in different positions and ranges, increasing the flexibility and expression ability of the network. Embedding data can map high-dimensional discrete data (such as words, categories, etc. in text) to low-dimensional continuous vector space. The present invention writes weight data, bias data and embedding data to the corresponding cache in sequence, so that weight data, bias data and embedding data have their specific cache locations, so that parameters can be quickly located and read as needed, data access speed can be improved, data transmission and search time overhead can be reduced, and efficient operation of reinforcement learning network can be ensured.
[0040] In the present invention, each computing component will calculate each computing layer in turn, and specifically, it can perform corresponding calculation processing according to the state vector, action sequence and reward sequence in the register group, the weight data, bias data and embedded data in the cache, combined with the read feature data, memory access parameters and calculation parameters, to obtain the calculation results of each computing layer. Among them, the calculation results of each computing layer can be transferred between each group of caches and register groups, so that the storage and access methods of data can be optimized, and there is no need to repeatedly calculate the results that have been obtained, but directly read from the cache or register group, which improves the reusability of data, reduces the amount of calculation, and improves the calculation efficiency.
[0041] Furthermore, in a specific implementation, the above-mentioned reinforcement learning accelerator provided in the embodiment of the present invention may also include: a cache cross selection component, which is respectively connected to the controller and the computing component, and may include an input cache, a cache for respectively storing weight data, bias data and embedded data, and a register group for respectively storing state vectors, action sequences and reward sequences; a cache cross selection component, which may be used to select and schedule data sources among the input cache, cache, and register group.
[0042] In implementation, Figure 2 As shown, the cache cross selection component may include an input cache (such as current input and historical input), a cache 1 for storing weight data, a cache 2 for storing bias data, a cache 3 for storing embedded data, and a register group 1 for storing the input state vector, a register group 2, and register groups 3 and 4 for storing the input action sequence and reward sequence. It may also include a cache 4 for storing intermediate calculation results and final calculation results, and a register group 5 for storing state features calculated by the current image. The cache cross selection component can flexibly select appropriate data (the data can be a feature source or result data with a bit width of 2048 bits) from data sources such as input caches, caches of different purposes, and register groups according to the current needs of the computing component, and reasonably schedule the transmission of data according to the priority and time sequence of the computing task. For example, when performing a certain layer of neural network calculation, if weight data is currently needed, it can quickly select data from the cache storing weight data and provide it to the computing component, avoiding blind searches in multiple data storage areas, thereby reducing the delay of data access and improving the speed of data acquisition. When the computing unit needs to process the state vector and action sequence in sequence, the cache cross-selection unit can provide the data in the corresponding register group to the computing unit in the correct order, ensuring the continuity and efficiency of the computing process.
[0043] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, the computing component can be used to store the calculation result data of each layer into a result memory in a debugging mode, so that the result memory returns the calculation result data of each layer to the controller; or, in a non-debugging mode, store the final calculation result data into the result memory, so that the result memory returns the final calculation result data to the controller.
[0044] In implementation, the present invention adopts a two-level control structure to obtain calculation results. In debugging mode, the computing component stores the calculation result data of each layer into the result memory, which can obtain the detailed data of each layer in the model calculation process, provide rich information for debugging, and facilitate accurate positioning of problems. In non-debugging mode, only the final calculation result data is stored, which reduces unnecessary data storage and transmission. Compared with storing each layer of data in debugging mode, it reduces the space occupied by the result memory and the amount of data transmission, so that system resources can be more efficiently used for core computing tasks, improving the efficiency of the system during normal operation. The computing component can switch between debugging and non-debugging modes, which not only meets the needs of detailed analysis of the model calculation process in the development and debugging stage, but also ensures that the system outputs the final results efficiently and quickly in the actual operation stage, enhancing the versatility and applicability of the computing component.
[0045] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, the computing component can be used to extract the intermediate calculation result data of the corresponding layer returned during the calculation process according to the pre-set layer number when the debugging mode is turned on and store it in the result memory; when the debugging mode is not turned on, calculations are performed on all calculation layers. After the calculations of all layers are completed, the final calculation results are stored in the result memory, and the feature data of the next set of input caches, the corresponding memory access parameters and calculation parameters are loaded to perform the next set of calculations.
[0046] In implementation, in debugging mode, the computing component can extract the intermediate calculation result data of the corresponding layer according to the pre-set layer number, which greatly improves the debugging efficiency. When the debugging mode is not turned on, the computing component can only store the final calculation result in the result memory, avoiding the storage of a large amount of intermediate result data and reducing the occupation of storage resources. After the computing component completes the calculation of all layers, the calculation result can be converted to the bus width through the random access memory (RAM), and the result can be read back after the completion signal triggers an interrupt or the host computer uses a query method to judge, and load the next set of feature data, memory access parameters and calculation parameters in the input cache to perform the next set of calculations. This continuous processing method reduces the idle time in the calculation process, improves the utilization of computing resources, speeds up the overall calculation speed, and enables the system to process multiple sets of data more efficiently.
[0047] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided by an embodiment of the present invention, the computing component may include: a general matrix multiply (GEMM) computing component connected to a cache cross selection component; the matrix multiplication computing component includes a front-end matrix multiplication calculator, an accumulator and a back-end selection calculator; the front-end matrix multiplication calculator is used to output a multiplication result after multiplying each group of corresponding elements; the accumulator is used to accumulate the multiplication result when in an accumulation enable cycle; and to prepare for the next round of accumulation when in a reset cycle; the accumulator includes a first accumulator and a second accumulator; the first accumulator and the second accumulator perform accumulation in a ping-pong manner; and the back-end selection calculator is used to receive data processed by the accumulator and selectively calculate the data.
[0048] In implementation, Figure 2 As shown, the weight cache, bias cache, and embedding cache receive data from the controller (such as data with a bit width of 512 bits), and after cache processing, output weight data, bias data, and embedding data (such as data with a bit width of 2048 bits) to provide data support for the matrix multiplication calculation component. In addition, the cache cross-selection component can provide the matrix multiplication calculation component with attention key value (KV) matrix data, etc. The attention KV matrix data and weight data can be input into the matrix multiplication calculation component after being processed by the selector. Since the calculation time of the accumulator part is divided into an accumulation enable cycle and a reset cycle, but the data of the front-end block matrix multiplication is pipelined, two sets of accumulators are used for ping-pong accumulation to ensure continuous and uninterrupted data processing.
[0049] Figure 3 Schematic diagram of the structure of the matrix multiplication calculation component provided by the embodiment of the present invention. Figure 3 As shown in the figure, matrix multiplication is the core operation of many computing tasks. By processing element multiplication in parallel through the front-end matrix multiplication calculator, the operation speed of matrix multiplication can be significantly accelerated. The accumulator uses the first accumulator and the second accumulator to perform accumulation in a ping-pong manner. In the accumulation enable cycle, one accumulator performs accumulation operations, while the other accumulator can prepare for the next round of accumulation in the reset cycle. In this way, the accumulation operation can be seamlessly connected, avoiding the time overhead of waiting for reset in the traditional accumulation method, and greatly improving the accumulation efficiency. The front-end matrix multiplication calculator, accumulator and back-end selection calculator work together in sequence to form an efficient calculation pipeline. The matrix multiplication results can be accumulated in time, and the accumulated data can quickly enter the back-end selection calculation. The close cooperation between the various components reduces the delay of data processing and improves the operation efficiency of the entire computing component.
[0050] Furthermore, in a specific implementation, in the above reinforcement learning accelerator provided in the embodiment of the present invention, if Figure 3 As shown, the back-end selection calculator may include: one or a combination of a scaling unit, a mask unit, an attention mask unit, a bias unit, a residual unit, an embedding unit and an activation unit; the scaling unit is used to perform a scaling operation on the data under the control of a scaling enable signal; the mask unit is used to perform a mask calculation on the data under the control of a mask enable signal; the attention mask unit is used to perform an attention mask calculation on the data under the control of an attention mask enable signal; the bias unit is used to perform a bias calculation on the data under the control of a bias enable signal; the residual unit is used to perform a residual calculation on the data under the control of a residual enable signal; the embedding unit is used to perform an embedding calculation on the data under the control of an embedding enable signal; and the activation unit is used to perform an activation calculation on the data under the control of an activation enable signal.
[0051] In implementation, the scaling unit can enlarge or reduce the data as needed to help optimize the training effect of the model and improve the convergence speed and accuracy. The mask unit can selectively shield or retain certain parts of the data through mask calculation, process data more flexibly, and highlight key information. The attention mask unit can assign different attention weights to different data elements according to specific task requirements, so that the model can focus on key information and ignore irrelevant information, thereby improving the model's ability to handle complex tasks. The bias unit can add a fixed offset to the data by performing bias calculation on the data, which helps the model better fit the data and increase the flexibility and expression of the model. The residual unit allows the model to directly learn the residual relationship between input and output, improving the performance and accuracy of the model. The embedding unit is used to map data to a low-dimensional vector space, which can capture semantic or structural information between data. The activation unit can introduce nonlinear factors to the neural network, enabling the model to learn more complex nonlinear relationships. The present invention uses a scaling part, a mask part, an attention mask part, a bias part, a residual part, an embedding part and an activation part in sequence, and each part has a corresponding enable signal to control whether to perform the calculation step. If not, it can be skipped directly to avoid additional processing cycles.
[0052] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided by an embodiment of the present invention, the computing component may also include: a flattening and concatenation (Flatten_Concat) component; the flattening and concatenation component includes a first input port, a second input port, a channel control input part, a data selection controller, a multiplexer and an output port; the first input port is connected to the channel control input part, and is used to read the feature data of each channel grouping one by one through the channel control input part; the data selection controller is connected to the multiplexer, and is used to rearrange the feature data of each channel grouping, and transmit the rearranged feature data to the multiplexer; the second input port is connected to the multiplexer, and is used to read the state vector and transmit it to the multiplexer; the output port is used to output the feature data and state vector transmitted through the multiplexer, and store them in the result memory.
[0053] In implementation, Figure 4 This is a schematic diagram of the structure of the flattened splicing assembly provided by an embodiment of the present invention. Figure 4 As shown, the flattened splicing component can use two-port input. First, the first input port reads the BEV feature data of each channel group from the BEV feature cache in sequence through the channel control input unit, where each group of channel data can be output to the result cache after obtaining the rearranged data through the data selection controller. Figure 5 Schematic diagram of data processing of flattened splicing components provided by an embodiment of the present invention. Figure 5 As shown, each group of channel data is rearranged accordingly. The second input port reads the state features from the state feature buffer and outputs them to the result buffer in sequence. After the BEV feature exists, the two groups of features are concatenated.
[0054] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, the computing component may also include: a stack component, a query key value split (QKV_Split) component, a split concatenation (Split_Concat) component and a port cross selector; the stack component, the query key value split component, the split concatenation component and the flattened concatenation component are all connected to the port cross selector; the port cross selector is connected to the cache cross selection component.
[0055] In implementation, Figure 2As shown, the stacking component can stack multiple data units or processing results. The query key value splitting component can split the data in the form of query, key, and value. Split splicing component: allows flexible splitting and re-splicing of data. The stacking component, query key value splitting component, split splicing component, and flatten splicing component are all connected to the port cross selector, and the port cross selector is connected to the cache cross selection component. This connection architecture forms an efficient data transmission network. It can ensure that data is transmitted quickly and accurately between components, reduce delays and errors in the data transmission process, improve the efficiency of data interaction, and ensure the smooth operation of the computing components. According to the computing requirements, the required data is obtained from the cache in a timely manner, and the computing results are stored in the cache in a timely manner, so as to realize efficient caching and access of data, and further improve the efficiency and performance of the entire computing process.
[0056] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, the computing component may also include: a normalized exponent (Softmax) component respectively connected to the cache cross selection component and the first selector; the first selector is respectively connected to the parameter distribution component and the second selector; the second selector is respectively connected to the parameter distribution component, the matrix multiplication calculation component and the result memory; a normalized exponent component, used to perform a normalized exponential operation on the read data.
[0057] In implementation, the normalized index component can perform normalized index operations on the read data and convert the data into a probability distribution form, making the calculation results more stable in terms of numerical value, improving the accuracy and reliability of the calculation, and reducing the occurrence of calculation errors or model instability caused by numerical problems. The above connection structure forms a flexible data path, and data can be efficiently transmitted and interacted between different components through various selectors according to different calculation requirements, so that the normalized component can obtain the required data for calculation at the right time, and accurately pass the results to the subsequent calculation links, optimizing the overall calculation process and improving calculation efficiency.
[0058] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, the computing component may also include: a layer normalization (LN) component connected to the cache cross selection component and the first selector respectively; a layer normalization component, used to receive weight data and bias data, and perform normalization processing in combination with the read data.
[0059] In implementation, the layer normalization component can make the model more adaptable to different input data by normalizing the data of each layer, reduce the model's dependence on specific data distribution, and thus improve the generalization ability of the model. The weight data (LN_Weight) and bias data (LN_Bias) in the layer normalization component are two important parameters in the layer normalization operation, which are used to scale and translate the input data respectively, improve the adaptability and processing ability of the computing component to different data, and improve the reliability and stability of the model in various practical application scenarios.
[0060] In practical applications, the present invention analyzes the reinforcement learning network structure. Tables 1 and 2 are the calculation order of the network, where C is the number of blocks in each row of the block matrix, R is the number of rows in the block matrix, H is the number of attention heads, X is the feature matrix, W is the weight matrix, E is the residual matrix, b is the bias, D is the embedding, S is the scaling, M is the mask, A is the attention mask, and are different elements in the feature matrix. Conv means convolution, i in RAMi stands for input, RAM stands for memory, ReLU stands for rectified linear, Softplus stands for an activation function, Att stands for attention score, and RegGrp stands for register group.
[0061] Table 1 Feature cache location example 1
[0062]
[0063] Table 2 Feature cache location example 2
[0064]
[0065] Tables 1 and 2 illustrate the switching order of the feature cache storage location in each step of the network. According to the data cache requirements, the K matrix and the V matrix are stored in the cache at the same time. .
[0066] It should be pointed out that matrix multiplication in reinforcement learning networks is mainly divided into the following cases:
[0067] ; (1)
[0068] ; (2)
[0069] ; (3)
[0070] ; (4)
[0071] ; (5)
[0072] ; (6)
[0073] ; (7)
[0074] Among them, formula (1) corresponds to numbers 1-8, 10, 11, 25, 29, 30, 33, and 34 in Tables 1 and 2; formula (2) corresponds to numbers 12, 13, and 14 in Table 1; formula (3) corresponds to numbers 18 and 35 in Table 2; formula (4) corresponds to number 20 in Table 2; formula (5) corresponds to number 22 in Table 2; formula (6) corresponds to numbers 23 and 26 in Table 2; and formula (7) corresponds to numbers 31 and 32 in Table 2.
[0075] Bias, weights, and embeddings use independent caches. The feature cache includes the input cache of the historical BEV and the current BEV. Four groups of feature calculation flow caches and register groups (RegGrp) 1-register group 8, a total of eight groups of dedicated register groups, of which register groups 1 and 2 are the input state vectors, register groups 3 and 4 are the input action sequences and reward sequences, register group 5 is the state feature calculated by the current BEV, registers 6 and 7 are the distribution mean and distribution variance of the calculated policy network, and register group 8 is the value scalar of the calculated value network. In addition, in Table 1, when the split operation No. 19 is performed, since three groups of matrices need to be generated, in order to reduce the total number of cache usage and since the K matrix and the V matrix will not participate in the calculation at the same time, the K matrix and the V matrix can be placed on the same group of cache and stored in two areas. This avoids the space waste caused by allocating independent caches for the K matrix and the V matrix respectively, and significantly reduces the overall cache usage. In the case of limited memory resources, this optimization can free up more cache space for other data or operations, improve the memory usage efficiency of the entire system, reduce the overall memory access scheduling overhead of the system, and make the system structure more compact. Compared with allocating a large cache space to each matrix, this method can make better use of every byte of the cache, making more efficient use of the cache space, reducing the problem of idle space due to fixed allocation, and enabling subsequent data access to hit the cache faster, reducing the number of cache misses and improving computing performance.
[0076] Furthermore, in a specific implementation, the above-mentioned reinforcement learning accelerator provided in the embodiment of the present invention may also include: a memory access control component, which is respectively connected to the cache cross selection component, the weight cache, the bias cache and the embedded cache; the weight cache, the bias cache and the embedded cache are connected to the controller; the memory access control component is used to receive data in the weight cache, the bias cache and the embedded cache, and transmit the data in the weight cache, the bias cache and the embedded cache to the cache cross selection component as needed.
[0077] In implementation, the memory access control component is responsible for receiving data in the weight cache, bias cache and embedded cache, and transmitting these data to the cache cross-selection component as needed. This can reasonably arrange the data transmission order and time according to the needs of the computing component, reduce data waiting time, improve the timeliness and accuracy of data access, avoid unnecessary data transmission, and improve the flexibility of the system and resource utilization.
[0078] Furthermore, in a specific implementation, in the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, the controller can also be used to read the calculation result data of each layer and / or the final calculation result data from the result memory for data comparison.
[0079] In implementation, the controller can read the calculation result data of each layer and compare it with the final calculation result data to fully understand the output of the model at different stages and as a whole, so as to accurately judge the performance of the model and provide a basis for further optimization of the model.
[0080] It should be noted that the present invention adopts an offline scheduling method. After the host computer preloads the current image, historical images and weight data into the accelerator cache, it sends a start calculation command to the instruction control component. After the instruction control component sends a start instruction to the parameter distribution component, the parameter distribution component controls each calculation component to complete the corresponding calculation process in sequence. Each component completes data loading, calculation and result writing back according to the calculation parameters.
[0081] Figure 6 This is a schematic diagram of a calculation process of a reinforcement learning accelerator provided in an embodiment of the present invention. Figure 7 The second schematic diagram of the calculation process of the reinforcement learning accelerator provided by the embodiment of the present invention. The calculation process may include: Figure 6 As shown, convolution feature extraction is performed on the current image to obtain the current image features; Figure 6N in it represents the number of loops. According to the current image features, the current state vector is obtained. Then, the current state vector is processed by matrix multiplication to obtain the current state features. The current state vector and the current image features are flattened and spliced to obtain the current spliced features, which are subjected to matrix multiplication to obtain the current state sequence. Next, according to the current state sequence and the historical image features, the historical image features are extracted. According to the historical image features, the historical state vector is obtained. The historical state vector is processed by matrix multiplication to obtain the historical state features. The historical state features and the historical image features are flattened and spliced to obtain the historical spliced features, which are subjected to matrix multiplication to obtain the historical state sequence. Figure 7 As shown in the figure, the historical state sequence is embedded to obtain the state feature embedding, and then the action sequence is obtained. The action sequence is embedded to obtain the action feature embedding. According to the action feature embedding, the reward sequence is obtained. The reward sequence is embedded to obtain the reward feature embedding. Then, the state feature embedding, action feature embedding, and reward feature embedding are stacked. After the stacked data is layer normalized and one-dimensional convolution operation is performed, the query key value is split to obtain the corresponding query, key, and value, and the attention score is obtained under the attention mask. Then, the calculation result is obtained after the normalized exponential function and other operations. If all data blocks are calculated, the layer normalization process is continued, and the current state sequence is split and spliced to obtain the state feature, and then a series of calculation steps are performed, including multiple fully connected layers (Fully-Connected Layer, FC), rectified linear units (Rectified Linear Unit, ReLU), normalized exponential functions, activation functions (such as Softplus) and other calculation links, to obtain the distribution mean and distribution variance of the policy network, and finally obtain the value scalar of the value network to complete the calculation process.
[0082] After the calculation process is completed, the host computer can be notified through an interrupt or the host computer can query the completion status through the instruction control module. After reading the results, the next set of feature data is loaded and a start calculation command is issued, and this cycle is repeated until all feature data calculations are completed. The calculation process is controlled by the instruction parameter distribution module according to the parameters obtained by decoding the pre-loaded instructions, without the participation of the host computer, which improves processing efficiency and avoids interaction delays.
[0083] In the above embodiments, the reinforcement learning accelerator is described in detail. Based on the same inventive concept, the embodiments of the present invention also provide an acceleration method of the reinforcement learning accelerator and a corresponding embodiment of the electronic device.
[0084] Figure 8 Flow chart of the acceleration method of the reinforcement learning accelerator provided by the embodiment of the present invention. The acceleration method of the reinforcement learning accelerator provided by this embodiment is as follows: Figure 8 As shown, the specific steps include:
[0085] S801. The controller writes the current image features and the historical image features into the corresponding input buffers respectively, and writes the instruction sequence into the instruction control component; the instruction control component is connected to the controller.
[0086] S802, the instruction control component sends an instruction sequence to the parameter distribution component; the parameter distribution component is connected to the controller and the instruction control component respectively.
[0087] S803, the parameter distribution component parses the instruction sequence, and distributes the obtained memory access parameters, computing parameters and startup instructions of each computing layer to the computing component; the computing component is connected to the parameter distribution component.
[0088] S804. After receiving the start instruction, the computing component sequentially reads the current image feature data and the historical image feature data in the input cache, the corresponding memory access parameters and the computing parameters, performs corresponding computing processing, obtains the result data and returns it to the controller.
[0089] In the acceleration method of the above-mentioned reinforcement learning accelerator provided by the embodiment of the present invention, the controller writes the current image and historical image features into the corresponding input cache respectively, so that the feature data is stored in order, which is convenient for subsequent computing components to read. At the same time, the instruction sequence is written into the instruction control component. This clear data and instruction transmission path ensures the efficient transmission of data and instructions, reduces the confusion and delay of data transmission, and improves the overall data processing efficiency of the system. The instruction control component receives the instruction sequence and sends it to the parameter distribution component. After the parameter distribution component parses the instruction sequence, it accurately distributes the computing layer access parameters, calculation parameters and startup instructions to the computing component. This process ensures that the computing component can accurately obtain the required parameters and instructions, avoids calculation errors or resource waste caused by improper instruction processing, and improves the accuracy and reliability of instruction execution. After receiving the startup instruction, the computing component reads the feature data in the input cache in turn, and performs calculation processing on the access parameters and calculation parameters, so that the computing component can flexibly perform various computing tasks according to different parameters, which can adapt to the diverse computing needs in reinforcement learning, enhance the versatility and flexibility of the accelerator, and help improve the operation efficiency and performance of the reinforcement learning algorithm. The entire process deeply analyzes the characteristics of the reinforcement learning model, and can fully, efficiently, and compactly support the calculations of each layer of the network, improve the processing efficiency of such models, do not require the participation of the host computer, avoid interaction delays, and reduce the resource utilization of the accelerator.
[0090] Since the embodiments of the acceleration method part correspond to the embodiments of the reinforcement learning accelerator part, the embodiments of the acceleration method part refer to the description of the embodiments of the reinforcement learning accelerator part, which will not be repeated here. And it has the same beneficial effects as the reinforcement learning accelerator mentioned above.
[0091] Furthermore, in the specific implementation, in the acceleration method of the above reinforcement learning accelerator provided by the embodiment of the present invention, while executing step S801, it can also include: the controller writes the state vector, action sequence and reward sequence into the corresponding register group, and writes the weight data, bias data and embedded data into the corresponding cache in sequence. Accordingly, when executing step S804, the calculation component performs corresponding calculation processing according to the state vector, action sequence and reward sequence in the register group, the weight data, bias data and embedded data in the cache, combined with the read feature data, memory access parameters and calculation parameters, to obtain the calculation results of each layer; the calculation results of each layer are transferred between each group of caches and register groups.
[0092] Furthermore, in a specific implementation, the acceleration method of the reinforcement learning accelerator provided in the embodiment of the present invention may further include: a cache cross selection component selects and schedules data sources among the input cache, the cache, and the register group. The cache cross selection component is connected to the controller and the computing component respectively.
[0093] Furthermore, in a specific implementation, in the acceleration method of the above-mentioned reinforcement learning accelerator provided in an embodiment of the present invention, when executing step S804, the computing component stores the calculation result data of each layer into the result memory in the debugging mode, so that the result memory returns the calculation result data of each layer to the controller; or, stores the final calculation result data into the result memory in the non-debugging mode, so that the result memory returns the final calculation result data to the controller.
[0094] For more specific working processes of the above steps, please refer to the corresponding contents disclosed in the above embodiments, which will not be repeated here.
[0095] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, including the above-mentioned reinforcement learning accelerator. Since the principle of solving the problem by the electronic device is similar to that of the above-mentioned reinforcement learning accelerator, the implementation of the electronic device can refer to the implementation of the reinforcement learning accelerator, and the repeated parts will not be repeated.
[0096] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0097] Finally, it should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by technicians in the technical field of the present invention; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the terms "include", "comprise" and "have" and any other variations thereof in the present invention are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. In the present invention, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0098] For the above-mentioned embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the order of the actions described, because according to the present invention, some steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0099] The above is a detailed introduction to the reinforcement learning accelerator, acceleration method and electronic device provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can refer to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can refer to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A reinforcement learning accelerator, characterized in that: include: A controller, used to write the current image features and the historical image features into the corresponding input buffers, and write the instruction sequence into the instruction control component; It is also used to write the state vector, action sequence and reward sequence into the corresponding register group, and write the weight data, bias data and embedding data into the corresponding cache in sequence; The instruction control component is connected to the controller and is used to send the instruction sequence to the parameter distribution component; The parameter distribution component is connected to the controller and the instruction control component respectively, and is used to parse the instruction sequence and distribute the obtained memory access parameters, computing parameters and startup instructions of each computing layer to the computing component; The computing component is connected to the parameter distribution component, and is used to read the current image feature data and the historical image feature data in the input cache, the corresponding memory access parameters and the computing parameters in sequence after receiving the start instruction, and perform corresponding computing processing according to the state vector, action sequence and reward sequence in the register group, the weight data, bias data and embedded data in the cache, combined with the read feature data, memory access parameters and computing parameters, to obtain the computing results of each layer; the computing results of each layer are transferred between each group of caches and register groups.
2. The reinforcement learning accelerator according to claim 1, characterized in that: Also includes: A cache cross-selection component, connected to the controller and the computing component respectively, including the input cache, a cache for storing weight data, bias data and embedded data respectively, and a register group for storing state vectors, action sequences and reward sequences respectively; The cache cross selection component is used to select and schedule data sources among the input cache, cache, and register group.
3. The reinforcement learning accelerator according to claim 1, characterized in that: The calculation component is used to store the calculation result data of each layer into the result memory in the debugging mode, so that the result memory returns the calculation result data of each layer to the controller; Or, in a non-debugging mode, the final calculation result data is stored in a result memory, so that the result memory returns the final calculation result data to the controller.
4. The reinforcement learning accelerator according to claim 3, characterized in that: The calculation component is used to extract the corresponding layer intermediate calculation result data returned during the calculation process according to the preset layer number when the debugging mode is turned on, and store it in the result storage; When the debugging mode is not turned on, calculations are performed on all calculation layers. After calculations on all layers are completed, the final calculation results are stored in the result memory, and the next set of feature data in the input cache, the corresponding memory access parameters and calculation parameters are loaded to perform the next set of calculations.
5. The reinforcement learning accelerator according to claim 2, characterized in that: The computing unit includes: a matrix multiplication computing component connected to the cache cross selection unit; The matrix multiplication calculation component includes a front-end matrix multiplication calculator, an accumulator and a back-end selection calculator; The front-end matrix multiplication calculator is used to output a multiplication result after multiplying each group of corresponding elements; The accumulator is used to accumulate the multiplication result when in an accumulation enable cycle; and to prepare for the next round of accumulation when in a reset cycle; the accumulator includes a first accumulator and a second accumulator; the first accumulator and the second accumulator perform accumulation in a ping-pong manner; The back-end selection calculator is used to receive the data processed by the accumulator and selectively calculate the data.
6. The reinforcement learning accelerator according to claim 5, characterized in that: The back-end selection calculator includes: one or a combination of a scaling part, a mask part, an attention mask part, a bias part, a residual part, an embedding part and an activation part; The scaling unit is used to perform a scaling operation on the data under the control of a scaling enable signal; The mask part is used to perform mask calculation on the data under the control of the mask enable signal; The attention mask unit is used to perform attention mask calculation on the data under the control of the attention mask enable signal; The biasing unit is used to perform bias calculation on the data under the control of the bias enabling signal; The residual part is used to perform residual calculation on the data under the control of the residual enable signal; The embedding unit is used to perform embedding calculation on the data under the control of the embedding enable signal; The activation unit is used to perform activation calculation on the data under the control of the activation enable signal.
7. The reinforcement learning accelerator according to claim 5, characterized in that: The computing component also includes: a flattening and splicing component; The flattening and splicing assembly includes a first input port, a second input port, a channel control input portion, a data selection controller, a multiplexer and an output port; The first input port is connected to the channel control input unit and is used to read the characteristic data of each channel group one by one through the channel control input unit; The data selection controller is connected to the multiplexer and is used to rearrange the characteristic data of each channel group and transmit the rearranged characteristic data to the multiplexer; The second input port is connected to the multiplexer and is used to read the state vector and transmit it to the multiplexer; The output port is used to output the feature data and state vector transmitted through the multiplexer and store them in the result memory.
8. The reinforcement learning accelerator according to claim 7, characterized in that: The computing component also includes: a stacking component, a query key value splitting component, a splitting and splicing component and a port cross selector; The stacking component, the query key value splitting component, the splitting and splicing component and the flattening and splicing component are all connected to a port cross selector; The port cross selector is connected to the cache cross selection component.
9. The reinforcement learning accelerator according to claim 5, characterized in that: The calculation component further includes: a normalized index component connected to the cache cross selection component and the first selector respectively; the first selector is connected to the parameter distribution component and the second selector respectively; the second selector is connected to the parameter distribution component, the matrix multiplication calculation component and the result memory respectively; The normalized index component is used to perform a normalized index operation on the read data.
10. The reinforcement learning accelerator according to claim 9, characterized in that: The computing component further includes: a layer normalization component connected to the cache cross selection component and the first selector respectively; The layer normalization component is used to receive weight data and bias data, and perform normalization processing in combination with the read data.
11. The reinforcement learning accelerator according to claim 2, characterized in that: Also includes: A memory access control component, connected to the cache cross selection component, the weight cache, the bias cache and the embedding cache respectively; The weight cache, the bias cache and the embedding cache are connected to the controller; The memory access control component is used to receive the data in the weight cache, the bias cache and the embedded cache, and transmit the data in the weight cache, the bias cache and the embedded cache to the cache cross selection component as needed.
12. The reinforcement learning accelerator according to claim 3, characterized in that: The controller is also used to read the calculation result data of each layer and / or the final calculation result data from the result storage for data comparison.
13. A method for accelerating a reinforcement learning accelerator, characterized in that: The reinforcement learning accelerator is the reinforcement learning accelerator according to any one of claims 1 to 12, and the acceleration method comprises: The controller writes the current image features and the historical image features into the corresponding input buffers respectively, writes the instruction sequence into the instruction control component; writes the state vector, the action sequence and the reward sequence into the corresponding register group, and writes the weight data, the bias data and the embedded data into the corresponding buffers in sequence; the instruction control component is connected to the controller; The instruction control component sends the instruction sequence to the parameter distribution component; the parameter distribution component is connected to the controller and the instruction control component respectively; The parameter distribution component parses the instruction sequence and distributes the obtained memory access parameters, computing parameters and startup instructions of each computing layer to the computing component; the computing component is connected to the parameter distribution component; After receiving the start-up instruction, the computing component sequentially reads the current image feature data and the historical image feature data in the input cache, the corresponding memory access parameters and the computing parameters, and performs corresponding computing processing according to the state vector, action sequence and reward sequence in the register group, the weight data, bias data and embedded data in the cache, in combination with the read feature data, memory access parameters and computing parameters, to obtain the computing results of each layer; the computing results of each layer are transferred between each group of caches and register groups.
14. An electronic device, characterized in that: Comprising a reinforcement learning accelerator as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Stream processor for extracting features of stream data in transmission process and implementation method thereof
CN113138804A
Instruction issuing and man-machine interaction method based on deep learning and sight tracking
CN118192805A