Matrix accelerator sharing chip system memory and operation method thereof
By sharing the matrix accelerator of the chip system memory, the problems of system bus load and memory waste during large matrix operations are solved, and efficient matrix operations are achieved, which is suitable for small-volume and low-power chips.
Patent Information
- Application Number
- CN202511316431.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing matrix accelerators frequently occupy system bus bandwidth when performing large matrix operations, resulting in system performance degradation and serious waste of memory resources, making them unsuitable for small-volume, low-power chips.
A matrix accelerator that uses shared chip system memory achieves memory sharing by converting part or all of the single-port SRAM into dual-port SRAM, and uses task queues and data access modules to optimize the interaction of computing tasks and reduce the number of interactions between the CPU and the matrix accelerator.
It improves CPU efficiency, reduces memory resource waste, and enhances chip system performance, making it suitable for small-sized and low-power navigation baseband chips.
Smart Images

Figure CN120832333A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a matrix accelerator sharing chip system memory and a running method thereof. BACKGROUND
[0002] In the navigation positioning solution algorithm, a large number of vector and matrix related operations are used, such as the commonly used single point weighted least square algorithm. The basic operation units of the algorithm are matrix multiplication, matrix inversion, matrix addition / subtraction, etc. These basic operation units can complete a weighted least square operation according to a specific combination order. Similar operations can also be extended to other algorithms such as Kalman filtering. These algorithms are stable and mature, have many operation steps, and the operation time increases significantly with the increase of the matrix dimension. Functionally, they can be completely separated from software, and hardware acceleration of matrix operation can be realized by using RTL (Register Transfer Level), so as to ultimately achieve the purpose of reducing CPU (Central Processing Unit) load and improving system performance.
[0003] The current general matrix accelerator can independently realize matrix operation, has no other interaction with the outside except the bus and the interrupt, and is convenient for verification and transplantation. However, data prefetching and result returning are required for each operation (such as matrix multiplication and matrix addition). With the increase of the matrix dimension, the data volume of matrix operation will increase by orders of magnitude. Moreover, the existing matrix accelerator uses the DMA (Direct Memory Access) mode to move data. With the increase of the matrix dimension, the DMA is prone to occupy the chip system bus and the system memory bandwidth for a long time, causing large system bus load and overall system performance decline. In addition, a large number of intermediate variables and results need to be stored in the matrix operation. Therefore, the capacity of the internal memory of the matrix accelerator needs to be designed according to the maximum possibility of the matrix dimension, which inevitably increases the chip area and power consumption. Moreover, when the matrix accelerator is not working, the internal memory of the module is idle. SUMMARY
[0004] Therefore, it is necessary to provide a matrix accelerator sharing chip system memory and a running method thereof in view of the above technical problems.
[0005] A matrix accelerator sharing chip system memory, comprising: a system memory module, a task management module, a data access module and a matrix operation module. The system memory module is configured to convert part or all of the single-port SRAM (Static Random Access Memory) in the chip system memory into double-port SRAM according to the maximum memory requirement of the matrix accelerator, so that the matrix accelerator shares the chip system memory through the double-port SRAM. The task management module is configured to coordinate the CPU to control the matrix accelerator to perform operation and interact with data, load operation tasks and task control instructions generated by the CPU based on the operation tasks, and send the operation tasks and the corresponding task control instructions to the data access module when the task queue is not empty, and generate an interrupt signal to inform the CPU to complete the operation when all the operation tasks are completed. The data access module is configured to send the matrix operation data read from the dual-port SRAM and the operation operation instructions synchronously generated to the matrix operation module according to the operation tasks and the corresponding task control instructions sent, and write the operation result data sent by the matrix operation module into the dual-port SRAM, and pull up the task completion signal to indicate that the operation task sent currently is completed when the operation result data is written. The matrix operation module is configured to select corresponding operation sub-modules for the input matrix operation data to perform matrix operation, and send the operation result data to the data access module when the operation is completed.
[0006] Further, the system memory module further comprises a protocol conversion submodule, one end of the dual-port SRAM is connected with the protocol conversion submodule through a MEM bus, the protocol conversion submodule is configured to interconnect with the system bus and the CPU through an AXI (Advanced eXtensible Interface) protocol, and ensure seamless connection between the system memory module and the CPU; the other end of the dual-port SRAM is connected with the data access module through a MEM (Memory) bus, so as to realize direct access of the matrix accelerator to the dual-port SRAM.
[0007] Further, the task management module further comprises a user interface submodule, the user interface submodule is composed of an operation start and stop control register, an interrupt enable and clear control register, a status register and a task queue interface, the user interface submodule is connected with the system bus through an AHB (Advanced High-performance Bus), and is interconnected with the CPU through an AXI protocol, and is configured to complete operation start and stop control, interrupt enable and clear control and task queue initialization of the matrix accelerator under the control of the CPU.
[0008] Further, the working process of the task queue is as follows: When the task queue receives a start signal sent by the CPU through the user interface submodule, the operation tasks and the corresponding task control instructions initialized by the CPU in the task queue interface are loaded into the task queue in an address-increment order; wherein, the task queue is managed in a first-in first-out manner; When the task queue is not empty, the operation task and the corresponding task control instruction located at the first position in the task queue are sent to the data access module; When the task queue detects that the data access module pulls up the task completion signal, it is determined that the current issued operation task is completed, the current issued operation task is deleted in the task queue, and the corresponding task number is uploaded to the state register of the user interface submodule for the CPU to query the task state information. Meanwhile, the task queue determines whether the user interface submodule generates an interrupt signal according to whether the interrupt reporting enable in the current corresponding task control instruction is high. If the task queue is empty, the matrix operation is stopped, and the interrupt signal generated by the user interface submodule notifies the CPU that the operation is completed. If the task queue is not empty, the operation task and the corresponding task control instruction are issued according to the task queue sorting to continue the matrix operation until the task queue is empty.
[0009] Further, the task control instruction is composed of a task number, an interrupt reporting enable, an operation type selection, a storage address of an operand matrix X, a storage address of an operand matrix Y, a storage address of an operation result, a row number of the operand matrix X, a column number of the operand matrix X and a row number of the operand matrix Y, a column number of the operand matrix Y, and an instruction validity identifier.
[0010] Further, the data access module includes an SRAM read / write control submodule, a matrix operation control submodule, a matrix multiplication / addition operation control submodule, and a matrix inversion operation control submodule. The matrix operation control submodule is configured to receive the operation task and the corresponding task control instruction issued by the task queue for parameter analysis, and send the analyzed operation parameters to the SRAM read / write control submodule, the matrix multiplication / addition operation control submodule, and the matrix inversion operation control submodule. The matrix operation control submodule is further configured to indicate that the current issued operation task is completed by pulling up the task completion signal after receiving the operation completion signal generated by the SRAM read / write control submodule. The operation parameters include an operation type, a matrix data storage starting address, and a matrix dimension parameter. The matrix multiplication / addition operation control submodule is configured to calculate, based on the operation parameters, an access starting address of matrix multiplication / addition operation data, a read / write cycle number, and an increment / decrement amplitude of an address in each read / write cycle, generate an operation operation instruction corresponding to the matrix multiplication / addition operation, and send the operation operation instruction to the matrix operation module. The matrix inversion operation control submodule is configured to generate an LDL decomposition inversion algorithm control flow based on the operation parameters, calculate, in the control flow, an access starting address of matrix inversion operation data, a read / write cycle number, and an increment / decrement amplitude of an address in each read / write cycle, generate an operation operation instruction corresponding to the matrix inversion operation, and send the operation operation instruction to the matrix operation module. The SRAM read-write control submodule is configured to generate SRAM read-write control signals based on the operation parameters and the access start address, the read / write cycle number, and the increment / decrement amplitude of each read / write cycle address of the matrix operation data provided by the matrix multiplication / addition operation control submodule and the matrix inversion operation control submodule, and to perform a read operation to read the matrix operation data from the dual-port SRAM and send the matrix operation data to the matrix operation module according to the SRAM read-write control signals, perform a write operation to write the operation result data sent by the matrix operation module into the dual-port SRAM after the operation is completed, and generate an operation completion signal and report the operation completion signal to the matrix operation control submodule after the last operation result data is written.
[0011] Further, the matrix operation module includes an operation input-output control submodule, a floating-point multiplication / addition operation unit, and a floating-point division operation unit. The operation input-output control submodule is configured to receive the operation operation instruction and the matrix operation data sent by the data access module, select a corresponding operation unit to perform a matrix operation according to the operation operation instruction, and send the matrix operation data to the corresponding operation unit, and obtain the operation result data output by the operation unit and send the operation result data to the data access module after the operation is completed. The floating-point multiplication / addition operation unit is configured to receive the matrix operation data sent by the operation input-output control submodule, complete multiplication, addition, and multiplication-accumulation operations according to the corresponding operation operation instruction, and return the operation result data to the operation input-output control submodule. The floating-point division operation unit is configured to receive the matrix operation data sent by the operation input-output control submodule, complete division and inversion operations according to the corresponding operation operation instruction, and return the operation result data to the operation input-output control submodule.
[0012] A running method of a matrix accelerator sharing a chip system memory, the method comprising the following steps: The CPU is configured to prepare matrix operation data in a dual-port SRAM of the chip system memory, initialize an operation task and a corresponding task control instruction, and start the matrix accelerator, wherein the matrix accelerator includes a system memory module, a task management module, a data access module, and a matrix operation module. After the CPU starts the matrix accelerator, the operation task and the task control instruction generated by the CPU based on the operation task are loaded by the task queue in the task management module, and when the task queue is not empty, the operation task and the corresponding task control instruction are sent to the data access module, the matrix operation data read from the dual-port SRAM and the operation operation instruction synchronously generated are sent to the matrix operation module by the data access module according to the operation task and the corresponding task control instruction sent, the matrix operation module selects the corresponding operation sub-module for the input matrix operation data according to the operation operation instruction to perform the matrix operation, and after the operation is completed, the operation result data is written into the dual-port SRAM by the data access module, and after the operation result data is written, the task completion signal is pulled high by the data access module to indicate that the current sent operation task is completed. The task queue is detected whether it is empty, if not, the operation task and the corresponding task control instruction are continuously sent by the task queue for matrix operation until the task queue is empty, if empty, it is determined that the matrix accelerator completes all operation tasks, an interrupt signal is generated by the task management module to inform the CPU to complete the operation, and the final matrix operation result is obtained from the dual-port SRAM by the CPU.
[0013] The above-mentioned matrix accelerator sharing the memory of the chip system and the operation method thereof, by converting part or all of the single-port SRAM in the memory of the chip system into dual-port SRAM according to the maximum memory requirement of the matrix accelerator, so that the matrix accelerator directly shares the memory of the chip system through the dual-port SRAM, without the need to design the capacity of the internal memory of the matrix accelerator according to the maximum possibility of the matrix dimension, so that the matrix accelerator can be applied to small-size and low-power navigation baseband chips, and at the same time, when the matrix accelerator is not working, the shared memory can still be used as the memory of the chip system, avoiding the idle waste of memory resources. In addition, through the task queue, multiple operation tasks and corresponding task control instructions can be loaded into the matrix accelerator at one time, and the matrix accelerator can be controlled to execute the sent operation tasks and corresponding task control instructions in sequence according to the programmed order to perform the matrix operation, and after all the operation tasks are completed, an interrupt signal is generated to complete the interaction between the CPU and the matrix accelerator. This batch processing method significantly reduces the interaction times between the CPU and the matrix accelerator, does not occupy the system bus bandwidth for a long time, maximizes the efficiency of the CPU, and significantly improves the performance of the chip system. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 It is a structural schematic diagram of a matrix accelerator sharing the memory of the chip system in one embodiment; Figure 2 It is a structural schematic diagram of a task management module in one embodiment; Figure 3Figure 1 is a schematic diagram of a task queue and a task control instruction coding format in one embodiment; Figure 3 (a) a task queue, Figure 3 (b) a task control instruction coding format; Figure 4 Figure 2 is a schematic diagram of a structure of a data access module in one embodiment; Figure 5 Figure 3 is a schematic diagram of a structure of a matrix operation module in one embodiment; Figure 6 Figure 4 is a flowchart of a running method of a matrix accelerator sharing a chip system memory in one embodiment; Figure 7 Figure 5 is a schematic diagram of a P0 port and a P1 port address interconnection relationship of SRAM3 in one embodiment. DETAILED DESCRIPTION
[0015] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0016] In one embodiment, as shown in Figure 1, a matrix accelerator sharing a chip system memory is provided, which includes a system memory module, a task management module, a data access module and a matrix operation module. Figure 1
[0017] The system memory module is configured to convert part or all of single-port SRAM in the chip system memory into double-port SRAM according to the maximum memory requirement of the matrix accelerator, so that the matrix accelerator shares the chip system memory through the double-port SRAM.
[0018] The task management module is configured to coordinate the operation control and data interaction of the CPU on the matrix accelerator, load operation tasks and task control instructions generated by the CPU based on the operation tasks through the task queue, and when the task queue is not empty, issue the operation tasks and the corresponding task control instructions to the data access module, and generate an interrupt signal to notify the CPU of the completion of the operation when all the operation tasks are completed.
[0019] The data access module is configured to send the matrix operation data read from the double-port SRAM and the operation operation instructions synchronously generated to the matrix operation module according to the issued operation tasks and the corresponding task control instructions, write the operation result data sent by the matrix operation module into the double-port SRAM, and pull up the task completion signal to indicate the completion of the current issued operation task after the writing of the operation result data is completed.
[0020] The matrix operation module is configured to select corresponding operation sub-modules for input matrix operation data according to operation operation instructions, perform matrix operation, and send operation result data to the data access module after the operation is completed.
[0021] Further, the system memory module further comprises a protocol conversion sub-module, one end of the dual-port SRAM is connected with the protocol conversion sub-module through the MEM bus, the protocol conversion sub-module is configured to interconnect with the system bus and the CPU through the AXI protocol, and seamless connection between the system memory module and the CPU is ensured; the other end of the dual-port SRAM is connected with the data access module through the MEM bus, so as to realize direct access of the matrix accelerator to the dual-port SRAM. Figure 1 SRAM3 in the matrix accelerator is a dual-port SRAM, SRAM3 has two independent read-write ports P0 and P1, the port P0 is connected with the protocol conversion sub-module, and the function of the chip system memory is realized; the port P1 is configured to connect the data access module in the matrix accelerator, and the matrix accelerator can directly access the SRAM3 memory through the P1 port.
[0022] Further, as shown in Figure 2 the task management module further comprises a user interface sub-module, the user interface sub-module comprises operation start and stop control registers, interrupt enable and clear control registers, a state register, and a task queue interface (the task queue interface is a group of control registers), and the like, the user interface sub-module is connected with the system bus through the AHB bus, and is interconnected with the CPU through the AXI protocol, and is configured to complete operation start and stop control of the matrix accelerator, interrupt enable and clear control, and task queue initialization and the like under the control of the CPU.
[0023] Further, compared with the general matrix accelerator, the task management module of the matrix accelerator of the present application adds a task queue as shown in Figure 3 (a), and the working process of the task queue is as follows: (1) when the task queue receives a start signal sent by the CPU through the user interface sub-module, the operation task and the corresponding task control instruction initialized by the CPU in the task queue interface are loaded into the task queue in the order of address increment; wherein, the task queue is managed in the form of first-in first-out (FIFO). For example, the FIFO depth is 32, that is, at most 32 matrix operation tasks can be temporarily stored.
[0024] (2) when the task queue is not empty, the operation task and the corresponding task control instruction located at the first position in the task queue are sent to the data access module.
[0025] (3) When the task queue detects that the data access module pulls up the task completion signal, it is determined that the current issued operation task is completed, the current issued operation task is deleted in the task queue and the corresponding task number is uploaded to the state register of the user interface submodule for the CPU to query the task state information. Meanwhile, the task queue determines whether the user interface submodule generates an interrupt signal according to whether the interrupt reporting enable in the current corresponding task control instruction is high.
[0026] (4) It is detected whether the task queue is empty. If it is empty, the matrix operation is stopped, an interrupt signal is generated by the user interface submodule to inform the CPU that the operation is completed. If it is not empty, steps (1) to (3) are repeated, that is, the operation task and the corresponding task control instruction are issued according to the task queue sorting to continue the matrix operation until the task queue is empty.
[0027] It should be understood that through the task queue, multiple operation tasks and corresponding task control instructions can be sent into the matrix accelerator at one time, the matrix accelerator will execute the matrix operation operation in the programmed order in turn, and an interrupt signal is generated after all operation tasks are completed to complete the interaction between the CPU and the matrix accelerator. This batch processing mode significantly reduces the interaction times between the CPU and the matrix accelerator, and maximizes the CPU efficiency.
[0028] Further, the encoding format of the task control instruction is as shown in Figure 3 (b), which is composed of a task number, an interrupt reporting enable, an operation type selection, a storage address of an operand matrix X, a storage address of an operand matrix Y, a storage address of an operation result, a number of rows of the operand matrix X, a number of columns of the operand matrix X and a number of rows of the operand matrix Y, a number of columns of the operand matrix Y, and an instruction validity identifier.
[0029] Further, the function of the data access module is to calculate the address parameters of reading and writing the SRAM in the operation flow according to the issued task control instruction, and to complete the reading and writing operation of the SRAM. As shown in Figure 4 the data access module includes an SRAM read-write control submodule, a matrix operation control submodule, a matrix multiplication / addition operation control submodule, and a matrix inversion operation control submodule.
[0030] The matrix operation control submodule receives operation tasks and corresponding task control instructions issued by the task queue, performs parameter parsing, and sends the parsed operation parameters to the SRAM read / write control submodule, the matrix multiplication / addition operation control submodule, and the matrix inversion operation control submodule. Upon receiving the operation completion signal generated by the SRAM read / write control submodule, the matrix operation control submodule pulls the task completion signal high to indicate the completion of the currently issued operation task. The task management module then updates the task queue, status register, and generates an interrupt signal accordingly. Operation parameters include the operation type, the starting address for storing matrix data, and the matrix dimension parameters.
[0031] The matrix multiplication / addition operation control submodule is used to calculate the access starting address of the matrix multiplication / addition operation data, the number of read / write cycles, and the increment / decrement amplitude of the address in each read / write cycle based on the operation parameters. At the same time, it generates the operation operation instructions corresponding to the matrix multiplication / addition operation and sends them to the matrix operation module to indicate which operation will be performed on the currently output operation data.
[0032] The matrix inversion operation control submodule is used to generate the LDL decomposition inversion algorithm control process based on the operation parameters. In the control process, the access starting address of the matrix inversion operation data, the number of read / write cycles, and the increment / decrement amplitude of the address in each read / write cycle are calculated. At the same time, the operation operation instructions corresponding to the matrix inversion operation are generated and sent to the matrix operation module, indicating whether the current output operation data will be subjected to multiplication and accumulation or differentiation operation.
[0033] The SRAM read-write control submodule is used to generate an SRAM read-write control signal based on the operation parameters and the access starting address of the matrix operation data provided by the matrix multiplication / addition operation control submodule and the matrix inversion operation control submodule, the number of read / write cycles, and the increment / decrement amplitude of the address of each read / write cycle. According to the SRAM read-write control signal, a read operation is first performed to read the matrix operation data from the dual-port SRAM and send it to the matrix operation module. After the operation is completed, a write operation is performed to write the operation result data sent by the matrix operation module into the dual-port SRAM. After the last operation result data is written, an operation completion signal is generated and reported to the matrix operation control submodule.
[0034] Furthermore, the function of the matrix operation module is to select the corresponding operator for the input matrix operation data according to the operation instruction and complete the specific calculation, such as Figure 5 As shown, the matrix operation module includes an operation input and output control submodule, a floating-point multiplication / addition operator, and a floating-point division operator.
[0035] The operation input and output control submodule is used to receive the operation instructions and matrix operation data sent by the data access module, select the corresponding operator to perform matrix operation according to the operation instruction, and send the matrix operation data to the corresponding operator. After the operation is completed, the operation result data output by the operator is obtained and sent to the data access module.
[0036] The floating-point multiplication / addition operator is used to receive matrix operation data sent by the operation input and output control submodule, complete multiplication, addition, and multiplication-accumulation operations according to the corresponding operation operation instructions, and return the operation result data to the operation input and output control submodule.
[0037] The floating-point division operator is used to receive the matrix operation data sent by the operation input and output control submodule, complete the division and inversion operations according to the corresponding operation instructions, and return the operation result data to the operation input and output control submodule.
[0038] In one embodiment, Figure 6 As shown, a method for operating a matrix accelerator sharing chip system memory is provided, comprising the following steps: First, the CPU stores the dual-port SRAM in the chip system memory (such as Figure 1 The matrix accelerator is configured to prepare matrix operation data, initialize operation tasks and corresponding task control instructions, and start the matrix accelerator in the SRAM3 of the CPU. The matrix accelerator includes a system memory module, a task management module, a data access module, and a matrix operation module. After the CPU starts the matrix accelerator, the task queue in the task management module loads the operation tasks and the task control instructions generated by the CPU based on the operation tasks. When the task queue is not empty, the operation tasks and corresponding task control instructions are sent to the data access module. The data access module, based on the issued operation tasks and corresponding task control instructions, sends the matrix operation data read from the dual-port SRAM and the synchronously generated operation operation instructions to the matrix operation module. The matrix operation module selects the corresponding operator for the input matrix operation data based on the operation operation instructions to perform matrix operations. After the operation is completed, the operation result data is written to the dual-port SRAM via the data access module. After the operation result data is written, the data access module pulls up the task completion signal to indicate the completion of the currently issued operation task. Check whether the task queue is empty. If it is not empty, the task queue continues to issue calculation tasks and corresponding task control instructions to perform matrix operations until the task queue is empty. If it is empty, it is determined that the matrix accelerator has completed all calculation tasks, and the task management module generates an interrupt signal to notify the CPU to complete the calculation, and the CPU obtains the final matrix calculation result from the dual-port SRAM.
[0039] It should be understood that the above-mentioned matrix accelerator and its operating method for sharing chip system memory converts part or all of the single-port SRAM in the chip system memory into dual-port SRAM according to the maximum memory requirement of the matrix accelerator, so that the matrix accelerator directly shares the chip system memory through the dual-port SRAM. There is no need to design the internal memory capacity of the matrix accelerator based on the maximum possibility of matrix dimensions, making the matrix accelerator suitable for small-volume, low-power navigation baseband chips. At the same time, when the matrix accelerator is not working, the shared memory can still be used as chip system memory, avoiding idle memory resources. In addition, through the task queue, multiple computing tasks and corresponding task control instructions can be loaded into the matrix accelerator at one time, and the matrix accelerator can be controlled to execute the computing tasks and corresponding task control instructions in the programmed order to perform matrix computing operations. After all computing tasks are completed, an interrupt signal is generated to complete the interaction between the CPU and the matrix accelerator. This batch processing method significantly reduces the number of interactions between the CPU and the matrix accelerator, does not occupy the system bus bandwidth for a long time, maximizes the efficiency of the CPU, and significantly improves the performance of the chip system.
[0040] Further, taking the matrix operation D=A*B+C as an example, the following Figure 1 The operation process of a matrix accelerator that shares chip system memory is described.
[0041] Table 1 SRAM3 P0 port, P1 port address mapping relationship
[0042] Specifically, Figure 1 The total storage capacity of SRAM3 is 1M bytes, the data bit width is 64 bits, and the interconnection relationship between the P0 port and the P1 port address is as follows: Figure 7 As shown. The system address space of SRAM3 in the chip is 0×300000~0×3fffff. Figure 7 The address mapping relationship between the P0 port and the P1 port under the interconnection mode shown is shown in Table 1.
[0043] Assume that the matrices involved in the operation are A (20×40), B (40×40), and C (20×40), and the output matrix is D (20×40), with matrix elements being 64-bit double-precision floating-point numbers. The operation process includes: 1. The CPU prepares matrix operation data in the system memory SRAM3.
[0044] The CPU places matrix A at the system memory address 0×00300000 through the P0 port, and stores it in order of element rows, with 8 bytes storing one element, and the total data length is 20*40*8 bytes.
[0045] CPU places matrix B at the memory address 0x00310000 of the system through the P0 port, sequentially stores in the element row order, 8 bytes store an element, the total data length is 40*40*8 bytes.
[0046] CPU places matrix C at the memory address 0x00320000 of the system through the P0 port, sequentially stores in the element row order, 8 bytes store an element, the total data length is 20*40*8 bytes.
[0047] The physical path of the data flow is: CPU→system bus→protocol conversion submodule→P0 port of SRAM3→storage unit of SRAM3.
[0048] 2. CPU initializes the operation task and the task control instruction through the task queue interface.
[0049] (1) Task 1 (complete C_tmp=A*B).
[0050] Task number: 1, indicating that the current task number is 1.
[0051] Interrupt reporting enable: 0, indicating that the task with the number 1 does not generate an interrupt after completion.
[0052] Operation type selection: 0x03, indicating that the multiplication operation is selected.
[0053] Operand matrix X storage address: 0x00000000, indicating the first address of the P1 port where the matrix A is stored.
[0054] Operand matrix Y storage address: 0x00002000, indicating the first address of the P1 port where the matrix B is stored.
[0055] Operation result storage address: 0x00006000, indicating the first address of the P1 port where the operation result matrix C_tmp is stored.
[0056] Row number of the operand matrix X: 20, indicating the row dimension of the matrix A.
[0057] Column number of the operand matrix X / row number of the operand matrix Y: 40, indicating the column dimension of the matrix A and the row dimension of the matrix B.
[0058] Column number of the operand matrix Y: 40, indicating the column dimension of the matrix B.
[0059] Instruction validity identifier: 1.
[0060] (2) Task 2 (complete D=C_tmp+C).
[0061] Task number: 2, indicating that the current task number is 2.
[0062] Interrupt report enable: 1, indicating that the task numbered 2 completes to generate an interrupt.
[0063] Operation type selection: 0x00, indicating that the addition operation is selected.
[0064] Operand matrix X storage address: 0x00006000, indicating that the P1 port first address of the matrix C_tmp is stored.
[0065] Operand matrix Y storage address: 0x00004000, indicating that the P1 port first address of the matrix C is stored.
[0066] Operation result storage address: 0x00008000, indicating that the P1 port first address of the operation result matrix D is stored.
[0067] Number of rows of the operand matrix X: 20, indicating the row dimension of the matrix C_tmp.
[0068] Number of columns of the operand matrix X / number of rows of the operand matrix Y: invalid, which can not be filled.
[0069] Number of columns of the operand matrix Y: 40, indicating the column dimension of the matrix C.
[0070] Instruction effective identification: 1.
[0071] 3、The CPU starts the matrix accelerator.
[0072] The user interface sub-module writes 1 to the operation start register, and loads the task control instruction to the task queue.
[0073] 4、Matrix operation.
[0074] (1) The task queue is not empty, and the operation task 1 is issued.
[0075] The data access module obtains the first row elements of the matrix A from the addresses 0x00000000, 0x00000001, …, 0x00000027 through the P1 port in turn.
[0076] The data access module obtains the first column elements of the matrix B from the addresses 0x00002000, 0x00002028, …, 0x00002618 through the P1 port in turn.
[0077] The matrix operation module performs the multiplication and accumulation operation on the first row elements of the matrix A and the first column elements of the matrix B, calculates the first column element of the first row of the matrix C_tmp, and saves the result to the P1 port address 0x00006000.
[0078] Repeat the above three steps until the 40th column element of the 20th row of the matrix C_tmp is calculated and written into the P1 port address 0x0000631f.
[0079] (2) The operation task 1 is completed, and the task queue discards the task control instruction of the task 1. It is checked that the task queue is not emptied, and the operation task 2 is continuously issued.
[0080] The data access module obtains the first row elements of the matrix C_tmp from the addresses 0x00006000, 0x00006001,..., 0x00006027 through the P1 port in sequence.
[0081] The data access module obtains the first row elements of the matrix C from the addresses 0x00004000, 0x00004001,..., 0x00004027 through the P1 port in sequence.
[0082] The matrix operation module performs addition operation on the first row elements of the matrix C_tmp and the matrix C in sequence, obtains the first row elements of the matrix D, and saves the results into the P1 port addresses 0x00008000, 0x00008001,..., 0x00008027 in sequence.
[0083] Repeat the above three steps until the 20th row elements of the matrix D are calculated and written into the P1 port addresses 0x000082f8, 0x000082f9,..., 0x0000831f in sequence.
[0084] The physical path of data reading is: the storage unit of the SRAM3→the P1 port of the SRAM3→the data access module.
[0085] The physical path of data storage is: the data access module→the P1 port of the SRAM3→the storage unit of the SRAM3.
[0086] 5、The matrix accelerator completes the calculation.
[0087] The operation task 2 is completed, the task queue is emptied, the current operation is completed, the matrix accelerator generates an interrupt signal to inform the CPU that the operation is completed.
[0088] 6、The CPU obtains the calculation result in the system memory SRAM3.
[0089] The CPU accesses the system memory address 0x00340000 to obtain the final calculation result matrix D. The system memory address 0x00330000 is accessed to also obtain the intermediate result, that is, the calculation result of A*B.
[0090] The physical path of the data stream is: memory unit of SRAM3→P0 port of SRAM3→protocol conversion submodule→system bus→CPU.
[0091] The technical features of the above embodiments can be combined in any manner. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not contradict, they should be considered as the scope of the present disclosure.
[0092] The above embodiments only express several implementation manners of the present application, the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the present application. It should be pointed out that, for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the protection scope of the present application.
Claims
1. A matrix accelerator sharing a system memory of a chip, characterized in that, The matrix accelerator comprises a system memory module, a task management module, a data access module and a matrix operation module; The system memory module is configured to convert part or all of the single-port SRAM in the chip system memory into double-port SRAM according to the maximum memory requirement of the matrix accelerator, so that the matrix accelerator shares the chip system memory through the double-port SRAM. The task management module is configured to coordinate the operation control and data interaction of the CPU on the matrix accelerator, load the operation task and the task control instruction generated by the CPU based on the operation task through the task queue, and send the operation task and the corresponding task control instruction to the data access module when the task queue is not empty, and generate an interrupt signal to inform the CPU of the completion of the operation when all the operation tasks are completed. The data access module is configured to send the matrix operation data read from the double-port SRAM and the operation operation instruction synchronously generated to the matrix operation module according to the operation task and the corresponding task control instruction sent, write the operation result data sent by the matrix operation module into the double-port SRAM, and pull up the task completion signal to indicate that the current operation task is completed after the operation result data is written. The matrix operation module is configured to select a corresponding operation sub-module for the input matrix operation data according to the operation operation instruction to perform matrix operation, and send the operation result data to the data access module after the operation is completed.
2. The matrix accelerator sharing the system memory of a chip according to claim 1, characterized in that, The system memory module further comprises a protocol conversion submodule, one end of the double-port SRAM is connected with the protocol conversion submodule through a MEM bus, the protocol conversion submodule is configured to interconnect with the system bus and the CPU through an AXI protocol to ensure seamless connection between the system memory module and the CPU, and the other end of the double-port SRAM is connected with the data access module through the MEM bus to realize direct access of the matrix accelerator to the double-port SRAM.
3. The matrix accelerator sharing on-chip system memory of claim 1, wherein, The task management module further comprises a user interface submodule, the user interface submodule comprises an operation start and stop control register, an interrupt enable and clear control register, a status register and a task queue interface, the user interface submodule is connected with the system bus through an AHB bus and interconnected with the CPU through the AXI protocol, and is configured to complete operation start and stop control, interrupt enable and clear control and task queue initialization of the matrix accelerator under the control of the CPU.
4. The matrix accelerator sharing the system memory of a chip according to claim 3, characterized in that, The working process of the task queue is as follows: When the task queue receives a start signal sent by the CPU through the user interface submodule, the operation task and the corresponding task control instruction initialized by the CPU in the task queue interface are loaded into the task queue in the order of address increment; wherein, the task queue is managed in a first-in first-out manner; When the task queue is not empty, the operation task and the corresponding task control instruction located at the first position in the task queue are sent to the data access module; When the task queue detects that the data access module pulls up the task completion signal, it is determined that the current issued operation task is completed, the current issued operation task is deleted in the task queue, and the corresponding task number is uploaded to the state register of the user interface submodule for the CPU to query the task state information. Meanwhile, the task queue determines whether the user interface submodule generates an interrupt signal according to whether the interrupt reporting enable in the current corresponding task control instruction is high. It is detected whether the task queue is empty. If it is empty, the matrix operation is stopped, and the interrupt signal generated by the user interface submodule notifies the CPU that the operation is completed. If it is not empty, the operation task and the corresponding task control instruction are issued according to the task queue sorting to continue the matrix operation until the task queue is empty.
5. The matrix accelerator sharing the system memory of a chip according to claim 4, characterized in that, The task control instruction is composed of a task number, an interrupt reporting enable, an operation type selection, a storage address of an operand matrix X, a storage address of an operand matrix Y, a storage address of an operation result, a number of rows of the operand matrix X, a number of columns of the operand matrix X and a number of rows of the operand matrix Y, a number of columns of the operand matrix Y, and an instruction validity identifier.
6. The matrix accelerator sharing on-chip system memory according to claim 4, wherein, The data access module includes an SRAM read / write control submodule, a matrix operation control submodule, a matrix multiplication / addition operation control submodule, and a matrix inversion operation control submodule. The matrix operation control submodule is configured to receive the operation task and the corresponding task control instruction issued by the task queue for parameter analysis, and send the analyzed operation parameters to the SRAM read / write control submodule, the matrix multiplication / addition operation control submodule, and the matrix inversion operation control submodule. The matrix operation control submodule is also configured to, after receiving the operation completion signal generated by the SRAM read / write control submodule, indicate that the current issued operation task is completed by pulling up the task completion signal. The operation parameters include an operation type, a matrix data storage starting address, and a matrix dimension parameter. The matrix multiplication / addition operation control submodule is configured to, based on the operation parameters, calculate the access starting address of the matrix multiplication / addition operation data, the read / write cycle number, and the increment / decrement amplitude of the address of each read / write cycle, generate the operation operation instruction corresponding to the matrix multiplication / addition operation, and send the operation operation instruction to the matrix operation module. The matrix inversion operation control submodule is configured to, based on the operation parameters, generate an LDL decomposition inversion algorithm control flow, calculate the access starting address of the matrix inversion operation data, the read / write cycle number, and the increment / decrement amplitude of the address of each read / write cycle in the control flow, generate the operation operation instruction corresponding to the matrix inversion operation, and send the operation operation instruction to the matrix operation module. The SRAM read-write control submodule is configured to generate an SRAM read-write control signal based on the operation parameters and the access start address, the read / write cycle number, and the increment / decrement amplitude of the address of each read / write cycle of the matrix operation data provided by the matrix multiplication / addition operation control submodule and the matrix inversion operation control submodule, perform a read operation first according to the SRAM read-write control signal, read the matrix operation data from the dual-port SRAM and send the matrix operation data to the matrix operation module, perform a write operation after the operation is completed, write the operation result data sent by the matrix operation module into the dual-port SRAM, and generate an operation completion signal and report the operation completion signal to the matrix operation control submodule after the last operation result data is written.
7. The matrix accelerator sharing on-chip system memory according to claim 6, wherein, The matrix operation module comprises an operation input-output control submodule, a floating-point multiplication / addition operation unit, and a floating-point division operation unit. The operation input-output control submodule is configured to receive the operation operation instruction and the matrix operation data sent by the data access module, select a corresponding operation unit to perform a matrix operation according to the operation operation instruction, and send the matrix operation data to the corresponding operation unit, obtain the operation result data output by the operation unit after the operation is completed, and send the operation result data to the data access module. The floating-point multiplication / addition operation unit is configured to receive the matrix operation data sent by the operation input-output control submodule, complete multiplication, addition, and multiplication-accumulation operations according to the corresponding operation operation instruction, and return the operation result data to the operation input-output control submodule. The floating-point division operation unit is configured to receive the matrix operation data sent by the operation input-output control submodule, complete division and reciprocal operations according to the corresponding operation operation instruction, and return the operation result data to the operation input-output control submodule.
8. A method of operating a matrix accelerator sharing a system memory of a chip according to any one of claims 1 to 7, characterized in that, The method comprises the following steps: The CPU prepares matrix operation data in the dual-port SRAM of the chip system memory, initializes an operation task and a corresponding task control instruction, and starts the matrix accelerator; wherein the matrix accelerator comprises a system memory module, a task management module, a data access module, and a matrix operation module. After the CPU starts the matrix accelerator, the task queue in the task management module loads the operation task and the task control instruction generated by the CPU based on the operation task, and when the task queue is not empty, the operation task and the corresponding task control instruction are sent to the data access module, the data access module sends the matrix operation data read from the dual-port SRAM and the operation operation instruction generated synchronously to the matrix operation module according to the operation task and the corresponding task control instruction sent, the matrix operation module selects a corresponding operation unit to perform a matrix operation for the input matrix operation data according to the operation operation instruction, writes the operation result data into the dual-port SRAM through the data access module after the operation is completed, and pulls up the task completion signal through the data access module to indicate that the current sent operation task is completed after the operation result data is written. The task queue is detected whether it is empty, if not empty, the operation task and the corresponding task control instruction are continuously issued by the task queue to carry out the matrix operation until the task queue is empty; if empty, it is determined that the matrix accelerator completes all operation tasks, an interrupt signal is generated by the task management module to inform the CPU to complete the operation, and the final matrix operation result is obtained from the dual-port SRAM by the CPU.
Citation Information
Patent Citations
Universal floating point matrix processor hardware structure based on FPGA (field programmable gate array)
CN104391820A
Instruction set device for reconfigurable deep neural network accelerator
CN116431214A
Universal computing accelerator based on shared tight coupling
CN117290279A
Data exchange method and device of shared storage pool based on SOC chip
CN117389767A
Visual self-attention accelerator optimization method based on FPGA
CN117610612A
Cited By
Matrix reduction operation method and device, equipment, storage medium and product
CN121143750A