Adaptive Reconfigurable Processing Array and Master Control Interaction Method and Device

By adding a control processing unit to the reconfigurable processing array, the clock cycle waste caused by the main control control operation is solved, and the computing performance is improved, which is suitable for data-intensive and computationally intensive applications.

CN113792009BActive Publication Date: 2025-07-22TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110861861.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-29
Publication Date
2025-07-22
Estimated Expiration
2041-07-29

AI Technical Summary

Technical Problem

In the prior art, the control operation of the main control on the reconfigurable processing array results in a large number of clock cycles waste, and the time overhead generated during task iteration increases orders of magnitude, affecting computing performance.

Method used

The control-type processing unit is added to the reconfigurable processing array, instead of the master's reading and writing of the global register GR in the coprocessor interface, data handling and array execution are realized, and the coupling between the array and the master is reduced.

Benefits of technology

It greatly shortens the execution time of the application, greatly improves computing power and computing performance, and is suitable for data-intensive and computing-intensive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113792009B_ABST
    Figure CN113792009B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive reconfigurable processing array and a main control interaction method and device. The device includes a control type processing unit, which is arranged on the reconfigurable processing array and is used to replace the main control to read and write the global register GR in the coprocessor interface, so as to realize data transfer and array execution. The present invention greatly reduces the coupling degree between the array and the main control, shortens the execution time of the application, greatly improves the computing power and computing performance, meets the requirements of the application computing performance, and is very suitable for hardware acceleration design for data-intensive and computing-intensive applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large-scale integrated circuit technology, and particularly to an adaptive reconfigurable processing array and a master control interaction method and device. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention described in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] The coprocessor interface module is the bridge for interaction between the master control and the reconfigurable processing array. In addition to being able to read and write the shared global registers on the reconfigurable processing array, the coprocessor interface also has 10 global registers for functional control of the reconfigurable processing array, and these 10 registers are invisible to the processing units in the reconfigurable processing array. The coprocessor interface module generates an enable signal by parsing the values of these 10 registers to control tasks such as calculation, data transfer, and configuration transfer of the reconfigurable processing array.

[0004] Before adding the control-type processing unit, the values of these special global registers for control were written by the master control. Each time the master control writes a value to a global register, a large number of clock cycles are wasted, and each time a reconfigurable processing array startup task is initiated or a data transfer task is initiated, values are written to multiple special global registers. When an application is executing, these tasks will iterate multiple times, so the resulting time overhead will increase exponentially. Summary of the Invention

[0005] An embodiment of the present invention provides an adaptive reconfigurable processing array and master control interaction method, the method comprising:

[0006] Adding a control-type processing unit on the reconfigurable processing array to replace the master control for reading and writing the global register GR in the coprocessor interface, so as to implement data transfer and array execution.

[0007] An embodiment of the present invention further provides an adaptive reconfigurable processing array and master control interaction device, the device comprising: a control-type processing unit, arranged on the reconfigurable processing array, used to replace the master control for reading and writing the global register GR in the coprocessor interface, so as to implement data transfer and array execution.

[0008] In the embodiments of the present invention, compared with the prior art in which all tasks such as controlling the calculation, data transfer, and configuration transfer of the reconfigurable processing array are completed by the main controller, by setting a control-type processing unit on the reconfigurable processing array to replace the main controller for reading and writing the global register GR in the coprocessor interface, data transfer and array execution are realized. It can greatly reduce the coupling degree between the array and the main controller, shorten the execution time of the application, greatly improve the computing power and computing performance, meet the requirements of the application computing performance, and is very suitable for hardware acceleration design for data-intensive and computing-intensive applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0010] Figure 1 is the architecture diagram of the reconfigurable processor unit (RPU);

[0011] Figure 2 is the schematic diagram of the operation process required by the main controller before the execution of the task of transferring the configuration;

[0012] Figure 3 is the schematic diagram of the operation process that the main controller needs to perform before transferring data to the shared memory;

[0013] Figure 4 is the architecture diagram of the interactive device between the adaptive reconfigurable processing array and the main controller in the embodiments of the present invention, that is, the architecture diagram of the reconfigurable array after adding the control-type PE;

[0014] Figure 5 is the architecture diagram of the control-type PE in the embodiments of the present invention;

[0015] Figure 6 is the schematic diagram of the configuration information format of the control-type PE in the embodiments of the present invention;

[0016] Figure 7 is the schematic diagram of the execution process of the original 512-point FFT;

[0017] Figure 8 is the schematic diagram of the execution process of the new 512-point FFT in the embodiments of the present invention;

[0018] Figure 9 is the schematic diagram of the configuration information executed by the control-type PE in the embodiments of the present invention. Detailed Implementation Modes

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0020] Figure 1 It is an architecture diagram of a reconfigurable processor unit (RPU). As Figure 1 shown, it is an architecture diagram of a reconfigurable processor unit (RPU). An RPU has four processing element arrays (PEAs). The main control RISCV (RISCV processor) interacts with the PEA through a coprocessor interface. The Data Cache stores data and sends the data to the shared memory module on the PEA array through the AHB bus. The Context Cache stores configuration information and distributes the configuration information to each processing element on the PEA array through the AHB bus (the bridge between the Cache and the PEA array) and the configuration control module (Config_Control) on the PEA array. GR: Global register. PE: Processing element. CM: Context memory, which stores the configuration information of the PE.

[0021] The global register (GR) for storing the function control of the PEA is located in the coprocessor interface (Corprocessor Interface). Among them, GR32 stores the start control signal, GR35 stores the base address of the data in the cache, GR36 stores the length of the data transferred from the Cache, GR33 stores the base address of the configuration information in the Cache, GR34 stores the length of the configuration information transferred from the Cache, GR37 stores the base address for reading and writing the shared memory, and GR39 stores the finish signal for the completion of data transfer, configuration transfer, and PEA calculation.

[0022] As Figure 2As shown, these are the operations required by the main controller before the task execution of the handling configuration. The main controller needs to write the base address of the configuration information in the cache to GR33 through the coprocessor interface, write the length of the handling configuration information to GR34, and write the enable signal for starting the handling of the configuration information to GR32. After writing these global registers, start the handling of the configuration information.

[0023] As Figure 3 shown, these are the operations that the main controller needs to perform before transferring data to the shared memory. The main controller needs to write the base address of the data in the cache to GR35 through the coprocessor interface, write the length of the transferred data to GR36, write the base address for reading and writing the shared memory to GR37, and write the enable signal for starting the data transfer to GR32. After writing these global registers, start the data transfer.

[0024] Each time the main controller writes data to a GR, it will take thousands of clock cycles, and a single task requires writing to multiple GRs. Thus, the time overhead for executing a single task is huge.

[0025] Due to the above problems, in order to reduce the coupling degree between the array and the main controller, as Figure 4 shown, a control-type PE (Ctrl_PE) is added to the PEA (processing element array) to replace the main controller for reading and writing GRs, automatically calculate information such as address offset and transfer length, and further control the execution of various tasks (Exec_PE).

[0026] Figure 5 is the architecture diagram of the control-type PE. The configuration information storage module (Config_memory) can store up to 16 pieces of configuration information at most, and the local register file (Local_Regfile) can store 8 data. The control-type PE supports a four-stage pipeline of read configuration (control), decode, execute, and writeback. Among them, the execute module includes common operations such as addition, subtraction, AND, OR, NOT, XOR, equality, and shift to support flexible address jump modes. It also includes a special operation "wait" operation for controlling the execution of the array. The "wait" operation continuously reads the signal in GR39 to detect whether tasks such as data transfer and array calculation are completed.

[0027] Figure 6The configuration information format of the control-type PE is shown. There are two inputs, input1 and input2, both of which support multiple data sources. Input1 can access the GR of the coprocessor interface, the previous operation result, and the local register. Input2 can access the shared register of the PEA array, the previous operation result, and the local register. Moreover, when the Imm immediate number enable field is set to 1, the Input2 field is an immediate number and participates in the operation. The output out1 also supports multiple output destinations, including the coprocessor interface register, the shared register of the PEA array, and the local register. The control-type PE supports iteration and pausing to support flexible and complex control functions.

[0028] Through the technology of the present invention, a control-type PE is added on the basis of the original PEA array to complete tasks such as data transfer and array startup instead of the main control, greatly reducing the coupling degree between the array and the main control, shortening the execution time of the application, greatly improving the computing power and computing performance, meeting the requirements of the application computing performance, and being very suitable for hardware acceleration design for data-intensive and computing-intensive applications.

[0029] The following analyzes the adaptive reconfigurable processing array and the main control interaction method and device proposed by the present invention through examples.

[0030] As Figure 7 shown is the execution process of a 512-point FFT, which is also the execution method without adding a control-type PE. It can be seen from the figure that for each layer iteration of the FFT, the main control needs to control the data transfer and the execution of the array, resulting in huge time overhead. The total execution time of the 9-layer FFT is 358,000 ns.

[0031] As Figure 8 shown, it is the execution process of the 512-point FFT after adding a control-type PE. Compared with the original version, except for pre-storing the data required for calculation in the GR of the coprocessor interface in advance, the entire execution process has been greatly streamlined. After transferring the configuration and writing the global register, the main control only needs to enable the control-type PE, and the control-type PE can control the completion of the 9-layer FFT iteration, instead of the main control being responsible for the data transfer of each layer and the start of the iterative calculation as frequently as before, reducing the coupling degree between the main control and the PEA array. The control-type PE completes the data transfer of each layer and the start of the iterative calculation through the configuration information as Figure 9 shown. The running time of the 512-point FFT is 15,600 ns. Compared with the previous version, the operation time is shortened by 95.648%.

[0032] An adaptive reconfigurable processing array and master control interaction method is also provided in an embodiment of the present invention, as described in the following embodiments. Since the principle of this method for solving problems is similar to that of the adaptive reconfigurable processing array and master control interaction device, the implementation of this device can refer to the implementation of the adaptive reconfigurable processing array and master control interaction device, and the repeated parts will not be elaborated.

[0033] The adaptive reconfigurable processing array and master control interaction method includes:

[0034] Add a control processing unit on the reconfigurable processing array to replace the master control for reading and writing the global register GR in the coprocessor interface, so as to realize data transfer and array execution.

[0035] In an embodiment of the present invention, the control processing unit includes a configuration information storage module, local registers, a configuration module, a decoding module, an execution module, and a write-back module;

[0036] Among them, the configuration information storage module is used to store configuration information;

[0037] The local registers are used to store data;

[0038] The configuration module, decoding module, execution module, and write-back module are used to implement a four-stage pipeline of read configuration, decoding, execution, and write-back.

[0039] In an embodiment of the present invention, the execution module is further used to implement a wait operation, where the wait operation detects whether tasks such as data transfer and array calculation are completed by reading the finish signal in the global register GR39.

[0040] In an embodiment of the present invention, the control processing unit configures the information format in the following manner:

[0041] The information format includes two inputs, input1 and input2, and supports multiple data sources;

[0042] Among them, input1 accesses the global register GR in the coprocessor interface, the result of the previous operation, and the local registers;

[0043] Input2 accesses the shared registers on the reconfigurable processing array PEA, the result of the previous operation, and the local registers;

[0044] It includes an immediate number enable field. When the immediate number enable field is set to 1, the input2 field is an immediate number and participates in the operation.

[0045] In an embodiment of the present invention, the control processing unit configures the information format in the following manner:

[0046] The information format includes output out1, which supports multiple output destinations, including global registers in the coprocessor interface, shared registers on the reconfigurable processing array (PEA array), and local registers.

[0047] In the embodiments of the present invention, compared with the prior art in which all tasks such as reconfigurable processing array calculation, data transfer, and configuration transfer are controlled by the main control, by setting a control processing unit on the reconfigurable processing array to replace the main control for reading and writing the global registers (GR) in the coprocessor interface, data transfer and array execution are realized. This can greatly reduce the coupling degree between the array and the main control, shorten the execution time of the application, greatly improve the computing power and computing performance, meet the requirements of the application computing performance, and is very suitable for hardware acceleration design for data-intensive and computing-intensive applications.

[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0049] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0050] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks Figure 1 of the functions specified in one block or a plurality of blocks.

[0052] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adaptive reconfigurable processing array and master control interaction method, characterized in that Comprising: Adding a control processing unit on the reconfigurable processing array to replace the master for reading and writing the global register GR in the coprocessor interface, so as to implement data transfer and array execution; The control processing unit includes a configuration information storage module, local registers, a configuration module, a decoding module, an execution module, and a write-back module; among them, the configuration information storage module is used to store configuration information; the local registers are used to store data; the configuration module, the decoding module, the execution module, and the write-back module are used to implement four-level pipelining of reading configuration, decoding, execution, and write-back; The control processing unit configures the information format in the following manner: The information format includes two inputs, input1 and input2, supporting multiple data sources; among them, input1 accesses the global register GR in the coprocessor interface, the previous operation result, and the local registers; input2 accesses the shared registers on the reconfigurable processing array PEA, the previous operation result, and the local registers; it includes an immediate number enable field, when the immediate number enable field is set to 1, the input2 field is an immediate number and participates in the operation.

2. The adaptive reconfigurable processing array and master control interaction method according to claim 1, wherein, The execution module is also used to implement waiting operations, where the waiting operations detect whether the data transfer task and the array calculation task are completed by reading the finish signal in the global register GR39.

3. The adaptive reconfigurable processing array and master control interaction method according to claim 1, characterized in that, The control processing unit configures the information format in the following manner: The information format includes an output out1, supporting multiple output destinations, including the global register in the coprocessor interface, the shared registers on the reconfigurable processing array PEA, and the local registers.

4. An adaptive reconfigurable processing array and master control interaction device, characterized in that Comprising: A control processing unit, arranged on the reconfigurable processing array, used to replace the master for reading and writing the global register GR in the coprocessor interface, so as to implement data transfer and array execution; The control processing unit includes a configuration information storage module, local registers, a configuration module, a decoding module, an execution module, and a write-back module; among them, the configuration information storage module is used to store configuration information; the local registers are used to store data; the configuration module, the decoding module, the execution module, and the write-back module are used to implement four-level pipelining of reading configuration, decoding, execution, and write-back; The control processing unit configures the information format in the following manner: The information format includes two inputs, input1 and input2, supporting multiple data sources; among them, input1 accesses the global register GR in the coprocessor interface, the previous operation result, and the local registers; input2 accesses the shared registers on the reconfigurable processing array PEA, the previous operation result, and the local registers; it includes an immediate number enable field, when the immediate number enable field is set to 1, the input2 field is an immediate number and participates in the operation.

5. The adaptive reconfigurable processing array and master control interaction device according to claim 4, characterized in that The execution module is also used to implement waiting operations, where the waiting operations detect whether the data transfer task and the array calculation task are completed by reading the finish signal in the global register GR39.

6. The adaptive reconfigurable processing array and master control interaction device according to claim 4, wherein The control processing unit configures the information format in the following manner: The information format includes output out1, supporting multiple output destinations, including global registers in the coprocessor interface, shared registers on the reconfigurable processing array PEA array, and local registers.

Citation Information

Patent Citations

  • Coprocessor based on reconfigurable computational array

    CN105630735A

  • Register file design method and device of reconfigurable processing unit array

    CN112486904A