Coarse-grained reconfigurable array and hardware modeling generation method thereof

By designing coarse-grained reconfigurable arrays and their hardware modeling and generation methods, and using the intermediate representation of hardware description to achieve target redirection, the problem of inefficiency of existing hardware automatic generation methods is solved, and efficient hardware design and software stack development are achieved.

CN119940250AActive Publication Date: 2025-05-06UNIV OF SCI & TECH OF CHINA

Patent Information

Application Number
CN202510016687.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Existing hardware automatic generation methods often introduce greater area and performance overhead when generating hardware, and are difficult to debug and verify, making it difficult to meet the efficiency needs of hardware design and software stack development.

Method used

Design a coarse-grained reconfigurable array and its hardware modeling generation method, realize the target redirection of driving the compilation tool through the hardware description intermediate representation, support the highly free Internet network of the processing unit array, and generate agile modeling and RTL-level code.

Benefits of technology

This greatly improves the efficiency of hardware accelerators in hardware design and software stack development, and maintains a good compromise in programming flexibility, performance, and power consumption of coarse-grained reconfigurable arrays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940250A_ABST
    Figure CN119940250A_ABST
Patent Text Reader

Abstract

The invention discloses a coarse-grained reconfigurable array and a hardware modeling generation method thereof, and belongs to the technical field of digital circuit design and processing. The coarseness reconfigurable array comprises a processing unit array of a reconfigurable array formed by M rows and N columns of processing units, a configuration memory used for temporarily storing operation configuration information of the processing units, a configuration bus used for connecting the configuration memory and the processing units, and an input / output interface. The modeling generation method comprises the steps that submodule design parameters are generated layer by layer from top to bottom through input top-layer hardware design parameters, and hardware RTL-level codes are generated from a bottom-to-top instantiation hardware generator. By the adoption of the coarse-grained reconfigurable array and the hardware modeling generation method thereof, agile modeling and RTL-level code generation of the coarse-grained reconfigurable array can be completed; the target redirection of the driving compiling tool is realized by designing hardware description intermediate representation; and the efficiency of hardware design and software stack development of the hardware accelerator is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital circuit design and processing, and in particular to a coarse-grained reconfigurable array and a hardware modeling generation method thereof. Background Art

[0002] In the related art, traditional computing hardware includes general-purpose processors and application-specific integrated circuits (ASICs). The former has good flexibility but cannot meet the real-time requirements of specific application fields. Domain-specific hardware accelerators can provide several orders of magnitude more acceleration and energy efficiency than general-purpose processors. However, domain-specific hardware accelerators require a lot of manual work in hardware design and software stack development, and cannot take flexibility into account. In recent years, coarse-grained reconfigurable arrays (CGRAs) have attracted widespread attention due to their good compromise between programming flexibility and performance and energy consumption.

[0003] Traditional hardware RTL design is directly mapped to the hardware structure. Designers can achieve extremely high performance through fine optimization, but the design efficiency is low. Existing hardware automatic generation methods such as high-level synthesis can convert high-level languages ​​(such as C, C++, SystemC, etc.) into hardware RTL-level languages, but the generated hardware often introduces larger area and performance overhead, and is difficult to debug and verify. Summary of the invention

[0004] The purpose of the present invention is to provide a coarse-grained reconfigurable array and a hardware modeling generation method thereof, which can complete agile modeling and RTL-level code generation of the coarse-grained reconfigurable array; realize target redirection of the driver compilation tool by designing an intermediate representation of the hardware description; and greatly improve the efficiency of hardware accelerators in hardware design and software stack development.

[0005] To achieve the above object, the present invention provides a coarse-grained reconfigurable array, comprising:

[0006] The processing unit array is a reconfigurable array formed by M rows and N columns of processing units, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1. The processing unit performs a calculation or access operation according to input operation configuration information and input data, and outputs a calculation or access result;

[0007] Configuration memory, used to temporarily store operation configuration information required for execution of each node in the processing unit array;

[0008] A configuration bus, used to connect the configuration memory and the processing units, and input operation configuration information to at least one processing unit;

[0009] The input and output interface connects the processing unit to external devices and is used to input data to the processing unit or output the calculation results of the processing unit.

[0010] Preferably, each column of processing units shares a configuration bus.

[0011] Preferably, the processing unit is an access unit or a computing unit;

[0012] The access unit is used to store and read data. When the coarse-grained reconfigurable array is working, the access unit stores input data before the calculation unit starts calculating, temporarily stores intermediate data when the calculation unit calculates, and stores output data after the calculation unit completes the calculation.

[0013] The computing unit is used to perform computing operations according to input operation configuration information and input data, and output execution results.

[0014] Preferably, the access unit includes:

[0015] The storage configuration register is used to temporarily store the storage operation configuration information input by the configuration bus;

[0016] Scratch pad memory, composed of SRAM, used to cache data;

[0017] A storage multiplexer connected to adjacent computing units selects corresponding input data according to storage operation configuration information in the storage register;

[0018] The logic unit is used to execute the logic of the access unit; according to the storage operation configuration information, when working in the loading state, the data is taken out from the note memory and output; when working in the storage state, the input data is stored in the note memory and the input data is directly output.

[0019] Preferably, the storage operation configuration information includes input selection information, address information of access data and access operation code.

[0020] Preferably, the calculation unit comprises:

[0021] A calculation configuration register, used to temporarily store the calculation operation configuration information input by the configuration bus;

[0022] A calculation multiplexer is connected to adjacent calculation units and selects corresponding input data according to calculation operation configuration information in the calculation register;

[0023] A synchronizer delays the input data selected by the multiplexer to synchronize the input data, and selects the corresponding data to be input to the arithmetic operation unit according to the delay selection information in the calculation operation configuration information;

[0024] The arithmetic operation unit performs arithmetic operations according to the calculation operation code in the calculation operation configuration information to obtain result data; or directly outputs the input data.

[0025] Preferably, the computing operation configuration information includes input selection information, immediate data, delay selection information and computing operation code.

[0026] Preferably, in the computing unit, a synchronizer is only provided after the computing multiplexer on one side, and the operation is immediately performed when the next clock arrives after the input data of the computing multiplexer on the other side arrives.

[0027] The hardware modeling generation method based on the above coarse-grained reconfigurable array includes the following steps:

[0028] S1. Modeling: divide the hardware structure of the coarse-grained reconfigurable array into top-level design primitives, middle-level design primitives and bottom-level design primitives. The top-level design primitive is the coarse-grained reconfigurable array. The middle-level design primitives include configuration memory, bus, computing unit and access unit. The bottom-level design primitives include computing configuration register, computing multiplexer, synchronizer, arithmetic operation unit and storage configuration register, storage multiplexer and memo memory. According to the input top-level hardware design parameters, the hardware of the top-level design primitive, the middle-level design primitive and the bottom-level design primitive are modeled from top to bottom and the sub-module hardware design parameters of the top-level design primitive, the middle-level design primitive and the bottom-level design primitive are generated. The hardware description intermediate representation is output to drive the target redirection of the compilation tool.

[0029] S2. Generation. After the modeling is completed, the hardware generator is instantiated from the bottom to the top according to the top-level hardware design parameters, the middle-level hardware design parameters and the bottom-level hardware design parameters, and the hardware RTL-level code is output.

[0030] Preferably, in S1, the hardware description intermediate representation includes a hardware architecture description and a hardware configuration description.

[0031] The advantages and positive effects of the coarse-grained reconfigurable array and the hardware modeling generation method thereof described in the present invention are: by designing the intermediate representation of the hardware description, the target redirection of the driver compilation tool is realized. The coarse-grained reconfigurable array and the hardware modeling generation method thereof described in the present invention can greatly improve the efficiency of hardware accelerators in hardware design and software stack development, support a highly free interconnection network of the processing unit array, and retain the good compromise of the coarse-grained reconfigurable array in programming flexibility and performance and power consumption.

[0032] 1. The present invention can perform agile modeling and RTL-level code generation for a coarse-grained reconfigurable array in a huge design space composed of a processing unit array and its homogeneity, an interconnection network, a configuration system, a memory, etc.

[0033] 2. The present invention makes it possible to design a redirectable compilation mapping method by designing a hardware description intermediate representation.

[0034] 3. The present invention greatly improves the efficiency of hardware accelerators in hardware design and software stack development, supports a highly free interconnection network of processing unit arrays, and retains the good compromise between programming flexibility, performance and power consumption of coarse-grained reconfigurable arrays.

[0035] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of a coarse-grained reconfigurable array according to an embodiment of the present invention;

[0037] Figure 2 A schematic diagram of a computing unit of a coarse-grained reconfigurable array according to an embodiment of the present invention;

[0038] Figure 3 A schematic diagram of an access unit of a coarse-grained reconfigurable array according to an embodiment of the present invention;

[0039] Figure 4 A schematic diagram of a synchronizer of a coarse-grained reconfigurable array according to an embodiment of the present invention;

[0040] Figure 5 This is a flow chart of a method for generating hardware modeling of a coarse-grained reconfigurable array according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] In the description of the present invention, it should be noted that the terms "upper", "lower", "inside", "outside", etc. indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, or the positions or positional relationships in which the invented product is usually placed when in use. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In the description of the present invention, it should also be noted that, unless otherwise clearly specified and limited, the terms "setting", "installation", and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0042] In this application, unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of this application. In the event of any inconsistency, the meaning described in this specification or the meaning derived from the contents recorded in this specification shall prevail. In addition, the terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit this application. In order to accurately describe the technical content in this application and to accurately understand the present invention, the following explanations or definitions are given to the terms used in this specification before describing the specific embodiments:

[0043] The embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.

[0044] Example

[0045] like Figure 1 A coarse-grained reconfigurable array includes a configuration memory, a processing unit array, a configuration bus, and an input-output interface.

[0046] The processing unit array is a reconfigurable array formed by M rows and N columns of processing units, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1. Each processing unit performs a specific calculation or access operation according to input operation configuration information and input data, and outputs the calculation or access result.

[0047] Configuration memory (Config Memory), composed of a register array, is used to temporarily store the operation configuration information required for each node in the processing unit array to execute. The operation configuration information includes the type of calculation performed by each processing unit in the processing unit array, circuit selection information, and address information for accessing data.

[0048] The configuration bus is used to connect the configuration memory and the processing unit. The configuration bus inputs operation configuration information to at least one processing unit.

[0049] Each column of processing elements shares a configuration bus.

[0050] The input and output interface directly connects the processing unit to external devices such as memory, and is used to input data to the processing unit or output the calculation results of the processing unit.

[0051] The processing unit array consists of two types of processing units: computing unit (PE, Process Element) and access unit (LSU, Load Store Unit).

[0052] The access unit is used to store and read data. When the coarse-grained reconfigurable array is working, it stores input data before the calculation unit starts calculating, temporarily stores intermediate data when the calculation unit calculates, and stores output data after the calculation unit completes the calculation.

[0053] The calculation unit is used to perform calculation operations or logical operations according to input operation configuration information and input data, and output the execution results.

[0054] The number of rows and columns of the processing unit array can be configured by the top-level hardware design parameters in the modeling method. The processing unit can be connected to any other processing unit or input and output interface in the array, and the interconnection network of the processing unit can be configured by the top-level hardware design parameters in the modeling method.

[0055] In this embodiment, the first and last columns of processing units in the processing unit array are configured as access units or computing units by the top-level hardware design parameters in the modeling method. The arithmetic or logical operation types supported by the computing unit are configured by the top-level hardware design parameters in the modeling method. The input and output interfaces can be connected to any processing unit, and the positions of the input and output interfaces are configured by the top-level hardware design parameters in the modeling method.

[0056] Figure 1 In the example, the interconnection network of the processing units is of N2N (neighbor to neighbor) type, the first column of the processing unit array is configured as an access unit, and the last column of computing units is configured with an input and output interface.

[0057] like Figure 2 The computing unit includes:

[0058] The calculation configuration register (Config) is used to temporarily store the calculation operation configuration information input by the configuration bus.

[0059] The calculation operation configuration information includes input selection information, immediate value (const), delay selection information and calculation operation code, etc.

[0060] A calculation multiplexer (MUX) is connected to a plurality of adjacent calculation units and selects corresponding input data according to calculation operation configuration information in a calculation register.

[0061] The synchronizer (Sync) delays the input data selected by the multiplexer to synchronize the input data.

[0062] like Figure 4 The synchronizer is composed of a cascaded register group, which can delay the input data selected by the multiplexer. The number of delay cycles is determined by the delay selection information in the calculation operation configuration information in the unit, so that the two input data of the calculation unit can be synchronized and the calculation operation can be performed in the same clock cycle.

[0063] The arithmetic logic unit (ALU) performs arithmetic operations according to the calculation operation code in the calculation operation configuration information to obtain result data. It can also directly output input data, in which case the calculation unit is used as a router.

[0064] The arithmetic operations supported by the arithmetic operation unit include addition, subtraction, multiplication, shift, comparison and other arithmetic and logical operations.

[0065] In the computing unit, there is only a synchronizer after a computing multiplexer on the right side. The input data that needs to be delayed will be selected from the computing multiplexer on the right side, and the operation will be performed immediately when the next clock arrives after the input data on the left side arrives. This single-sided synchronizer design can effectively reduce the hardware area while supporting data synchronization.

[0066] The arithmetic operation types supported by the arithmetic operation unit and the depth of the synchronizer, i.e., the maximum number of delay cycles that can be supported, are configured by the top-level hardware design parameters in the modeling method. The number of input ports of the computing unit and the bit width of the configuration register are further calculated from the top-level design parameters in the modeling method.

[0067] like Figure 3 The access unit includes:

[0068] The storage configuration register is used to temporarily store the storage operation configuration information input by the configuration bus.

[0069] The storage operation configuration information includes input selection information, address information of access data, access operation code, etc.

[0070] The scratch pad memory, which is composed of SRAM (Static Random Access Memory), is used to cache data.

[0071] The storage multiplexer is connected to a plurality of adjacent computing units and selects corresponding input data according to the storage operation configuration information in the storage register.

[0072] The logic unit is used to execute the logic of the access unit. According to the storage operation configuration information, when working in the loading state, the data is taken out from the note memory and output; when working in the storage state, the input data is stored in the note memory and the input data is directly output.

[0073] The word width and word length of the scratchpad memory, that is, the capacity of the memory, are configured by the top-level design parameters in the modeling method. The number of input ports of the access unit and the bit width of the access configuration register are further calculated from the top-level design parameters in the modeling method.

[0074] like Figure 5The hardware modeling generation method of the above coarse-grained reconfigurable array includes the following steps:

[0075] S1. Modeling. The coarse-grained reconfigurable array is abstracted into a series of hardware primitives. According to the hardware structure of the coarse-grained reconfigurable array, the hardware primitives include three levels: top-level design primitives, middle-level design primitives, and bottom-level design primitives. The top-level design primitive is the coarse-grained reconfigurable array. The middle-level design primitives include configuration memory, bus, computing unit, and access unit. The bottom-level design primitives include computing configuration registers, computing multiplexers, synchronizers, arithmetic operation units, storage configuration registers, storage multiplexers, and scratch pad memory. Each hardware primitive corresponds to a hardware generator, and the input of the hardware generator is the design parameters of the hardware.

[0076] Verify the hardware design parameters according to the input top-level hardware design parameters, confirm the legality of the input parameters, and then generate the hardware architecture description (.json file). After the hardware architecture description is generated, the sub-module hardware design parameters corresponding to the top-level design primitives, middle-level design primitives, and bottom-level design primitives are generated from top to bottom. After the sub-module hardware design parameters are determined, the hardware configuration description will be generated.

[0077] The hardware description intermediate representation ADIR (Architecture Description Intermediate Representation) output during the modeling process of the coarse-grained reconfigurable array includes the hardware architecture description and the hardware configuration description (.json file), which can be further used to drive the target redirection of the compilation tool.

[0078] S2, generation, is started after modeling is completed. After modeling is completed, the hardware generator is instantiated from bottom to top according to the top-level hardware design parameters, the middle-level hardware design parameters, and the bottom-level hardware design parameters, and the hardware RTL-level code is output.

[0079] The specific operation process is as follows: The top-level hardware design parameters are shown in Table 1.

[0080] Table 1 Top-level hardware design parameters

[0081]

[0082]

[0083] Table 1 shows the top-level hardware design parameters and the description and optional values ​​of each parameter. From the top-level hardware design parameters, the design parameters used to instantiate the hardware generators at each level can be directly or indirectly obtained. Among them, connect_topology and additional_connections together constitute the interconnection network of the complete processing unit array. Connect_topology represents the basic interconnection network, which can support systolic, N2N (n2n), diagonal N2N (n2n_diag), torus, and multi-hop. Additional_connections represents the additional routes added on the basis of the basic interconnection network. When connect_topology is "none", the interconnection network is uniquely determined by additional_connections.

[0084] The operation types of the arithmetic operation unit in the computing unit are shown in Table 2.

[0085] Table 2 Operation types supported by the arithmetic operation unit

[0086]

[0087] Table 2 shows that the arithmetic operation unit in this embodiment supports a total of 14 types of operations, covering arithmetic operations, bit operations and comparison operations. op0 and op1 represent two operands input to the arithmetic operation unit, where op0 represents the operand selected by the synchronizer in the calculation unit, and op1 represents another operand directly selected by the calculation multiplexer in the calculation unit. The arithmetic operation unit determines the type of operation during execution according to the calculation operation configuration information corresponding to the arithmetic operation unit in the calculation configuration register of the calculation unit, that is, the calculation operation code, to achieve reconfigurable calculation requirements.

[0088] The total bit width of the calculation multiplexer configuration bit width mux0 and mux1 of the calculation unit, the configuration bit width alu of the arithmetic operation unit, the immediate bit width const and the synchronization unit bit width sync and the calculation configuration register pe_cfg_width can be calculated by the following formula:

[0089]

[0090] const = data_width

[0091]

[0092] The configuration bit width alu of the arithmetic operation unit has nothing to do with the number of operations supported in the top-level hardware design parameters during hardware instantiation, but is only related to the number of operations supported by the arithmetic operation unit. More specifically, according to an embodiment of the present disclosure, as shown in Table 2, the number of operations supported by the arithmetic operation unit is 14, then alu=4, and the correspondence between the calculation operation code and the operation type remains unchanged.

[0093] for Figure 1 The interconnection network is of N2N type. The first column is configured as access units, and the last column is configured with input and output ports. Each processing unit is connected to the four processing units above, below, left and right. The access unit has 2 read ports and 1 write port, and the input and output interface has 2 inputs and 1 output. Then max(inputs_num) = 5. If max_delay = 3, sync = 2 can be calculated according to the above formula. The complete configuration information of the computing unit is shown in Table 3. It can be represented by the following splicing, with the right side representing the low bit and the left side representing the high bit, and the high bit is filled with 0.

[0094] Table 3 Complete configuration information of the computing unit

[0095] 20 bits, padded with 0 2bit, sync 3bit, mux1 3bit, mux0 4bit, alu 32bit, const

[0096] The configuration bit width mux of the storage multiplexer of the access unit, the storage address selection bit width sel_addr and the total bit width lsu_cfg_width of the configuration register can be calculated by the following formula:

[0097]

[0098] for Figure 1 The interconnect network is of N2N type. The first column of the processing unit array is configured as an access unit with only 1 input, so the access unit multiplexer is configured with bit width mux=0. When sram_depth=128, sel_addr=8. The lowest bit WEN indicates write enable, which is valid at low level. Therefore, the complete configuration information of the access unit is shown in Table 4. The right side indicates the low bit, the left side indicates the high bit, and the high bit is filled with 0.

[0099] Table 4 Complete configuration information of access unit

[0100] 7 bits, fill with 0 8bit, sel_addr 1bit, WEN

[0101] The hardware description intermediate representation output during the modeling process of the coarse-grained reconfigurable array includes the hardware architecture description and the hardware configuration description. The hardware architecture description is an extension of the input top-level hardware design parameters, which is used to specify the hardware's computing, storage, and routing resources. The hardware configuration description saves the one-to-one correspondence between the different multiplexer gating states, delay clock numbers, access addresses, read and write enables of the computing units and access units, and the reconfigurable working states and operation codes that need to be configured clearly. The generated hardware description intermediate representation is used to drive the subsequent retargetable compilation mapping design.

[0102] Therefore, by adopting the coarse-grained reconfigurable array and the hardware modeling generation method thereof described in the present invention, agile modeling and RTL-level code generation of the coarse-grained reconfigurable array can be completed; by designing the intermediate representation of the hardware description, the target redirection of the driver compilation tool can be realized; and the efficiency of the hardware accelerator in hardware design and software stack development can be greatly improved.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A coarse-grained reconfigurable array, characterized in that: include: The processing unit array is a reconfigurable array formed by M rows and N columns of processing units, where M is an integer greater than or equal to 1, and N is an integer greater than or equal to 1. The processing unit performs a calculation or access operation according to input operation configuration information and input data, and outputs a calculation or access result; Configuration memory, used to temporarily store operation configuration information required for execution of each node in the processing unit array; A configuration bus, used to connect the configuration memory and the processing units, and input operation configuration information to at least one processing unit; The input and output interface connects the processing unit to external devices and is used to input data to the processing unit or output the calculation results of the processing unit.

2. A coarse-grained reconfigurable array according to claim 1, characterized in that: Each column of the processing units shares a configuration bus.

3. The coarse-grained reconfigurable array according to claim 1, characterized in that: The processing unit is an access unit or a computing unit; The access unit is used to store and read data. When the coarse-grained reconfigurable array is working, the access unit stores input data before the calculation unit starts calculating, temporarily stores intermediate data when the calculation unit calculates, and stores output data after the calculation unit completes the calculation. The computing unit is used to perform computing operations according to input operation configuration information and input data, and output execution results.

4. A coarse-grained reconfigurable array according to claim 3, characterized in that: The access unit comprises: The storage configuration register is used to temporarily store the storage operation configuration information input by the configuration bus; Scratch pad memory, composed of SRAM, used to cache data; A storage multiplexer connected to adjacent computing units selects corresponding input data according to storage operation configuration information in the storage register; The logic unit is used to execute the logic of the access unit; according to the storage operation configuration information, when working in the loading state, the data is taken out from the note memory and output; when working in the storage state, the input data is stored in the note memory and the input data is directly output.

5. The coarse-grained reconfigurable array according to claim 4, characterized in that: The storage operation configuration information includes input selection information, address information of access data and access operation code.

6. The coarse-grained reconfigurable array according to claim 3, characterized in that: The computing unit comprises: A calculation configuration register, used to temporarily store the calculation operation configuration information input by the configuration bus; A calculation multiplexer is connected to adjacent calculation units and selects corresponding input data according to calculation operation configuration information in the calculation register; A synchronizer delays the input data selected by the multiplexer to synchronize the input data, and selects the corresponding data to be input to the arithmetic operation unit according to the delay selection information in the calculation operation configuration information; The arithmetic operation unit performs arithmetic operations according to the calculation operation code in the calculation operation configuration information to obtain result data; or directly outputs the input data.

7. The coarse-grained reconfigurable array according to claim 6, characterized in that: The calculation operation configuration information includes input selection information, immediate data, delay selection information and calculation operation code.

8. The coarse-grained reconfigurable array according to claim 6, characterized in that: In the computing unit, a synchronizer is only arranged after the computing multiplexer on one side, and the operation is immediately performed when the next clock arrives after the input data of the computing multiplexer on the other side arrives.

9. A method for generating hardware modeling of a coarse-grained reconfigurable array based on any one of claims 1 to 8, characterized in that: The following steps are involved: S1. Modeling, the hardware structure of the coarse-grained reconfigurable array is divided into top-level design primitives, middle-level design primitives and bottom-level design primitives. The top-level design primitive is the coarse-grained reconfigurable array. The middle-level design primitives include configuration memory, bus, computing unit and access unit. The bottom-level design primitives include computing configuration register, computing multiplexer, synchronizer, arithmetic operation unit and storage configuration register, storage multiplexer and scratch pad memory. According to the input top-level hardware design parameters, the hardware of the top-level design primitive, the middle-level design primitive and the bottom-level design primitive are modeled from top to bottom and the sub-module hardware design parameters of the top-level design primitive, the middle-level design primitive and the bottom-level design primitive are generated; Output hardware description intermediate representation, which is used to drive target redirection of compilation tools; S2. Generation. After the modeling is completed, the hardware generator is instantiated from the bottom to the top according to the top-level hardware design parameters, the middle-level hardware design parameters and the bottom-level hardware design parameters, and the hardware RTL-level code is output.

10. The method for generating hardware modeling of a coarse-grained reconfigurable array according to claim 9, characterized in that: In S1, the hardware description intermediate representation includes a hardware architecture description and a hardware configuration description.

Citation Information

Patent Citations

  • High-efficiency coarse granularity reconfigurable computing system

    CN105468568A

  • Executable file static instrumentation technical framework supporting multiple architectures

    CN112882701A

  • Coarse-grained reconfigurable array system for deep learning and calculation method

    CN115168284A

  • Instructions to configure coarse-grained reconfigurable array to execute program of data stream

    CN118922824A

  • Microcomputer for low power efficient baseband processing

    US20140137123A1

Cited By

  • Multi-core simulation method, system and equipment and storage medium

    CN120297209A