Accelerator design method based on generator modularization

Through the modular design method based on the generator, explicitly defining the interface and unified memory management and performance models, the problems of low performance and poor coupling between modules in the existing accelerator design technology are solved, and the efficient and flexible design of the accelerator is achieved.

CN116305817BActive Publication Date: 2025-06-06PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310103771.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-06-06
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

The existing accelerator design technology lacks a unified performance model and cannot handle complex memory system integration, resulting in low accelerator performance and inseparable modules, which makes it less effective.

Method used

By establishing a modular design method based on generators, explicitly defining interfaces and unified memory management and performance models, the high communication efficiency and flexibility design of the accelerator are achieved. This method includes pipeline-aware optimization algorithm and hierarchical memory management, which supports the design generation of accelerators in different architectures.

Benefits of technology

It improves the communication efficiency of the accelerator and the flexibility of the modular design, saves accelerator design time, and promotes the agile design and performance improvement of the accelerator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116305817B_ABST
    Figure CN116305817B_ABST
Patent Text Reader

Abstract

The present invention discloses an accelerator design method based on generator modularization, and establishes a two-stage process for designing an accelerator, including a generation and selection stage, and an integration stage; including pre-selecting the modules required for generator generation, and obtaining the factors and parameters required for generator generation to specify optimized functions and constraints; and then reducing the impact of the integrated modules on the accelerator performance through a hierarchical memory management method. The present invention develops an accelerator by integrating and building generator modules, so that the accelerator has high communication efficiency, and the modular design has flexibility and high productivity, which can promote the efficiency and performance of agile design of accelerators in specific fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of electronic information technology, relates to module integration and accelerator design technology, and in particular to an accelerator design method based on generator modularization and design space exploration. Background Art

[0002] A generator is a tool that uses high-level and simple descriptions to automatically generate low-level and detailed target hardware. Parameterized IP, high-level synthesis HLS, and hardware construction languages ​​(such as chisel) are all included in the generator design technology.

[0003] In existing approaches, creating chip generators, to create customized heterogeneous designs, it is best to start with a flexible homogeneous architecture because it makes verification and software easier. To enable customization, the internal structure of the architecture must be built from highly parameterized and extensible modules so that application designers can easily shape the components to create heterogeneous systems that suit specific application needs. The design technology of generators is to initially create these flexible and customizable modules.

[0004] In generator design, design space exploration plays an important role in optimization and trade-offs, among which performance modeling is a key element in design space exploration and focusing on different indicators for different applications. The modular design approach includes three stages: module decomposition, module implementation, and module integration. Module abstraction includes module interface and module information description, timing and loaded data, and registered inputs and outputs. With the development of automated module implementation, these concepts have been transformed into interface definition, performance model, and memory management.

[0005] However, existing accelerator integration technologies can only integrate IP that conforms to standard interfaces and can only perform simple integration. There is no unified performance model to guide the optimization goals of integration, and it is also unable to handle complex memory system integration.

[0006] The defects of existing accelerator design technologies are: 1. Most of them adopt manual description methods and lack an accurate module performance model. Moreover, the performance models between generators are not unified, and most of them are designed to work alone; 2. For the interface between modules, there are some standards and extensions that can be used to describe the information of integrated IP, but there is no available interface to tightly couple the modules generated by the generator 3. Memory management: The memory system is usually designed manually, but because the memory system is generated at the same time as the generated module, this may cause conflicts between the hierarchical structure of the global shared memory system and the local private memory, affecting the performance of the accelerator in memory-constrained application scenarios. Therefore, the existing accelerator design technology has low effectiveness due to the inconsistent performance model and the inability to tightly couple between modules, and the generated accelerator has low performance due to memory constraints. Summary of the invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides an accelerator design method based on generator modularization (named Weave in the present invention), including explicitly defining interfaces and establishing unified memory management and performance models, and developing accelerators by integrating and building generator modules, so that the accelerator has high communication efficiency, and the modular design is flexible and productive, which can promote the efficiency and performance of agile design of accelerators in specific fields.

[0008] Accelerators in common fields can be divided into two architectures, instruction-based accelerators and data flow-based accelerators. The performance of instruction-based accelerators is determined by instructions including data dependency and branch control, while it is determined by the rate of data consumption and generation in data flow-based accelerators. In the present invention, the module of modular abstract design includes the interface, memory access model, timing model and resource model of the generator. The generator modular design method eliminates the differences between different accelerator architectures by unifying the interface and performance analysis model with modular abstraction, thereby supporting the design and generation of accelerators with different architectures.

[0009] The technical solution provided by the present invention is:

[0010] An accelerator design method based on generator modularization is implemented by establishing a two-stage process (Weave). First, the modules required for generator generation are pre-selected, and the factors and parameters required for generator generation are obtained to specify the optimized functions and constraints. The Weave two-stage process includes a generation and selection stage and an integration stage; the Weave two-stage integration process is used to design accelerators, including instruction-based accelerators and data flow-based accelerators. The Weave two-stage integration process includes the following steps:

[0011] 1) In the generation and selection stage, the design architecture of the target accelerator to be designed, i.e., the optimal module set of the target accelerator, is obtained by designing a pipeline-aware optimization method;

[0012] In the generation and selection stage, a pipeline-aware optimization algorithm is designed to coordinate the contradiction between the design space of the accelerator and the accelerator system performance (model), so as to retain as many optional modules as possible and avoid excessive development time by exploring too large a design space; the pipeline-aware optimization algorithm is used to determine the effective sub-design space of the accelerator and iteratively explore the space to obtain the optimal module set for the accelerator design.

[0013] The generator design method is to design the accelerator by specifying optional modules generated by the designer and selecting the best combination according to the optimization function and resource constraints. On the one hand, for systematic considerations, it is necessary to retain as many optional modules as possible to avoid missing the best choice as much as possible. On the other hand, retaining as many optional modules as possible may lead to excessive exploration of the accelerator design space, resulting in too long accelerator development time. Therefore, a pipeline-aware optimization algorithm is designed to solve this problem.

[0014] The pipeline-aware optimization algorithm includes the following steps:

[0015] 2) In the integration stage, a hierarchical memory management method is designed to group the memories of the multiple accelerator optimal modules generated in step 1), and manage the global on-chip and off-chip data access of the integrated circuit of the chip, thereby completing the design of the target accelerator; the hierarchical memory management method includes the following steps:

[0016] 21) Define the interface of the target accelerator; including the interface type and interface protocol of the target accelerator;

[0017] 22) grouping the memories of multiple accelerator optimal modules generated in step 1);

[0018] 23) Manage the chip's integrated circuit global on-chip and off-chip data access;

[0019] After optimization in the generation and selection phases, the architecture of the target accelerator design is obtained. The designed architecture is used as the main framework of the target accelerator. The interface is a key element for interconnection in module integration, and the interface of the target accelerator is defined in the modular abstraction of the generated module. According to the interface type and protocol used in the accelerator design, in order to properly connect the modules, we classify the interfaces by the control signals (information) and data transmission of the interfaces used in the accelerator design, that is, including control signal interfaces and data transmission interfaces. Among them, the interface type is the type of connection mode, and the interface protocol (such as AXI-4 protocol) is the mode of data transmission. For the architecture of tightly coupled accelerators, the memory subsystem is kept in a harmonious structure, because it will affect the performance when there are conflicts between different modules. In addition, a unified global memory access module will reduce the development cost of managing the global. Therefore, based on the abstract interface and memory access, we propose a hierarchical memory management method to reduce the impact on accelerator performance when integrating modules.

[0020] The present invention designs memory management from three levels:

[0021] Accelerator level: The present invention handles memory access between on-chip and off-chip at the accelerator level, treating the entire accelerator as a general DMA module. The key design at this level is the integration of DMA modules of different types of functional modules. The support of different access modes from the generator is retained, the general DMA module is implemented and the global memory portion is provided for data cache.

[0022] Inter-module level: When an independent global DMA module and global memory are implemented, the original DMA module is replaced and connected with the global memory access. In addition, there are multiple modes of data transmission requests between modules. The present invention implements a suitable memory module to cache and reorganize data by designing a suitable memory management method at the inter-module level, such as BRAM (block random-access memory) for batch processing of memory information based on modular abstraction, and FIFO for stream processing of memory information based on modular abstraction.

[0023] Internal module level: The present invention does not change the local memory inside the module and treats it as the original private type.

[0024] The hierarchical memory management method includes the following steps:

[0025] A) Handle memory access between on-chip and off-chip at the accelerator level, treating the entire accelerator as a general-purpose Direct Memory Access (DMA) module.

[0026] B) Process data transfer requests in various modes at the inter-module level, design appropriate memory management methods at the inter-module level for caching and reorganizing data, and implement appropriate memory modules to cache and reorganize data.

[0027] C) At the internal module level, local memory is not altered and is preserved as its original private type.

[0028] Through the above steps, an accelerator design method based on generator modularization is implemented for agile design of accelerators.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] The present invention provides an accelerator design method based on generator modularization. An accelerator is developed by integrating generated modules. The generated accelerator improves communication efficiency, enhances the flexibility of accelerator component modules, saves accelerator design time, and effectively promotes agile design of accelerators in application fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1Schematic diagram of the modules of the two accelerator architectures.

[0032] Figure 2 It is a flowchart of the method of the present invention.

[0033] Figure 3 It is a two-stage integration flow chart. DETAILED DESCRIPTION

[0034] The present invention will be further described below by way of embodiments in conjunction with the accompanying drawings, but the scope of the present invention is not limited in any way.

[0035] Accelerators in common fields can be divided into two architectures: instruction-based accelerators and data flow-based accelerators. The performance of instruction-based accelerators is determined by instructions including data dependency and branch control, while the performance of data flow-based accelerators is determined by the rate at which data is consumed and generated. The modules included in the two architectures of accelerators are as follows: Figure 1 As shown. In the present invention, the module of modular abstract design includes the interface, memory access model, timing model and resource model of the generator. The generator modular design method eliminates the differences between different accelerator architectures by unifying the interface and performance analysis model with modular abstraction, thereby supporting the design and generation of accelerators with different architectures.

[0036] Accelerators can be divided into two architectures: instruction-based architecture and data flow-based architecture. Specific modules are as follows: Figure 1 As shown. The accelerator design method based on generator modularization can be implemented by establishing a two-stage integration process (Weave). First, the modules required for generator generation are pre-selected, and the factors and parameters required for generator generation are obtained to specify the optimized functions and constraints. The Weave two-stage integration process includes the generation and selection stage and the integration stage; the Weave two-stage integration process is used to design accelerators, including instruction-based accelerators and data flow-based accelerators.

[0037] When the present invention is implemented, the Weave two-stage integration process includes the following steps:

[0038] As shown in Table 1, the modular design module of the accelerator Weave designed by the present invention includes an interface, a memory access model, a timing model, and a resource model. The present invention adopts an optimization method and a memory management method. Figure 2 As shown, the present invention designs a two-stage integration process based on two algorithms, builds a timing model and a resource model to guide the generator and the design space explorer to generate and select the most effective accelerator pipeline module, and designs a hierarchical memory management method based on the memory access model and interface. Figure 2 The present invention proposes a two-stage integrated process for the optimization problem in the problem statement of Figure 3shown.

[0039] Table 1 Module names, definitions and contents of the generator-based modular accelerator design

[0040]

[0041] First, the pipeline-aware optimization algorithm is introduced in detail. From the perspective of design space, the execution frequency and II initiation interval are affected by the modules with comparison functions, and the delay nature is cumulative. If we set a suitable limit Freq sys and II sys , we only need to explore the generated modules with delays that satisfy the constraints of suitable boundaries (suitable boundaries Freqsys and IIsys). In addition, the choice of boundaries defines a valid sub-design space, which is intuitively important for the search process and the optimal solution. The present invention sets a boundary selection factor, and the exploration generated module method performed on the sub-design space will improve the search efficiency through appropriate boundary selection (i.e. setting suitable boundaries Freqsys and IIsys). Considering that the characteristics between different generators and modules vary greatly, a model-independent learning algorithm is suitable. In addition, flexible trade-offs in performance (throughput and latency) and resource usage should be considered according to different application scenarios. Based on these considerations, the pseudo code of the pipeline-aware optimization algorithm for the accelerator design of the present invention is as follows.

[0042] Algorithm 1 Pipeline-aware optimization algorithm

[0043] Requirements: factors, generators, constraints

[0044] Guarantee: optimal sets of modules

[0045]

[0046] The pipeline-aware optimization algorithm for accelerator design is used to determine the effective sub-design space of the accelerator and iteratively explore the space to obtain the optimal module set for the accelerator design. The pipeline-aware optimization algorithm for accelerator design is an active learning process, which determines the effective sub-design space of the accelerator and iteratively explores the space through the pipeline-aware optimization algorithm. The algorithm requires throughput and delay factors α and β, generation factor G, and resource capacity B. In addition, the present invention introduces three boundary selection factors, ratio_f, decay_f and decay_ii. Among them, ratio_f and decay_f are frequency scaling factors and updated frequency scaling factors for boundary calculation, respectively, and decay_ii is the update factor of the upper boundary of II. The algorithm obtains the initial state and settings of G according to multiple generation factors G (lines 1-5 of the pseudo code). After multiple iterations (lines 6-18), the generation function calls G to generate multiple modules under different constraints, which include different trade-offs of frequency constraints and startup interval II. In addition, two additional parameters, lwb_opt_freq and upb_opt_ii, are used to help G generate modules that are more likely to be selected. These generated modules form a valid sub-design space, which is used as the design space explored by the selection function. In the selection function, the DSE engine (Design Space Exploration) will select valuable choices for the optimal strategy. This process attempts to solve the NP-hard combinatorial optimization problem. Therefore, we use the classic model-free simulated annealing algorithm as the DSE engine. At the same time, in order to support multi-objective optimization (resource usage, throughput and latency), the goal of the DSE engine is to explore the Pareto frontier as close to the real Pareto frontier as possible. Therefore, we choose IPV (the Improvement of Pareto Hypervolume) as the optimization goal. The use of IPV converts the sum of different types of factors in multi-objective optimization into a multiplication relationship, eliminating the influence of factors in classical multi-objective optimization. The Selection function is used to explore the Pareto optimal set under the objective function in the design space. The output of the Selection function is a series of Pareto sets under the current sub-design space. According to the number and quality of the output set, the boundary selection factor and the cumulative optimal module set M are updated. At the end of the process, we select the optimal module set M_opt by calculating the Pareto frontier of the cumulative optimal module set M. In this scheme, we improve the overall DSE search efficiency by generating sub-design spaces, and also obtain good quality of generated modules through Pareto-driven exploration.

[0047] In the generation and selection phase, the design architecture of the target accelerator to be designed, i.e., the optimal module set of the target accelerator, is obtained by designing a pipeline-aware optimization method;

[0048] In the generation and selection stage, a pipeline-aware optimization algorithm is designed to coordinate the contradiction between the design space of the accelerator and the accelerator system performance (model), so as to retain as many optional modules as possible and avoid excessive development time by exploring too large a design space; the pipeline-aware optimization algorithm is used to determine the effective sub-design space of the accelerator and iteratively explore the space to obtain the optimal module set for the accelerator design.

[0049] The generator design method is to design the accelerator by specifying optional modules generated by the designer and selecting the best combination according to the optimization function and resource constraints. On the one hand, for systematic considerations, it is necessary to retain as many optional modules as possible to avoid missing the best choice as much as possible. On the other hand, retaining as many optional modules as possible may lead to excessive exploration of the accelerator design space, resulting in too long accelerator development time. Therefore, a pipeline-aware optimization algorithm is designed to solve this problem.

[0050] The pipeline-aware optimization algorithm for accelerator design includes the following steps:

[0051] 1) obtaining the initial state and setting of the accelerator according to the multiple generation factors G;

[0052] 2) Perform multiple iterations to generate multiple modules under different constraints according to the generation factor G. Each time a different constraint is given and the generator is called, it is an iteration process;

[0053] 3) The generated modules form an effective sub-design space of an accelerator; the DSE engine is used to select the optimal strategy; the classic model-free simulated annealing algorithm is used as the DSE engine;

[0054] The goal of the DSE engine is to make the Pareto frontier as close to the real Pareto frontier as possible;

[0055] IPV is selected as the optimization objective. The use of IPV converts the sum of different types of factors in multi-objective optimization into multiplication.

[0056] 4) Update the boundary selection factor and the cumulative optimal module set M;

[0057] The limitation of the boundary selection factor will affect the speed of the Generation function process and the quality of the generated module M; the appropriate selection factor can generate better modules faster, so it is necessary to iteratively update the boundary selection factor to achieve the optimal one.

[0058] 5) Select the optimal module set by calculating the Pareto frontier of the cumulative optimal module set M.

[0059] The pipeline-aware optimization algorithm specifically:

[0060] The initial state and setting of the target accelerator to be designed are set by setting the generation factor G;

[0061] The generation factor (G) is a parameter of the generator, which determines the initial state and settings of the accelerator.

[0062] Generate multiple modules of different accelerators by calling G iteratively multiple times; organize the multiple accelerator modules generated each time into an effective sub-design space, and obtain a series of Pareto sets under the sub-design space;

[0063] Design a generation function, which generates multiple modules of different accelerators by calling G iteratively multiple times; the multiple accelerator modules generated each time form an effective sub-design space, which is represented as a Selection function; the output of the Selection function is a series of Pareto sets under the sub-design space. For a given set of optimal solutions, if the solutions in the set are mutually non-dominated, that is, the solutions in the set are not in a dominating relationship, this optimal solution set is the Pareto Set.

[0064] Obtain the optimal module set for the target accelerator to be designed;

[0065] For example, there are many ways to implement the matrix multiplication calculation module (multiple modules are generated), and the one with the smallest resource usage that meets the frequency and II requirements is the optimal matrix multiplication module; the modules with other functions are similar; selecting the optimal module set is to calculate the boundary of the Pareto set;

[0066] Update the boundary selection factors ratio_f, decay_f and decay_ii and the cumulative accelerator optimal module set M. At the end of the update process, the optimal module set M_opt is selected by calculating the Pareto frontier of the cumulative optimal module set M. Specifically, the Pareto frontier calculation algorithm is used, that is, the boundary is formed according to the frequency of the module set, II and resource usage.

[0067] By setting boundary selection factors, the exploration-generated modular approach performed on the sub-design space will improve the search efficiency through boundary selection (i.e., setting appropriate limits of Freqsys and IIsys).

[0068] The boundary selection factors include: frequency scaling factor for boundary calculation, updated frequency scaling factor, and update factor of the upper boundary of the startup interval (II); the initial value setting generally adopts empirical values, for example, 0.75, 0.1, 0.1 can be taken; the boundary selection factor value is generally not updated, and the two variables of boundary selection (boundary Freqsys and IIsys) are updated.

[0069] Specific process: The initial state and settings of the target accelerator are obtained according to multiple generation factors G. After multiple iterations, G is called to generate multiple modules under different constraints (including frequency constraints and startup interval constraints). Two additional parameters (the lowest possible frequency limit and the highest possible II limit, the corresponding parameters are lwb_opt_freq and upb_opt_ii) are used to help G generate modules that are more likely to be selected. These generated modules constitute an effective sub-design space; the optimal strategy is selected by designing a design space exploration engine DSE engine. The present invention adopts the classic model-free simulated annealing algorithm as the DSE engine. At the same time, in order to support multi-objective optimization (resource utilization, throughput and latency), the goal of the DSE engine is to explore the Pareto frontier as close to the real Pareto frontier as possible, specifically selecting IPV (the Improvement of Pareto Hypervolume) as the optimization target.

[0070] In the accelerator integration stage, the obtained modules are integrated, and the generated interfaces and memories are connected to obtain the accelerator. The hierarchical memory management in the accelerator integration stage is introduced in detail below. An inappropriate integration method will lead to low efficiency of the pipeline, and the interface and memory management methods are key elements in the integration stage. In the integration stage, global memory management is essential for the resource utilization and timing performance of the accelerator. We propose a hierarchical memory management method to divide the memory of the grouped modules into shared and private, and manage on-chip and off-chip data access in the global view. The present invention designs memory management from three levels:

[0071] Accelerator level: This level handles memory access between on-chip and off-chip, treating the entire accelerator as a general DMA module. The key design of this level is the integration of DMA modules of different types of functional modules. We retain the support of different access modes from the generator, implement the general DMA module and provide a global memory section for data cache.

[0072] Inter-module level: When an independent global DMA module and global memory are implemented, the present invention replaces the original DMA module and connects with the global memory access. In addition, there are multiple modes of data transfer requests between modules. We implement appropriate memory modules to cache and reorganize data, such as BRAM for batch processing of memory information based on modular abstraction, and FIFO for stream processing of memory information based on modular abstraction.

[0073] Internal module level: The present invention does not change the local memory inside the module and treats it as the original private type.

[0074] Under the guidance of this hierarchical memory management design, the present invention not only integrates modules with little impact, but also increases data reuse and reduces the cost of memory access. From the perspective of memory management implementation, we designed a general DMA module with off-chip memory, using the AXI-4 protocol to support active random and sequential access modes. At the same time, the size of the global shared memory is determined by the available remaining resources and the module size defined in the memory information. The hierarchical design supports the design space exploration of shared memory size and partitioning methods, but it is not the focus of the present invention. For this hierarchical memory management design, the present invention implements the key memory modules of the accelerator in the integration stage, allowing all modules to work correctly.

[0075] It should be noted that the purpose of publishing the embodiments is to help further understand the present invention, but those skilled in the art can understand that various substitutions and modifications are possible without departing from the scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the contents disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined in the claims.

Claims

1. An accelerator design method based on generator modularization establishes a two-stage process for accelerator design, including a generation and selection stage and an integration stage; firstly, the modules required for generator generation are pre-selected, and the factors and parameters required for generator generation are obtained to specify the optimized functions and constraints; then, the impact of the integrated modules on the accelerator performance is reduced through a hierarchical memory management method; the steps include: 1) In the generation and selection stage, the design architecture of the target accelerator to be designed, i.e., the optimal module set of the target accelerator, is obtained by designing a pipeline-aware optimization method; The pipeline-aware optimization algorithm includes the following steps: 11) Setting the initial state and settings of the target accelerator to be designed by setting the generation factor G; 12) Generate multiple modules of different accelerators by calling G iteratively multiple times ; The modules of multiple accelerators generated each time are combined into an effective sub-design space, and a series of Pareto sets under the sub-design space are obtained; 13) Obtain the optimal module set of the target accelerator to be designed by calculating the boundary of the Pareto set; Specifically, the Pareto boundary calculation algorithm is used to set the boundary selection factor, and the generated modules form an effective sub-design space through boundary selection, that is, setting appropriate boundaries Freqsys and IIsys; the exploration and generation module method is executed on the sub-design space; the Pareto boundary calculation algorithm takes the improvement of Pareto hypervolume IPV as the optimization goal and is executed on the sub-design space; The boundary selection factors include: frequency scaling factor for boundary calculation, frequency scaling factor for update, update factor for upper boundary of start interval; the initial value setting adopts empirical value; The generated modules form an effective sub-design space. By constructing the design space exploration engine DSE, the boundary of the Pareto set is calculated to obtain the optimal module set of the target accelerator to be designed. 2) In the integration stage, a hierarchical memory management method is designed to group the memories of the multiple accelerator optimal modules generated in step 1), and manage the global on-chip and off-chip data access of the integrated circuit of the chip, so as to complete the design of the target accelerator; the hierarchical memory management levels include: accelerator level, inter-module level and internal module level; the hierarchical memory management method includes the following steps: 21) defining the interface of the target accelerator, including the interface type and interface protocol of the target accelerator; 22) grouping the memories of multiple accelerator optimal modules generated in step 1); 23) Manage the chip's integrated circuit global on-chip and off-chip data access; After optimization in the generation and selection stages, the architecture of the target accelerator design is obtained; Step 22) grouping includes: A) Handle memory access between on-chip and off-chip at the accelerator level, treating the entire accelerator as a general-purpose direct memory access DMA module; B) Processing data transfer requests of various modes at the inter-module level, designing the memory management method at the inter-module level for caching and reorganizing data, and implementing appropriate memory modules to cache and reorganize data; C) At the internal module level, local memory is not changed and is kept as its original private type; Through the above steps, an accelerator design method based on generator modularization is implemented for agile design of accelerators.

2. The accelerator design method based on generator modularization as claimed in claim 1, Its characteristics are: The accelerator includes an instruction-based accelerator and a data flow-based accelerator.

3. The accelerator design method based on generator modularization as claimed in claim 1, Its characteristics are: Step 12) specifically includes: Design a generation function, which generates multiple modules of different accelerators by calling G multiple times through iterations; The multiple accelerator modules generated each time are combined into a valid sub-design space, which is represented as a Selection function; The output of the Selection function is a series of Pareto sets under the sub-design space.

4. The accelerator design method based on generator modularization as claimed in claim 1, Its characteristics are: The memory module cache and data reorganization are realized through the inter-module level, including: using BRAM for batch processing of memory information based on modular abstraction, and using FIFO for stream processing of memory information based on modular abstraction.

Citation Information

Patent Citations

  • Multi-target swarm intelligence algorithm parallel optimization method based on cloud computing

    CN113010316A

  • Staged efficient constraint optimization method and device based on Kriging agent model

    CN114595577A