Automatic configuration of pipeline module in electronics system
The automatic configuration of pipeline modules in SoCs through an optimized RTL generation process addresses the inefficiencies of manual RTL generation, reducing time and computational intensity while ensuring timely constraint compliance.
Patent Information
- Application Number
- JP2025063102
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-11
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
AI Technical Summary
Conventional RTL generation for SoCs with reconfigurable and parameterizable hardware components is time-consuming and requires expensive iterations due to manual configuration of pipeline modules, with timing and area constraints often not identified until late in the synthesis flow.
A method for automatically configuring pipeline modules by generating an optimized RTL description using a database of elements and configurable components, prioritizing timing and area optimization through a search algorithm that reduces register usage while meeting constraints.
Significantly reduces RTL generation time from hours to minutes, enhances predictability, and ensures timely identification of configuration issues, allowing for efficient synthesis without extensive computational overhead.
Smart Images

Figure 2025102988000001_ABST
Abstract
Description
Technical Field
[0001] Field of the Invention The present invention generally relates to the design of electronic systems, and more specifically to the automatic optimization of pipeline configurations.
Background Art
[0002] Background Register Transfer Level (RTL) typically refers to a design abstraction that models digital circuits as the flow of data signals between hardware registers and the logical operations performed on those signals. That is, it describes how data is manipulated and moved between registers. RTL can be used in the design and verification flow of electronic systems. For example, RTL can be used in the design and verification flow of a System-on-Chip (SoC).
[0003] Conventional RTL generation for an SoC is extremely time-consuming for systems that utilize reconfigurable and parameterizable hardware components. For example, an initial RTL description is generated and sent to the SoC integrator to determine whether certain constraints are met. If there is a violation of the constraints, a new RTL description is generated and verification is repeated. It may take several hours to perform multiple iterations. The problem is not only the time required for RTL generation but also the time to generate the final acceptable RTL. At present, designers or users manually create the configuration of the pipeline module and generate RTL. If timing and / or area criteria / constraints are not met, problems in the configuration settings may not be found until the latter half of the synthesis flow. As a result, the user has to go back (expensive iteration) and change the pipeline configuration and keep trying this process until it works. Therefore, there is a need for a system and method for automatically configuring pipeline modules in an electronic system.
Summary of the Invention
Means for Solving the Problems
[0004] Summary Disclosed are various embodiments and methods for automatically configuring pipeline modules in an electronic system. The method implemented by embodiments of the present invention includes generating a complete register transfer level (RTL) description of an electronic device system. The method includes generating an optimized pipeline configuration from an input that includes a database of RTL elements and a list of configurable pipeline components, and generating a complete RTL description with pipeline components configured according to the optimized pipeline configuration. Generating the configuration includes performing a search for a configuration that optimizes area and timing. As disclosed herein, various advantages result from the embodiments and methods according to the present invention. Further, the methods disclosed herein are general and are not limited to where the pipeline belongs or is located.
[0005] Brief Description of the Drawings To more fully understand the present invention, reference is made to the accompanying drawings or figures. The present invention is described in accordance with the aspects and embodiments in the following description with reference to the drawings or figures, in which like numbers represent the same or similar elements. It is understood that these drawings should not be regarded as limiting the scope of the present invention, and the aspects and the best mode presently understood of the present invention are described in further detail with the use of the accompanying drawings.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0007] Detailed Description In the following detailed description, reference is made to the accompanying drawings which form a part hereof and which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention, and other embodiments may be utilized and mechanical, procedural, and other changes may be made without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the invention is defined only by the appended claims and includes all equivalents thereof given by such claims.
[0008] Throughout this specification, references to "one embodiment", "an embodiment", "an example", or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in one embodiment", "in an embodiment", "an example", or "an example" in various places throughout this specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, databases, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Additionally, the drawings provided herein are for illustrative purposes for those skilled in the art and it is to be understood that the figures are not necessarily drawn to scale. In one or more embodiments or examples, they can be combined in any suitable combination and / or partial combination. Further, it is to be understood that the drawings provided in this specification are for the purpose of explanation to those skilled in the art, and the figures are not necessarily drawn to scale.
[0009] As used herein, "source", "master", and "initiator" refer to similar intellectual property (IP) blocks, modules, or units, and these terms are used interchangeably within the scope and embodiments of the present invention. As used herein, "sink", "slave", and "target" refer to similar IP modules or units, and these terms are used interchangeably within the scope and embodiments of the present invention. As used herein, a transaction can be a request transaction or a response transaction. Examples of request transactions include write requests and read requests.
[0010] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each step in the flowchart or block diagram can represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that each step in the block diagrams and / or flowchart diagrams, as well as combinations of steps or blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions. Further, these computer program instructions may be stored in a tangible computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instruction means for implementing the function / operation specified in the steps or blocks of the flowchart and / or block diagram.
[0011] Referring to FIG. 1, FIG. 1 shows a computer-implemented method for generating a complete RTL description of an electronic device system. The electronic device system utilizes reconfigurable and / or parameterizable hardware blocks. In FIG. 1, the blocks include reconfigurable pipeline components. Examples of electronic device systems include, but are not limited to, artificial intelligence (AI), system-on-chip (SoC), and system-in-package (SiP).
[0012] In step 100, an “optimized” pipeline configuration is generated. The optimized configuration is generated from an input that includes a database of RTL elements and a list of reconfigurable and / or parameterizable hardware components. The list of configurable pipeline elements identifies different configuration modes for each of the pipeline elements. The list refers to the target pipeline instances within the system being designed. According to some aspects and embodiments of the present invention, the target list can include all possible pipeline instances or only a subset of the instances that exist within some specific subsystems. Logical operation and timing path information are known for each configuration mode. The terms “timing path” and “timing arc” are related in that a timing arc is one of the components of a timing path. A timing arc refers to a path between ports of the same library component. A timing path of a path that crosses multiple instances of a library component.
[0013] Generating an optimized configuration involves performing a search for a configuration that optimizes area and timing. According to some aspects and embodiments of the present invention, the process prioritizes timing over area. In the search, the process reduces the number of registers (and thus saves area) as long as the timing constraints are met. When the process hits a pipeline configuration on a path that leads to a timing violation, it stops reducing the slack on the path and uses the last configuration without a violation. Thus, according to some aspects and embodiments of the present invention, some optimized configurations are not necessarily optimal and are not necessarily the only possible configurations. There may be multiple configurations that satisfy both the timing constraints and the area constraints. However, the search quickly finds a configuration that balances timing and area.
[0014] In step 110, a complete RTL description is generated with pipeline components configured according to an optimized pipeline configuration. As used herein, a complete RTL description refers to an RTL description that is accurately synthesized from a very large library of primitives. The method of FIG. 1 has several advantages over conventional methods of generating a complete RTL description. The method of FIG. 1 has greater predictability and significantly reduces the generation time of a complete RTL generation that meets timing and area constraints. The method of FIG. 1 is not as computationally intensive. A complete RTL description can be generated only once without modification. Instead of taking hours to generate, a complete RTL description can be generated in a short time, such as a few minutes.
[0015] The method of FIG. 1 is not limited to any particular intellectual property (IP) block, but is particularly useful for configurable pipeline components. In the following paragraphs, examples of configurable pipeline components, as well as examples of computing platforms and computer-implemented methods for generating an optimized pipeline configuration are described.
[0016] Referring now to Figure 2, Figure 2 shows an example of the architecture of a configurable pipeline component 210. The component 210 includes a control unit (Ctl) 220 and a data unit (dp) 230. Here, specific signals related to the pipeline component 210 will be described in the context of the configuration mode.
[0017] Referring further to Figure 3, Figure 3 shows examples of various configuration modes of the pipeline component 210. This particular example shows four modes, namely P00, P01, P10, and P11.
[0018] Mode P00 reflects the transparent or "invalidation" mode of the pipeline component 210. There are timing paths between out_ready and in_ready, and between in_valid and out_valid. There are no timing paths that end within the pipeline component or start within the pipeline component. Mode P00 has no logic cost.
[0019] Mode P01 has a timing path between out_valid and in_valid, but no timing path between in_ready and out_ready. All timing paths entering from in_ready end within the pipeline module. All timing paths originating from out_ready start within the pipeline component.
[0020] Mode P10 has no timing path between out_valid and in_valid. However, it has a timing path between in_ready and out_ready.
[0021] Mode P11 reflects the full activation mode. Mode P11 has no timing path between out_ready and in_ready and no timing path between in_valid and out_valid. Mode P11 has the highest logic cost.
[0022] The pipeline component 210 can be characterized for each mode of configuration by a look-up table (LUT). For each mode, the path between the output port and the register, the path from the input port to the register, and the path between the input and the output are described. According to various aspects and embodiments of the present invention, the configurable pipeline element 210 has the same port interface regardless of its mode. According to some aspects and embodiments of the present invention, it is important that the relaxation-based algorithm has the same port interface for the pipeline instance in order to prevent the design from having to be resynthesized every time different implementation modes of the pipeline are tried.
[0023] In the LUT, the modes are preferably sorted in descending order of the number of registers used. Mode P11 is considered first because it is the mode in which the most registers are enabled. Mode P00 is considered last because it is the mode with the fewest enabled registers. This order is called the "relaxation order". For example, the modes in FIG. 3 are considered in the following order: mode P11 → mode P01 → mode P10 → mode P00. Moving from mode P11 and towards mode P00 is called "progressive relaxation".
[0024] Referring now to FIG. 4A, FIG. 4A depicts LUTs 410, 420, and 430 that characterize an example of a pipeline component having three modes, namely, disabled, Mode 1, and Mode 2, according to various aspects and embodiments of the present invention. Each path within each of the LUTs 410, 420, and 430 is described in terms of registers and combinational logic. The path type represents the type of timing path. According to some aspects and embodiments of the present invention, there are multiple types of paths: 1) Pi2Po is a direct path between an input port and an output port, 2) Pi2Reg is a path between an input port and a register, and 3) Reg2Po is a path from a register to an output port. The information within LUT 420 exists for documentation / reporting purposes and is not required by the search algorithm.
[0025] Referring now to FIGS. 4B and 4C, FIG. 4B depicts a pipeline component configured in Mode 1 according to various aspects and embodiments of the present invention. Three registers (Reg1, Reg2, and Reg3) are utilized. FIG. 4C depicts a pipeline component configured in Mode 2 according to various aspects and embodiments of the present invention. Two registers are utilized. Disabled (not shown) utilizes a minimum number (zero) of registers, and Mode 1 utilizes the most registers. Thus, the order of relaxation is Mode 1 → Mode 2 → Disabled.
[0026] Referring now to FIG. 5, FIG. 5 depicts a system 510 that includes a plurality of configurable pipeline components 520, 530, 540, and 550 between block instances B1, B2, B3, and B4. According to some aspects and embodiments of the present invention, B1, B2, B3, and B4 are instances of any general-purpose data processing module. FIG. 5 emphasizes the fact that it may be necessary to configure a pipeline between these blocks to assist the system in operating at the required frequency.
[0027] The pipeline components of FIGS. 2 - 4C, and the system 510 of FIG. 5 are provided only to facilitate understanding of the following platform and method. Pipeline components having different architectures and modes and LUTs can also be utilized.
[0028] Referring now to FIGS. 6 and 7, FIGS. 6 and 7 illustrate an example of a method of using a module 610 to generate an optimized pipeline configuration for an electronic device system. A finite set of pipeline configuration options is available to the system.
[0029] FIG. 6 shows input data processed by module 610. The input data includes a complete RTL description 620 of a design according to various aspects and embodiments of the present invention. The RTL description (block 620) in the input includes all components within the system being designed prior to synthesis. The pipeline modules are not yet configured. Thus, the design description includes all components and has not yet been synthesized, or the pipeline components have not yet been configured.
[0030] The input data further includes a list 630 of existing configurable pipeline components within the RTL design, including their names and locations within the design, and their parameters used to configure them, and the values of the existing parameters for each. According to some aspects and embodiments of the present invention, the list of configurable pipeline elements (block 630) refers to pipeline modules and their instantiations within the system being designed. Configurable pipeline components are instantiated as modules within the RTL description.
[0031] The input data further includes logic operation and timing path information 640 for each configuration mode according to various aspects and embodiments of the present invention. This information can be provided by a LUT that sorts the configuration modes by their relaxation order.
[0032] The input data further includes synthesis primitives 650 and 660 regarding delay and area. These synthesis primitives include a basic set such as logic gates and flip - flops. These primitives are mapped to the RTL description for calculating area and delay.
[0033] Figure 6 further shows the output data generated by module 610. The output data includes a report 670 regarding the timing paths based on the combinational delay paths extracted between I / Os, between I / O and register, and between registers. According to some aspects and embodiments of the present invention, report 670 provides the evaluated timing paths after all the pipelines are configured.
[0034] The delay is synthesized according to various aspects and embodiments of the present invention for their delay through logic primitives rather than for the wires connecting them. This greatly simplifies the logic synthesis process because it does not require physical information about how components are arranged on the system. The output data further includes a report 680 of area numbers for each cell instance base. According to some aspects and embodiments of the present invention, the cell instance base refers to primitive cells, and since the final synthesis with optimization is not performed at this stage, it includes how many instances exist in the design, such as the number of gates, the number of muxes, the number of registers, etc., and reports only the area numbers regarding primitive instances (primitive cells).
[0035] The output data finally further includes, most importantly, a report 690 of all the configured pipeline components. Report 690 is for configurable pipeline components For each component, it includes the values of its respective configuration parameters. The information within this report 690 is used to generate a complete RTL description.
[0036] Referring further to FIG. 7, FIG. 7 shows an example of a method for generating an optimized pipeline configuration. At step 710, user options are provided for selecting all pipeline components to be considered for configuration. Not all configurable pipeline components are configured. For example, a user may want to retain the user-explicit configuration of a pipeline component. This allows the user to implement their own ideas regarding the configuration space, such as reusing blocks from a previous version of the system. As another example, certain pipeline components may be optional, and some of those optional components may not be selected. All selected pipeline components are considered for configuration.
[0037] At step 720, all pipeline components considered for configuration are fully enabled according to various aspects and embodiments of the present invention. A configurable pipeline component is considered to be fully enabled when configured in the mode with the most registers. A fully enabled pipeline component achieves the best timing but utilizes the most area.
[0038] In block 730, a baseline RTL description having a fully enabled pipeline is synthesized to generate a set of flow paths that achieve the best timing but utilize the maximum area. The synthesis process includes mapping the RTL representation to primitive cells including logic gates and registers. It generates a netlist of connected instances of those primitive cells, which is used by the method of the inventors of the present invention for pipeline configuration and related timing and area evaluation. A large library of complete logic primitives can be used to perform accurate synthesis. However, it has been shown that using small basic logic primitives significantly reduces the processing time while still obtaining accurate results. The technology library used by the synthesis tool can include thousands of cells that differ in type, transistor size, drive strength, power consumption, etc. Therefore, synthesis mapping and optimization can take several hours to map the RTL to an appropriate gate-level representation. According to some aspects and embodiments of the present invention, the process uses a very small set of cells, namely inverters, AND gates, OR gates, Muxes, and registers, and there are no changes because no optimization is required. Various aspects of the present invention simply require a quick mapping to this small set for a quick evaluation of the area and timing required for the pipeline configuration.
[0039] In block 740, the path and area delays are calculated from the RTL synthesized using primitives 650 and 660. This step gives the best timing baseline since the pipeline is configured in the mode with the most registers, but the area and leakage are the worst.
[0040] In block 750, although it is the influence of the worst timing, the best area is determined for the entire design. This can be done by disabling the pipeline component under consideration. If all paths still meet the timing constraints, a pipeline configuration is found and all pipeline components are disabled. According to some aspects and embodiments of the present invention, step 750 refers to a baseline where all timing paths meet the required frequency when all pipeline modules are "disabled". This is a corner case that may not actually occur, but still worth checking. According to some aspects and embodiments of the present invention, the process executes step 760 when 750 is not called at all. According to some aspects and embodiments of the present invention, the process executes step 760 when step 750 fails to meet the timing requirements after all pipeline modules are disabled.
[0041] According to some aspects and embodiments of the present invention, the processing of step 760 is detailed in FIG. 8 and starts by enabling all pipeline modules. In step 760, if the timing constraints are not met, the valid flow paths in the baseline RTL description are repeatedly modified until an optimized pipeline configuration according to various aspects and embodiments of the present invention is found. Generally, repeatedly modifying the valid flow paths includes reducing the total number of registers while still meeting the timing constraints. With each iteration, the number of registers is further reduced and the area further decreases (however, the timing increases). An example of such iterative modification is shown in FIG. 8. The modification can be performed progressively according to various aspects and embodiments of the present invention. That is, the modification is performed in the order of relaxation.
[0042] In step 770, the configured pipeline settings are reported. These settings are used for generating a complete RTL description that refers to the final RTL description in which all pipeline module instances are configured. Area and timing are also reported according to various aspects and embodiments of the present invention. This is related to blocks 670 and 680 that report the impact of timing and area after all pipelines are configured. This accompanying result gives a reference point that may be useful for the designer to know.
[0043] FIG. 8 shows an example of a method of iteratively modifying a flow path to find an optimized pipeline configuration according to various aspects and embodiments of the present invention. After step 740 is completed, the flow paths are sorted in descending order of their timing lengths (step 810). Then, the analysis of each path is performed in descending order as follows.
[0044] In steps 820 and 830, a pipeline component instance is selected and a more relaxed configuration mode of that instance is selected. In steps 840 and 850, the timing path across the selected instance is recalculated and analyzed against the timing constraints.
[0045] If a constraint is violated (step 860), for the selected instance, a previously less relaxed configuration mode is selected (step 870), and the next pipeline instance in descending order of timing length is selected (step 820).
[0046] If the selected configuration mode does not violate the target frequency (step 860), if there is a more relaxed configuration mode (step 880), the next configuration mode for that instance is selected (step 830).
[0047] For the selected instance, there is no less restrictive configuration mode (step 880). If there are further pipeline instances to consider (step 890), another pipeline instance is selected (step 820).
[0048] If there are no more pipeline instances to consider (block 890), the search selects the next timing path to work on. When all timing paths have been processed (all pipelines along those paths have been configured), the final pipeline configuration is reported (step 770).
[0049] Referring now to FIG. 9, FIG. 9 shows an example of components of a computing platform 910 for performing the method of this specification. The computing platform 910 includes a memory 920 and a processing unit 930. The memory 920 stores instructions 940 that, when executed, cause the processing unit 620 to generate an optimized pipeline configuration and optionally generate a complete RTL description having the optimized pipeline configuration. Examples of the computing platform 910 include, but are not limited to, workstations, laptops (Windows®, MacOS), servers, and cloud computing.
[0050] The methods and platforms disclosed herein are not limited to any particular electronic device system. Examples of possible systems include, but are not limited to, any electronic system made of reconfigurable pipeline components.
[0051] Consider an example of an SoC 1010 that includes a NoC 1020 as shown in FIG. 10 according to various aspects and embodiments of the present invention. The SoC includes a plurality of initiators and targets such as video, a central processing unit (CPU), a camera, direct memory access (DMA), random access memory (RAM), dynamic random access memory (DRAM), input / output (IO), and a hard disk drive (HDD). The NoC 1020 provides packet-based communication between the initiator and the target.
[0052] The NoC 1020 of FIG. 10 includes a network interface unit (NIU) 1030 responsible for converting and converting from several supported protocols and data sizes to and from the packet transport architecture. The NoC further includes switches, width adapters, firewalls, clock adapters, and individual pipeline registers. These components enable the creation of various network topologies (mesh, ring, etc.), accommodate data widths and packet styles, and are parameterized to enable / disable specific functions based on the user's requirements. Configurable pipeline components are available in many locations within the NIU unit as individual blocks. The parameterizable and configurable pipeline components of the NoC 1020 can be configured according to the methods of this specification.
[0053] Referring to FIG. 11, FIG. 11 shows an example of a layered NIU 1030 having a plurality of pipeline components according to some aspects and embodiments of the present invention. The pipeline components are used in the native layer 1110, the common layer 1120, and the packet layer 1130. According to some aspects and embodiments of the present invention, the common layer 1120 includes a partial address map (PAM) that defines an address space for the NIU initiator from the perspective of which target NIU to communicate with, and other related routing information. Each of these pipeline components can be configured according to the methods described herein.
[0054] Embodiments according to the present invention can be embodied as an apparatus, method, or computer program product. Accordingly, the present invention can take the form of an embodiment completely constituted by hardware, an embodiment completely constituted by software (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which can be broadly referred to herein as "circuit", "module", or "system". Furthermore, embodiments of the present invention can take the form of a computer program product embodied in any tangible medium.
[0055] One or more computer-usable media or any combination of computer-readable media can be utilized. For example, the computer-readable media can include one or more of a portable computer diskette, hard disk, random access memory (RAM) device, read-only memory (ROM) device, erasable programmable read-only memory (EPROM or flash memory) device, portable compact disk read-only memory (CDROM), optical storage device, and magnetic storage device. The computer program code for carrying out operations of the present invention may be written in any combination of one or more programming languages. Such code may be compiled from source code into a computer-readable assembly language or machine code suitable for a device or computer that executes the code.
[0056] Embodiments may be implemented in a cloud computing environment. As used herein and in the following claims, "cloud computing" can be defined as a model for enabling ubiquitous and convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned via virtualization with minimal management effort or service provider interaction and then scaled accordingly.
[0057] Although specific examples have been described herein, it should be noted that different combinations of different components from different examples may be possible. Distinctive features have been presented to better illustrate the examples, but it is clear that specific features can be added, changed, and / or omitted without altering the functional aspects of these examples as described.
[0058] Various examples are ways of using the behavior of any or a combination of machines. An example of a method is completed no matter where in the world most of the configuration steps are performed. For example, according to various aspects and embodiments of the present invention, an IP element or unit may include a processor (e.g., a CPU or GPU), a random access memory (e.g., a RAM such as off-chip dynamic RAM or DRAM), a network interface for wired or wireless connections such as Ethernet®, WiFi, 3G, 4G Long Term Evolution (LTE), 5G, and other wireless interface standard radios. The IP may optionally include various I / O interface devices for various peripheral devices such as touch screen sensors, geographical location information receivers, microphones, speakers, Bluetooth® peripherals, and USB devices (notably, keyboards and mice, etc.). By executing instructions stored in the RAM device, the processor executes the steps of the methods described herein.
[0059] Some examples are one or more non-transitory computer-readable media configured to store such instructions for the methods described herein. Any machine holding a non-transitory computer-readable media containing any of the required code can implement an example. Some examples may be implemented as physical devices such as semiconductor chips, hardware description language representations of the logical or functional behavior of such devices, and one or more non-transitory computer-readable media configured to store such hardware description language representations. The description herein listing principles, aspects, and embodiments encompasses both their structural and functional equivalents. Elements described herein as "coupled" have an effective relationship that may be realized by direct connection or indirectly through one or more other intervening elements.
[0060] Those skilled in the art will understand that various other modifications can be made to the device without departing from the spirit and scope of the present invention (notably, various programmable features). It will be so. All such modifications and changes are included within the technical scope of the claims and are intended to be protected by the claims. Further, those skilled in the art will recognize numerous modifications and variations. The modifications and variations include any relevant combination of the disclosed features. The description in this specification listing principles, aspects, and embodiments encompasses both their structural and functional equivalents. Elements described herein as "coupled" or "communicatively coupled" have an effective relationship that can be realized by direct connection or indirect connection using one or more other intervening elements. Embodiments described herein as "communicating" or "interacting" with another device, module, or element include any form of communication or link and include an effective relationship. For example, a communication link can be established using a wired connection, a wireless protocol, a short-range wireless protocol, or radio frequency identification (RFID).
[0061] All illustrations in the drawings are for the purpose of explaining selected versions of the present invention and are not intended to limit the scope of the present invention. Accordingly, the scope of the present invention is not intended to be limited to the exemplary embodiments illustrated and described herein. Rather, the scope and spirit of the present invention are embodied by the claims.
Claims
1. A method for generating a complete register transfer level (RTL) description of an electronic system, comprising: generating an optimized pipeline configuration from an input including a database of RTL elements and a list of configurable pipeline components, wherein generating the configuration includes performing a search for a configuration that optimizes area and timing; generating the complete RTL description with the pipeline components configured according to the optimized pipeline configuration.
2. The method of claim 1, wherein the list of configurable pipeline elements identifies different configuration modes for each pipeline element, and the input further includes timing path information for each configuration mode.
3. Performing the search includes: generating a baseline RTL description with the pipeline configured as fully enabled to generate a set of flow paths with best timing; eliminating flow paths in the description that do not meet timing constraints; iteratively modifying the valid flow paths in the baseline RTL description until an optimized pipeline configuration is found.
4. The method of claim 3, wherein the baseline RTL description is generated from a library of basic logic primitives, and the flow path delay and pipeline area are synthesized by mapping the baseline RTL description to logic synthesis primitives.
5. The method of claim 4, wherein the flow path delay takes into account the delay by logic primitives.
6. Iteratively modifying the valid flow paths includes modifying the flow paths by progressive relaxation of the configuration, wherein the state where all the pipeline components are fully enabled has the lowest degree of relaxation, and the state where all the pipeline components are in the mode of being fully disabled has the highest degree of relaxation.
7. The method of claim 3, wherein generating the baseline RTL description includes maintaining a user-explicit configuration of pipeline components and configuring the remaining pipeline components.
8. The method according to claim 3, wherein repeatedly modifying the effective flow path includes reducing the total number of registers while still satisfying timing constraints.
9. The target parameter includes a frequency, Modifying the flow path by progressive relaxation includes Sorting the paths in descending order of their timing lengths, and Starting from the path with the longest timing length, selecting each of the paths for analysis, wherein the analysis of each selected path includes Analyzing the selected path against a target clock frequency, and if the selected path does not violate the target frequency, then Selecting a pipeline component instance and selecting a more relaxed configuration mode for the selected instance, Recalculating the timing path across the selected instance, If the selected path violates the target frequency, reinstantiating a previously less relaxed configuration mode, If the selected configuration mode does not violate the target frequency and a more relaxed configuration mode exists, saving the selected instance The method according to claim 8, including looping through instances of pipeline components along the selected path.
10. The electronic device system includes a network-on-chip (NoC), at least some of the configurable pipeline components are associated with the NoC, and the complete RTL description includes a NoC pipeline configured according to the optimized configuration, the method according to claim 1.
11. A computer-implemented method for generating an optimized pipeline configuration for an electronic device system, Accessing a database of register transfer level (RTL) elements, a list of configurable pipeline components related to the electronic device system, and timing path information for each configuration mode of each configurable pipeline component, Generating a baseline RTL description of the system from the RTL elements with all of the configurable pipeline components fully enabled, Eliminating flow paths that do not meet timing constraints in the description, A computer-implemented method including repeatedly modifying the configuration of the pipeline components in a valid flow path in order of relaxation until an optimized pipeline configuration is found.
12. The method according to claim 11, wherein the timing is optimized with respect to the area by the specific criterion including an area.
13. Repeatedly modifying the valid flow path includes modifying the flow path by progressive relaxation of the configuration until target parameters are met, and the state where all the pipeline components are in a fully enabled mode has the lowest degree of relaxation, and the state where all the pipeline components are in a fully enabled mode has the highest degree of relaxation. The method according to claim 11.
14. The electronic device system includes a network-on-chip (NoC), at least some of the configurable pipeline components are related to the NoC, and the complete RTL description includes a NoC pipeline configured according to the optimized configuration. The method according to claim 11.
15. A processing unit, when executed, cause the processing unit to access a database of register transfer level (RTL) elements, a list of configurable pipeline components related to the electronic device system, and timing path information for each configuration mode of each configurable pipeline component, generate a baseline RTL description of the system from the RTL elements with all of the configurable pipeline components fully enabled, exclude flow paths that do not meet the timing constraints among the descriptions, repeatedly modify the configuration of the pipeline components in a valid flow path in order of relaxation until an optimized pipeline configuration is found a memory storing executable instructions A computing platform comprising.
16. The computing platform according to claim 15, wherein the timing is optimized with respect to the area by the specific criterion including an area.
17. Repeatedly modifying the effective flow path includes modifying the flow path by progressive relaxation of the configuration until the target parameters are met, where the state in which all the pipeline components are in a fully enabled mode has the lowest degree of relaxation and the state in which all the pipeline components are in a fully enabled mode has the highest degree of relaxation. The computing platform according to claim 15.
18. The electronic device system includes a network-on-chip (NoC), at least some of the configurable pipeline components are related to the NoC, and the complete RTL description includes a NoC pipeline configured according to the optimized configuration. The computing platform according to claim 15.
19. The executable instructions further cause the computing platform to generate a complete RTL description of the electronic device system with the pipeline components configured according to the optimized pipeline configuration. The computing platform according to claim 15.
Citation Information
Patent Citations
Method for designing pipeline stage in computer-aided design system
JP1995036968A
Method, apparatus, and program for detecting clock gating opportunities in pipelined electronic circuit design
JP2010537293A
Operation synthesizer, operation synthesis method, data processing system including operation synthesizer, and operation synthesis program
JP2014006650A
Automatic pipelining of noc channels to meet timing and / or performance
JP2017500810A
Methods and tools for designing integrated circuits with auto-pipelining capabilities
US20150121319A1