Design technology co-optimization of power performance area optimization in a flow
Patent Information
- Application Number
- CN202580017437.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-01-31
- Publication Date
- 2026-09-22
Smart Images

Figure CN122804233A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to power, performance, and area (PPA) optimization of integrated circuits. More specifically, this disclosure relates to PPA optimization in the Design Technology Co-optimization (DTCO) flow. Background Technology
[0002] Power, performance, and area (PPA) optimization is a holistic approach in semiconductor design aimed at achieving the optimal balance between power consumption, performance, and chip area. PPA provides an important consideration in the development of integrated circuits (ICs) to ensure that the final product meets desired specifications and requirements. Attached Figure Description
[0003] This disclosure will be more fully understood from the following detailed description and the accompanying drawings illustrating embodiments of the present disclosure. The drawings are provided to give knowledge and understanding of embodiments of the present disclosure and are not intended to limit the scope of the disclosure to these specific embodiments. Furthermore, the drawings are not necessarily drawn to scale.
[0004] Figure 1 The illustration shows an example design technique called Co-optimization of Power, Performance, and Area (PPA) flow.
[0005] Figure 2A The illustration shows a first example method for achieving fast and high-quality design-scale process optimization using a domain-specific search algorithm.
[0006] Figure 2B The diagram illustrates PPA points or samples on a graph when using a domain-specific search algorithm.
[0007] Figure 3 The illustration shows a second example of a method to achieve rapid and high-quality design-scale process optimization by performing collaborative optimization on multiple groups.
[0008] Figure 4A The diagram illustrates an example flowchart of a third approach for achieving rapid and high-quality design-scale process optimization using the Pareto frontier of PPA.
[0009] Figure 4B The diagram illustrates PPA points or samples on a graph when using the Pareto frontier of PPA.
[0010] Figure 4C The illustration shows an example set of the best samples from all bins according to the first embodiment.
[0011] Figure 4D The illustration shows an example set of the optimal samples from all bins according to the second embodiment.
[0012] Figure 5The diagram illustrates an example chart showing how designers can select PPA points based on design requirements.
[0013] Figure 6 An example computer system in which embodiments of the present disclosure may operate is illustrated.
[0014] Figure 7 The illustrations depict a collection of example processes used during the design, verification, and manufacturing of articles such as integrated circuits to transform and verify design data and instructions representing integrated circuits. Detailed Implementation
[0015] Various aspects of this disclosure relate to power, performance, and area (PPA) optimization in Design Technology Co-optimization (DTCO) flows. Further aspects of this disclosure relate to maximizing PPA gain.
[0016] PPA optimization is a holistic approach in semiconductor design aimed at achieving the optimal balance between power consumption, performance, and chip area. PPA optimization is a crucial consideration in integrated circuit development to ensure the final product meets desired specifications and requirements.
[0017] Power consumption is a critical issue in electronic devices. Lower power consumption translates to longer battery life, less heat dissipation, and enhanced environmental sustainability. Power consumption analysis (PPA) assesses power usage at various stages of chip operation, from idle to peak performance, to ensure efficient power management.
[0018] Performance is typically measured in clock frequency or execution speed, and determines how quickly a chip can process instructions and deliver results. Higher-performance chips enable faster computation and improved user experience. PPA analysis aims to achieve the highest possible performance while adhering to power and area constraints.
[0019] Area refers to the area of a chip, that is, the physical space occupied by the chip on a silicon (Si) die. Smaller chips allow for higher density, thus saving costs and increasing functionality. PPA analysis strives to minimize the physical footprint of the chip while maintaining desired power and performance characteristics.
[0020] Therefore, PPA can refer to a metric used in the semiconductor industry to compare, for example, processor cores or other electronic components.
[0021] By simultaneously optimizing power, performance, and area, semiconductor designers aim to find a balanced solution that meets the specific requirements of an application. Achieving optimal power-to-area (PPA) balance is crucial for designing efficient and competitive integrated circuits (ICs) across various industries, including consumer electronics, automotive, and industrial applications.
[0022] As Moore's Law slows, DTCO (Technical Development and Computation) becomes crucial for continuing the rapid scaling of silicon transistors. DTCO involves jointly considering manufacturing processes and system performance to guide development and evaluate emerging materials / devices at advanced technology nodes. Unlike traditional technology development, where process optimization and system performance evaluation are separate phases, DTCO integrates system performance and reliability considerations early in the process and device optimization phase.
[0023] DTCO is a concept in the semiconductor industry that involves simultaneously optimizing both design and manufacturing technologies. In the context of semiconductor manufacturing, DTCO refers to the synergistic optimization of chip design and manufacturing process technologies to achieve better performance, power efficiency, and cost-effectiveness.
[0024] Traditionally, chip design and manufacturing have been considered separate phases in the semiconductor development process. However, with technological advancements, close collaboration between design and manufacturing teams can lead to better overall results. DTCO aims to address the challenges and constraints associated with the complex interplay between design considerations and manufacturing process capabilities.
[0025] By collaboratively optimizing design and technology, semiconductor manufacturers can achieve improvements across various aspects, such as performance, power consumption, and yield. This approach becomes particularly important as semiconductor technology advances and reaches smaller process nodes, where the interaction between design and manufacturing becomes more complex. Overall, DTCO is a holistic approach that involves considering both design and manufacturing simultaneously to achieve optimal results in semiconductor chip development.
[0026] Traditional design PPA optimization involves two orthogonal approaches: process optimization and design optimization. Process optimization is performed by foundries, such as technology computer-aided design (TCAD) users in the early stages of process design kit (PDK) definition. There are many choices among various technology variants (such as material, device, layer, physical, and electrical constraints). To achieve rapid turnaround, process optimization is typically performed on small assemblies of cells. Design optimization, on the other hand, is performed by the design team on the actual production design after process definition. Design optimization uses a longer design cycle time and involves multiple steps (such as library characterization, placement and routing, sign-off analysis, and multiple iterations of Engineering Change Orders (ECOs)) to meet PPA objectives.
[0027] In earlier technology nodes, independent process and design development may have been sufficient to meet PPA requirements. However, for more advanced technology nodes, process constraints and design techniques become more complex, leading to challenging optimization problems. To maximize PPA, closer interaction between process and design becomes more important.
[0028] Attempts have been made to perform both process and design optimization at the design level. However, the design-process co-optimization feedback loop is prohibitively expensive in terms of runtime. This design-process co-optimization feedback loop constitutes a high-dimensional optimization problem, resulting in millions to billions of design-level runs. Each design-level run involves costly library characterization and parasitic parameter extraction. Production libraries can take weeks to months to characterize because the cell library includes timing, power, and layout information for various standard cells based on a given timing arc, output load, and input transition. The cell library can vary significantly for different process-voltage-temperature (PVT) corners. Therefore, during system-level DTCO iterations, especially for process development of emerging materials, it is necessary to characterize the cell library across a wide range of voltages, temperatures, and threshold voltages whenever the material or device structure is updated. Furthermore, characterizing the cell library across a wide range of aging-induced threshold voltage shifts is required when evaluating reliability caused by aging of new processes / devices (such as hot carrier injection (HCI) and negative bias temperature instability (NBTI)). As the number of analyses for advanced technology nodes and emerging technologies increases dramatically, this approach becomes increasingly expensive and time-consuming. The prohibitively high runtime limits the success of design-process co-optimization.
[0029] A common compromise is to use small integrated circuit blocks with a limited number of library cells for PPA evaluation. However, optimization quality cannot be guaranteed in actual production design.
[0030] To address these issues, three exemplary approaches are proposed.
[0031] The first approach to achieving fast and high-quality design-scale process optimization for different types of DTCO streams involves using domain-specific search algorithms.
[0032] The second approach for achieving rapid and high-quality design-scale process optimization for different types of DTCO flows involves performing rapid co-optimization of multiple parameter sets. In a practical application, these parameter sets include a combination of front-end process (FEOL) and back-end process (BEOL) parameters.
[0033] A third approach for achieving rapid and high-quality design-scale process optimization for different types of DTCO streams involves the rapid generation of Pareto fronts through PPA (Performance-Based Analysis).
[0034] These three methods enable fast DTCO and can be applied to different types of DTCO flows (such as early DTCO, late DTCO), as well as static timing analysis (STA) driven or implementation tool driven DTCO for process selection to maximize PPA gain.
[0035] Figure 1 The diagram illustrates an example DTCO flow used to implement PPA optimization.
[0036] As semiconductor components continue to shrink, the challenges associated with design-to-manufacturing (DFM) and DTCO increase. The complexity of IC design and manufacturing processes requires extending traditional DFM and DTCO techniques to overcome systematic failures associated with complex design-process interactions.
[0037] The IC flow from design to manufacturing has well-defined modules (such as physical design, mask synthesis, mask writing, wafer fab processes, and inspection and testing). Each module includes industry-standard verification processes (such as physical verification, optical process correction and mask proximity correction (OPC / MPC) verification, metrological inspection, and physical failure analysis (PFA)).
[0038] DTCO is crucial for yield breakthroughs and rapid product scaling throughout the lifecycle of new technology nodes. DTCO involves simultaneous optimization of chip design and manufacturing processes, taking into account the interdependencies between the two. DTCO uses iterative design and manufacturing simulations to identify optimal design and manufacturing parameters.
[0039] Later in the process node lifecycle, this co-optimization is accomplished by traditional technologies such as DFM and lithography-friendly design (LFD). However, as manufacturing technologies shrink and design complexity increases, greater challenges arise that existing DFM and DTCO technologies cannot overcome.
[0040] Traditional DTCO, DFM, and LFD methods have proven their value, but many designs require more. Systematic defects evade traditional detection methods and emerge during yield improvements and eventual high-volume manufacturing (HVM). To improve the efficiency and effectiveness of DFM and DTCO processes for complex, advanced node designs, the industry needs more methodologies for performing rapid and high-quality PPA searches to maximize PPA gains. An exemplary embodiment presents such a methodology for maximizing PPA gains.
[0041] Back Figure 1 The DTCO flow 100 includes millions to billions of states 102 provided to the production design phase 130. In one example, state 102 may include FEOL process variant 110 and BEOL process variant 120. FEOL process variant 110 may involve device variable 112, while BEOL process variant 120 may involve layer variable 122.
[0042] During the production design phase 130, placement 132, routing 134, and the STA / ECO process 136 are performed during the domain-driven artificial intelligence (AI) search 138. After the domain-driven AI search 138 is completed, the PPA search results are generated and represented as the PPA Pareto front 140. The PPA Pareto front 140 allows the designer to select the optimal PPA point. This can be targeted at the optimal F... max Optimal power and / or F max The optimal set of solutions is generated by weighing trade-offs between power and efficiency.
[0043] Layout 132 is part of the physical design flow, where the layout assigns precise locations to various circuit components within the chip's core region. The placeholder performs the assignment while optimizing multiple objectives to ensure the circuit meets its performance requirements. The placeholder receives a given synthesized circuit netlist along with a technology library and generates an efficient layout. The layout is optimized according to the objectives and prepared for cell size adjustments and buffering.
[0044] Routing 134 builds upon placement, which determines the location of each active component on an IC or printed circuit board (PCB). Following placement, the routing step adds the necessary traces to properly connect the placed components while adhering to all IC design rules. Routing 134 involves providing pre-existing polygons consisting of pins (or terminals) on components, and optionally some pre-existing wiring referred to as pre-routing. Each of these polygons is associated with a net, typically by name or number. The primary task of the router is to create geometry such that all terminals assigned to the same net are connected, terminals assigned to different nets are not connected, and all design rules are followed.
[0045] Regarding the STA / ECO process 136, STA is a method for verifying the timing performance of a design by examining all possible paths of timing violations. STA decomposes the design into timing paths, calculates the signal propagation delay along each path, and checks for timing constraints violations within the design and at input / output interfaces.
[0046] Regarding the STA / ECO process 136, in a system-on-chip (SoC) workflow environment, a functional ECO is a method for directly patching or modifying the design at the gate level or in the synthesized version. The reasons for such modifications may include fixing bugs found in the register-transfer-level (RTL) (pre-synthesis) version of the design, applying optimizations to the design, or updating the design based on new customer requirements. ECOs provide an important late-stage optimization step for any design. This is a step for correcting bugs, applying optimizations, and meeting customer requests in later stages.
[0047] The PPA Pareto Frontier 140 provides a collection of optimal PPA solutions covering the entire frequency or power range. The PPA Pareto Frontier 140 is a collection of processor solutions that provide optimal power or optimal frequency. For example, given a frequency value, the optimal achievable power can be determined. Alternatively, given a power value, the optimal achievable frequency can be determined. The PPA solution collection provides, for example, a trade-off between frequency and power.
[0048] Figure 2A The illustration shows a first example of a method for achieving fast and high-quality design-scale process optimization using a domain-specific search algorithm.
[0049] The first approach for rapid and high-quality design-scale process optimization for different types of DTCO flows involves using domain-driven search algorithms.
[0050] In previous methodologies, the search algorithm could be a machine learning (ML) algorithm based on artificial intelligence (AI) search. AI search based on ML algorithms can use general optimization algorithms. These previous search algorithms provide different optimization qualities for different applications. In other words, these previous search algorithms may be effective for some practical applications, but ineffective for others. Domain-driven search algorithm 210 uses the physical and electrical characteristics of process variants to drive a more targeted PPA search. Domain-driven search algorithm 210 explores the relationship between process variants and design quality and derives an analytical PPA model, which enables fast PPA evaluation with high accuracy. Domain-driven search algorithm 210 allows a full scan to be performed and evaluate all samples or targets throughout the entire design optimization space to obtain the optimal sample set with top samples or targets. Typically, a full scan refers to the measured output value varying with progressive input parameters. In this paper, a full scan refers to searching for PPA points within the design optimization space. The optimal sample set or target set provides the most desired or top or preferred or ideal or best samples or targets. The best samples or top samples (or targets) provide the highest quality PPA points or samples or targets. The top sample can also be referred to as a subset of the samples.
[0051] Before delving into the details of Domain-Driven Search Algorithm 210, let's make two observations. The first observation concerns the nonlinearity between PPA gain and variables. The second observation concerns the correlation between variables.
[0052] Regarding the first observation, different types of variables can have different effects on timing and power. Some variables have a dominant effect on timing and a negligible effect on power. Alternatively, some variables have a dominant effect on power and a negligible effect on timing. The relationship between process variables and design quality was examined. An examination of how several parameters affect PPA gain revealed a common trend: a slight nonlinearity was observed between PPA gain and the variables. Slight nonlinearity can refer to a small amount of nonlinearity. Slight nonlinearity can refer to nonlinearity that is slightly noticeable, slightly detectable, or slightly perceptible. Slightly perceptible nonlinearity can refer to nonlinearity that is detected by the user in only a few instances, for example, less than an integer (such as 5 or 10). This observation of slight nonlinearity allows for the determination of a PPA model for each parameter. The PPA effect of each parameter can be modeled with reasonable accuracy using quadratic fitting.
[0053] Regarding the second observation, weak or non-substantial correlations between process parameters and cross-parameter terms can be ignored. A weak correlation can refer to a non-substantial relationship or correlation between parameters. A weak correlation can refer to a weak linear relationship between two quantitative variables. A weak correlation can also refer to a statistical relationship between two variables where the values can be very close but not proportional. For weak correlations, the correlation coefficient is typically less than 0.3. The correlation coefficient can be a measure of how closely two variables are related. Thus, the combined PPA gain can be obtained by linearly superimposing the effects of each individual parameter. Therefore, linearly superimposing the effects of each variable provides reasonable precision.
[0054] Domain-driven search algorithm 210 is based on first and second observations. In the first method 200A, the conclusions of the two observations 202 are combined. In other words, the first conclusion 204 regarding some mild nonlinearity is combined with the second conclusion 206 regarding weak variable correlation to develop the domain-driven search algorithm 210. The domain-driven search algorithm 210 performs the following steps: constructing a surrogate model 220, performing an exhaustive scan throughout the design or optimization space 230, and performing an analysis of the top or optimal M candidates 240.
[0055] The surrogate model 220 is constructed by establishing quadratic models to represent the PPA effect of each process parameter or variable. A quadratic model is a mathematical model represented by a quadratic equation or a group of quadratic equations. The quadratic equation representing the effect from each parameter is shown. The PPA effect of each process parameter or variable is represented as a quadratic model. The quadratic models of all process parameters or variables are summed or added together. The construction of the quadratic model uses 2N runs, where N is the number of process parameters or variables. In one example, there are 6 process parameters or variables. This results in 2 × 6 = 12 runs because each variable includes 2 coefficients. The overall PPA gain is obtained from the sum of the quadratic models corresponding to all variables.
[0056] Typically, constructing a surrogate model 220 (also known as an alternative or proxy model) involves creating a simplified representation of a more complex system or process. Surrogate models are used across various fields, including engineering, optimization, machine learning, and simulation, to approximate the behavior of more complex and computationally expensive models. The goal of constructing a surrogate model 220 is to obtain a faster and less resource-intensive way to perform predictions or conduct analyses.
[0057] Proxy models are typically used when dealing with computationally expensive or time-consuming complex systems or simulations. These systems can include physical experiments, numerical simulations, or other processes. To construct the surrogate model 220, a set of training data is used. The training data is typically obtained by running the complex model or simulation at various input points. The input-output pairs of these simulations are then used to train the surrogate model 220.
[0058] The surrogate model 220 captures the fundamental relationship between input and output without replicating the full complexity of the original model. Once trained, the surrogate model 220 can be used for prediction or analysis at a significantly lower computational cost compared to the original model. This is particularly useful in scenarios requiring rapid evaluation or optimization. The surrogate model 220 can be iteratively refined as more data becomes available or as the complexity of the problem is better understood. This iterative process helps improve the accuracy of the surrogate model 220 over time. Therefore, building the surrogate model 220 allows practitioners to balance the trade-off between computational cost and model accuracy, making it a valuable tool in scenarios such as maximizing PPA gains, where efficiency and speed are paramount.
[0059] After constructing the surrogate model 220, an exhaustive scan of the entire optimization space 230 is performed. All combinations of variables are enumerated very quickly, for example, within a few seconds.
[0060] After performing an exhaustive scan of the entire optimization space 230, analysis 240 is performed on the top, optimal, or most relevant M candidates. After obtaining the top or optimal M candidates or sample candidates from the surrogate model 220, analysis 240 is performed by initiating M runs to determine the PPA gain. In one example, M=50. In other words, the method selects the top M sample runs estimated by the surrogate model 220 (e.g., M=50). Thus, the PPA gain of the top 50 samples or targets is obtained. Of course, M can be set to any other integer based on desired design criteria. The top M candidates or optimal M candidates are those that exhibit the strongest statistical relationship between process parameters, i.e., high linearity and strong correlation.
[0061] The domain-driven search algorithm 210 was tested on a central processing unit (CPU) block, and the results were plotted, as shown below. Figure 2B As shown.
[0062] Figure 2B The diagram illustrates PPA points or samples on graph 200B when using a domain-specific search algorithm.
[0063] In one example, the Domain-Driven Search algorithm 210 was tested on a central processing unit (CPU) block with six FEOL process parameters or variables. In Figure 200B, the x-axis represents the normalized F... max 250, where the y-axis represents the normalized leakage 252. When using the traditional AI algorithm, a 3.91% PPA gain is achieved after 200 runs. Sample 260 is the result using the traditional AI algorithm. Sample 260 is scattered across the entire graph. In contrast, when using the Domain-Driven Search algorithm 210, a 4.43% PPA gain is achieved after only 62 runs. The 62 runs are derived from 2N + M = (2 × 6) + 50 = 12 + 50 = 62. Therefore, the Domain-Driven Search algorithm 210 provides a better PPA gain with fewer runs. Sample 270 is the result using the Domain-Driven Search algorithm 210. Sample 270 represents the top-ranked process points or PPA points and is shown as a cluster along the curve.
[0064] Figure 3 The illustration shows a second example of a method to achieve rapid and high-quality design-scale process optimization by performing collaborative optimization on multiple groups.
[0065] A second approach for rapid and high-quality design-scale process optimization for different types of DTCO flows involves performing rapid collaborative optimization across multiple groups.
[0066] The optimization space 310 comprises multiple groups. In this example, there are a first group 312, a second group 314, and a third group 316. The first group 312 may include N1 samples, the second group 314 may include N2 samples, and the third group 316 may include N3 samples. The total number of samples is N1 × N2 × N3. The correlation between variables between groups is weak or non-substantial. Note that each group may include a very large number of samples or points. Here, the optimization space of the first group 312 is very large, and the optimization space of the second group 314 is also very large. If the first group 312 is combined with the second group 314 (or even multiple groups), the resulting co-optimization space becomes extremely large, potentially generating millions to billions of samples. Execution on such a large dataset would lead to runtime issues during sample evaluation, as this processing is time-consuming and cost-inefficient.
[0067] At position 320, key samples are selected in each group, and collaborative optimization is performed on the key samples across all groups. Key samples are considered top candidates.
[0068] The optimization space 330 has been adjusted to provide co-optimization. The first group 332 depicts key samples 334, the second group 336 depicts key samples 338, and the third group 340 depicts key samples 342. The total number of key samples is M1×M2×M3. Here, only key samples or top candidates from each group are used to generate the co-optimization space. This significantly reduces runtime because instead of processing millions to billions of samples, only hundreds or thousands of samples (key samples or top candidate samples) are processed. Key samples can also be referred to as important samples, meaningful samples, significant samples, purposeful samples, useful samples, or relevant samples. A sample is considered key, significant, meaningful, or purposeful if it has a significant or substantial or serious or important impact on the PPA gain. Key samples can also be referred to as target samples, candidate samples, or dominant samples.
[0069] In a practical example, FEOL and BEOL processes can be synergistically optimized. The manufacturing process of Very Large Scale Integration (VLSI) ICs consists of a set of basic steps, from crystal growth, wafer fabrication, epitaxy, dielectric and polysilicon thin film deposition, oxidation, photolithography to dry etching. Different patterns are developed using resists and masks, and after patterning, the resist is stripped from the wafer. During the manufacturing process, devices are created on the chip, and many of these basic steps are repeated multiple times. The process threshold voltage can vary depending on the dose or energy of ion implantation. Furthermore, the chip's location on the wafer determines the threshold voltage and mobility. Therefore, all chips manufactured on the same wafer may differ in performance. The manufacturing process is divided into FEOL, MOL, and BEOL. FEOL involves transistor-level layout design, MOL involves transistor-level interconnects, and BEOL involves netlist recovery rate (PnR) level interconnects.
[0070] FEOL encompasses the processing of the active portion of a chip, namely the transistors located at the bottom of the chip. Transistors function as electrical switches and use three electrodes for their operation: the gate, source, and drain. Current in the conduction channel between the source and drain can be switched on and off, and operation is controlled by the gate voltage.
[0071] BEOL (Block Interconnect) is the final stage of processing and refers to the interconnects located on top of the chip. Interconnects are complex wiring schemes that distribute clock and other signals, provide power and ground, and transmit electrical signals from one transistor to another. BEOL is organized in different metal layers, local (Mx), intermediate, semi-global, and global wiring. The total number of layers can be up to 15, however, the typical number of Mx layers ranges from 3 to 6. Each of these layers contains (unidirectional) metal lines organized in regular tracks, along with dielectric material. They are vertically interconnected through metal-filled via structures.
[0072] FEOL and BEOL are linked together via MOL. As device scaling continues to 3nm and below, processing each of these modules presents numerous challenges. This has forced chip manufacturers to turn to new device architectures in FEOL and new materials and integration schemes in BEOL. FEOL process variations have a dominant impact on chip performance and leakage power, while BEOL process variations can more effectively affect chip performance and dynamic power. Therefore, synergistic optimization of FEOL and BEOL can lead to a significant maximization of PPA gain.
[0073] However, as mentioned above, different types of transistors and interconnect layers have different process variations, even for the same type of transistor or interconnect layer. Process parameters can differ from die to die due to variations in masks, photolithography, chemical mechanical polishing (CMP), etc. The combination of these two types of process variations can result in very large sample sizes, leading to prohibitively high runtimes.
[0074] The above text is about Figure 2B In the example, if we assume the FEOL process has 6 process parameters or variables, and each parameter varies in the range [-30mV, 30mV] with a resolution of 5mV, the resulting search space will have approximately 4.8 million samples. For the BEOL process, if the same design has 12 metal layers, and the thickness of each layer varies in the range [-20%, +20%] with a resolution of 10%, the resulting search space will have approximately 244 million samples. Therefore, the number of samples in the co-optimization space (both FEOL and BEOL) will be approximately 4.8 million × 244 million = 1.17 × 10⁻⁶. 15 A sample size of 100 samples. Processing such a large number of samples will introduce runtime issues during sample evaluation.
[0075] The second approach proposes an effective and efficient combination of FEOL and BEOL processes for optimization. For example... Figure 3 As shown, there are multiple groups for collaborative optimization, and each group represents a type of optimization space. Within each optimization space, not all samples are critical, important, or beneficial to the PPA gain. The second method involves selecting critical samples from each group using a domain-driven search algorithm 210, and performing collaborative optimization only among the critical samples in each group. When the correlation between variables in a group is not strong, weak, or substantial, the second method can significantly reduce the number of sample evaluations while maintaining high optimization quality.
[0076] Figure 4A The diagram 400 illustrates an example flowchart of a third approach to achieving rapid and high-quality design-scale process optimization using the Pareto frontier of PPA.
[0077] At position 402, the domain-driven search algorithm 210 performs a PPA evaluation to evaluate all samples in the entire optimization space. Here, the PPA evaluation model is obtained using the domain-driven search algorithm 210.
[0078] At position 404, sampled data is obtained. Sampled data can be, for example, process data, F... max Data, power data, etc. Therefore, it is used to evaluate F... max The PPA model derived from the power was used for each process point or each sample.
[0079] At position 406, select a leading-edge sample. Obtain F-based... max And an updated PPA frontier sample set of power data.
[0080] At point 408, the selected sample is analyzed. An analysis is performed on the frontier sample by initiating a full design run to generate an accurate Pareto front for the PPA. The analysis includes analyzing the top M candidates. In one example, the top M candidates are 50 candidates. The top M candidates are selected or picked by a surrogate model. The analysis involves determining the nonlinearities and correlations between process parameters. Specifically, the analysis involves determining whether the nonlinearities are mild and whether the correlations are weak. Therefore, the analysis involves determining various relationships between process parameters (such as statistical relationships between process parameters). This involves discovering, for example, how strong the statistical relationships between process parameters are. Determining the strong or weak relationships between process parameters allows the user to find the optimal PPA point.
[0081] At position 410, the Pareto front of the PPA is generated.
[0082] The third method generates the PPA Pareto front. The PPA Pareto front is an optimal set of processes covering the entire performance and power range. The PPA Pareto front can determine the optimal F max What is the power, and determine F? max The trade-off between power and PPA. The Pareto front determines F. max The trade-off between power Figure 4B As shown in the image.
[0083] Figure 4B The diagram illustrates the PPA points or samples on chart 420 when using the Pareto frontier of PPA.
[0084] In Figure 420, the x-axis represents the normalized F. max 422, the y-axis represents the normalized leakage 424. The curve depicts the set of samples. The optimal power solution set is located at point 430, where 25% of the leakage power is achieved. The optimal performance solution set is located at point 432, where 4.2% of the F... max Gain. The optimal power and performance combination solution can be found at point 434, where a 3% gain is achieved. max Gain and 14% leakage gain.
[0085] Figure 4C An example diagram 450 is shown, illustrating the collection of the best samples from all boxes according to the first embodiment.
[0086] Given a metric (e.g., F) maxGain), select the metric resolution. The entire metric range is divided into multiple bins. The bin size is the specified resolution. Once a new process point (e.g., process, F) has been evaluated... max (Power), then find its corresponding F max The optimal process point with the maximum power gain within each bin is updated. After evaluating all process points, the optimal process point from each bin is collected and used as a frontier sample.
[0087] See Figure 4C Chart 450 includes representations of F max The x-axis represents 452 and the y-axis represents power 454. Bins are designated as 464. Each bin 464 includes a sample 460. Samples 460 can also be referred to as data points. The optimal sample 462 from each bin 464 is ultimately selected. Bins 464 are represented as those used to evaluate F. max The vertical line.
[0088] F max It has a specified range. F max This range can be divided into multiple bins (e.g., bin 464). In one example, there might be 100 bins, and F max The resolution can be, for example, 0.05%. Samples 460 are distributed within each bin 464. Each bin 464 may contain a different number of samples 460. Here, multiple samples in each bin 464 may have the same frequency. However, the samples 460 in each bin 464 have different power values. The goal is to determine the optimal power value in each bin 464 for the samples 460 in that bin. The optimal sample in each bin 464 is designated as 462. As samples 460 are added to bins 464, the optimal sample 462 can be continuously updated. Therefore, the optimal sample 462 provides the optimal power value for that frequency. In other words, the optimal sample 462 in each bin 464 represents the optimal combination of power value and frequency value, and this is how the frontier samples are generated. The optimal sample 462 can be referred to as the Pareto front sample utilized by the user.
[0089] Figure 4D An example diagram 470 illustrates the collection of the best samples from all boxes according to the second embodiment.
[0090] Chart 470 includes representations of F max The x-axis represents 452 and the y-axis represents power 454. Bins are designated as 476. Each bin 476 includes a sample 472. Samples 472 can also be referred to as data points. The optimal sample 474 from each bin 476 is ultimately selected. Bins 476 are represented as horizontal lines used to evaluate power.
[0091] The power has a specified range. The power can be divided into multiple bins (e.g., bin 476) along this range. In one example, there might be 100 bins, and the power resolution could be, for example, 0.05%. Samples 472 are distributed within each bin 476. Each bin 476 can contain a different number of samples 472. Here, multiple samples in each bin 476 can have the same power. However, the samples 472 in each bin 476 have different frequency values. The goal is to determine the optimal frequency value in each bin 476 for the samples 472 in that bin. The optimal sample in each bin 476 is designated as 474. As samples 472 are added to bins 476, the optimal sample 474 can be continuously updated. Therefore, the optimal sample 474 provides the optimal frequency value for that power level. In other words, the optimal sample 474 in each bin 476 represents the optimal combination of power and frequency values; this is how the frontier samples are generated. The optimal sample 474 can be referred to as the Pareto frontier sample utilized by the user.
[0092] Figure 5 The illustration in Figure 500 shows an example of how designers can select PPA points based on design requirements.
[0093] Chart 500 includes representations of F max The x-axis represents 502 and the y-axis represents power 504. The curves depict the set of samples. The optimal power solution set lies at point 410, where a 17% total power gain and a 1% Fg are achieved. max Gain. The optimal performance solution set lies at point 514, where a 5% F is obtained. max Gain, and a 2% power gain. The optimal power and performance combination solution can be found at point 512, where a 4% F gain is achieved. max Gain and a 10% power gain. Figure 500 can relate to a second approach. Specifically, this could be 4nm production of a CPU core including 6 FEOL process parameters or variables and 12 BEOL process parameters or variables. Therefore, when performing co-optimization on both FEOL and BEOL, the designer can choose between points 510, 512, and 514 to meet desired system performance criteria. Of course, the designer can choose any point along the curve to meet desired system performance criteria. Design runs have been performed from 10... 15 The number of runs has been reduced to several hundred. In addition, the turnaround time (TAT) is less than 30 hours.
[0094] In summary, methods for fast and high-quality design-scale process optimization are proposed for different types of DTCO flows. The first method for fast and high-quality design-scale process optimization for different types of DTCO flows involves using a domain-driven search algorithm employing a surrogate model. The second method for fast and high-quality design-scale process optimization for different types of DTCO flows involves performing fast co-optimization across multiple groups. In a practical application, these groups include a combination of FEOL and BEOL processes. The third method for fast and high-quality design-scale process optimization for different types of DTCO flows involves performing rapid generation of the PPA Pareto front. The PPA Pareto front allows designers to select the optimal PPA point. This can be achieved by optimizing the optimal F... max Optimal power and / or F max The optimal set of solutions is generated by weighing trade-offs between power and efficiency. These three approaches achieve fast DTCO and can be applied to different types of DTCO flows (such as early DTCO, late DTCO), as well as STA-driven or implementation tool-driven DTCO for process selection.
[0095] Figure 6 An example computer system in which embodiments of the present disclosure may operate is illustrated.
[0096] Figure 6 The illustration depicts an example machine of computer system 600, within which a set of instructions can be executed to cause the machine to perform any or more methodologies discussed herein. In alternative implementations, the machine can be connected (e.g., networked) to other machines on a local area network (LAN), intranet, extranet, and / or the Internet. The machine can operate as a server or client machine in a client-server network environment, as a peer-to-peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0097] A machine can be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, network device, server, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequentially or otherwise) specifying the actions to be taken by that machine. Furthermore, while a single machine is illustrated, the term "machine" should also be understood to include any collection of machines that, individually or jointly, execute a set (or more) of instructions to perform any or more methodologies discussed herein.
[0098] Example computer system 600 includes processing device 602, main memory 604 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM)), static memory 606 (e.g., flash memory, static random access memory (SRAM) etc.)) and data storage device 618, which communicate with each other via bus 630.
[0099] Processing device 602 represents one or more processors (such as microprocessors, central processing units, etc.). More specifically, the processing device may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor that implements other instruction sets, or a combination of instruction sets. Processing device 602 may also be one or more special-purpose processing devices (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc.). Processing device 602 may be configured to execute instructions 626 to perform the operations and steps described herein.
[0100] The computer system 600 may also include a network interface device 608 for communication via a network 620. The computer system 600 may also include a video display unit 610 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), a graphics processing unit 622, a signal generation device 616 (e.g., a speaker), a graphics processing unit 622, a video processing unit 628, and an audio processing unit 632.
[0101] Data storage device 618 may include machine-readable storage medium 624 (also known as non-transitory computer-readable medium) on which one or more sets of instructions 626 or software embody any one or more methodologies or functions described herein. Instructions 626 may also reside wholly or at least partially in main memory 604 and / or processing device 602 during execution by computer system 600, which also constitute machine-readable storage media.
[0102] In some implementations, instruction 626 includes instructions that implement functions corresponding to this disclosure. While machine-readable storage medium 624 is shown as a single medium in the example implementation, the term "machine-readable storage medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of instructions. The term "machine-readable storage medium" should also be understood to include any medium capable of storing or encoding a set of instructions for machine execution, and the medium enabling the machine and processing device 602 to perform any one or more methodologies of this disclosure. The term "machine-readable storage medium" should be accordingly understood to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0103] Figure 7 The illustration depicts a set 700 of example processes used during the design, verification, and manufacturing of articles such as integrated circuits to transform and verify design data and instructions representing integrated circuits. Each of these processes can be structured and enabled as multiple modules or operations. The term "EDA" stands for "Electronic Design Automation." These processes begin with the creation of a product concept 710, accompanied by information provided by the designer, which is transformed to create a manufactured article using a set of EDA processes 712. When the design is completed, it is tape-out 734, which occurs when patterns (e.g., geometric patterns) for the integrated circuit are sent to a manufacturing facility to create a mask set, which is then used to manufacture the integrated circuit. After tape-out, semiconductor dies are manufactured 736 and packaging and assembly processes 738 are performed to produce a finished integrated circuit 740.
[0104] Specifications for circuits or electronic structures can range from low-level transistor material layouts to high-level description languages. High-level representations can be used to design circuits and systems using hardware description languages (“HDLs”) such as VHDL, Verilog, SystemVerilog, SystemC, MyHDL, or OpenVera. HDL descriptions can be transformed into logic-level register-transfer-level (“RTL”) descriptions, gate-level descriptions, layout-level descriptions, or mask-level descriptions. Each lower level of representation, as a more detailed description, adds more useful details to the design description, such as more details for modules that include the description. Lower levels of representation, as more detailed descriptions, can be computer-generated, derived from design libraries, or created by another design automation process. An example of a lower-level specification language used to specify a more detailed representation language is SPICE, which is used for detailed descriptions of circuits with many analog components. The description at each level of representation can be used by the corresponding tools for that layer (e.g., formal verification tools). The described processes can be enabled by EDA products (or tools).
[0105] During system design phase 714, the functionality of the integrated circuit to be manufactured is specified. The design can be optimized for desired characteristics such as power consumption, performance, area (physical and / or lines of code), and cost reduction. At this stage, the design can be divided into different types of modules or components.
[0106] During logic design and functional verification 716, modules or components in a circuit are specified using one or more description languages, and the functional accuracy of the specifications is checked. For example, components of a circuit can be verified to generate outputs that match the requirements of the specifications of the circuit or system being designed. Functional verification can be performed using simulators and other programs such as test bench generators, static HDL checkers, and formal verifiers. In some embodiments, a special component system, referred to as a “simulator” or “prototype system,” is used to accelerate functional verification.
[0107] During the synthesis and design phase 718 for testing, HDL code is transformed into a netlist. In some embodiments, the netlist may be a graph structure, where edges of the graph structure represent components of the circuit, and nodes of the graph structure represent how the components are interconnected. Both HDL code and netlist are hierarchical fabricated articles that EDA products can use to verify that integrated circuits perform according to a specified design during manufacturing. The netlist can be optimized for a target semiconductor manufacturing technology. Furthermore, the completed integrated circuit can be tested to verify that the integrated circuit meets the specifications.
[0108] During netlist verification (720), the netlist is checked to ensure it meets timing constraints and corresponds to the HDL code. During design planning (722), the overall planar layout of the integrated circuit is constructed, and timing and top-level routing are analyzed.
[0109] During layout or physical implementation 724, physical placement (positioning of circuit components such as transistors or capacitors) and wiring (connecting circuit components through multiple conductors) occur, and elements can be selected from a library to enable specific logic functions. As used herein, the term "element" can specify a set of transistors, other components, and interconnections that provide Boolean logic functions (e.g., AND, OR, NOT, XOR) or storage functions (such as flip-flops or latches). As used herein, a circuit "block" can refer to two or more elements. Both elements and circuit blocks can be referred to as modules or components and are implemented in both physical structure and simulation. Parameters (such as dimensions) are specified for the selected elements (based on "standard cells") and are available in a database for use in EDA products.
[0110] During Analysis and Extraction 726, circuit functionality is verified at the layout level, allowing for refinement of the layout design. During Physical Verification 728, the layout design is checked to ensure that manufacturing constraints such as DRC constraints, electrical constraints, and lithographic constraints are correct, and that the circuit functionality matches the HDL design specifications. During Resolution Enhancement 730, the geometry of the layout is transformed to improve how the circuit design is manufactured.
[0111] During the tape-out process, data is created for the production of the photomask (if appropriate, after the application of photolithography enhancement). During mask data preparation 732, the "tape-out" data is used to produce the photomask, which is then used to produce the finished integrated circuit.
[0112] The storage subsystem of a computer system (such as computer system 700 in Figure 10) can be used to store programs and data structures used by some or all of the EDA products described herein, as well as products for the development of library units and for physical and logical design using the library.
[0113] This disclosure relates to methods for rapid and high-quality design-scale process optimization for different types of DTCO flows. A first method for rapid and high-quality design-scale process optimization for different types of DTCO flows involves using a domain-driven search algorithm employing a surrogate model. A second method for rapid and high-quality design-scale process optimization for different types of DTCO flows involves performing rapid co-optimization of multiple groups. In a practical application, multiple groups include a combination of FEOL and BEOL processes. A third method for rapid and high-quality design-scale process optimization for different types of DTCO flows involves performing rapid generation of the PPA Pareto front. The PPA Pareto front allows designers to select the optimal PPA point. This can be achieved for the optimal F... max Optimal power and / or F max The optimal set of solutions is generated by weighing trade-offs between power and efficiency. These three approaches achieve fast DTCO and can be applied to different types of DTCO flows (such as early DTCO, late DTCO), as well as STA-driven or implementation tool-driven DTCO for process selection.
[0114] In one example, the method includes: constructing a surrogate model representing the impact of multiple metrics on multiple process parameters; performing a scan to determine multiple samples in an optimization space including the multiple process parameters; selecting a subset of sample candidates from the surrogate model; and generating a PPA model based on the subset of sample candidates using a processing device to output an improved sample set. The subset of sample candidates consists of the top M candidates, or the best M candidates, or the optimal M candidates that exhibit the strongest statistical relationship between the process parameters, i.e., high linearity and strong correlation.
[0115] In another example, the method includes: creating multiple groups in the optimization space, each group including samples of different process parameters, selecting a dominant sample in each group, and performing collaborative optimization using the dominant sample from each group.
[0116] In yet another example, the method includes: using a domain-driven search algorithm to generate a power, performance, and area (PPA) model; using the PPA model to evaluate a first metric and a second metric for each of a plurality of process parameters; updating a PPA front sample set comprising a plurality of samples based on the first metric data and the second metric data; and performing analysis on the PPA front sample set to generate a PPA Pareto front.
[0117] Some parts of the foregoing detailed description have been presented in the form of symbolic representations of algorithms and operations on data bits within computer memory. These algorithmic descriptions and representations are the most effective way for those skilled in the art of data processing to communicate the essence of their work to others skilled in the art. An algorithm can be a sequence of operations that leads to a desired result. An operation is one that requires physical manipulation of a physical quantity. This quantity can take the form of an electrical or magnetic signal that can be stored, combined, compared, and otherwise manipulated. Such a signal can be referred to as a bit, value, element, symbol, character, term, number, etc.
[0118] However, it should be remembered that all these and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless explicitly stated otherwise in this disclosure, it should be understood that throughout the description, certain terms refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within computer system registers and memories into other data similarly represented as physical quantities within computer system memory or registers or other such information storage devices.
[0119] This disclosure also relates to means for performing the operations described herein. The means may be specifically constructed for the intended purpose, or it may comprise a computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards or optical cards, or any type of medium suitable for storing electronic instructions, each medium being coupled to a computer system bus.
[0120] The algorithms and displays presented herein do not inherently relate to any particular computer or other device. Various other systems may be used in conjunction with the programs taught herein, or it may be convenient to construct more specialized devices to execute the methods. Furthermore, this disclosure is not described with reference to any particular programming language. It will be understood that the teachings disclosed herein can be implemented using various programming languages.
[0121] This disclosure can be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, which can be used to program a computer system (or other electronic device) to perform processes according to this disclosure. Machine-readable media include any mechanism for storing information in a machine-readable (e.g., computer-readable) form. For example, machine-readable (e.g., computer-readable) media include machine-readable (e.g., computer-readable) storage media such as read-only memory (“ROM”), random access memory (“RAM”), disk storage media, optical storage media, flash memory devices, etc.
[0122] In the foregoing disclosure, implementations of this disclosure have been described with reference to specific examples thereof. It is apparent that various modifications may be made thereto without departing from the broader spirit and scope of the disclosure as set forth in the following claims. Where elements are referred to in the singular tense in this disclosure, more than one element may be depicted in the figures, and the same elements are labeled with the same numerals. Therefore, the disclosure and the figures should be considered illustrative rather than restrictive.
Claims
1. A method comprising: Construct a proxy model that represents the impact of multiple metrics on multiple process parameters; Perform a scan to determine the number of samples in the optimization space, which includes the plurality of process parameters; Select a subset of sample candidates from the proxy model; as well as The processing device generates a PPA model based on a subset of the sample candidates to output an improved sample set.
2. The method according to claim 1, wherein the surrogate model is the summation of multiple quadratic models for each process parameter.
3. The method of claim 2, wherein a slight nonlinearity between the PPA gain and the plurality of process parameters enables the use of the quadratic model.
4. The method of claim 2, wherein the weak correlation between the plurality of process parameters enables the use of the quadratic model.
5. The method of claim 1, wherein the scanning involves enumerating all combinations of the plurality of process parameters.
6. The method of claim 1, wherein each process parameter involves two run times.
7. The method of claim 1, wherein a subset of the sample candidates is displayed in a compressed manner along the curve.
8. The method of claim 1, wherein the subset of the sample candidates exhibits strong statistical correlation among the subsets of the process parameters of the plurality of process parameters.
9. The method of claim 1, wherein the proxy model is executed by a domain-driven search algorithm, and wherein the domain-driven search algorithm is used for different types of Design Technique Co-optimization (DTCO) streams.
10. A method comprising: Multiple groups are created in the optimization space, each group including samples of different process parameters; Select a dominant sample in each group; as well as Collaborative optimization is performed using the dominant samples from each group.
11. The method of claim 10, wherein the first group of the plurality of groups includes samples related to front-end process (FEOL) process changes, and the second group of the plurality of groups includes samples related to back-end process (BEOL) process changes.
12. The method of claim 10, wherein the dominant sample is displayed in a compressed manner along the curve.
13. The method of claim 10, wherein the dominant sample comprises several hundred samples.
14. The method of claim 10, wherein the dominant sample in each group is selected using a domain-driven search algorithm.
15. A method comprising: Domain-driven search algorithms are used to generate power, performance, and area (PPA) models; The PPA model is used to evaluate the first and second metrics for each of the multiple process parameters. The PPA frontier sample set is updated based on the first metric data and the second metric data; the PPA frontier sample set includes multiple samples; and An analysis is performed on the PPA front sample set to generate the PPA Pareto front.
16. The method of claim 15, wherein the first metric is frequency and the second metric is power.
17. The method of claim 16, further comprising selecting a resolution for the first metric.
18. The method of claim 17, wherein the analysis of the PPA frontier sample set comprises: The range of the first metric is divided into multiple bins, and the size of each bin in the multiple bins of the first metric is determined by the resolution of the first metric; The plurality of samples are assigned to the plurality of bins associated with the first metric; as well as Update the optimal set of samples in each bin associated with the first metric.
19. The method of claim 18, further comprising collecting each optimal sample set from each bin to create a group of optimal frontier samples.
20. The method of claim 16, further comprising selecting a resolution for the second metric such that the analysis of the PPA frontier sample set includes: The range of the second metric is divided into multiple bins, and the size of each bin in the multiple bins of the second metric is determined by the resolution of the second metric; The plurality of samples are assigned to the plurality of bins associated with the second metric; Update the optimal set of samples in each bin associated with the second metric; as well as Collect each optimal sample set from each bin to create a set of optimal frontier samples.