Predicting optimal parameters for physical design synthesis

By combining a machine learning model of variational autoencoder and regression network, the resource and time problems of optimizing design flow parameter prediction in IC design are solved, achieving efficient and accurate design flow parameter prediction and optimizing power, congestion and timing objectives.

CN120937012APending Publication Date: 2025-11-11INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480025238.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-28
Filing Date
2024-04-18
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing IC design physical design synthesis, the prediction of optimization design flow parameters is usually a nondeterministic polynomial-time (NP) complete problem, resulting in high resource costs and long running time. Traditional methods require a large number of synthesis and build jobs to be performed on multiple machines.

Method used

A machine learning neural network model combining variational autoencoder (VAE) and regression network is adopted. Through training and interpolation, predictions of optimization design process parameters are generated, reducing dimensionality and constraining the potential space. The input gradient is used to search for the optimization design objective.

Benefits of technology

It effectively reduces runtime, improves the accuracy and efficiency of parameter prediction, avoids the need for multiple machines to run online in traditional methods, and achieves efficient prediction of optimization design goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937012A_ABST
    Figure CN120937012A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide enhanced systems and methods for optimal design process parameters for predicting optimal output objectives for physical design synthesis of a given IC design. A variational automatic encoder (VAE) is trained along with a regression network using a dataset that includes an integrated design build stream from a historical IC design to provide a training data representation of the dataset constrained to a potential space of the VAE. The system generates a feature vector based on the training data representation of the dataset and updates the feature vector with initial design characteristics of the given IC design. The system iteratively performs an input gradient search of the updated feature vectors to optimize an objective function of the design objective to identify locally optimized design parameters. The system identifies global optimization design process parameters for optimization design objectives based on the local optimization design parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This invention relates to electronic design automation (EDA) for integrated circuit (IC) design, and more specifically, to the prediction of optimal design flow parameters for power, congestion, and timing in the physical design synthesis of IC designs.

[0002] Industrial digital implementation of the construction process involves thousands of parameters that control algorithms to optimize circuits. Generating an optimized set of design parameters based on given design characteristics presents a significant challenge in delivering optimized power, performance, and area at the end of the construction process, given a legal setup and a routeable design. Currently, predicting and optimizing design process parameters is typically a nondeterministic polynomial-time (NP)-complete problem, or a computational problem for which efficient solutions have not yet been found.

[0003] Parameters for EDA physical design synthesis used in IC design are determined using various techniques and algorithms, and typically include trial and error handling. One technique for determining design flow parameters for physical design synthesis involves running numerous synthesis build jobs with different parameter values, collecting metrics on the results of parameter combinations, and generating recommendations for the parameters. This technique is exceptionally resource-intensive, consuming considerable runtime and processing power, requiring the design to be executed on many machines, and is constrained by the runtime of the build flow. Summary of the Invention

[0004] Embodiments of this disclosure provide systems and methods for predicting optimized design flow parameters for physical design synthesis of integrated circuit (IC) designs to achieve optimized output targets. The disclosed systems and methods enable the prediction of optimized design flow parameters for physical design synthesis to achieve optimized output targets, resulting in enhanced runtime performance. Optimized output targets may include power, congestion, and timing output targets.

[0005] According to one aspect, a method is provided, comprising: receiving initial design characteristics, design flow parameters, and design objectives from a design construction flow of a given integrated circuit (IC) design; receiving a historical dataset, the historical dataset including synthesis design construction flows from historical IC designs; using the historical dataset to train a variational autoencoder (VAE) along with a regression network to provide a training data representation of the dataset constrained to a latent space of the VAE; generating feature vectors based on the training data representation of the dataset, and updating the feature vectors using the initial design characteristics; iteratively performing an input gradient search on the updated feature vectors to optimize an objective function of the design objectives, thereby identifying locally optimized design parameters, wherein the training data samples are constrained to the initial design characteristics; and obtaining predictions of globally optimized design flow parameters based on the identified locally optimized design parameters for optimizing the design objectives.

[0006] This method effectively and efficiently predicts the optimization design process parameters for an optimization design objective that is trained offline using a VAE and a regression network. This method effectively eliminates the need for traditional parameter prediction systems to run multiple comprehensive design construction processes online across many machines to determine the design process parameters.

[0007] According to one or more preferred embodiments, the present invention effectively and efficiently achieves the prediction of optimized design flow parameters, such as power, congestion, and timing targets for physical design synthesis used to optimize IC design; reduces runtime; and overcomes some shortcomings of traditional parameter prediction systems.

[0008] According to one or more preferred embodiments, the present invention performs interpolation training of the VAE along with the regression network to smooth the constraints on training data points in the latent space and minimize reconstruction error, thereby achieving more accurate parameter prediction. Preferably, the present invention performs interpolation training of the VAE along with the regression network, generating interpolation vectors for the dataset in each period and combining the interpolation vectors with the training vectors to produce an enhanced dataset.

[0009] According to one or more preferred embodiments, the present invention decodes a random sample set from the training data representation of the dataset to generate feature vectors including design characteristics and design process parameters. The system updates the generated feature vectors by replacing the design characteristics of the generated feature vectors with initial design characteristics from the design construction process of a given design. Preferably, the present invention iteratively performs an input gradient search on the updated generated feature vectors to optimize the objective function of the design objective, thereby identifying locally optimized design parameters.

[0010] According to another aspect, a system is provided, comprising: a processor; and a memory, wherein the memory includes a computer program product configured to perform operations for predicting optimized design flow parameters for physical design synthesis of a given integrated circuit (IC) design, the operations including: receiving initial design characteristics, design flow parameters, and design objectives from a design construction flow of the given IC design; receiving a historical dataset including synthesis design construction flows from historical IC designs; using the historical dataset to train a variational autoencoder (VAE) along with a regression network to provide a training data representation of the dataset constrained to a latent space of the VAE; generating feature vectors based on the training data representation of the dataset, and updating the feature vectors using the initial design characteristics; iteratively performing an input gradient search on the updated feature vectors to optimize an objective function of the design objective, thereby identifying locally optimized design parameters, wherein the training data samples are constrained to the initial design characteristics; and obtaining predictions of globally optimized design flow parameters based on the identified locally optimized design parameters for optimizing the design objective.

[0011] According to another aspect, a computer program product is provided for predicting optimized design flow parameters for physical design synthesis of a given integrated circuit (IC) design. The computer program product includes: a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations including: receiving initial design characteristics, design flow parameters, and design objectives from a design construction flow of the given IC design; receiving a historical dataset including synthesis design construction flows from historical IC designs; using the historical dataset to train a variational autoencoder (VAE) along with a regression network to provide a training data representation of the dataset constrained to a latent space of the VAE; generating feature vectors based on the training data representation of the dataset and updating the feature vectors using the initial design characteristics; iteratively performing an input gradient search on the updated feature vectors to optimize an objective function of the design objective, thereby identifying locally optimized design parameters, wherein the training data samples are constrained to the initial design characteristics; and obtaining predictions of globally optimized design flow parameters based on the identified locally optimized design parameters for optimizing the design objective.

[0012] Other disclosed embodiments include a computer system and computer program product for predicting optimized design flow parameters for physical design synthesis of IC designs, thereby implementing the features of the disclosed methods. Attached Figure Description

[0013] Preferred embodiments of the invention will now be described by way of example only and with reference to the following figures:

[0014] Figure 1 This is a block diagram of an example computer environment used in conjunction with one or more disclosed embodiments of prediction of optimized design flow parameters for achieving optimized output targets of physical design synthesis of IC designs.

[0015] Figure 2 It is a block diagram of an example system for predicting optimized design flow parameters for implementing physical design synthesis of one or more disclosed embodiments;

[0016] Figure 3 An exemplary variational autoencoder (VAE) of one or more disclosed embodiments and Figure 2 A schematic diagram combining the system's regression network model;

[0017] Figure 4 A flowchart of example operations is provided for an example method of predicting optimization design process parameters for achieving the optimization design objective of physical design synthesis of one or more disclosed embodiments;

[0018] Figure 5 Two-dimensional example images of potential spaces for congestion, timing, and power output targets according to one or more disclosed embodiments are shown;

[0019] Figure 6 The example data stream training steps of a VAE and a combined regression network of one or more disclosed embodiments are illustrated schematically.

[0020] Figure 7 The diagram illustrates one or more disclosed embodiments of generating interpolation vectors for interpolation training of VAEs and combined regression networks.

[0021] Figure 8 An example interpolation training of a VAE and a combined regression network, representing one or more disclosed embodiments, is illustrated schematically.

[0022] Figure 9 An example initialization for optimizing gradient search according to one or more disclosed embodiments is illustrated schematically;

[0023] Figure 10 Exemplary optimization operations of one or more disclosed embodiments are schematically illustrated to obtain predictions of optimized design process parameters; and

[0024] Figure 11 Exemplary optimization operations for fine-tuning design process parameters are illustrated in one or more of the disclosed embodiments. Detailed Implementation

[0025] Existing parameter prediction systems suffer from limitations including running numerous synthesis build jobs, collecting metrics on the results of parameter combinations, and the inherent time constraints of generating parameter recommendations and the considerable runtime consumed by synthesis build jobs. Embodiments of this disclosure provide efficient and effective techniques for predicting optimized design flow parameters to achieve optimized output targets for physical design synthesis of IC designs. The disclosed systems and methods can perform the prediction of optimized design flow parameters for optimized output targets, resulting in a significant speedup and reduced computational resources compared to conventional parameter identification arrangements.

[0026] In one embodiment, the disclosed system uses an offline machine learning neural network model that combines a variational autoencoder (VAE) providing dimensionality reduction with a regression network to generate predictions of optimized design flow parameters for an optimized output objective of physical design synthesis for a given integrated circuit (IC) design. By combining the VAE with the regression network, the system is trained to model those features that influence the regression, while gaining the benefits of the VAE, including constructing a constrained latent space and generalizing input features. Because the latent space is centered at the origin, we can use the radius of a point in the latent space to compare with the radius of our training points in the latent space to determine whether we are inside or outside our training domain. Neural networks cannot extrapolate, and they can only interpolate where they have been trained; therefore, if interpolation is attempted in a latent space that has not yet been trained, the predictions are unreliable. If a simple autoencoder is used to reduce dimensionality, it is not easy to determine whether a data point is in a part of the trained latent space because the data point can be mapped anywhere. On the other hand, the VAE constrains the latent space such that all data points are mapped close to each other within the latent space.

[0027] In one embodiment, the system receives a historical dataset comprising features and metrics from the synthesis and construction flow of historical IC designs. Target metrics may include selected design objectives used at the end of the design construction flow to identify and optimize objectives such as congestion, timing, and power. For example, target labels and metrics may include target names and values ​​such as congestion (weighted average congestion estimate), timing (sum of latch-to-latch margins), and total power (sum of leakage power and dynamic power). The system uses the historical dataset to perform training on a VAE along with a regression network to provide a training data representation of the dataset constrained to the latent space of the VAE.

[0028] In one embodiment, the disclosed system performs interpolation training to smooth training data points in the latent space and minimizes the reconstruction error of the interpolated values ​​to achieve more accurate parameter predictions. The system can interpolate sampling points from the latent space representation of the dataset used in machine learning.

[0029] In one embodiment, the system generates feature vectors based on the training data representation of the dataset and updates these feature vectors with initial design characteristics. The system iteratively performs an input gradient search on the updated feature vectors to optimize the objective function of the design objective, thereby identifying locally optimized design parameters, where the updated feature vectors are constrained to the initial design characteristics of a given IC design. Prediction of global optimization design flow parameters is based on the identified locally optimized design parameters for optimizing design objectives, such as congestion, timing, and power output objectives. The system classifies samples of locally optimized constraint latent space based on their selected contributions to the objective function and selects optimization latent points from the constraint samples for prediction. During optimization, the radius of each point in the latent space is calculated based on the standard deviation of each dimension. Collapsed dimensions are treated as if they have a standard deviation of one. A loss function equal to the square of the radius is added; this is called drift. This causes the gradient search to remain within the latent space covered by the training data. The system obtains predictions of optimized design flow parameters for optimizing power, congestion, and timing objectives based on the selected optimization latent points. This method effectively predicts optimized design flow parameters using offline machine learning and the latent space of training, without requiring multiple integrated design construction flows to be run online on many machines as traditional parameter prediction systems do.

[0030] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies on the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0031] The following reference is made to embodiments presented in this disclosure. However, the scope of this disclosure is not limited to the specifically described embodiments. Rather, any combination of the following features and elements is contemplated for implementing and practicing the intended embodiments, regardless of whether different embodiments are involved. Furthermore, while the embodiments disclosed herein may achieve advantages over other possible solutions or prior art, whether a given embodiment achieves a particular advantage does not limit the scope of this disclosure. Therefore, the following aspects, features, embodiments, and advantages are merely illustrative and should not be considered elements or limitations of the appended claims unless expressly stated in the claims. Similarly, references to “the invention” should not be construed as a generalization of any inventive subject matter disclosed herein and should not be considered elements or limitations of the appended claims unless expressly stated in the claims.

[0032] Various aspects of this disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). Regarding any flowchart, depending on the technology involved, operations may be performed in a different order than that shown in a given flowchart. For example, again according to the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.

[0033] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively included in a collection of one or more storage devices, the collection of one or more storage devices collectively including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions used by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or pits / platforms formed in the main surface of the disk), or any suitable combination of the foregoing. Computer-readable storage media, as used in this disclosure, should not be construed as storing transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As those skilled in the art will understand, data is typically moved at certain incidental points in time during the normal operation of the storage device, such as during access, defragmentation, or garbage collection; however, this does not make the storage device transient, because the data is not transient when it is stored.

[0034] refer to Figure 1The computing environment 100 includes examples of environments for executing at least some of the computer code involved in performing the methods of the present invention, such as parameter prediction control unit 182, feature target metric dataset 184, and hyperparameters and values ​​186 in block 180. In addition to block 180, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end-user equipment (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuitry 120 and a cache 121), a communication structure 111, volatile memory 112, persistent storage device 113 (including an operating system 122 and block 180, as described above), a peripheral device set 114 (including a user interface (UI) device set 123, a storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. Public cloud 105 includes gateway 140, cloud coordination module 141, host physical machine set 142, virtual machine set 143, and container set 144.

[0035] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future capable of running programs, accessing networks, or querying databases such as remote database 130. As is well known in the field of computer technology, and depending on the technology, the performance of a computer-implemented method can be distributed across multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, specifically computer 101, to keep the presentation as simple as possible. Computer 101 can reside in the cloud, even... Figure 1 It is not shown in the cloud, and on the other hand, computer 101 does not need to be in the cloud unless it can be indicated with certainty to any extent.

[0036] Processor assembly 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, such as multiple cooperating integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be readily accessible by the threads or cores running on processor assembly 110. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor assembly may be located “off-chip.” In some computing environments, processor assembly 110 may be designed to work with qubits and perform quantum computing.

[0037] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowcharts and / or descriptive descriptions of the computer-implemented method included in this document (collectively, the “method of the invention”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the method of the invention. In computing environment 100, at least some of the instructions for performing the method of the invention may be stored in permanent storage device 113, within block 180.

[0038] Communication structure 111 is a signal transmission path that allows the various components of computer 101 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths that form buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.

[0039] Volatile memory 112 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 101.

[0040] The persistent storage device 113 is any form of non-volatile memory known now or developed in the future for use with a computer. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 may be a read-only memory (ROM), but typically at least a portion of the persistent memory allows data to be written, deleted, and rewritten. Some common forms of persistent storage include hard disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems or operating systems employing an open-source portable operating system interface type with a kernel. The code included in box 180 typically includes at least some of the computer code involved in performing the methods of the present invention.

[0041] Peripheral device set 114 includes a set of peripheral devices for computer 101. Data communication connections between peripheral devices and other components of computer 101 can be implemented in various ways, such as Bluetooth connectivity, near field communication (NFC) connectivity, connections made by cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 124 can be permanent and / or volatile. In some embodiments, storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 requires substantial storage (e.g., where computer 101 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed for storing very large amounts of data, such as a Storage Area Network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 comprises sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.

[0042] Network module 115 is a collection of computer software, hardware, and firmware that allows computer 101 to communicate with other computers via WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in network module 115.

[0043] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology known now or developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced by and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area (e.g., a Wi-Fi network). WANs and / or LANs typically include computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0044] End User Equipment (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in conjunction with computer 101. EUD 103 typically receives useful and available data from the operation of computer 101. For example, assuming computer 101 is designed to provide recommendations to the end user, these recommendations are typically transmitted from network module 115 of computer 101 to EUD 103 via WAN 102. In this way, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 can be client equipment, such as a thin client, heavy client, mainframe, desktop computer, etc.

[0045] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 can be controlled and used by the same entity operating computer 101. Remote server 104 represents a machine that collects and stores useful and available data used by other computers, such as computer 101. For example, if computer 101 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 101 from a remote database 130 of remote server 104.

[0046] Public cloud 105 is any computer system that can be used by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (particularly data storage (cloud storage) and computing power) without the need for direct, active management by users. Cloud computing typically leverages resource sharing to achieve scalability consistency and economy. Direct and active management of the computing resources of public cloud 105 is performed by the computer hardware and / or software of cloud coordination module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers constituting the host physical machine set 142, which is the entirety of physical computers in and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud coordination module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allow public cloud 105 to communicate via WAN 102.

[0047] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically appear as actual computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a characteristic known as containerization.

[0048] Private cloud 106 is similar to public cloud 105, except that computing resources are available only to a single enterprise. While private cloud 106 is depicted as communicating with WAN 102, in other embodiments, private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by standardization or proprietary technology that enables coordination, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 105 and private cloud 106 are both part of a larger hybrid cloud.

[0049] Embodiments of this disclosure provide systems and methods for predicting design flow parameters to achieve optimized output targets for physical design synthesis of an IC design, avoiding the generally extreme runtime of traditional parameter prediction deployments. In the disclosed embodiments, the system receives initial design characteristics, design flow parameters, and design targets to be optimized for a given IC design's physical design synthesis construction flow. The system uses a machine learning neural network model that combines a variational autoencoder (VAE) providing dimensionality reduction with a regression model to predict optimized design flow parameters for the optimized design targets of a given IC design. The system uses a dataset of design characteristics, design parameters, and target metrics extracted from a database of historical synthesis design construction flows to fundamentally train the VAE and the combined regression network together. The VAE and the combined regression network generate predictions of optimized design flow parameters for optimizing design targets such as optimized power, congestion, and timing targets for the physical design synthesis construction flow of a given IC design.

[0050] Figure 2 An example system 200 for implementing enhanced parameter prediction in one or more of the disclosed embodiments is shown. System 200 uses a machine learning neural network model 300 that includes a variational autoencoder (VAE) combined with a regression network model, such as... Figure 3 As shown, this is used to achieve enhanced parameter prediction.

[0051] System 200 can use offline machine learning with machine learning neural network model 300 to efficiently and effectively predict the parameters of the optimization design process, avoiding the online operation of multiple integrated design and construction processes on many machines, as is done by traditional parameter prediction systems.

[0052] System 200 includes controller 202 and parameter prediction control unit 182, for example with... Figure 1Together with the processor computer 101, it controls the operation of the machine learning neural network model 300 to achieve prediction of optimized design flow parameters. In the disclosed embodiments, the system 200 includes a feature target metric dataset 184 and hyperparameters and values ​​186, which are used with the controller 202 and the processor computer 101 to achieve enhanced prediction of optimized design flow (physical synthesis) parameters for optimized congestion, timing, and power targets for the disclosed embodiments.

[0053] System 200 uses a feature target metric dataset 184 to train a machine learning neural network model 300. This dataset includes design characteristic features, design parameter features, and target metrics from historical synthesis design construction flow files. System 200 receives initial design characteristics, design flow parameters, and design objectives to be optimized for a given IC design's design construction flow. The design objectives to be optimized include selected design objective outputs, such as congestion, timing, and power. In one embodiment, system 200 uses hyperparameters and values ​​186, such as weights for regression, reconstruction, and regularization of the machine learning neural network model 300.

[0054] Also refer to Figure 3 The example structure of the offline machine learning neural network model 300 shown includes a variational autoencoder (VAE) 302 providing dimensionality reduction, combined with a regression network 306 of the disclosed embodiment. The VAE 302 and the combined regression network 306 predict optimized design flow parameters for optimized design objectives (e.g., congestion, timing, and power objectives) for the synthesis construction of a given integrated circuit in the disclosed embodiment. An example dataflow operation of the VAE 302 and the combined regression network 306 is referenced. Figure 6-11 Show and describe.

[0055] In the disclosed embodiments, VAE 302 provides dimensionality reduction, constrains the latent space 307, and allows interpolation of sampled points in the latent space 307 from the machine learning latent space representation of the feature objective metric dataset 184. VAE 302 can be implemented, for example, using TensorFlow and Keras layers. TensorFlow is an open-source end-to-end platform for executing libraries for multiple machine learning tasks, while Keras is a high-level neural network library that runs on top of TensorFlow. Both provide high-level application programming interfaces (APIs) for building and training models.

[0056] The VAE 302 of the machine learning neural network model 300 includes a neural network encoder 304, a neural network decoder 305, and a latent space generally indicated by 307. A regression network 306, together with the VAE 302, is used for training, inference, and optimization operations according to the disclosed embodiments. The encoder 304, decoder 305, and latent space 307 of the VAE 302, combined with the regression network 306 of the disclosed embodiments, produce predictions of optimized design flow parameters for an optimized synthesis and construction flow target for a given IC design.

[0057] The latent space 307 of the VAE 302 implements a machine learning latent space representation of the encoder 304 of the combined VAE 302 and regression network 306 of the disclosed embodiment, which inputs the training feature metric dataset 184. The latent space 307 includes, for example, a random sample layer consisting of multiple nodes (e.g., 35 nodes). Each of the 35 latent space nodes in the example represents one dimension of the latent space 307. The latent space 307 can be sampled using a first diagonal multivariate Gaussian #1, 308 or a second diagonal multivariate Gaussian #2, 309. The first diagonal multivariate Gaussian #1, 308 is used for basic training and interpolation training of the regression network 306 of the combined VAE 302 and the disclosed embodiment. Using the feature target metric dataset 184, the neural network learns the mean Zmean 342 and standard deviation Zstddev 344 corresponding to each point in the latent space 307, which the diagonal multivariate Gaussian #1 308 uses to generate a sample vector Zsample 345. The second diagonal multivariate Gaussian #2,309 is used to generate interpolation vectors for the interpolation training and initialization operations of the optimized gradient search in the disclosed embodiment. Using the feature target metric dataset 184, the diagonal multivariate Gaussian #2,309 is fitted to all points in the latent space 307 corresponding to all vectors in the feature target metric dataset 184 and is used to generate the sample vector Zrandom 346.

[0058] like Figure 3 As shown in the machine learning neural network model 300, the data input layer X 336 of the encoder 304 represents the design features 310 and design flow parameters 312 of the feature target metric dataset 184, where the target output of the regression network 306 includes the target output Y' 318. The decoded or reconstructed output X' 348 of the decoder 305 includes the reconstructed design features 314 and the reconstructed design flow parameters 316 (e.g., names and values). For example, Figure 2 The feature target metric dataset 184 was obtained from the historical synthesis design construction flow file of historical IC designs. The target output Y' 318 of the regression network 306 includes selected targets, such as congestion, timing, and power.

[0059] In one disclosed embodiment of VAE 302 in System 200, an example neural network encoder 304 includes two fully connected dense layers 320, 322 with rectified linear unit (ReLU) activation functions. Encoder 304 includes multiple nodes (e.g., 256 nodes) on the outer layer 320 and multiple nodes (e.g., 128 nodes) on the inner layer 322. The ReLU activation function in neural network encoder 304 can define how a weighted sum of inputs is transformed into the output of the neural network encoder. For example, the ReLU activation function of encoder 304 can include a piecewise linear function that outputs the input directly if it is positive, and zero otherwise. The example model encoder 304 using the ReLU activation function can be trained relatively easily to achieve efficient performance of VAE 302.

[0060] In the disclosed embodiments, the example decoder 305 of VAE 302 can be configured as the opposite of encoder 304. As shown, decoder 305 includes, for example, two fully connected layers 324, 326 with ReLU activation functions, a first plurality of nodes (e.g., 128 nodes) on inner layer 324 and a second plurality of nodes (e.g., 256 nodes) on outer layer 326, wherein a plurality of sub-output nodes (e.g., 200 nodes) of decoder output X' 348 of decoder 305 are reconstructed using linear activation.

[0061] For example, the data input layer X 336 of encoder 304 includes multiple sub-input nodes, such as 200 nodes, including 26 nodes for design features 310 and 176 nodes for design flow parameters 312. Similarly, the decoding output X'348 of decoder 305 includes multiple sub-nodes, such as 200 nodes, including 26 nodes for decoding design features 310 and 176 nodes for decoding design flow parameters 312.

[0062] In the disclosed embodiments, the regression network 306 may include two fully connected layers 330, 332, each layer having 64 nodes using the Tanh activation function and 3 nodes using linear activation on the target output Y' 318. For example, the Tanh activation function is a hyperbolic tangent sigmoid function ranging from -1 to 1, which can model nonlinear boundaries. The Tanh activation function can be applied at different scales without losing its effectiveness. In system 200, the target output Y' 318 of the regression network 306 may include congestion (weighted average congestion estimate), timing (zero-based latching to latch the sum of negative margins), and total power (the sum of leakage power and dynamic power).

[0063] like Figure 3As shown, VAE 302 offers multiple configurations, including individual configurations for encoder 304, decoder 305, and regression network 306, depending on the data flow used for training, inference, and optimization operations. As illustrated, the individual paths shown by dashed lines provide corresponding connections from Zmean 342, Zsample 345, or Zrandom 346 to latent space 307 or connection block Z347. Block Z347 is not a layer; it represents a connection for three alternative VAE configurations: Zmean 342 for inference operations, Zsample 345 for training operations, or Zrandom 346 for interpolation training and optimization operations. As shown, VAE 302 includes one of the alternative connections via connection block Z347 to decoder 305 and regression network 306: Zmean 342, Zsample 345, or Zrandom 346.

[0064] Zsample 345 is provided by sampling with the first diagonal multivariate Gaussian #1 308, which is used during basic training and interpolation training of VAE 302, but not during inference or optimization (i.e., the first diagonal multivariate Gaussian #1 is not used during optimization initialization or optimization operations). Zsample 345 of the first diagonal multivariate Gaussian #1 308 is centered on each of the Zmeans of the training data, including some standard deviation Zstddev 344 relative to the Zmean. The output Zrandom 346 of the second diagonal multivariate Gaussian function #2 309 is used to generate the interpolation vector and for optimization initialization of the optimization operations. The output Zrandom 346 of the second diagonal multivariate Gaussian #2 309 is a single distribution with zero mean and standard deviation (equal to the root mean square of the Zmean 342 of the training data).

[0065] The combination of VAE 302 and regression network 306 is trained together using the feature target metric dataset 184 (e.g., trained simultaneously) to constrain the latent space 307 and provide a latent space representation of the dataset. VAE 302 constrains the latent space 307 such that all training data points are mapped close to each other in the latent space 307. During the basic and interpolation training of VAE 302, a first diagonal multivariate Gaussian #1, 308 samples the Zmean 342 and Zstddev 344 of the output encoder 304, providing the output Zsample 345.

[0066] During training, the data stream output Zmean 342 represents the mean of the encoded latent space sample points, while Zstddev 344 represents the variance or standard deviation of the sample points. The encoded data Zmean 342 and Zstddev 344 are applied to a first diagonal multivariate Gaussian #1, which provides conventional VAE sampling and data stream output Zsample 345 during basic and interpolation training of VAE 302. The 308 data points from the first diagonal multivariate Gaussian #1 are then applied to the decoder 305 and the regression network 306.

[0067] The example schematic illustration of the first training data sample distribution in latent space 307 shown below the first diagonal multivariate Gaussian #1, 308 represents the example output Zsample 345 of the first diagonal multivariate Gaussian #1, 308. These points represent the Zmean 342 position in latent space 307, and the curve represents the Gaussian distribution around the Zmean 342 position with a standard deviation Zstddev 344. Although only four distributions are shown, there are actually distributions around each Zmean 342 position. For clarity, we only show four. The example schematic illustration of the second training data sample distribution in latent space 307 shown above the second diagonal multivariate Gaussian #2, 309 represents the output Zrandom 346 of the second diagonal multivariate Gaussian #2, 309. Both the first and second illustrations include the same training data points representing the encoder output Zmean 342. Each training data point represents the synthesis construction process, including design characteristics 310 and design flow parameters 312 of a specific historical IC design obtained from the feature target metric dataset 184.

[0068] For the first diagonal multivariate Gaussian #1,308, each training data point has a distribution in each dimension of the latent space 307; however, only the four training data distributions illustrated are shown in one dimension. In the first diagonal multivariate Gaussian #1,308, each of the four training data distributions is centered at one of the four Zmeans and has a standard deviation Zstddev relative to the corresponding Zmean. The training data sample distribution shown represents the sampled output Zsample 345 in the latent space 307 of the first diagonal multivariate Gaussian #1,308.

[0069] For the second diagonal multivariate Gaussian #2,309, only one distribution exists in the latent space 307. The single distribution of the output Zrandom 346 has multiple dimensions of the latent space 307, but the training distribution illustrated is shown in one dimension. As shown, the output Zrandom 346 of the second diagonal multivariate Gaussian #2,309 is used to generate the interpolation vectors for interpolation training and the initialization for the optimization gradient search of the optimization operation. The output Zrandom 346 is a single distribution with zero mean and the standard deviation of each dimension is equal to the root mean square (RMS) of the Zmean of the training data.

[0070] In system 200, design characteristics 310 include design-specific features that are typically static or unchanging. Design characteristics 310 may include the height and width of the design's placeable area; area utilization (the amount of the total placeable area used in the design); the number of gates (the number of instances, i.e., the amount of cells used); the number of networks (connecting instances); and the number of connections between instances (receive pins); design characteristics 310 also include the number of metal layers used for routing; the number of routing blocks in the design; the number of ports (main inputs / outputs) in the design; the average 1% placement pin density in the design; the average 1% pin density of the placement pins in the design; and the voltage of the design; design characteristics 310 also include the worst-case margin of latch-to-latch paths; the number of fault timing paths; clock gating switching factor (dynamic power); system clock cycle; latch output switching factor (dynamic power), etc.

[0071] Design flow parameters 312, representing design flow parameters, may include a set of parameter names and combinations of values. In the disclosed embodiments, design flow parameters are converted into binary features called runids. Design runids make manual parameter selection easier and continue the practice of automatic parameter optimization to maintain consistency. A runid is a small subset of parameters to which specific values ​​are applied, and together they define a useful optimization function. This additional layer of abstraction allows the function to be transparent to changes in the underlying programming. If the appropriate runid is selected, the input to the neural network is set to 1; otherwise, it is set to 0. A runid can represent a logic synthesis parameter with no, downward, flowing, flattening, crushing, destroying, or selected values ​​or ranges of values ​​or ranges; an adder configuration with fluctuating, fast, and fastest values ​​or ranges; and an area effort with values ​​or ranges of values ​​or ranges of 0-99. Other runids can represent placement type parameters with pin expansion factors in the range of values ​​or ranges of 0.0-2.0, target density in the range of values ​​or ranges of 30-80, and true or false timing-driven attraction. Other runids can represent clock tree synthesis, such as latch fanout, latch configuration, and multipliers, each with its own value or range. Other runids can represent extended type algorithm parameters. Other runids can represent global route diffusion iterations with values ​​or ranges of 0, 5, 10, 15, and 20, and pin-aware density layout diffusions with values ​​or ranges of 0, 5, 10, 15, and 20. Other runids can represent late-stage design flow optimization parameters, such as coarse selection iterations with values ​​or ranges of 1, 2, 3, and 4, fine selection iterations with values ​​or ranges of 1, 2, 3, and 4, coarse selection diffusions with true / false values ​​or ranges, fine selection diffusions with true / false values ​​or ranges, and particularly detailed layouts with true / false values ​​or ranges. Other runids can represent power optimization type parameters, such as low voltage (Vt) percentage with values ​​or ranges of 0.01, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30, and 0.60, and area timing margin with values ​​or ranges of 5 and 10.

[0072] System 200 can use Kendall correlation on the feature objective metric dataset 184 to isolate important design flow parameters 312 for the output Y' 318 of the congestion, time series, and power objectives to be optimized. Hyperparameters and values ​​186 provide, for example, regression weights, reconstruction weights, and L2 regularization weights; training data batch size, such as 256, 512, or 1024; multiple epochs, such as 15,000 epochs; and Adam optimizers with step learning rates of 1e-3 and 1e-4 at 25% and 95% of the epochs.

[0073] refer to Figure 4System 200, for example, uses feature target metric dataset 184, hyperparameters and values ​​186, and utilizes controller 202 and parameter prediction control unit 182. Figure 1 The processor computer 101 implements the operation of example method 400 in Figure 4 In the figures, the same reference numerals are used Figure 3 The same or similar components.

[0074] As shown in box 402, system 200 accesses a database of historical synthesis design construction processes and receives a feature target metric dataset 184 to simultaneously train VAE 302 and regression network 306. The feature target metric dataset 184 used to train VAE 302 and regression network 306 includes design features 310 and design parameters 312, including a set of parameter names and value combinations (i.e., runids), and a target metric or target output Y' 318, such as... Figure 3 As shown, in addition to accessing a database of historical integrated design and construction processes to identify datasets, system 200 can also identify other data such as designer builds and regression development data to provide datasets. At box 404, system 200 identifies feature vectors and outputs target vectors from the datasets for training the VAE 302 and regression network 306.

[0075] At box 406, system 200 provides or constructs a machine learning neural network model 300, which includes a dimensionality-reduced VAE 302 combined with a combined regression network 306 to obtain predictions of optimized design flow parameters of the disclosed embodiments. As described above. Figure 3 The structure of VAE 302 and the combined regression network 306 is schematically shown. In box 408, system 200 can use the vectors identified at box 404 to perform optional basic training of VAE 302 together with the combined regression network 306 to provide a training data representation of the dataset constrained to the latent space 307 of VAE 302. System 200 can use unsupervised machine learning with feature vectors of the training data and supervised machine learning with target vectors of the training data output to perform training of VAE 302 together with the regression network 306 to provide a training data representation of the dataset constrained to the latent space 307 of VAE. In box 408, basic training may include information about... Figure 6 The example data stream of the basic training operation 600 is schematically shown and described.

[0076] At box 410, system 200 alternatively performs interpolation training on VAE 302 along with regression network 306. Note that basic training at box 408 is optional, not a prerequisite for interpolation training at box 410. In the disclosed embodiments, interpolation training can provide improved results for optimizing basic training. At box 410, system 200 can use the identified vectors at box 404 to provide a training data representation of the dataset constrained to the latent space 307 of VAE 302. At box 410, system 200 can generate interpolated vectors by running the training data representation of the sampled dataset per period through decoder 305 to produce an augmented dataset for interpolation training operations. At box 410, interpolation training can include generating interpolated vectors during the vector generation phase, such as... Figure 7 As illustrated in the diagram, such as Figure 8 The diagram illustrates the interpolation training phase.

[0077] In box 410, interpolation training provides a form of regularization. The purpose of regularization is to allow encoder 305 to be generalized, that is, interpolated. In one embodiment, additional L2 regularization of encoder 304 is useful.

[0078] In box 412, system 200 accesses initial design features from a design build flow for a given IC design. The design build flow is an early build flow following initial placement and early optimization of a given IC design. The design build flow is implemented prior to the synthesis build flow for the given IC design. The design build flow includes features such as design features to be optimized and design flow parameters, as well as output targets (e.g., congestion, timing, and power) to be optimized at the end of the build flow. Design features from the design build flow for a given IC design are used for optimization to predict optimized design flow parameters for the disclosed embodiments.

[0079] In box 414, system 200 selects a set of random samples from latent space 307 and decodes the random samples to generate a complete feature vector including initial design characteristics and design flow parameters. In box 416, system 200 replaces the design characteristics in the generated feature vector with the received initial design characteristics for a given IC design to provide an updated feature vector set. System 200 uses this updated feature vector set to perform optimization to predict the optimized design flow parameters for the output objective of the given IC design. Figure 9 The operations at boxes 414 and 416 are illustrated schematically.

[0080] In box 418, system 200 performs an input gradient descent search on an updated set of feature vectors from a dataset constrained to initial design characteristics to optimize the objective function of the design objective, thereby identifying locally optimal design parameters. For example, system 200 optimizes the objective function of the output objective using relative weights with the importance of the constraints to keep the optimization within a latent space defined by the training data and to reliably predict locally optimal design parameters. In the disclosed embodiments, importance weights are assigned to contributions to the objective function, including, for example, congestion, timing and power, as well as drift components and reconstruction errors. During the input gradient descent search, the initial design characteristics remain fixed with constraints to keep the optimization within a latent space defined by the training data to generate reliable recommended design parameters. The operation in box 418 is, for example, in… Figure 10 As schematically shown, in one embodiment, at block 422, system 200 classifies local optimization design parameters based on the objective function of the design objective to identify global optimization of the local optimization design parameters used to optimize the output objective. At block 424, system 200 optionally performs threshold quantization of the global optimization design parameters. Alternatively, at block 426, system 200 optionally performs fine-tuning of the global optimization design parameters. The operation at block 426 is schematically shown, for example, in... Figure 11 In box 428, system 200 obtains a prediction of the optimized output target based on the global optimization design parameters.

[0081] Figure 5 A latent space sample image 500 of example projections for each output target, including congestion 502, timing 504, and power 506, is shown, representing one or more disclosed embodiments. The example image 500 includes a two-dimensional latent space for each output target, while the latent space 307 includes, for example, 35 dimensions. Corresponding first images 508, 510, and 512 for congestion 502, timing 504, and power 506 illustrate latent space projections and rotated images, where each point or data point represents a physical synthesis of design characteristics and parameters / values ​​from the feature target metric dataset 184. In congestion image 508, the right-hand shading legend represents the scale of the projection of the scaled and normalized congestion output values ​​rotated with respect to the illustrated vertical axis Z(1) and horizontal axis Z(0), approximately from about -4 to +4. In timing image 510, the right-hand shading legend represents the scale of the projection of the scaled and normalized timing output values ​​rotated with respect to the illustrated vertical axis Z(1) and horizontal axis Z(0), approximately from about -0.5 to +2.5. In power image 510, the shaded legend on the right represents the scale of the projection of the scaled and normalized power output values ​​rotated with respect to the illustrated vertical axis Z (1) and horizontal axis Z (0), approximately from -3 to +3.

[0082] The corresponding second images 514, 516, and 518 show how the values ​​of output target congestion 502, timing 504, and power 506 change in the latent space plane, respectively. Each of the individual images 514, 516, and 518 shows a different plane because the principal axes are different for congestion 502, timing 504, and power 506. The corresponding third images 520, 522, and 524 show the reconstruction errors in the axis planes for the indicated vertical axis Z(1) and horizontal axis Z(0), respectively. The corresponding right-hand shading legend indicates the error increasing upwards from zero (0). The corresponding third images 520, 522, and 524 can show where the effective features are located. Low error values ​​in images 520, 522, and 524 indicate good reconstruction; high errors indicate that the features mapped to that region of the latent space are unreliable for parameter prediction.

[0083] Figure 6 An example data stream of the basic training operation 600 of the VAE 302 and the combined regression network 306 of the disclosed embodiment is illustrated schematically. In one embodiment, system 200 trains the combination of VAE 302 and regression network 306 to constrain the latent space 307 and provide a latent space representation of the feature target metric dataset 184.

[0084] In the disclosed embodiments, system 200 identifies feature vectors (Xs) and output target vectors (Ys) from a feature target metric dataset 184 to train one or more VAEs 302 and combined regression networks 306 of the disclosed embodiments. Feature vectors Xs include, for example, design features 310 and design flow parameters 312. The input layer X 336 of encoder 304 receives feature vector data input 602 Xs. The output X' 346 of decoder 305 provides decoded or reconstructed feature vectors 608 X's. The output Y' 318 of regression network 306 provides a predicted output target vector 614 Y'. In the disclosed embodiments, system 200 calculates a reconstruction error, which is a measure of the difference between reconstructed features 608 X's and 610 Xs, where 610 Xs is the same data as the feature vector data input 602 Xs; and calculates a regression error, which is a measure of the difference between the actual target 612 Ys and the predicted output target 614 Y's. System 200 performs basic training of VAE 302 and combined regression network 306 using forward and backward propagation of training feature vector data input 602 Xs together to minimize Kulbeck-Leibler divergence and mean squared error between the training data points of reconstruction and regression.

[0085] Basic training occurs in multiple periods. Each period represents one iteration of the dataset using updated neural network weights and biases through the machine learning neural network model 300, although the dataset is typically divided into multiple batches, with updates occurring in each batch. Basic training involves encoding by encoder 304 by providing Zmean 342 and Zstddev 344 to the first diagonal multivariate Gaussian #1, 308, which provides input to decoder 305 and regression network 306 as Zsample 345. Zsample 345 is decoded by decoder 305 by providing the reconstructed feature vector 608X. The output Y' 318 of regression network 306 provides the predicted output target vector 614 Y'.

[0086] During training operation 600, the KL divergence loss term pushes each sampling distribution toward a multivariate Gaussian distribution with zero mean and unit variance. For each dimension of the latent space 307, these distributions tend toward one of two steady states. The first state distributes the sample means with small variances tending toward zero across the latent space 307. While the sample distributions may overlap during the early stages of training operation 600, this overlap decreases during the later stages of training.

[0087] The second state can move all sampling units to zero and all sampling variances to one. When the sampling units are moved to zero or close to zero, the dimension can no longer encode any information about the features, and the dimension of the latent space is said to have collapsed. An interesting behavior of latent space dimension collapse is that if latent space 307 is defined as having more dimensions than the dimensions used by the VAE to encode features, the additional dimensions collapse. The number of collapsed dimensions can be calculated, and latent space 307 can be reduced accordingly. To determine if a dimension has collapsed, the mean of the variances is compared to the variance of that mean. When the mean of the variances is large, this indicates that the dimension has collapsed. To prevent latent space 307 from completely collapsing, the effective KL weights can be adjusted indirectly, and the reconstruction weights can be adjusted, where latent space collapse is caused, for example, by excessively high KL weights or excessively low reconstruction weights.

[0088] Reference Figure 7 and Figure 8 , Figure 7 The diagram schematically illustrates an example operation of a regression network 306 combining VAE 302 and one or more disclosed embodiments for generating interpolation vectors for each period, and Figure 8 The diagram illustrates the interpolation training performed for each period using interpolation vectors.

[0089] exist Figure 7In the example operation 700 for generating the set of interpolated vectors, a random set of points is selected in the latent space at each period, and the decoder 305 of VAE 302 is used to reconstruct a vector for the random set of points. The output Zrandom 346 of the second diagonal multivariate Gaussian #2, 309 is used to generate the interpolated vectors. As shown in the figure, the latent space vector Zrandom 346 is applied to the decoder 305 and the regression network 306. The reconstructed output layer 348 X' of the decoder 305 provides the generated interpolated vector 702 Xi. The regression target output Y' 318 of the regression network 306 provides the corresponding output target.

[0090] Figure 8 The illustration shows the method of use. Figure 7 An exemplary interpolation training operation 800 involves interpolating the interpolated vectors generated in each period to VAE 302 and the combined regression network 306. In one embodiment, system 200 performs interpolation training to improve the reliability of predictions or recommendations for the design flow parameters to be optimized and to reduce regression error. Figure 8 In the process, sampling with the first diagonal multivariate Gaussian #1 308 provides Zsample 345, which is used for interpolation training of VAE 302 (e.g., compared with that used for...). Figure 6 The basic training data flow path is the same as the data flow path.

[0091] exist Figure 8 In this process, interpolation training operation 800 begins with training vectors 802 Xs and the output target vector 818 Ys of regression network 306. In interpolation training operation 800, in each period, interpolation vector 803 Xi generated by VAE 302 is combined (e.g., concatenated) with training vector 802 Xs, and in each period, interpolation vector 820 Yi is combined with the output target vector 818 Ys. In the disclosed embodiment, system 200 calculates a reconstruction error based on the difference between the combined training vector 810 Xs and interpolation vector 812 Xi, and the difference between the combined reconstructed vector 804 X's and the reconstructed interpolation vector 806 X'i. In the disclosed embodiment, system 200 calculates a regression error, which is a measure of the difference between the combined actual target vector 818 Ys and interpolation vector 820 Yi and the combined predicted output target vector 814 Y's and interpolation vector 816 Y'i. Each iteration combines the training vector 802 Xs and the interpolation vector 803 Xi with the regression output target vector 818 Ys and the interpolation vector 816 Yi to generate an augmented dataset, as shown by combining the reconstructed vectors 804 X's and 806 X'i with the output target vectors 814 Y' and 816 Y'. The augmented dataset is dynamic and is generated by using VAE 302 and regression network 306 under training.

[0092] refer to Figure 9 and 10 System 200 uses VAE 302 and combined regression network 306 to schematically illustrate the use of the disclosed embodiments. Figure 9 Example initialization operations 900 and 900 for optimizing input gradient search. Figure 10 The optimization operation 1000 includes iterative optimization of the input gradient search.

[0093] exist Figure 9 In example initialization operation 900, system 200 obtains a random sample set in latent space 307 from Zrandom 346 of the second diagonal multivariate Gaussian #2, 309. Output Zrandom 346 is processed by decoder 305 and regression network 306. Decoded output X' 348 includes decoded design characteristic value Xo 902 and decoded design flow parameter Xo 904. Regression network target output Y' 318 includes decoded target Yo 906. System 200 obtains input 907, which includes initial design characteristic 908 and design flow parameter 910, as well as decoded design characteristic value Xo 902 and decoded design flow parameter Xo 904, wherein the initial design characteristic 908 and design flow parameter 910 will be optimized to predict optimized design flow parameters for the physical design synthesis construction flow of a given IC design. System 200 combines input 907 (initial design characteristic 908) and decoded design flow parameter Xo 904 from decoder 305 for use in... Figure 10 The next iteration of the optimized gradient search.

[0094] Figure 10 An exemplary optimization operation 1000 is schematically illustrated to predict the optimized design process parameters for one or more of the optimization output targets of the disclosed embodiments. In such... Figure 10 In the disclosed embodiment shown, initial design feature 1002 is used ( Figure 9 The system 200 performs optimization using initial design characteristics 908 and modified design flow parameters 1024 for multiple iterations of the optimization operation. The system iteratively performs an input gradient search constrained to a latent space sample set of initial design characteristics 1002 to optimize an objective function of a design objective with relative weights of importance, thereby identifying local optimization design parameters for optimizing the design objective. In the disclosed embodiments, importance weights are assigned to contributions to the objective function, including, for example, congestion, timing, and power; reconstruction error; and drift components, to keep the optimization within the latent space defined by the training data and to enable reliable prediction of local optimization design parameters.

[0095] In the disclosed embodiments, system 200 receives data input design features 1002, which include initial design features 1002 implemented for physical design flow synthesis of a given IC design, and optimized design flow parameters 1004 (e.g., Figure 9 The decoder 305 decodes the design flow parameters X0 (904) for the input Zrandom 346. The encoder 304 encodes the initial design characteristic 1002 and the design flow parameters 1004, providing the encoder output Zmean 342 to the decoder 305 and the regression network 306. The output X' 348 of the decoder 305 provides the iterative design characteristic value 1006 and the iterative design flow parameters 1008 for each iteration of optimization. The target output Y' 318 of the regression network 306 provides the iterative target 1010. As shown, the initial design characteristic 1002 remains unchanged (modified by zero), and the design flow parameters 1004 being optimized are modified by the input gradient 1022. The modified design flow parameters 1024 provided by the modification and the same initial design characteristic 1002 are used for the next iteration of the optimization operation.

[0096] Figure 11 An exemplary fine-tuning operation 1100 for optimizing design flow parameters of one or more disclosed embodiments is schematically illustrated. Input 1102, or fine-tuning, includes design characteristics 1104 and design flow parameter values ​​1006 provided by gradient descent optimization. A vector set 1110 of multiple single vectors includes the invariant design characteristics 1104 and design flow parameters 1112 including some unquantized parameter features (e.g., some binary features of the optimized design parameter 1106) that should be quantized to zero or one. For each binary parameter feature to be quantized (e.g., the five binary parameter features with steps 0 and 1 shown in the steps), one feature is fixed for quantization at a time. As shown, the quantized design flow parameter 1114 includes the same design characteristics 1104 and binary parameter features 1116 quantized, for example, with a second parameter value of 0 to be fixed. For all binary features, zero or one is tried, and an objective function is computed for each vector using each parameter block or slice. Parameter values ​​are flipped to zero or one to identify the optimal parameter value (e.g., 0 or 1) based on the objective function. The quantization of eigenvalues ​​continues until all binary eigenvalues ​​are quantized to the optimal parameter values ​​of zero and one.

[0097] While the foregoing relates to embodiments of the present invention, other and further embodiments of the present invention may be designed without departing from the basic scope of the present invention, and the scope of the present invention is defined by the appended claims.

Claims

1. A method comprising: Receive initial design characteristics, design flow parameters, and design objectives from the design construction flow for a given integrated circuit (IC) design; Receive historical datasets, which include synthesis design construction flows from historical IC designs; The historical dataset is used to train a variational autoencoder (VAE) along with a regression network to provide a training data representation of the dataset constrained to the latent space of the VAE. Feature vectors are generated based on the training data representation of the dataset, and the feature vectors are updated using the initial design features; Iteratively perform an input gradient search on the updated feature vector to optimize the objective function of the design objective, thereby identifying locally optimized design parameters, wherein the training data samples are constrained to the initial design characteristics; as well as Based on the identified local optimization design parameters used to optimize the design objectives, predictions of global optimization design process parameters are obtained.

2. The method according to claim 1, further comprising: Decode a random sample set from the training data representation of the dataset to generate a feature vector that includes design characteristics and design process parameters; Furthermore, updating the feature vector using the initial design features includes: replacing the design features in the generated feature vector with the initial design features to provide the updated feature vector.

3. The method according to claim 1, wherein, Iteratively performing the input gradient search further includes: calculating the input gradient for each optimization period, and applying the input gradient to modify the next design flow parameters for the next optimization iteration.

4. The method according to claim 1, wherein, The prediction of obtaining global optimization design process parameters based on the local optimization design parameters for the optimization design objective further includes: quantizing at least one binary parameter element of the global optimization design process parameters.

5. The method according to claim 1, further comprising: The VAE is constructed including an encoder, the latent space, and a decoder; wherein the encoder includes one or more neural network layers, having a first plurality of nodes on the inner layer and a second plurality of nodes on the outer layer; wherein the decoder includes one or more neural network layers, having the second plurality of nodes on the inner layer of the decoder and the first plurality of nodes on the outer layer of the decoder; and wherein the latent space includes a random sample layer, the random sample layer including a plurality of nodes.

6. The method according to claim 1, wherein, The prediction of obtaining global optimization design process parameters based on the identified local optimization design parameters for the optimization design objective further includes: sorting the local optimization design parameters based on the objective function to identify the global optimization design process parameters.

7. The method according to claim 1, wherein, Training the VAE together with the regression network further includes: performing interpolation training of the VAE together with the regression network using a combination of interpolation vectors and training vectors to minimize the reconstruction error of the training data representation of the dataset.

8. The method according to claim 1, wherein, Training the VAE together with the regression network further includes: performing training of the VAE together with the regression network using unsupervised machine learning with vectors having features of the training data and supervised machine learning with vectors having output targets of the training data, to provide the training data representation of the dataset constrained to the latent space of the VAE.

9. The method according to claim 1, wherein, Training the VAE together with the regression network further includes: generating an interpolation vector for the dataset in each period and combining the interpolation vector with the training vector to generate an enhanced dataset.

10. The method according to claim 1, wherein, Iteratively performing the input gradient search further includes using hyperparameter values, which include at least one of regression weights, reconstruction weights, the number of periods for training, or the training batch size.

11. The method according to claim 1, wherein, The regression network is connected to a neural network layer of the VAE representing the latent space.

12. A system comprising: processor; as well as A memory, wherein the memory includes a computer program product configured to perform operations for predicting optimized design flow parameters for physical design synthesis of a given integrated circuit (IC) design, the operations including: Receive initial design features, design flow parameters, and design objectives from the design construction flow for the given integrated circuit (IC) design; Receive historical datasets, which include synthesis design construction flows from historical IC designs; The historical dataset is used to train a variational autoencoder (VAE) along with a regression network to provide a training data representation of the dataset constrained to the latent space of the VAE. Feature vectors are generated based on the training data representation of the dataset, and the feature vectors are updated using the initial design features; Iteratively performing an input gradient search on the updated feature vectors to optimize the objective function of the design objective, thereby identifying locally optimized design parameters, wherein the training data samples are constrained to the initial design characteristics; and Based on the identified local optimization design parameters used to optimize the design objectives, predictions of global optimization design process parameters are obtained.

13. The system of claim 12, further comprising: Decode a random sample set from the training data representation of the dataset to generate a feature vector that includes design characteristics and design process parameters; Furthermore, updating the feature vector using the initial design features includes: replacing the design features in the generated feature vector with the initial design features to provide the updated feature vector.

14. The system according to claim 12, wherein, Training the VAE together with the regression network further includes: performing training of the VAE together with the regression network using unsupervised machine learning with vectors having features of the training data and supervised machine learning with vectors having output targets of the training data, to provide the training data representation of the dataset constrained to the latent space of the VAE.

15. The system according to claim 12, wherein, Iteratively performing the input gradient search further includes: calculating the input gradient for each optimization period, and applying the input gradient to modify the next design flow parameters for the next optimization period.

16. The system according to claim 12, wherein, The prediction of obtaining the global optimization design process parameters further includes: quantizing at least one binary parameter element of the global optimization design process parameters.

17. A computer program product for predicting optimized design flow parameters for physical design synthesis of a given integrated circuit (IC) design, the computer program product comprising: A computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being executable by one or more computer processors to perform operations including: Receive initial design characteristics, design flow parameters, and design objectives from the design construction flow for a given integrated circuit (IC) design; Receive historical datasets, which include synthesis design construction flows from historical IC designs; The historical dataset is used to train a variational autoencoder (VAE) along with a regression network to provide a training data representation of the dataset constrained to the latent space of the VAE. Feature vectors are generated based on the training data representation of the dataset, and the feature vectors are updated using the initial design features; Iteratively performing an input gradient search on the updated feature vectors to optimize the objective function of the design objective, thereby identifying locally optimized design parameters, wherein the training data samples are constrained to the initial design characteristics; and Based on the identified local optimization design parameters used to optimize the design objectives, predictions of global optimization design process parameters are obtained.

18. The computer program product of claim 17, further comprising: The local optimization design parameters are sorted based on the objective function to obtain predictions of global optimization design flow parameters for the given IC design.

19. The computer program product according to claim 17, wherein, Training the VAE together with the regression network further includes: generating an interpolation vector for the dataset in each period and combining the interpolation vector with the training vector to generate an enhanced dataset.

20. The computer program product according to claim 17, wherein, Training the VAE together with the regression network further includes: performing training of the VAE together with the regression network using unsupervised machine learning with vectors having features of the training data and supervised machine learning with vectors having output targets of the training data, to provide the training data representation of the dataset constrained to the latent space of the VAE.

21. The computer program product of claim 16, further comprising: Decode a random sample set from the training data representation of the dataset to generate the feature vector that includes design features and design process parameters; Furthermore, updating the feature vector using the initial design features includes: replacing the design features in the generated feature vector with the initial design features to provide the updated feature vector.

22. A computer program comprising program code means, wherein when the program is run on a computer, the program code means is adapted to perform the method of any one of claims 1 to 11.