Domain randomization-based value decomposition multi-agent reinforcement learning method for analog circuit design migration

By introducing a domain-randomized value decomposition multi-agent reinforcement learning method into analog circuit design, the problems of low automation and poor efficiency in analog circuit design transfer are solved, and rapid automatic adaptation and efficient transfer are achieved when circuit conditions change.

CN121503398APending Publication Date: 2026-02-10FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510895977.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-30
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing analog circuit design migration methods require redesign when faced with changes in circuit conditions, resulting in wasted time and high computational complexity. They are particularly difficult to efficiently and automatically adapt to new operating conditions in complex circuits.

Method used

A value decomposition-based multi-agent reinforcement learning method based on domain randomization is adopted. By assigning an agent to each sub-circuit and using a hybrid network to model the relationship between the system-level circuit and the sub-circuit, the system is trained to automatically adapt and quickly transfer under different circuit conditions.

Benefits of technology

It enables rapid and automatic adaptation to changes in circuit conditions, improving the efficiency and automation of analog circuit design migration, and can efficiently solve the optimization and migration problems of complex system-level circuits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005475986250000031
    Figure BDA0005475986250000031
  • Figure BDA0005475986250000034
    Figure BDA0005475986250000034
  • Figure BDA0005475986250000035
    Figure BDA0005475986250000035
Patent Text Reader

Abstract

The invention belongs to the field of integrated circuit computer aided design / electronic design automation, and particularly relates to a value decomposition multi-agent reinforcement learning and domain randomization method which is used for solving the problem of design migration of a constrained complex system-level analog circuit. According to the method, a layered multi-agent reinforcement learning architecture is adopted, and an agent network and a hybrid network are respectively used for modeling the value of a sub-circuit and the overall value of a system-level circuit; multi-agent training is carried out under a randomized circuit working condition by adopting a domain randomization technology, so that the method can be quickly migrated to different circuit working points. The invention belongs to a fast migration algorithm for an analog circuit. The algorithm has relatively high simulation efficiency and operation time efficiency for migration optimization of a complex system-level analog circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of computer-aided design / electronic design automation (CAD / EDA) for integrated circuits, specifically involving a value decomposition-based multi-agent reinforcement learning method for the analog circuit design transfer problem. Background Technology

[0002] Design migration in analog circuits is crucial for quickly addressing different application scenarios by adjusting design parameters without altering the circuit topology. In applications such as analog-to-digital converters (ADCs), changes in operating conditions (e.g., sampling frequency) are common. However, analog circuit design is a time-consuming process, especially when circuit conditions change, often requiring designers to redesign, resulting in significant time wastage. Popular approaches formulate design migration as an optimization problem [1, 2], requiring numerous iterations to achieve convergence. As analog modules evolve towards system-level integration, the increasing complexity and scale lead to more complex designs and exponentially increasing optimization costs.

[0003] Currently, there has been considerable research on analog circuit design problems internationally, such as population-based optimization methods [3], gradient-based optimization methods [4], and Bayesian optimization methods [5]; and single-agent reinforcement learning methods based on graph neural networks (GCN-RL) [6] and two-layer deep learning network reinforcement learning methods (DNN-Opt) [7]. However, in terms of analog circuit design transfer, the convergence speed of analog circuit design transfer based on optimization methods largely depends on the circuit size. In the case of complex analog circuits, high-dimensional design parameters pose a challenge to computational complexity, which leads to the need for a large number of simulations to thoroughly explore and utilize the entire design space. Moreover, once the circuit conditions change, the optimization method needs to start iterating from scratch, resulting in a significant waste of time. Reinforcement learning-based methods either have extremely high computational complexity and cannot handle complex circuits, or still need to re-explore the design space in the transfer problem, resulting in unnecessary costs.

[0004] Multi-agent reinforcement learning methods for analog circuit design are still in their early stages. Existing MARL methods [8] face challenges in accurately and automatically decomposing the contributions between sub-blocks, which requires manually defining the specifications of complex analog system sub-modules based on the design experience of human experts. In terms of design transfer, the model cannot automatically and accurately adapt to new operating conditions when conditions change. Moreover, the transferability of multi-agent reinforcement learning in the automation of analog design has not been thoroughly studied.

[0005] Based on the current state of the technology, the inventors of this application aim to address the problems of low automation and poor efficiency in analog circuit design transfer methods. They utilize a hierarchical multi-agent deep neural network structure to model the behavior of sub-circuits and the overall system-level circuit. Each sub-circuit is assigned an agent to be responsible for its design and evaluation. Simultaneously, a hybrid network is designed at the top level to model the relationship between the behavior of each sub-circuit and the behavior of the system-level circuit, evaluating the impact of changes made by each sub-circuit on the system-level circuit. To improve the transfer capability of the algorithm, this method uses circuit conditions as random variables, training the multi-agent reinforcement learning algorithm under randomized circuit conditions to enhance its adaptability when encountering new circuit conditions. This invention will contribute to solving the complex system-level analog circuit optimization transfer problem.

[0006] References:

[0007] [1] Morteza Fayazi, Zachary Colter, Ehsan Afshari, and Ronald Dreslinski. 2021. Applications of artificial intelligence on the modeling and optimization for analog and mixed-signal circuits: A review. IEEE Transactionson Circuits and Systems I: Regular Papers 68, 6 (2021), 2418–2431.

[0008] [2]Georges GE Gielen and Rob A Rutenbar.2000.Computer-aided design ofanalog and mixed-signal integrated circuits.Proc.IEEE 88,12(2000),1825–1854.

[0009] [3]Po-Cheng Pan,Hung-Ming Chen,Chien-Chih Lin,et al.2013.PAGE:parallel agile genetic explorationtowards utmostperformance for analogcircuit design..InDATE.1849–1854.

[0010] [4]Zhaori Bi,Dian Zhou,Sheng-Guo Wang,and Xuan Zeng.2017.Optimizationand quality estimation of circuit design via random region coveringmethod.ACM Transactions on Design Automation ofElectronic Systems(TODAES)23,1(2017),1–25.

[0011] [5]Tianchen Gu,Wangzhen Li,Aidong Zhao,Zhaori Bi,Xudong Li,Fan Yang,Changhao Yan,Wenchuang Hu,Dian Zhou,Tao Cui,et al.2023.BBGP-sDFO:BatchBayesian and Gaussian Process Enhanced Subspace Derivative Free Optimizationfor High-Dimensional Analog Circuit Synthesis.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems(2023).

[0012] [6] Hanrui Wang, Kuan Wang, Jiacheng Yang, Linxiao Shen, Nan Sun, Hae-SeungLee, and Song Han. 2020. GCN-RL circuit designer: Transferable transistor sizing with graph neural networks and reinforcement learning. In 202057th ACM / IEEEDesign Automation Conference (DAC). IEEE, 1–6.

[0013] [7]Ahmet F Budak,Prateek Bhansali,Bo Liu,Nan Sun,David Z Pan,andChandramouli V Kashyap.2021.Dnn-opt:An rl inspired optimization for analogcircuit sizing using deep neural networks.In 2021 58thACM / IEEEDesignAutomation Conference(DAC).IEEE,1219–1224

[0014] [8]Jinxin Zhang,Jiarui Bao,Zhangcheng Huang,Xuan Zeng,andYeLu.2023.Automated Design of Complex Analog Circuits with Multiagent basedReinforcement Learning.In 202360th ACM / IEEE Design Automation Conference(DAC).1–6.https: / / doi.org / 10.1109 / DAC56929.2023.10247909 Summary of the Invention

[0015] The purpose of this invention is to provide a multi-agent reinforcement learning method based on domain randomization for value decomposition, addressing the design transfer problem in analog circuits.

[0016] Specifically, this method first assigns an agent to each sub-circuit based on the functional division of the system-level circuit. During training, a circuit condition is randomly generated. In each step of the Markov chain, each sub-circuit adjusts its parameters according to the actions proposed by each agent. Each agent proposes actions according to the following method:

[0017]

[0018] Where θ i These are the circuit conditions in this round of training. The function modeled for a single agent means that the corresponding sub-circuit proposes an action in the current state (circuit parameters). The expected benefit of parameter changes (whether the change in this sub-circuit improves the overall circuit performance) is considered. The final action taken by the agent is the one that maximizes the expected benefit among all possible actions, i.e., argmax. After each agent takes an action, the actual expected benefit (value) of that action is derived through the agent network.

[0019]

[0020] The outputs of each agent network are integrated into the upper-level mixing network as input. The mixing network, as the top-level network in the overall architecture, is responsible for estimating the expected return (value) of the entire system-level circuit after taking this action based on the expected return (value) measured by each agent.

[0021]

[0022] In the simulator, the modified system-level circuit is simulated to obtain its FOM (Form of Object) value. The algorithm calculates the reward as the difference in FOM values ​​before and after the modification. End-to-end training is performed on the networks of all agents and the mixing network using the TD (Digital Transformation Error) method.

[0023]

[0024] Use this function as the loss function for gradient backpropagation to update the parameters of all neural networks. Repeat the above steps until convergence.

[0025] After training convergence, during execution, a new circuit condition θ is given. newThe multi-agent system is directly applied to new circuit conditions and undergoes several training and adjustment steps, during which the system automatically adapts to the new circuit conditions. Subsequently, each agent independently selects actions and moves forward, generating a path to the vicinity of the optimal solution within the new design space. Once all circuit parameters meet the requirements, the algorithm stops executing and returns a feasible solution under the new circuit conditions.

[0026] The method proposed in this invention is applicable to general constrained single-objective analog circuit optimization transfer problems. The design space is normalized to [0,1]. d The optimization problem is then defined as follows:

[0027]

[0028] stc i (x)≤0,

[0029] i = 1, ..., N c

[0030] Input parameters: Circuit optimization target performance f1, f2, ... f m Circuit constraint performance Optimization range S, circuit operating conditions θ;

[0031] Output: A feasible solution that satisfies the design constraints.

[0032] Step 1: Divide the system-level circuit into multiple sub-circuits according to their functions, and assign an agent to each sub-circuit.

[0033] Step 2: Randomly generate circuit operating conditions θ i And a set of random transistor sizes, and initialization of the proxy network.

[0034] and hybrid networks

[0035]

[0036] Step 3: Each agent network uses the cross-entropy method (CEM) to suggest actions to be taken to the corresponding sub-circuit parameters based on the transistor size (defined as state) of the current circuit. Then, the transistor parameters (state) of the entire circuit are updated based on the actions of each agent.

[0037] Step 4: Each agent network evaluates the expected return (value) of its sub-circuits and then sends it to the hybrid network to calculate the value of the system-level circuit.

[0038] Step 5: The updated circuit is sent to the simulator for simulation, and a reward is calculated based on the difference in FOM values ​​before and after the update. This reward is used to calculate the TD error, which is used as the loss function and for backpropagation to train each network. Steps 2 through 5 are repeated during the training phase until the EVDMARL model converges.

[0039] Step 6: In the transfer phase, under the given target circuit conditions, each agent operates according to the network model trained in the training phase. Data generated under the new circuit conditions is input into the neural network to adapt to the new operating conditions. Hybrid networks can also automatically adjust their models to adapt to changes in operating conditions with minimal steps. When the generated circuit parameters satisfy the given circuit constraints, the design point is returned, and the algorithm terminates.

[0040] Specifically, the following conceptual explanations are given: A state refers to a vector representing system-level circuit design parameters, including transistor width, channel length, etc., represented as follows: [W1, L1, W2, L2, ... Wn, Ln], where Wi represents the width of the i-th transistor and Li represents the channel length of the i-th transistor. Merging global states allows each agent to consider the actions of other agents, thereby improving decision-making ability. An action is represented by a vector, indicating the adjustments made by each agent to the sub-circuit design parameters. The reward is calculated based on the change in FOM caused by the action, reflecting the improvement in system-level circuit performance. The reward for a single-step action is defined as...

[0041]

[0042] A substantial reward (e.g., +10) is provided to incentivize the agent to explore the domain when a specific constraint is achieved in a step. In the context of the migration problem, we avoid directly introducing rewards based on actual FOM values, as these design metrics can fluctuate with changes in circuit operating conditions. Therefore, we utilize a standardized reward formula. The value is defined as the expected future reward for taking a specific action in a given circuit state. This value is defined as follows:

[0043]

[0044] in,

[0045]

[0046] It is the cumulative reward starting from time t, with a discount factor of γ. i At each time step, we select the action with the highest value to maximize the potential reward, consistent with the objective of optimizing the circuit's FOM, as shown in the formula:

[0047]

[0048] In step 1, for example, a system-level ADC circuit is divided into four main sub-circuits based on function: a driver, a comparator, a DAC logic circuit, and a comparator logic circuit. Each sub-circuit is assigned an agent that provides suggestions for improving transistor parameters for each sub-circuit.

[0049] In step 2 of the method of the present invention, the individual agent network and the hybrid network have the following structures respectively:

[0050] Each agent network is an artificial neural network, consisting of an input encoder, an output layer, and five hidden layers. These layers store the state of the entire system-level circuitry. and the possible actions of the corresponding sub-circuits As input, the output is the value Q of the agent taking this action in this state. a To determine the optimal action for the sub-circuit based on the current circuit state, we sample an action from the sub-circuit's action space using the cross-entropy (CEM) method. This action has the highest value from the agent network. This demonstrates that, given the current circuit conditions, this action is the best choice.

[0051] Hybrid Networks: Considering that the design goal of each sub-circuit is to optimize the FOM of the system-level circuit, the relationship between agents can be modeled as a purely cooperative relationship, i.e., the value of the system-level circuit is positively correlated with the value of each agent. The relationship between the value of the sub-circuit and the value of the overall circuit should satisfy the constraints of the equations.

[0052]

[0053] To ensure a positive correlation between the value of the sub-circuit and the value of the system-level circuit, we introduce a hybrid network with a supernetwork architecture, where the weight parameters of the hybrid network need to be positive. A supernetwork with an absolute activation function is used to generate the weight parameters and biases of the hybrid network.

[0054] Specifically, in step 3, the method of the present invention selects the sub-circuit operation using the following steps.

[0055] Step 3.1: Initialize random variables using a multivariate normal distribution

[0056]

[0057] Step 3.2: Initialize the storage variable B = {}, where B is a container responsible for temporarily storing all actions of this trial.

[0058] Step 3.3: From the distribution random sampling

[0059] Step 3.4: Input the current state of the circuit into the agent network to obtain the value of the action in this state. Will Store the samples in container B. Repeat steps 3.3-3.4 until N samples have been collected. N is a hyperparameter.

[0060] Step 3.5: Take N from B best The action sample with the highest Q value is called the new Population. N best This is a hyperparameter.

[0061] Step 3.6: Calculation The mean and variance are calculated, and the distribution is updated. Repeat steps 3.2-3.6 M times, where M is a hyperparameter.

[0062] Step 3.7: From the distribution One action is randomly selected from the samples and chosen as the final action.

[0063] Specifically, in step 4, the method of the present invention utilizes the following steps to calculate the sub-circuit value function and the system-level circuit value.

[0064] Step 4.1: For each agent, the network input circuit state and the action taken by the agent are used to obtain the value function of that sub-circuit;

[0065]

[0066] Step 4.2: The hybrid network takes the value functions of each sub-circuit as input to obtain the value function of the entire circuit;

[0067]

[0068] In step 5, the method of the present invention utilizes the following steps to update the parameters of the agent network and the hybrid network.

[0069] Step 5.1: Input the updated circuit netlist into the simulator for simulation to obtain the FOM value.

[0070] Step 5.2: Calculate the reward based on the difference in FOM values ​​before and after.

[0071]

[0072] Step 5.3 Calculate the TD error based on the reward. The specific form of the TD error is as follows:

[0073]

[0074] Where r is the reward obtained after completing this step, which is the actual circuit performance improvement obtained through simulation by the simulator; To anticipate future rewards starting from the next step, we input the next state and the best action for that agent to take, thus obtaining the expected future reward for each agent's output. This reward is then obtained by integrating the results using a hybrid function to get the expected reward for the entire circuit. The expected future reward based on the current state. r+ and The difference lies in the fact that, having gone through this step, the reward *r* for that step is known. By calculating the square of the distance between the two, we can obtain the gap between them, thus making the future reward (value) estimated by the agent and the hybrid network closer to the actual reward.

[0075] Step 5.4: We use the TD error as the loss function to perform gradient backpropagation, thereby training the agent network and the hybrid network as a whole end-to-end.

[0076] In step 6 of the method of the present invention, the specific technical details of the domain randomization are as follows:

[0077] Since the Bendo agent reinforcement learning framework is an offline policy approach, the training database can be populated with data from different circuit operating conditions. This helps to train a more comprehensive and operationally-oriented situation-aware policy. This method uses the circuit operating point θ... i The search strategy is determined based on the condition, which is treated as a random variable, as shown below.

[0078] π tot =π(a t |s t ,θ)

[0079] In each training epoch, we randomly initialize the operating conditions and generate a design trajectory based on those conditions. This trajectory data is then merged into the training database. By enabling the reinforcement learning network to converge to different operating conditions, we expect it to have the ability to generalize to unseen operating conditions.

[0080] The advantages of the method of the present invention are:

[0081] (1) A hierarchical multi-agent network architecture is used to introduce a hybrid network to model the relationship between system-level circuits and sub-blocks as a higher-level structure, thereby automatically combining the value function of each sub-block.

[0082] (2) Domain randomization technique is used to enhance the generalization and adaptability of multi-agent networks under different working conditions. Therefore, during migration, the algorithm can quickly migrate to a new working point according to the change of working conditions.

[0083] (3) Compared with traditional optimization and reinforcement learning algorithms, this method can solve more complex system-level circuit migration optimization problems and has a higher degree of automation and migration efficiency. Attached Figure Description

[0084] Figure 1 A framework for a multi-agent reinforcement learning circuit transfer method based on domain randomization-based value decomposition.

[0085] Figure 2 ADC circuit diagram

[0086] Figure 3 A schematic diagram comparing the migration effects of different ADC circuit algorithms. Detailed Implementation

[0087] The method of the present invention is described below through specific examples.

[0088] Implementation example

[0089] We designed a complex system-level ADC circuit, the schematic of which is shown below. Figure 2 As shown. In the process of designing ADC circuits, it is usually necessary to create circuits for different frequency points. When encountering a new frequency point, design parameters need to be adjusted, leading to a complex and iterative process. To address this challenge, we use domain randomization in the 400-600MHz range to train our algorithm, and then use it to efficiently generate designs for 200MHz and 800MHz. In the ADC design process, four sub-circuits need to be carefully designed—the driver, comparator, DAC logic, and comparator logic. To optimize the design, we assign a proxy to each sub-circuit. The optimization objectives and constraints are shown below.

[0090]

[0091] st TD=THD<-49.5dB

[0092] ENOB represents the effect bit depth, and THD represents the total harmonic distortion. For each case, we compare the proposed EVDMARL method with TuRBO, DDPG, and MARL.

[0093] The results are as follows Figure 3 As shown, at 800MHz, EVDMARL using the DR algorithm is the only method that meets the specifications in 250 simulations. At 200MHz, although all algorithms meet the design specifications, our method is 6.8–8.3 times faster than the slowest algorithm. These results demonstrate that at the new operating frequency, our algorithm requires at least 34 simulations to generate a well-designed ADC circuit that meets the design specifications.

[0094] Table 1: Transfer results of different methods (800MHz)

[0095] method FOM Simulation times acceleration ratio This method 3.46(±0.03) 60 fail TuRBO 9.68(±0.08) 250 fail DDPG 11.68(±0.51) 250 fail MARL 6.96(±0.04) 250 -

[0096] Table 2: Transfer results of different methods (200MHz)

[0097] method FOM Simulation times acceleration ratio This method 1.23(±0.03) 34 6.8 TuRBO 1.23(±0.05) 181 1.3 DDPG 1.23(±0.01) 231 1 MARL 1.23(±0.02) 164 1.4

Claims

1. A multi-agent reinforcement learning method based on domain randomization for value decomposition in analog circuit design transfer learning, characterized in that, A hierarchical multi-agent reinforcement learning architecture is adopted, using agent networks and hybrid networks to model the value of sub-circuits and the overall value of the system-level circuit, respectively. Domain randomization is employed for multi-agent training under randomized circuit operating conditions, enabling the method to quickly transfer to different circuit operating points. The steps include: Input parameters: Circuit optimization target performance f1, f2, ... f m Circuit constraint performance Optimization range S, circuit operating conditions θ; Output: Feasible solutions that satisfy the design constraints; Step 1: Divide the system-level circuit into multiple sub-circuits according to their functions, and assign an intelligent agent to each sub-circuit; Step 2: Randomly generate circuit operating conditions θ i And a set of random transistor sizes, and initialization of the proxy network. and hybrid networks Step 3: Each agent network uses the cross-entropy method (CEM) to suggest actions to be taken to the corresponding sub-circuit parameters based on the transistor size (defined as state) of the current circuit. Then, the transistor parameters (state) of the entire circuit are updated based on the actions of each agent. Step 4: Each agent network evaluates the expected return (value) of its sub-circuits and then sends it to the hybrid network to calculate the value of the system-level circuit. Step 5: The updated circuit is sent to the simulator for simulation. A reward is calculated based on the difference in FOM values ​​before and after the update. This reward is used to calculate the TD error, which is used as the loss function and for backpropagation to train each network. Steps 2 through 5 are repeated during the training phase until the EVDMARL model converges. Step 6: In the transfer phase, under the given target circuit conditions, each agent operates according to the network model trained in the training phase, and the data generated under the new circuit conditions is input into the neural network to adapt to the new operating conditions; the hybrid network can also automatically adjust its model to adapt to changes in operating conditions with minimal steps; when the generated circuit parameters meet the given circuit constraints, the design point can be returned and the algorithm ends.

2. The method according to claim 1, characterized in that, In step 2, each agent network and hybrid network are initialized before each training session, and the structures of the agent network and hybrid network are as follows: Each agent network is an artificial neural network, consisting of an input encoder, an output layer, and five hidden layers, which store the state of the entire system-level circuit. and the possible actions of the corresponding sub-circuits As input, the output is the value Q of the agent taking this action in this state. a To determine the optimal action for the sub-circuit based on the current circuit state, we use the cross-entropy method (CEM) to sample an action from the sub-circuit action space that has the highest value from the agent network, indicating that the action is the best choice given the current circuit conditions. Hybrid Networks: Considering that the design goal of each sub-circuit is to optimize the FOM of the system-level circuit, the relationship between agents can be modeled as a purely cooperative relationship, that is, the value of the system-level circuit is positively correlated with the value of each agent, and the relationship between the value of the sub-circuit and the value of the overall circuit should satisfy the constraints of the equation. In order to ensure a positive correlation between the value of the sub-circuit and the value of the system-level circuit, we introduce a hybrid network with a super-network architecture inspired by QMIX[12], where the weight parameters of the hybrid network need to be positive, and a super-network with an absolute activation function is used to generate the weight parameters and bias of the hybrid network.

3. The method according to claim 1, characterized in that, Given the current state of the circuit, each agent network uses the cross-entropy method (CEM) to suggest actions to be taken to the corresponding sub-circuit parameters based on the current circuit state. Then, the transistor parameters (state) of the entire circuit are updated based on the actions of each agent. The specific steps are as follows: Input parameter: Current circuit state S t Output: Action A taken this time. t and the state S at the next moment t , Step 3.1: Initialize random variables using a multivariate normal distribution Step 3.2: Initialize the storage variable B = {}, where B is a container responsible for temporarily storing all actions of this trial. Step 3.3: From the distribution random sampling Step 3.4: Input the current state of the circuit into the agent network to obtain the value of the action in this state. Will Store the samples in container B; repeat steps 3.3-3.4 until N samples are collected; N is a hyperparameter. Step 3.5: Take N from B best The action sample with the highest Q value is called the new Population, N best For hyperparameters; Step 3.6: Calculation The mean and variance are calculated, and the distribution is updated. Repeat steps 3.2-3.6 M times, where M is a hyperparameter; Step 3.7: From the distribution Randomly sample an action As the final choice; Step 3.8: Based on the current circuit state and actions, calculate the state at the next moment.

4. The method according to claim 1, characterized in that, In step 4, given the global state of each agent and the action suggested in step 3, the agent network outputs its value. The value output by each agent is then input into the hybrid network, which outputs the value of the entire circuit. The specific steps are as follows: Input parameter: Current circuit state S t and the actions taken by each agent Output: The value function Q of the entire circuit tot Step 4.1: For each agent, the network input circuit state and the action taken by the agent are used to obtain the value function of that sub-circuit; Step 4.2: The hybrid network takes the value functions of each sub-circuit as input to obtain the value function of the entire circuit; 5. The method according to claim 1, characterized in that, In step 5, the updated circuit is fed into a simulator to obtain a new FOM value, which is then used to calculate the reward and TD error. The TD error is used as the loss function to update the parameters of the agent network and the hybrid network. The sub-steps are as follows: Input parameters: Updated circuit netlist; Output: Updated agent network and hybrid network parameters; Step 5.1: Input the updated circuit netlist into the simulator for simulation to obtain the FOM value. Step 5.2: Calculate the reward based on the difference in FOM values ​​before and after. Step 5.3 Calculate the TD error based on the reward. The specific form of the TD error is as follows: Where r is the reward obtained after completing this step, which is the actual circuit performance improvement obtained through simulation by the simulator; To anticipate future rewards starting from the next step, we input the next state and the best action for that agent to take the next step into the expected future reward output by each agent. This is then obtained by integrating the expected reward of the entire circuit through a hybrid function. The expected future reward for achieving the current state; r+ and The difference lies in the fact that, after this step, the reward r of this step is known; by calculating the square of the distance between the two, the difference between the two can be obtained, thus making the future reward (value) estimated by the agent and the hybrid network closer to the real reward. Step 5.4: We use the TD error as the loss function to perform gradient backpropagation, thereby training the agent network and the hybrid network as a whole end-to-end.

6. The method according to claim 1, characterized in that, in During the transfer phase, under the given target circuit conditions, each agent operates according to the network model trained in the training phase, and the data generated under the new circuit conditions is input into the neural network to adapt to the new operating conditions; the hybrid network can also automatically adjust its model to adapt to changes in operating conditions with minimal steps; when the generated circuit parameters meet the given circuit constraints, the design point can be returned and the algorithm can be terminated.