Analog circuit netlist division method and system based on reinforcement learning
Through the method based on reinforcement learning, feature extraction, classification and function annotation of analog circuits is solved, and the complexity and accuracy of analog circuit division in the prior art is achieved, and efficient and automatic circuit division is achieved.
Patent Information
- Application Number
- PCT/CN2024/137360
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art is difficult to meet the high-precision requirements when processing analog circuits, and the circuit division algorithm is complex and computational complex, resulting in inefficient calculations or inability to process large-scale circuits.
The analog circuit netlist division method based on reinforcement learning is adopted to achieve efficient division of analog circuits through feature extraction, layered reinforcement learning algorithm classification and function annotation.
It improves the quality and efficiency of analog circuit division, can automatically identify and classify different circuit types, meet high-precision requirements, and reduces the computational complexity.
Smart Images

Figure CN2024137360_19062025_PF_FP_ABST
Abstract
Description
A method and system for analog circuit netlist partitioning based on reinforcement learning Technical Field
[0001] The present application relates to the field of circuit partitioning technology, and in particular to a method and system for analog circuit netlist partitioning based on reinforcement learning. Background Art
[0002] Circuit partitioning involves dividing unit devices into two or more subsets. The goal is to keep the size or area of each subset similar and minimize the number of interconnections between subsets, thereby reducing circuit design complexity and improving the readability of the partitioned circuits. The complex structures and variable parameters of analog circuits increase the complexity and computational complexity of circuit partitioning algorithms. Circuit partitioning algorithms in related technologies need to handle large-scale graph structures and high-dimensional parameter spaces, which can lead to inefficient computation or even an inability to handle large-scale circuits. Furthermore, analog circuits require high precision, which related circuit partitioning algorithms may not be able to meet.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a method and system for analog circuit netlist partitioning based on reinforcement learning, which can improve the quality and efficiency of analog circuit partitioning.
[0005] To achieve the above objectives, an embodiment of the present application provides a method for partitioning an analog circuit netlist based on reinforcement learning, the method comprising:
[0006] Get the analog circuit netlist;
[0007] Performing feature extraction processing on the analog circuit netlist to obtain circuit features;
[0008] Classifying the circuit features based on a hierarchical reinforcement learning algorithm to obtain a classification result;
[0009] Function labeling is performed on the classification results to obtain a modular circuit netlist.
[0010] In some embodiments, performing feature extraction processing on the analog circuit netlist to obtain circuit features includes:
[0011] Parsing the analog circuit netlist to obtain parsing information;
[0012] Performing data structure construction processing on the analog circuit netlist according to the parsing information to obtain a circuit data structure;
[0013] Performing component extraction and signal path extraction processing on the circuit data structure to obtain circuit node features;
[0014] Performing time domain and frequency domain simulation analysis on the circuit data structure to obtain simulation behavior characteristics;
[0015] The circuit node characteristics and the simulation behavior characteristics are determined as circuit characteristics.
[0016] In some embodiments, the classifying the circuit features based on the hierarchical reinforcement learning algorithm to obtain the classification results includes:
[0017] The analog circuit netlist partitioning problem is defined as a semi-Markov decision process, and a deep reinforcement learning model is constructed;
[0018] Solving the deep reinforcement learning model based on a hierarchical reinforcement learning algorithm to obtain a classification model;
[0019] The circuit features are input into the classification model to perform circuit division processing to obtain a classification result.
[0020] In some embodiments, the problem of partitioning the analog circuit netlist is defined as a semi-Markov decision process, and a deep reinforcement learning model is constructed, including:
[0021] The analog circuit netlist partitioning problem is defined as a semi-Markov decision process;
[0022] constructing a scheduling strategy and an association strategy based on the semi-Markov decision process;
[0023] According to the scheduling strategy, the components of the analog circuit netlist are divided and scheduled at different time steps to obtain a scheduling network;
[0024] Allocating and processing the components of the analog circuit netlist according to the association strategy to obtain an associated network;
[0025] A deep reinforcement learning model is constructed based on the scheduling network and the association network.
[0026] In some embodiments, the processing of solving the deep reinforcement learning model based on a hierarchical reinforcement learning algorithm to obtain a classification model includes:
[0027] Obtaining a reinforcement learning space set and a reinforcement learning function set according to the deep reinforcement learning model definition;
[0028] Initialize a value function and a hierarchical strategy according to the reinforcement learning space set and the reinforcement learning function set;
[0029] The value function and the stratification strategy are updated based on the rainbow algorithm to obtain a classification model.
[0030] In some embodiments, updating the value function and the hierarchical strategy based on the rainbow algorithm to obtain a classification model includes:
[0031] Initializing a neural network and an experience replay buffer of the deep reinforcement learning model, wherein parameters of the neural network include the value function and the hierarchical strategy;
[0032] Performing interactive processing between the deep reinforcement learning model and the environment, collecting interactive experience and storing the interactive experience in the experience playback buffer;
[0033] Performing random sampling processing on the experience playback buffer to obtain experience samples;
[0034] The parameters of the neural network are updated according to the experience samples to obtain a classification model.
[0035] In some embodiments, performing function labeling on the classification results to obtain a modular circuit netlist includes:
[0036] Performing component recognition processing on the classification result to obtain a recognition result;
[0037] The circuit modules in the classification results are named and labeled according to the recognition results to obtain a modular circuit netlist.
[0038] To achieve the above objectives, another aspect of the present application provides a reinforcement learning-based analog circuit netlist partitioning system, the system comprising:
[0039] The first module is used to obtain the analog circuit netlist;
[0040] The second module is used to perform feature extraction processing on the analog circuit netlist to obtain circuit features;
[0041] The third module is used to classify the circuit features based on a hierarchical reinforcement learning algorithm to obtain a classification result;
[0042] The fourth module is used to perform function labeling processing on the classification results to obtain a modular circuit netlist.
[0043] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0044] To achieve the above objectives, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0045] The embodiments of the present application include at least the following beneficial effects: The present application provides a method and system for analog circuit netlist partitioning based on reinforcement learning, which obtains circuit features by performing feature extraction processing on the analog circuit netlist; classifies the circuit features based on a hierarchical reinforcement learning algorithm to obtain classification results; and can combine machine learning methods to partition the analog circuit netlist, and can automatically identify and classify different circuit types, thereby improving the quality and efficiency of analog circuit partitioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG1 is a flow chart of a method for partitioning an analog circuit netlist based on reinforcement learning provided in an embodiment of the present application;
[0047] FIG2 is an interaction diagram of an agent and an environment in reinforcement learning provided by an embodiment of the present application;
[0048] FIG3 is a flow chart of a rainbow algorithm provided in an embodiment of the present application;
[0049] FIG4 is a flowchart of an implementation of an application scenario provided by an embodiment of the present application;
[0050] FIG5 is a schematic diagram of a deep reinforcement learning model provided in an embodiment of the present application;
[0051] FIG6 is a schematic diagram of the structure of an analog circuit netlist partitioning system based on reinforcement learning provided in an embodiment of the present application;
[0052] FIG7 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0054] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0055] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0057] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0058] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0059] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0060] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning (deep learning) typically includes techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0061] Reinforcement Learning (RL), also known as reinforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning. It is used to describe and solve the problem of how an agent can maximize rewards or achieve specific goals by learning strategies during its interaction with the environment.
[0062] Integrated circuits (ICs) play a crucial role in modern society. ICs are composed of multiple electronic components (such as transistors, capacitors, and resistors) connected by tiny wires to form a fully functional circuit. Circuit partitioning plays a crucial role in this process, and digital circuit partitioning is significantly more mature than analog circuit partitioning. This is primarily because the design and partitioning of digital circuits can be approached from a formal, algorithm-driven perspective, while analog circuits face greater challenges and complexity. Graph partitioning is a classic combinatorial optimization problem, widely used in fields such as hardware-software co-design, very large-scale integrated circuit (VLSI) design, and parallel computing. However, applying graph partitioning algorithms to analog circuits faces complexity and accuracy issues.
[0063] For example, graph partitioning algorithms operate based on discrete graph models and cannot be directly applied to analog circuits processing continuous signals. These algorithms require solving continuity issues, such as sampling and quantizing continuous signals. Analog circuits require high precision, and graph partitioning algorithms may not meet this high precision requirement. Algorithmic errors and approximations can lead to inaccurate descriptions of circuit behavior, which in turn affects circuit performance prediction and optimization. Analog circuits contain nonlinear components and characteristics, while graph partitioning algorithms operate based on linear models. This limits the algorithm's ability to capture nonlinear characteristics and restricts its application in analog circuit optimization. Analog circuits have complex structures and highly variable parameters, which increase the complexity and computational complexity of graph partitioning algorithms. The algorithms need to process large-scale graph structures and high-dimensional parameter spaces, which may lead to inefficient computation or even inability to handle large-scale circuits. Graph partitioning algorithms may also lack robustness to uncertainty and noise. Analog circuits are subject to factors such as noise, temperature variations, and device non-idealities, all of which can affect the algorithm's accuracy and stability.
[0064] In view of this, an embodiment of the present application provides a method and system for analog circuit netlist partitioning based on reinforcement learning. This solution obtains circuit features by performing feature extraction processing on the analog circuit netlist; classifies the circuit features based on a hierarchical reinforcement learning algorithm to obtain classification results; and performs function annotation processing on the classification results to obtain a modular circuit netlist. This can combine machine learning, optimization algorithms, and simulation technology to improve the analog circuit diagram partitioning algorithm. It can automatically identify and classify different circuit types by training and learning the input and output data of the analog circuit. Moreover, due to the combination of this solution with machine learning, the graph partitioning algorithm can better adapt to different types of circuits and problems, improving the quality and efficiency of partitioning.
[0065] The embodiment of the present application provides a method for partitioning an analog circuit netlist based on reinforcement learning, which relates to the field of artificial intelligence technology. The embodiment of the present application provides a method for partitioning an analog circuit netlist based on reinforcement learning, which can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a method for partitioning an analog circuit netlist based on reinforcement learning, etc., but is not limited to the above forms.
[0066] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0067] FIG1 is an optional flowchart of a method for analog circuit netlist partitioning based on reinforcement learning provided in an embodiment of the present application. The method in FIG1 may include but is not limited to steps S101 to S104 .
[0068] Step S101, obtaining an analog circuit netlist;
[0069] Step S102, performing feature extraction processing on the analog circuit netlist to obtain circuit features;
[0070] Step S103, classifying the circuit features based on a hierarchical reinforcement learning algorithm to obtain a classification result;
[0071] Step S104 , performing function labeling processing on the classification results to obtain a modular circuit netlist.
[0072] In steps S101 to S104 shown in the embodiment of the present application, by extracting features from the analog circuit netlist, the circuit elements, connection methods, analog behaviors and other features in the analog circuit are extracted to obtain circuit features, and then the circuit features are classified and processed based on the hierarchical reinforcement learning algorithm, and the analog circuit is divided by the reinforcement learning model to obtain classification results; finally, the classification results are functionally labeled to obtain a modular circuit netlist. The embodiment of the present application focuses on the application of machine learning in the graph partitioning algorithm in terms of feature extraction, parameter optimization, search strategy and algorithm combination. By combining with machine learning, the graph partitioning algorithm can better adapt to different types of circuits and problems, and improve the quality and efficiency of partitioning.
[0073] In some embodiments, step S101 may be performed by obtaining a pre-prepared analog circuit netlist file. Alternatively, the analog circuit may be identified by other means, such as crawlers, image recognition, or other techniques, to obtain the relevant analog circuit netlist file, without limitation. The analog circuit netlist is a text file containing circuit components, connections, and other related information.
[0074] In step S102 of some embodiments, performing feature extraction processing on the analog circuit netlist to obtain circuit features includes:
[0075] Parsing the analog circuit netlist to obtain parsing information;
[0076] Performing data structure construction processing on the analog circuit netlist according to the parsing information to obtain a circuit data structure;
[0077] Performing component extraction and signal path extraction processing on the circuit data structure to obtain circuit node features;
[0078] Performing time domain and frequency domain simulation analysis on the circuit data structure to obtain simulation behavior characteristics;
[0079] The circuit node characteristics and the simulation behavior characteristics are determined as circuit characteristics.
[0080] In an embodiment of the present application, an analog circuit netlist can be parsed using methods such as natural language processing to obtain parsed information, which includes circuit elements, connections, and other related information in the analog circuit. The analog circuit netlist is then subjected to data structure construction based on the parsed information, organizing the parsed information into an appropriate data structure for subsequent processing. This data structure can be a topological structure or other representation of the constructed analog circuit. Component extraction and signal path extraction are performed on the circuit data structure to obtain circuit node features, where circuit node features include circuit element information and circuit signal paths. This is specifically achieved by identifying the circuit data structure and extracting component information in the circuit, such as resistors, capacitors, inductors, and amplifiers. For each component, key information such as its value and connection is recorded. By identifying the circuit data structure and extracting the signal path in the circuit, including the propagation path of current and voltage, the signal path is understood, which is crucial for understanding the behavior of the circuit. The circuit data structure is then subjected to time domain and frequency domain simulation analysis to obtain simulated behavior features. In analog circuits, the focus is on the continuous-time behavior of the circuit. Time domain analysis, frequency domain analysis, and other methods are used to obtain information such as the circuit's dynamic response and steady-state response. The embodiment of the present application performs time domain simulation to obtain the response of the circuit at different time points, which helps to understand the transient behavior, response time, etc. of the circuit. Steady-state analysis is also performed to understand the performance of the circuit in a stable operating state, which involves DC analysis, etc. The embodiment of the present application determines the circuit node characteristics and analog behavior characteristics as circuit characteristics. By extracting specific features from the circuit, the features may include voltage and current values, frequency response, amplitude-frequency characteristics, phase-frequency characteristics, etc. at key nodes, which helps to quantify the performance and behavior of the circuit. It is conceivable that the embodiment of the present application can also save the extracted features into an appropriate data structure or file for future use. These features can be used for further analysis, optimization, or other electronic design automation (EDA) tasks.
[0081] In step S103 of some embodiments, the classifying process of the circuit features based on the hierarchical reinforcement learning algorithm to obtain the classification result includes:
[0082] The analog circuit netlist partitioning problem is defined as a semi-Markov decision process, and a deep reinforcement learning model is constructed;
[0083] Solving the deep reinforcement learning model based on a hierarchical reinforcement learning algorithm to obtain a classification model;
[0084] The circuit features are input into the classification model to perform circuit division processing to obtain a classification result.
[0085] In the embodiment of the present application, the analog circuit partitioning problem is defined as a semi-Markov decision process, thereby constructing a deep reinforcement learning model. The deep reinforcement learning model can use a convolutional neural network (CNN) to generate decisions through observation. The decision is made through two strategies: scheduling and association. The scheduling strategy is π S , the association strategy is π a . Then, the deep reinforcement learning model is solved based on the hierarchical reinforcement learning algorithm to obtain a classification model. In the embodiment of the present application, the hierarchical reinforcement learning algorithm adopts the rainbow algorithm to solve the deep reinforcement learning model, so that the performance of the model can be regularly evaluated and the parameters of the algorithm can be adjusted during the training process. The effect of the algorithm is measured by evaluation, and the training strategy and parameter settings are adjusted as needed. After the training is completed, the classification model is obtained, and the classification model is tested and applied by using the value function and strategy obtained by training. The hierarchical strategy obtained by training is deployed to the actual analog circuit partitioning problem to achieve the best partitioning decision and obtain the classification result. The embodiment of the present application can improve the precision and accuracy of analog circuit partitioning by combining the deep reinforcement learning model to divide the analog circuit.
[0086] In some embodiments, the problem of partitioning the analog circuit netlist is defined as a semi-Markov decision process, and a deep reinforcement learning model is constructed, including:
[0087] The analog circuit netlist partitioning problem is defined as a semi-Markov decision process;
[0088] constructing a scheduling strategy and an association strategy based on the semi-Markov decision process;
[0089] According to the scheduling strategy, the components of the analog circuit netlist are divided and scheduled at different time steps to obtain a scheduling network;
[0090] Allocating and processing the components of the analog circuit netlist according to the association strategy to obtain an associated network;
[0091] A deep reinforcement learning model is constructed based on the scheduling network and the association network.
[0092] In an embodiment of the present application, the analog circuit partitioning problem is defined as a semi-Markov decision process. The semi-Markov decision process is an extension of the Markov decision process (MDP) that allows the time of state transitions and rewards to not be completely dependent on the previous action and state. Both the state transitions and reward functions of the environment contain non-deterministic factors. The decision-making of the intelligent agent needs to take this non-determinism into account in order to select the optimal action strategy. In this definition, a scheduling policy and an association policy decision are constructed based on the semi-Markov decision process, and the decision is made through scheduling and association strategies. Among them, the scheduling policy specifies which components or sub-circuits are selected for partitioning and scheduling at different time steps. The association policy indicates how to allocate components to appropriate sub-circuits or processing units. Therefore, by selecting the components of the analog circuit netlist at different time steps according to the scheduling policy, a scheduling network is obtained, and the components of the simulated circuit netlist are allocated and processed according to the association policy to obtain an association network. A deep reinforcement learning model is constructed based on the scheduling network and the association network. In analog circuit partitioning, the scheduling policy can determine which components are selected for processing at each time step and determine the order in which they are divided. The scheduling strategy can be formulated based on the characteristics, performance requirements and optimization goals of the components to obtain the best possible partitioning results. The association strategy mainly involves allocating components to different sub-circuits or processing units to meet circuit performance requirements and constraints. The association strategy can take into account factors such as the interdependence between components, communication overhead, power consumption, etc. to achieve the optimal circuit partitioning result. By defining scheduling and association strategies, the semi-Markov decision process can formally describe the analog circuit partitioning problem and provide a framework to solve the optimal partitioning scheme. Various reinforcement learning algorithms or planning methods can be used to search for optimal scheduling and association strategies to maximize performance indicators (such as delay, power consumption, area, etc.) or meet specific constraints. The embodiment of the present application mainly designs the analog circuit partitioning as a non-deterministic partially observable Markov decision process problem under deep reinforcement learning (DQL), so as to reduce the complexity by considering the non-deterministic factors of circuit partitioning, in order to achieve the partitioning result.
[0093] Reinforcement learning involves two key elements: the agent and the environment. A diagram of the agent-environment interaction is roughly depicted in Figure 2. At each discrete time step t = 0, 1, 2, ..., the environment provides the agent with an observation St. The agent responds by choosing an action At, and the environment then provides the next reward Rt+1, a discount γt+1, and a new state St+1. This interaction is formalized as a Markov decision process (MDP), which is a tuple (S, A, T, r, γ), where S is a finite set of states, A is a finite set of actions, T(S, a, S') = P[St+1 = S' | St = S, At = a] is the (random) transition function, r(S, a) = E[Rt+1 | St = S, At = a] is the reward function, and γ∈[0, 1] is the discount factor. In the experiments, the MDP will be episodes with a constant γt = γ, except for the episode where γt = 0, but this is the general form of the algorithm. On the agent side, the action selection is given by the policy network π, which defines the probability distribution of the action in each state. Starting from the state St encountered at time t, the discounted return is defined as The agent's goal is to maximize the expected reward by finding a good policy. The policy can be learned directly or constructed as a function of some other learned quantity. In value-based reinforcement learning, the agent learns an estimate of the expected discounted payoff, or value, when following a policy π (vπ(s) = Eπ[Gt|St = s]) or state-action pair (qπ(s, a) = Eπ[Gt|St = s, At = a]) starting from a given state. A common approach to deriving new policies from state-action value functions is to take actions that are ρ-greedy with respect to action values. This corresponds to taking the action with the highest value (the greedy action) with probability (1-ρ) and taking other actions uniformly and randomly with probability ρ. This strategy is used to introduce a form of exploration: by randomly choosing suboptimal actions based on its current estimate, the agent can discover and correct its estimate when appropriate.
[0094] In summary, the embodiment of the present application defines the analog circuit partitioning problem as a semi-Markov decision process, which can establish a comprehensive decision-making model while taking into account time dependence as well as scheduling and associated strategy selection. This can better optimize the effect and performance of circuit partitioning, and decompose the analog circuit partitioning problem, which can effectively reduce the complexity of the analog circuit partitioning problem.
[0095] In some embodiments, the processing of solving the deep reinforcement learning model based on a hierarchical reinforcement learning algorithm to obtain a classification model includes:
[0096] Obtaining a reinforcement learning space set and a reinforcement learning function set according to the deep reinforcement learning model definition;
[0097] In an embodiment of the present application, the Rainbow algorithm is implemented in combination with a hierarchical strategy to solve the deep reinforcement learning model to obtain a classification model. Through the reinforcement learning space set, the reinforcement learning space set includes defining the state space, action space and observation space. The embodiment of the present application defines the state space S, action space A and observation space O according to the analog circuit partitioning problem. The state space may include the current circuit layout, the load conditions of the sub-circuit or processing unit, the communication requirements, etc.; the action space may include the action of selecting which operations to perform and which components to divide; the observation space may be an incomplete observation of the environmental state, such as data obtained by sensors or other available information. The definition of the reinforcement learning function set includes the transfer function, observation function and reward function. The transfer function T(s,a,s'), observation function Z(s,a,o) and reward function R(s,a) are set according to the specific problem. The transfer function describes the probability of reaching the next state from one state through a certain action; the observation function describes the information observed under a specific state and action; the reward function defines the immediate reward obtained by the intelligent agent when it is in a specific state and takes an action.
[0098] Initialize a value function and a hierarchical strategy according to the reinforcement learning space set and the reinforcement learning function set;
[0099] In this embodiment, the value function and hierarchical strategy are initialized based on a set of reinforcement learning spaces and a set of reinforcement learning functions. A value function and strategy are initialized for each state-action pair. The value function represents the expected reward when taking an action in a specific state, while the strategy determines which action to choose in a specific state. For hierarchical strategies, value functions and strategies can be set for the scheduling layer and the association layer, respectively.
[0100] The value function and the hierarchical strategy are updated based on the Rainbow algorithm to obtain a classification model. In this embodiment of the present application, the reinforcement learning algorithm selected is Rainbow, which is an enhanced version of the DQN algorithm that combines multiple reinforcement learning techniques. DQN is a deep reinforcement learning algorithm based on a neural network. It is used to learn the Q-value function, which is a function that maps states and actions to expected rewards. In the framework of semi-nondeterministic partially observable Markov decision process (ND-POMDP), the Rainbow algorithm can be selected as the main reinforcement learning algorithm. The Rainbow algorithm includes multiple components, such as priority experience replay, dual Q network, distributed Q value and n-step reward. This embodiment of the present application uses the selected reinforcement learning algorithm for training, collects sample data through interaction with the environment, and uses this data to update the value function and strategy. At each time step, an action is selected based on the current state and observation information, the action is executed and the new state and reward are observed, and then the parameters are updated using the collected data to obtain a classification model. Referring to Figure 3, this embodiment of the present application uses the Rainbow algorithm to update the reinforcement learning model. First, in DQN, by using convolutional neural networks to approximate the action value of a, deep networks and reinforcement learning are successfully combined. Given a state S t (Input is fed to the network in the form of a stack of raw pixel frames). At each step, the agent selects an action using a ρ-greedy strategy based on the current state relative to the action value and writes a data strip (S t ,A t ,R t+1 ,γ t+1 ,S t+1 ) is added to the replay buffer, which stores the last million transformations. Stochastic gradient descent is used to optimize the neural network parameters to minimize the loss.
[0101] In some embodiments, updating the value function and the hierarchical strategy based on the rainbow algorithm to obtain a classification model includes:
[0102] Initializing a neural network and an experience replay buffer of the deep reinforcement learning model, wherein parameters of the neural network include the value function and the hierarchical strategy;
[0103] Performing interactive processing between the deep reinforcement learning model and the environment, collecting interactive experience and storing the interactive experience in the experience playback buffer;
[0104] Performing random sampling processing on the experience playback buffer to obtain experience samples;
[0105] The parameters of the neural network are updated according to the experience samples to obtain a classification model.
[0106] In an embodiment of the present application, a neural network and an experience replay buffer of a reinforcement learning agent in a deep reinforcement learning model are initialized, wherein the parameters of the neural network include a value function and a hierarchical strategy. Then, the deep reinforcement learning model interacts with the environment, collects interaction experience by selecting actions and observing rewards and next states, and stores the interaction experience in the experience replay buffer. A batch of experience samples are randomly sampled from the experience replay buffer, and the loss is calculated based on the current strategy and the estimated value of the target network. The weights of the neural network are updated using the gradient descent method, thereby regularly updating the parameters of the neural network to obtain a classification model. In the embodiments of the present application, enhancement technologies such as Prioritized Experience Replay and N-step learning are introduced to improve the processing and learning of reward signals, and multi-step or distributed Q-value estimation methods such as n-step or n-step distributed Q-learning are used to reduce reward signal deviation and improve performance. Double Q-learning is used to alleviate the problem of overestimation, and distributed deep Q learning (Dueling DQN) and Noisy Nets are introduced to increase the exploration and randomness of the agent by using the current network for action selection and the target network for value estimation. By integrating all these technologies together to achieve better performance and stability, the above steps are repeated until the preset number of training steps is reached or the stopping condition is met. It should be noted that in order to avoid falling into a local optimum, simulated annealing technology can also be used in the embodiments of the present application.
[0107] In step S104 of some embodiments, performing function annotation processing on the classification results to obtain a modular circuit netlist includes:
[0108] Performing component recognition processing on the classification result to obtain a recognition result;
[0109] The circuit modules in the classification results are named and labeled according to the recognition results to obtain a modular circuit netlist.
[0110] In an embodiment of the present application, the classification results can be used to divide the circuit into different modules based on its structure and function. Modules can be divided based on component type (e.g., amplifier module, filter module, oscillator module, etc.) or function (e.g., signal amplification, filtering, frequency modulation, etc.). The components contained in each module in the classification results are then identified and their types and parameters recorded. This includes resistors, capacitors, inductors, transistors, etc. In an embodiment of the present application, each component can also be named and numbered for easy reference in subsequent documents. Each module is then assigned a clear name that reflects its primary function. For example, if a module has a low-pass filtering function, it can be named "Low-Pass Filter." Describe the primary function of each module. This helps others understand the design purpose and key characteristics of the circuit. For example, the function of an amplifier module may be signal amplification, while the function of a filter module may be signal filtering within a specific frequency range. If there are interfaces between modules, the characteristics of these interfaces are annotated, including signal levels, impedance matching, etc. This is important for ensuring good connectivity and interoperability between modules. Annotations can be added directly to the netlist file as comments or described in a separate document. The embodiment of the present application performs functional annotation on the classification results, so that other people or tools can easily understand and use the modular simulation netlist.
[0111] The following is a detailed description of the embodiments of the present application with reference to specific application examples:
[0112] 4 , the embodiment of the present application inputs an analog circuit netlist. By analyzing and extracting features from the analog circuit netlist, analog circuit features such as performance requirements, component characteristics, number of devices, power consumption, and communication overhead can be extracted. Then, based on the deep reinforcement learning model, the performance features and component features are input into the scheduling strategy, and the number of devices, power consumption, and communication overhead are input into the association strategy. The scheduling and association processes are performed continuously, by defining the joint problem as a semi-ND-POMDP, with the scheduler as an option of the ND-POMDP of the association process, and then attempting to use a hierarchical reinforcement learning algorithm to solve the problem. The idea of hierarchical reinforcement learning is to update two networks, one for the scheduling strategy and the other for the association strategy. The scheduling strategy is updated by using a fixed association strategy, and vice versa. Note that the scheduling and association networks share the same environmental rewards at different time scales, which ensures monotonic improvement of the environment. The agent continuously adjusts its association decisions based on its strategy, and the agent gains to improve the scheduling and association strategies. Referring to Figure 5, the deep reinforcement learning model includes a scheduling network and an association network, wherein the scheduling network is used to execute the scheduling strategy, and the association network is used to execute the association strategy. In the embodiment of the present application, the rainbow algorithm, simulated annealing algorithm and hierarchical reinforcement learning algorithm are used to update the scheduling strategy and the association strategy. The scheduling strategy can determine which components to select for processing at each time step and determine the order of division between them. The association strategy can be formulated based on the characteristics, performance requirements and optimization goals of the components to obtain the best possible division results. Finally, the classification results are functionally labeled and output to obtain a modular circuit netlist.
[0113] Referring to FIG. 6 , an embodiment of the present application further provides an analog circuit netlist partitioning system based on reinforcement learning, which can implement the above-mentioned analog circuit netlist partitioning method based on reinforcement learning. The system includes:
[0114] The first module 601 is used to obtain an analog circuit netlist;
[0115] The second module 602 is configured to perform feature extraction processing on the analog circuit netlist to obtain circuit features;
[0116] The third module 603 is used to classify the circuit features based on a hierarchical reinforcement learning algorithm to obtain a classification result;
[0117] The fourth module 604 is configured to perform function labeling on the classification results to obtain a modular circuit netlist.
[0118] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0119] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned reinforcement learning-based analog circuit netlist partitioning method. The electronic device can be any intelligent terminal, such as a tablet computer or an in-vehicle computer.
[0120] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0121] Please refer to FIG7 , which illustrates a hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0122] The processor 701 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0123] The memory 702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called by the processor 701 to execute the analog circuit netlist partitioning method based on reinforcement learning in the embodiments of this application.
[0124] Input / output interface 703, used to implement information input and output;
[0125] Communication interface 704, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0126] Bus 705 , which transmits information between various components of the device (e.g., processor 701 , memory 702 , input / output interface 703 , and communication interface 704 );
[0127] The processor 701 , the memory 702 , the input / output interface 703 and the communication interface 704 are connected to each other in communication within the device via a bus 705 .
[0128] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned reinforcement learning-based analog circuit netlist partitioning method.
[0129] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0130] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0131] The embodiment of the present application provides a method and system for analog circuit netlist partitioning based on reinforcement learning, which performs netlist partitioning on analog circuits through a centralized deep reinforcement learning (DRL) association method based on rainbow agents. In the deep reinforcement learning model, the intelligent agent needs to select the optimal action strategy to achieve the circuit partitioning goal and meet the design requirements based on the uncertainty of the observation information and environmental state. Thus, by combining machine learning, optimization algorithms and simulation technology, the analog circuit diagram partitioning algorithm is improved. By training and learning the input and output data of the analog circuit, different circuit types can be automatically identified and classified, which can improve the partitioning performance and efficiency of the analog circuit.
[0132] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0133] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0135] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0136] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0137] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0138] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0139] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0142] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for partitioning analog circuit netlists based on reinforcement learning, characterized in that: The method comprises: Get the analog circuit netlist; Performing feature extraction processing on the analog circuit netlist to obtain circuit features; Classify the circuit features based on a hierarchical reinforcement learning algorithm to obtain a classification result; Functional annotation is performed on the classification results to obtain a modular circuit netlist.
2. The method according to claim 1, characterized in that The step of performing feature extraction processing on the analog circuit netlist to obtain circuit features includes: Parsing the analog circuit netlist to obtain parsing information; Performing data structure construction processing on the analog circuit netlist according to the parsing information to obtain a circuit data structure; Performing component extraction and signal path extraction processing on the circuit data structure to obtain circuit node features; Performing time domain and frequency domain simulation analysis on the circuit data structure to obtain simulation behavior characteristics; The circuit node characteristics and the simulated behavior characteristics are determined as circuit characteristics.
3. The method according to claim 1, characterized in that The circuit features are classified based on the hierarchical reinforcement learning algorithm to obtain classification results, including: The analog circuit netlist partitioning problem is defined as a semi-Markov decision process, and a deep reinforcement learning model is constructed; Solving the deep reinforcement learning model based on a hierarchical reinforcement learning algorithm to obtain a classification model; The circuit features are input into the classification model to perform circuit division processing to obtain a classification result.
4. The method according to claim 3, characterized in that The problem of partitioning the analog circuit netlist is defined as a semi-Markov decision process, and a deep reinforcement learning model is constructed, including: The partitioning problem of the analog circuit netlist is defined as a semi-Markov decision process; constructing a scheduling strategy and an association strategy according to the semi-Markov decision process; According to the scheduling strategy, the components of the analog circuit netlist are divided and scheduled at different time steps to obtain a scheduling network; Allocate and process the components of the analog circuit netlist according to the association strategy to obtain an associated network; A deep reinforcement learning model is constructed based on the scheduling network and the association network.
5. The method according to claim 3, characterized in that: The step of solving the deep reinforcement learning model based on the hierarchical reinforcement learning algorithm to obtain a classification model includes: Obtaining a set of reinforcement learning spaces and a set of reinforcement learning functions according to the deep reinforcement learning model definition; Initialize the value function and the hierarchical strategy according to the reinforcement learning space set and the reinforcement learning function set; The value function and the stratification strategy are updated based on the rainbow algorithm to obtain a classification model.
6. The method according to claim 5, characterized in that The updating of the value function and the stratification strategy based on the rainbow algorithm to obtain a classification model includes: Initializing a neural network and an experience replay buffer of the deep reinforcement learning model, wherein parameters of the neural network include the value function and the hierarchical strategy; Performing interactive processing with the environment through the deep reinforcement learning model, collecting interactive experience and storing the interactive experience in the experience playback buffer; Performing random sampling processing on the experience playback buffer to obtain experience samples; The parameters of the neural network are updated according to the experience samples to obtain a classification model.
7. The method according to any one of claims 1 to 6, characterized in that: The function labeling process is performed on the classification result to obtain a modular circuit netlist, including: Performing component recognition processing on the classification result to obtain a recognition result; The circuit modules in the classification results are named and labeled according to the recognition results to obtain a modular circuit netlist.
8. An analog circuit netlist partitioning system based on reinforcement learning, characterized in that: The system comprises: The first module is used to obtain the analog circuit netlist; The second module is used to perform feature extraction processing on the analog circuit netlist to obtain circuit features; The third module is used to classify the circuit features based on a hierarchical reinforcement learning algorithm to obtain a classification result; The fourth module is used to perform function labeling processing on the classification results to obtain a modular circuit netlist.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Man-machine collaborative labeling system and method for analog circuit netlist
CN114662434A
Analog integrated circuit netlist labeling method based on VF3 algorithm
CN115906734A
Gate-level circuit component identification method and system, storage medium and equipment
CN115984633A
Analog circuit netlist division method and system based on reinforcement learning
CN117709271A
Cited By
Water affair system operation data model optimization scheduling method based on reinforcement learning
CN121212713A
Cable accessory intelligent maintenance strategy optimization method based on reinforcement learning
CN122021271A
Netlist generation method, device, equipment, storage medium and product
CN122549311A