Chip design system and method based on multiple agents
Through a chip design system based on multiple agents, and collaborating with multiple agents for chip design processing, the application limitations and complexity of LLM in chip design in the existing technology is solved, and the automation and efficiency improvement of the entire chip design process is achieved.
Patent Information
- Application Number
- CN202510314033.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
When using generative artificial intelligence (LLM) for chip design, the existing technology has problems such as limited design complexity and limited application to specific stages, and has failed to effectively explore the entire process of chip design.
Using a chip design system based on multiple agents, through the collaboration of multiple agents, it receives user design requests, performs chip design processing, and generates design files containing chip layout information to realize the automation of the entire chip design process.
It has realized the application of LLM in the entire chip design process, reduces human participation, improves chip design efficiency, significantly shortens the design cycle, and improves design efficiency and performance.
Smart Images

Figure CN120217995A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chip technology, and in particular, to a chip design system and method based on multi-agent. Background Art
[0002] With the development of generative artificial intelligence, large language models (LLMs) have demonstrated remarkable "intelligent" capabilities, which can improve the efficiency and productivity of various industries. For example, in the field of chip design, the ability to use LLMs for chip design has been studied and explored to reduce chip design complexity and shorten the design cycle. However, in the existing research on using LLMs for chip design, there are limitations in chip design, and only some chip designs with relatively low complexity can be carried out, such as pipeline design or simpler finite state machine (FSM) design; in addition, the application of LLMs is often limited to specific stages of the chip design process, ignoring the exploration of applying LLMs to the entire chip design process. Summary of the Invention
[0003] In view of the problems mentioned in the above background art, embodiments of this application propose a chip design system, method, device, and storage medium based on multi-agent to solve the above problems or at least partially solve the above problems. Among them,
[0004] In the first embodiment, this application provides a chip design file generation system. The system includes: multiple agents;
[0005] Among them, the multiple agents cooperate to perform the following operations:
[0006] Receive a design request from the user, where the design request includes chip design requirement information;
[0007] Based on the chip design requirement information, perform chip design processing to generate a chip design file; the chip design file contains chip layout information.
[0008] In the second embodiment, this application provides a method for generating a chip design file. The method includes:
[0009] Use multiple agents to cooperate to perform the following operation steps:
[0010] Receive a design request from the user, where the design request includes chip design requirement information;
[0011] Based on the chip design requirement information, perform chip design processing to generate a chip layout file.
[0012] In the third embodiment, this application provides an electronic device. The electronic device includes the chip design file generation system provided in the first embodiment of this application.
[0013] Fourthly, the present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a computer, it can implement the chip design file generation method provided in the second embodiment of the present application.
[0014] In the technical solution provided by the embodiments of the present application, multiple agents cooperate to receive a user's design request, and based on the chip design requirement information included in the design request, execute the entire chip design processing flow, thereby generating a chip design file (including chip layout information) for subsequent actual chip manufacturing. It can be seen that the solution of the present application realizes the application of agents based on LLM in the full-process exploration of chip design, which can reduce human participation and improve chip design efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1 It is a schematic flowchart of an existing chip design solution provided by an embodiment of the present application;
[0017] Figures 2a to 2d It is a schematic diagram of the structure of a chip design system provided by an embodiment of the present application and the principle of the full-process exploration of chip design by multiple agents collaborating in the system;
[0018] Figure 3 It is a schematic diagram of the principle of establishing an open-source hardware core library provided by an embodiment of the present application;
[0019] Figure 4a and Figure 4b It is a schematic diagram of the principle of an agent designing a hardware core module provided by an embodiment of the present application;
[0020] Figure 5 It is a schematic diagram of the principle of an agent generating module code for a hardware core module provided by an embodiment of the present application;
[0021] Figure 6 It is two examples of chip layouts designed according to the embodiments of the present application;
[0022] Figure 7 It is a bar chart showing the improvement of chip energy efficiency and performance designed by using the solution of the present application provided by an embodiment of the present application;
[0023] Figure 8A bar chart showing the energy efficiency improvement of the chip designed using the solution of this application compared with the corresponding existing chip provided by an embodiment of this application;
[0024] Figure 9 A schematic diagram showing the execution workflow of the NPU provided by an embodiment of this application and the generation process of the systolic array code that changes with the change in the number of agents in the workflow. Detailed implementation manners
[0025] In the field of chip design, a chip design often spans a long-stage process, which includes from chip architecture definition and related code development and testing to physical implementation. Especially for today's new generation of chip products, their complexity requirements are higher and the design cycle is longer. Therefore, this has led to the problems of design complexity and long design cycle faced by chip design, which has become the main bottleneck in the design of new generation of chip products. Therefore, it is crucial to explore an agile design method for chips to accelerate the design cycle of increasingly complex chips.
[0026] Currently, with the development of generative artificial intelligence and the emergence of large language models (LLMs), such as GPT-4 and Llama3, it provides a new perspective and technical support for the agile design of the new generation of increasingly complex chips. As shown in Figure 1 The typical flowchart of using an LLM for chip design shows that an LLM is used to design a chip (ChatCPU) for a chatbot. The ChatCPU is a RISC-V CPU designed using an LLM and has been taped out. Among them, RISC-V CPU refers to a central processing unit (CPU) designed and implemented based on the RISC-V instruction set architecture (ISA), which demonstrates how to accelerate the chip design process through generative artificial intelligence technology. Specifically, when implementing the design of ChatCPU, in the EDA (Electronic Design Automation) physical implementation stage, script generation, task decomposition and execution are performed based on the LLM. However, the existing chip design solutions using LLMs have the following limitations:
[0027] 1) The chip design complexity generated by the LLM is relatively low and limited. For example, it can only perform pipeline design or simpler finite state machines (FSMs).
[0028] 2) The application of the LLM is often limited to specific stages of the chip design process, while ignoring the exploration of applying the LLM to the entire chip design process. For example, the application is limited to RTL (Register Transfer Level) code writing.
[0029] To solve the above problems, an embodiment of the present application provides a chip design solution, and the basic idea is as follows:
[0030] Considering that open-source hardware (OSH) is another important innovative approach to support agile chip design, and in recent years, many open-source hardware designs have been released, which have significantly shortened the chip development cycle. By leveraging existing open-source hardware designs, designers can accelerate the development process and reuse existing modules instead of designing a module from scratch. Based on this, in the solution of the present application, the advantages of LLM and open-source hardware are combined to explore an end-to-end chip design solution to enable the design of complex chips, such as those for complex domain-specific architectures (DSAs) or domain-specific systems-on-chip (DSSoCs). A DSA is a computing architecture optimized for a specific class of applications or tasks, which improves the execution efficiency of specific tasks by integrating dedicated hardware accelerators. A DSSoC is a chip design tailored for specific tasks, usually adopting a heterogeneous architecture that includes multiple accelerators to handle heavy workloads. The traditional DSSoC design process is long and complex, requiring in-depth expertise and complex design tools at each design stage. To solve these problems, in the solution of the present application, specifically: a multi-agent system based on LLM is adopted to implement guiding chip architecture exploration and related code generation and verification, etc.; at the same time, open-source hardware IPs obtained from online available resources are utilized to accelerate the integration of chips (such as DSSoCs), and the design cycle is shortened through LLM.
[0031] Among them, the above-mentioned related code includes hardware description language (HDL, Hardware Description Language), HDL. HDL code can describe circuits in various ways, including behavioral level, RTL level, gate level, etc. RTL code is a specific type of HDL code dedicated to describing the register transfer level behavior of circuits. And each agent based on LLM in the present application is given a specific role, memory ability, planning ability, and tool usage ability, enabling the agent to execute the specified task. Additionally, in the present application, the end-to-end chip design process is segmented and assigned to different agents, and these agents are responsible for specific tasks. And, regarding the solution provided by the present application, it has been verified through, for example, two DSSoC design examples. The design cycles of these example chips are relatively short. For example, for the Internet of Things and mobile devices, the design cycle is only 2 - 4 weeks; through the architecture design of this specific domain DSSoC, compared with existing system-on-chips (SoCs) with similar performance, the design efficiency has been increased by approximately 23.81 to 32.43 times.
[0032] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0033] In addition, in some processes described in the specification, claims, and the above-mentioned accompanying drawings of this application, a plurality of operations that appear in a specific order are included. These operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish each different operation, and the serial numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types. And the term "or / and" in this application is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A or / and B means that A can exist alone, A and B can exist simultaneously, and B can exist alone. The character " / " in this application generally represents an "or" relationship between the associated objects before and after. It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to this commodity or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or system including the said element. In addition, the following embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0034] The technical solutions provided in the embodiments of this application will be introduced and described below.
[0035] First, the vocabulary involved in the embodiments of this application will be explained. It can be understood that this explanation is for a clearer understanding of the embodiments of this application and does not necessarily constitute a limitation to the embodiments of this application.
[0036] Agent: In this application, an agent refers to a software-based agent designed based on LLM and is used to handle a series of specific tasks.
[0037] Chip design: that is, IC (Integrated Circuit) design.
[0038] EDA (Electronic Design Automation) physical implementation: In the electronic design automation (EDA) process, the physical implementation stage refers to the process of converting the logical design into a physical layout. The goal of this stage is to generate a layout file that can be actually produced in the semiconductor manufacturing process. The layout file can be, but is not limited to, the GDSII (Graphic Data System II) file format. The GDSII file format is a standard file format (binary file format) widely used in the semiconductor industry. It is dedicated to storing the physical layout data exchange of integrated circuit (IC) designs. It contains all the geometric information and hierarchical structure of the chip design and is a key step in transferring the design from the EDA tool to the manufacturing plant.
[0039] Hardware core: refers to a module that can realize specific functions. It can be used directly or slightly modified to be integrated into a new chip design, such as a larger system-on-chip (SoC) design. A hardware core can be a simple logic unit or a complex processing unit. The hardware can be developed and designed in-house or obtained from outside.
[0040] Hardware IP (Intellectual Property Core): Hardware IP is a special form of hardware core, which includes not only design files and technical documents, but also related intellectual property protection. Hardware IP is usually developed by a third party and authorized to users through a licensing agreement.
[0041] It can be understood from the above that: the hardware core includes hardware IP (also called hardware IP core or hard IP core or hard core), which refers to some pre-designed (already physically designed and laid out and wired, etc.) standardized hardware core modules, which are usually provided in the form of RTL code and can be reused in different chip designs. Specifically, the hardware core can be a completed processor core, or it can also be a dedicated circuit module that performs a specific function. For example, the hardware core can be but is not limited to a processor core (such as ARM, RISC-V), a storage controller, an input and output interface (such as USB, HDMI), etc. These modules have been verified and optimized, and their functions are stable. They can be used directly or with slight modifications to be integrated into new chip designs.
[0042] Hardware core library: refers to a collection of hardware cores. In the hardware core library, each hardware core is equivalent to an independent "building block", and designers can build complex chip systems by combining these "building blocks".
[0043] In the embodiments of the present application, an open-source hardware core library is used, such as the Open Source Hardware (OSH) IP library. The OSH IP library is developed for the public and allows anyone to view, modify, distribute, and use the hardware IP in the library. In addition, when introducing the various embodiment solutions provided by the present application below, the above-mentioned "hardware core" is referred to as a "hardware core module".
[0044] Figure 2a The structural schematic diagram of the chip design system provided by the present application is shown. This system is a multi-agent system based on the LLM. Specifically, when implemented, this system is deployed on a corresponding electronic device, and the electronic device can be a terminal device or a server device. The terminal device can be, for example, a smart phone, a desktop computer, a laptop computer, a smart wearable device, etc., and the server device can be, for example, a single server, a cluster composed of multiple servers, a cloud server, a virtual server, etc. Among them, whether the system is deployed on a terminal device or a server device, a corresponding system entrance is provided to enable users to interact with this system, so as to enable users to access various functions and services of the system (such as automated services for chip design, etc.). The above-mentioned system entrance can be, but is not limited to, a web page, an application page, or a mini-program page, etc. In practical applications, the web page is usually used as the main entrance for users to interact with the system. This is because using the web page as the system entrance allows users to access the system without installing a specific application, only need to use a browser, and the web page can be used on various devices such as desktop computers, laptop computers, tablets, and smart phones. In addition, since all users access the same web version, it is also convenient for unified management and update, etc.
[0045] As shown in Figure 2a As shown, the chip design system provided in this embodiment includes: multiple agents. The collaboration of these multiple agents can effectively and automatically complete the entire chip design process with high quality.
[0046] During the chip design process, each agent is responsible for different design tasks, and the design tasks assigned to each agent can be determined by the identity information configured for it. Specifically, the identity information of each agent mainly includes the following components: role specification, behavior guidelines, partners, and output format, which are used to ensure that the behavior and functions of the agent meet the design requirements. The above role specification defines the responsibilities and functions of the agent in the entire system, and it describes the main tasks, goals, and its position and relationships in the system. The behavior guidelines stipulate the behavioral norms and strategies that the agent should follow when performing tasks, so as to ensure that the agent's behavior is consistent and meets expectations. Partners refer to other agents or entities that the agent needs to collaborate with during the task execution process. The output format defines the output content and format that the agent should generate after completing the task, which ensures that the output of the agent can be correctly parsed and used by other agents or systems.
[0047] Exemplarily, assume that in this system for the provided automated chip design service, there is a second agent 22 (i.e., Figure 2a the "architecture expert" agent shown in
[0048] 1) The role specification includes the following: the main task is to select appropriate hardware IPs according to the task flow graph (TFG) and optimize the architecture; the goal is to ensure that the design meets the requirements of PPA (power consumption, performance, area); the permission is to be able to access all hardware core libraries in the system and modify the design documents; and, the relationship with other roles is: to collaborate with the third agent 23 (a writing agent, specifically the "RTL writer" agent) and the fourth agent 24 (a testing agent, specifically the "RTL tester" agent) to provide a detailed module description document of the hardware core module.
[0049] Among them, the task flow graph (TFG) is a directed acyclic graph, and each node in it represents the key information of a task block. The task block is called a functional module in other embodiments below. Through the task flow graph, the decomposition structure, execution order, and dependency relationships of tasks (such as an algorithm task) can be clearly shown.
[0050] 2) The behavior guidelines include the following: the decision rule is to select the most suitable hardware core module according to the task flow graph (TFG) and optimize it according to the PPA metrics; the priority setting is to give priority to tasks that directly affect the system performance, such as module selection on the critical path; the exception handling is to find an alternative or notify the project manager (management agent 10) if a certain hardware core module is unavailable; the communication protocol is to use RESTful API to exchange data with the third agent 23 and the fourth agent 24.
[0051] 3) The partners include: the third agent 23 (which is responsible for converting the module description document into Verilog code, and the second agent 22 needs to provide the third agent 23 with a detailed module description document), and the management agent 10 (which is responsible for coordinating the entire chip design progress, and other agents need to regularly report progress and problems to the management agent).
[0052] 4) The output formats include: the output type is that the generated module description document is in JSON format, structured data format, metadata format, and interface specification (such as sending the module description document to the third agent 23 and the fourth agent 24 through RESTful API).
[0053] In this embodiment, the cooperation of multiple agents specifically performs the following steps:
[0054] 101. Receive the design request of the user, where the design request includes chip design requirement information;
[0055] 102. Based on the chip design requirement information, perform chip design processing to generate a chip design file; the design file contains chip layout information.
[0056] Specifically in implementation, the above-mentioned multiple agents include the management agent 10 (the "manager" agent shown in Figure 2a etc.) and at least one execution agent. The management agent is responsible for coordinating and managing various tasks and subtasks in the entire design process to ensure that the tasks are completed efficiently and orderly. For example, the management agent 10 can decompose the entire chip design task into multiple subtasks, allocate the subtasks to the most suitable other agents, and can monitor the situation of other agents completing tasks. At least one execution agent includes the second agent 22, the third agent 23, the fourth agent 24, etc. mentioned above.
[0057] Considering that having too many agents in a system will increase communication complexity, while having too few agents will lead to unclear responsibilities and a decline in output quality. To balance this, six different agents are defined and deployed in this system. The six agents include one management agent 10 and five execution agents. The five execution agents include the first agent 21 (a functional agent, such as an "algorithm expert" agent), the second agent 22 (the "architecture expert" agent), the third agent 23 (the "RTL writer" agent), the fourth agent 24 (the "RTL tester" agent), and the fifth agent 25 (the "EDA expert" agent, used for EAD physical implementation).
[0058] The detailed functions of each agent will be described in detail in other embodiments below and will not be specifically elaborated here.
[0059] Moreover, in this embodiment, different agents can be equipped with different LLMs. For example, for the third agent 23, a fine-tuned Llama 3.1 70B model can be used, which uses the open-source Verilog dataset and LoRA (Low-Rank Adaptation) tuning. LoRA is an efficient tuning method for large-scale pre-trained models, aiming to reduce the number of parameters and computational resources required for tuning through low-rank matrix factorization. For knowledge-intensive tasks, such as the first agent 21, the Chain of Thought (CoT) method can be adopted to solve them. Among them, the CoT method is a method for enhancing the reasoning ability of LLMs. For labor-intensive tasks, such as the third agent 23 and the fourth agent 24, the ReAct (Reasoning and Acting) method can be used for planning. Among them, ReAct is a framework that combines reasoning and acting, aiming to enhance the ability of LLMs to solve complex problems. For agents that require a memory module to support decision-making, such as the third agent 23, relevant information can be stored in the memory module therein for quick access.
[0060] In addition, in some agents, the required tools are also equipped / embedded. And to improve the utilization efficiency of the tools, a dedicated API (Application Programming Interface) interface is designed for each tool. Each API is a configuration or script file for a specific operation toolset, serving as a bridge for the agent to call these tools. For example, the second agent 22 collaborates with the management agent 10 and the first agent 21, and a hardware core library (such as the OSH IP library) is embedded in its memory module for efficient access; in addition, a simulation tool is also equipped in the second agent 22 to support informed decision-making and design optimization.
[0061] Combined with the above content, in the implementation of the automated chip design process, the functions of the six agents included in this system are as follows:
[0062] 1. The management agent 10 (i.e., the "manager" agent shown in Figure 2a The management agent is used to: receive the design request from the user, and perform chip design task decomposition based on the chip design requirement information included in the design request to obtain multiple subtasks; and assign each subtask to a suitable other agent (the second agent).
[0063]
[0064] In specific implementation, the user can submit a design request through the system entry (such as a web page) of the system. For example, the user can input chip design prompts (such as inputting "Please design a DSSoC suitable for the IOT side...") through the web page of the system. This prompt contains chip design requirement information, and then the user clicks to submit, thereby sending a design request to the system. The design request will be received by the management agent in the system. After receiving the design request, the management agent will parse out the chip design requirement information provided by the user from the request, and then decompose the complex entire chip design task based on this chip design requirement information, and output multiple decomposed subtasks to allocate these subtasks to corresponding other agents for processing. In addition to decomposing tasks for requests, the management agent will also perform other operations, such as recording relevant information such as timestamps and user identifiers (such as user IDs).
[0065] Among them, the above chip design requirement information includes but is not limited to: design goals such as required function information, and constraint conditions. Constraint conditions include, for example, chip application scenarios (such as Internet of Things (IOT)), chip types (such as DSSoc), chip PPA (Performance, Power, and Area requirements), and interface standards. Among them, performance requirements include, for example, clock frequency, throughput, latency, etc. Function information includes algorithm information (such as algorithm code, algorithm documents, etc.). In addition, other information can also be included in the function information, such as the type of data to be processed (such as processing images, videos, or audio, etc.).
[0066] Moreover, the multiple decomposed subtasks include: function relationship diagram generation task, architecture generation task, code writing task, testing task, and physical implementation task. The above function relationship diagram generation task can be generated based on the function information included in the chip design requirement information.
[0067] When the management agent executes the allocation of multiple subtasks, specifically: it allocates the function relationship diagram generation task to the first agent 21, the architecture generation task to the second agent 22, the code writing task to the third agent 23, the testing task to the fourth agent 24, and the physical implementation task to the fifth agent 25.
[0068] It should be supplemented and explained here that: in addition to executing the above-mentioned functions, the management agent can also execute other functions. For the other functions that can be executed, reference can be made to the relevant content in other embodiments.
[0069] II. Five execution agents (including the first agent 21, the second agent 22, the third agent 23, the fourth agent 24, and the fifth agent 25)
[0070] Continue to refer to Figure 2a As shown, the entire chip design mainly includes the following three stages: the architecture definition stage, the code generation and verification stage, and the EDA physical implementation stage. In the chip architecture definition stage, it is implemented through two LLM-based agents, namely: as Figure 2b shown, the two agents, the first agent 21 and the second agent 22, cooperate to execute the following four steps to explore and define the chip architecture: 1) task analysis, 2) infrastructure generation, 3) architecture evaluation, 4) architecture definition. The above 1) task analysis mainly refers to analyzing the provided algorithm information, which is completed by the first agent 21. The above steps 2)-4) are mainly implemented by the second agent 22 based on the output of the first agent 21. The code generation and verification stage is achieved through the cooperation of the third agent 23 and the fourth agent 24, and the EDA physical implementation is achieved through the fifth agent 25.
[0071] The specific functions of the above-mentioned execution agents are implemented as follows:
[0072] 2.1. The first agent 21 (such as the "algorithm expert" agent)
[0073] The first agent 21, when executing the function relationship diagram generation task, is specifically used to: generate a function relationship diagram according to the input algorithm information, and send the function relationship diagram to the second agent 22. Among them, each node in the function relationship diagram represents a function module, each edge represents the data dependency relationship between two function modules, and each of the function modules is responsible for completing part of the algorithm tasks in the algorithm information.
[0074] Specifically, the above algorithm information includes algorithm code. Of course, in some other embodiments, the algorithm information may also include other information, such as algorithm documents. An algorithm document refers to a collection of files or materials that describe in detail the design, implementation, usage method, and performance characteristics of an algorithm, and it aims to help developers, users, and other relevant personnel understand the working principle of the algorithm. And, the above first agent 21 can analyze the input algorithm code (or algorithm code and algorithm document), and use software analysis tools to generate the corresponding function relationship diagram. The function relationship diagram can be a task flow graph (TFG).
[0075] In order to realize the architecture exploration of the workload, the steps of generating the function relationship diagram in this embodiment are specifically as follows:
[0076] First, the first intelligent agent 21 parses the input algorithm code from the functional layer. Specifically, the algorithm code can be parsed in combination with the algorithm documentation to comprehensively understand the function and structure of the algorithm, so as to divide the algorithm code into multiple functional modules. Each functional module can include a single functional unit or an aggregation of multiple functional units. Each functional unit corresponds to at least part of the code in the algorithm code. Thus, each functional module and / or functional unit is responsible for the execution task of a specific piece of code.
[0077] Example 11, in the vision Transformer algorithm code, the Transformer code is regarded as a single functional unit, and this single functional unit is used as a functional module. Among them, Vision Transformer (ViT) is a deep learning model based on the Transformer architecture.
[0078] Example 12, a piece of code in the algorithm code that implements the matrix multiplication algorithm. In this piece of code, the inputs are two matrices A and B, and the output is matrix C. The steps involved are triple-loop calculations for the product and accumulation of each element. Then this piece of code can be decomposed into the following multiple functional units, and each functional unit is responsible for the execution task of a part of the specific code: an initialization unit (for initializing the result matrix C), a calculation unit (for implementing the specific calculation logic of matrix multiplication), an auxiliary function unit (for verifying the legality of the input matrices), and a post-processing unit (for performing necessary processing on the result matrix (such as normalization)). The aggregation of the above initialization unit, auxiliary function unit, and post-processing unit forms a functional module m1, and this functional module m1 is responsible for the algorithm tasks of initializing the result matrix C, verifying the legality of the input matrices, and further processing the result matrix (such as normalization) in the matrix multiplication algorithm. And the single calculation unit is determined as a functional module m2, and this functional module m2 is responsible for implementing the specific calculation logic task in the matrix multiplication algorithm.
[0079] Then, the first intelligent agent 21 uses built-in performance analysis tools (such as including gprof or PyTorch Profiler) to analyze the execution time, call frequency, basic operation types and their proportions, and data flow dependencies of the functional modules. Some reports will be generated through the analysis. According to these reports, a functional relationship diagram in JSON format (a task flow graph TFG) can be generated.
[0080] From the above content, it can be seen that each node in the functional relationship diagram can be understood as a workload node, and the workload node represents at least part of the algorithm tasks in the algorithm information (specifically, the execution tasks of part of the code in the algorithm code).
[0081] 2.2. The second intelligent agent 22 ("architecture expert" intelligent agent)
[0082] The second intelligent agent 22 has a memory module. The main function of the memory module is to endow the LLM with memory capabilities, enabling it to store and remember past interaction content. A memory module usually consists of multiple memory components, and each memory component is responsible for different memory tasks. In this embodiment, a hardware core library (such as an OSH IP library) is embedded in the memory module, and each hardware core module in the hardware core library has been pre-developed.
[0083] When the second intelligent agent 22 executes the architecture generation task, specifically, it is used to: allocate a hardware core module to each function module in the function relationship diagram to obtain module information of at least one hardware core module; perform an evaluation operation based on the module information of the at least one hardware core module; and find a matching hardware core module from the hardware core library according to the module information of the at least one hardware core module that passes the evaluation. The hardware module is used for integration into the system-on-chip design.
[0084] During specific implementation, as shown in Figure 2b , for step 2) infrastructure generation shown for the second intelligent agent 22, it can be understood that: allocate a hardware core module to each function module in the function relationship diagram to obtain module information of at least one hardware core module. The module information of the hardware core module includes hardware configuration information and the corresponding function module. The hardware configuration information includes key attribute information, such as storage attributes (cache size, memory hierarchy, etc.), computing attributes, bus bandwidth, etc. The corresponding function module is the workload of the hardware core module (i.e., the task load, which is a node in the function relationship diagram), indicating the algorithm task that the hardware core module is responsible for executing (such as the algorithm task is at least part of the code in the algorithm code). In addition, the module information of the hardware core module may also include other information, such as a hardware identifier (such as a hardware name).
[0085] In this embodiment, when the second agent implements the above step 2), it selects at least one appropriate hardware core module according to the functional relationship graph (TFG), such as one or more of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an NPU (Neural Network Processing Unit), etc., to generate a basic chip architecture (such as a basic SoC architecture) document. Specifically, the second agent uses the functional relationship graph as the input to its internal LLM model. The LLM model internally performs matching of functional modules with hardware core modules according to the characteristics of each functional module in the functional relationship graph (which can be understood as the matching of tasks and architectures shown in the graph), so as to select an appropriate hardware core module for each functional module to implement, and outputs a basic chip architecture document according to the selected hardware core modules. Among them, when selecting an appropriate hardware core module for a functional module, it will specifically determine what configuration (attributes) the hardware core module that this functional module should match. For example, a matching GPU needs to have a large storage capacity and / or strong computing power, etc., so as to generate a hardware table containing hardware configuration information and task loads (i.e., the corresponding functional modules), and this hardware table is the basic chip architecture document.
[0086] For example, assume that the functional relationship graph includes the functional module m1 and the functional module m2 in the foregoing example 12. According to the characteristics of the functional module m1, the hardware core module selected for this functional module m1 can be a CPU, because the tasks of this part of the functional module m1 mainly involve initialization, logical judgment, data verification, and post-processing, which are relatively simple and do not require a large amount of computing, and are suitable for execution on a CPU. According to the characteristics of the functional module m2, the hardware core module selected for this functional module m2 can be a GPU or a dedicated accelerator (such as a DSP, a TPU, etc.), because the tasks of this part of the functional module m2 require a large amount of parallel computing and require the hardware to have strong computing power, so it is suitable for execution on a GPU or a dedicated accelerator. According to the above selected CPU, GPU (or dedicated accelerator), a hardware table will be output as the basic chip architecture document.
[0087] It should be supplemented here that different functional modules can be assigned to different hardware core modules, or they can also be assigned to the same hardware core module.
[0088] From the above content, that is, in this embodiment, the basic chip architecture document can be understood to be a hardware table containing module information of hardware core modules (such as hardware core module identifiers, configuration information of hardware core modules, and corresponding functional modules, etc.). Exemplarily, this hardware table can be seen as shown in Table 1 below:
[0089] Table 1
[0090]
[0091] The above-mentioned Acc (Accumulator) is a special register used to store intermediate results of calculations, and it is commonly used in arithmetic and logical operations.
[0092] Furthermore, the second intelligent agent will execute step 3) architecture evaluation shown in the figure. Specifically, this step 3) can be understood as: performing architecture evaluation based on the basic chip architecture document (i.e., the module information of at least one hardware core module). Specifically, the second intelligent agent performs DSE (Design Space Exploration) using the simulation tool (simulator) configured therein according to the basic chip architecture document to evaluate the PPA (power consumption, performance, area) of the basic chip architecture and optimize the architecture parameters, so as to achieve a relatively optimal balance among architecture performance, function, and area, and then determine that the basic chip architecture meets the requirements.
[0093] Still further, after the evaluation passes, the second intelligent agent will execute step 4) architecture definition shown in the figure based on the simulation report (a PPA report) to generate a chip architecture and output the architecture document of the chip architecture. This step 4) can be understood as including: retrieving and finding matching hardware core modules from the hardware core library according to the module information of at least one hardware core module for subsequent processing steps.
[0094] Specifically in implementation, the matching hardware core modules can be retrieved and found from the hardware core library according to the hardware identifier (such as the hardware name), hardware configuration information, etc. included in the module information of the hardware core module.
[0095] The above-mentioned hardware core library is, for example, an open source (OSH) IP library. Open source IP is a key method to accelerate chip design. In order to achieve flexible retrieval and modification of existing IP, in the embodiments of the present application, an LLM is used to parse and integrate existing open source IP to form a structured OSH IP library (such as an IP library containing hardware IP) for intelligent agents (such as the second intelligent agent 22, etc.) to query and further modify and develop.
[0096] Figure 3 Shows an example flowchart for establishing an OSH IP library. The process of establishing the OSH IP library includes the following steps:
[0097] (1) Collect documents and open source code, and use an LLM to generate organized code and documents to generate structured text.
[0098] First, for each target open-source IP (such as the CVA6 CPU), collect its relevant content, which includes relevant IP documents and open-source IP code. The IP documents can include technical documents such as user manuals, API documents, design specifications, etc. The open-source IP code can include hardware description language codes such as Verilog, VHDL, SystemC, etc.
[0099] Then, use the LLM to parse the collected documents and open-source code to generate organized code and documents. The specific steps can but are not limited to the following:
[0100] 11) Identification and function extraction: Extract the module identification and its corresponding function description from the documents and open-source code. The identification can be a name. For example, for the CVA6 CPU, in addition to identifying the module name, it also includes identifying the names and functions of each unit (such as ALU, FPU, Cache, etc.) in the module.
[0101] 12) Interfaces: Extract the interface definitions between modules, and clarify the input and output signals and their data types. For example, the interface definition between the CPU and the memory controller.
[0102] 13) State machine description: For modules containing state machines, extract the state transition diagram and its description, which helps to understand the working process and behavior of the module.
[0103] 14) Relationships: Such as including the relationships between each unit in the module and the relationships with other modules (external modules), etc.
[0104] 15) Code: Such as the code of each unit in the module.
[0105] Finally, based on the information obtained through the above 11) - 15), generate structured text. The content of the structured text includes: name (such as module name, names of each unit in the module), function description, interface description, state machine description, code (such as module code and / or code of each unit in the module), relationships (such as relationships with other modules, relationships between each unit in the module), etc.
[0106] (2) Construct a graph-structured document (i.e., the module graph-structured document)
[0107] Based on the above-generated structured text, form a graph-structured document (such as the structured IP document shown in Figure 4). In this graph-structured document, a node can represent a unit and its attributes in the current target open-source IP (such as the CVA6 CPU) module, and the edges represent the relationships between different units.
[0108] The same methods given in (1) to (2) above can be applied to different open-source IP modules to generate corresponding organized documents and code, thus finally forming a document in a graph structure.
[0109] (3) Create a concise summary description (understandably, for the key functions of the module)
[0110] Next, create a concise summary description for each IP module to capture its core functions and features.
[0111] For example, the summary description of the CVA6 CPU is: A high-performance RISC-V processor that supports the RV64IMAFDC instruction set and has multi-stage pipelining and dynamic branch prediction functions.
[0112] (4) Classification and integration
[0113] Classify and integrate the IP library according to functions and purposes, and group each IP module into corresponding categories (such as categories like CPU, GPU, DSA, and IO, etc.), thus forming a comprehensively structured open-source IP library that covers a wide range of IP modules.
[0114] Exemplarily: Each IP module can be grouped into the following corresponding categories, but not limited to:
[0115] CPU: Such as IP modules including CVA6, Rocket Chip, etc.
[0116] GPU: Such as IP modules like OpenCL GPU, NVIDIA CUDA Core, etc.
[0117] DSA (Domain-Specific Accelerator): Such as IP modules like TPU for machine learning, dedicated accelerators for encryption algorithms, etc.
[0118] IO (Input / Output): Such as IP modules like USB controllers, PCIe interfaces, etc.
[0119] This structured OSH IP library formed through the above (1) to (4) can provide great support for the subsequent rapid query and flexible modification of the agents.
[0120] For example, when the second agent 22 uses a CPU module (such as the module module2 shown in Figure 3 ) in this OSH IP library, it can, according to actual needs, delete redundant units therein, and / or modify one or more units, and / or can also add new units, etc.
[0121] Furthermore, continue to refer toFigure 2b As shown, when the second intelligent agent 22 is executing step 4) of the architecture definition process, if a matching module cannot be retrieved from the hardware core library for an expected required hardware core module (such as a special NPU design), in this case, the second intelligent agent will use the LLM configured within it to design and generate this new hardware core module.
[0122] The above design and generation method includes: selecting a hardware core module of the same type from the hardware core library as the initial hardware core module, and then modifying the initial hardware core module to design this new hardware core module. Specifically, the design of this new hardware core module is achieved through hierarchical driving and data flow driving. Hierarchical driving means adopting a hierarchical method to refine and modify layer by layer in a top-down manner. Data flow driving means that for any hardware unit on each layer, its internal refinement and modification are based on the data transmission path. Conducting data flow driving design within the hardware unit means that after data is transmitted into a hardware unit, it will sequentially pass through multiple processing units according to a predetermined data path. Each processing unit is responsible for a specific task and passes the processing result to the next processing unit. Also, in order to achieve data flow driving within the unit, data interfaces are defined between each processing unit to ensure that data can be efficiently transmitted from one processing unit to the next processing unit.
[0123] Exemplarily, assume that the functional module represented by a node in the functional relationship graph (TFG) is responsible for tasks that include a large number of matrix multiplication operations with high performance requirements (such as CNN algorithms, sparse computations, etc.). It is determined that this functional module is to be assigned to an NPU for execution. That is: According to the functional relationship graph, what ultimately needs to be generated is a baseline architecture with a dedicated NPU domain-specific architecture (DSA), and a corresponding node in the functional relationship graph (the functional module represented by this node is responsible for tasks that include a large number of matrix multiplication operations with high performance requirements) will be assigned to this NPU intellectual property (IP). However, the expected NPU module cannot be found in the existing hardware core library, but an NPU module of the same type can be found. In this case, this found NPU module of the same type can be used as the reference initial NPU module. Then, a hierarchical method is adopted to gradually refine and modify this initial NPU module in a top-down manner, ultimately enabling the modified reference NPU module to support operations such as CNN algorithms and sparse computations. Specifically, the second agent 22 will first determine the design requirements for the module information of this NPU module. The design requirements are, for example, to design an NPU that supports convolutional neural network (CNN) algorithms and sparse computations. Then, according to the design requirements, the modification plan is determined layer by layer in a top-down manner, and the modification plan can be fed back to the management agent for the management agent to evaluate the modification plan. After the evaluation passes, the second agent 22 executes the modification operation according to the modification plan. Among them, the feedback to the management agent can be made after determining the modification plan for one level, or it can also be made after determining the modification plans for all levels. And, the modification operations included in the modification plan include deleting some existing redundant functional units, and / or adding new functional units, and / or modifying one or more functional units.
[0124] As Figure 4a shown, the initial NPU module is divided into three levels: the first level Level1, the second level Level2, and the third level Level3. Among them, the next level corresponds to a certain unit in the previous level Level1. The modification process for these three levels can be as follows:
[0125] First, according to the tasks to be responsible for, such as the need to support CNN algorithms and sparse computations, the modification plan for the first level Level1 is determined as: deleting redundant units. Among them, except for the computing units, storage units, and interconnection buses, other units are redundant. After this modification plan is fed back to the management agent 10 for evaluation and passes, the second agent 22 will delete the other units among them, completing the modification of the first level Level1.
[0126] Then, modify the second level Level2. The second level level2 corresponds to the computing units on the first level Level1, that is, each unit in the computing units is included on the second level Level2. For this second level Level2, in order to implement the CNN algorithm, the determined modification scheme is as follows: It is necessary to add an Img2col unit in this level to support the img2col technology; and by performing data flow-driven design inside the Img2col unit, it is also determined that after the data is transmitted into the Img2col unit, it needs to pass through the input configuration unit, padding unit, image expansion unit, and matrix control unit in sequence according to the corresponding data path. After this modification scheme is fed back to the management agent 10 for evaluation and approval, the second agent 22 will add an Img2col unit newly on the second level Level2, and this Img2col unit includes the four processing units of the input configuration unit, padding unit, image expansion unit, and matrix control unit. The data dependency relationship between these four processing units is: input configuration unit -> padding unit -> image expansion unit -> matrix control unit. The arrow "->" that appears in this application indicates the data flow direction. Img2col is a commonly used matrix transformation technology, widely used in convolutional neural networks (CNNs). Especially when implementing the convolutional layer, it unfolds the local area of the input image into matrix columns (columns), enabling the convolutional operation to be converted into matrix multiplication, thereby simplifying the calculation process and being able to utilize highly optimized linear algebra libraries (such as BLAS) for acceleration more efficiently. Also, the functions of each subunit included in the Img2col unit are as follows: The input configuration is responsible for receiving and parsing external input parameters, such as image size, convolutional kernel size, stride, and padding, etc. These parameters will be used to guide the subsequent image expansion process; the padding unit is responsible for adding zero-valued pixels around the input image to ensure that the convolutional operation does not lose edge information; the image expansion is responsible for unfolding the input image data according to the specified convolutional window and storing it into the matrix column; the matrix control is responsible for managing the timing and status of the entire img2col process. It ensures that each unit executes in the correct order and handles possible state transitions and error conditions.
[0127] Finally, modify the third level Level3. Specifically, to support sparse computing, based on the data flow drive of sparse computing, the modification plan for the PE unit is determined as follows: a zero detection unit and a MAC unit need to be added to the PE unit, and the data dependency relationship is zero detection unit -> MAC unit. After this modification plan is fed back to the management agent 10 for evaluation and approval, the second agent 22 will execute adding a zero detection unit and a MAC unit to the PE unit, thus forming a new PE unit. Among them, the PE array refers to a set of parallel PE units, and each PE unit is an independent computing unit (such as a floating-point arithmetic unit (FPU) or a dedicated accelerator (such as a matrix multiplication unit, etc.)), which can perform basic arithmetic and logical operations. These PE units are usually used to execute specific types of computing tasks. The MAC unit mainly performs multiplication and accumulation operations and is widely used in various computing tasks, such as matrix multiplication, convolution operations in convolutional neural networks (CNNs), filters in signal processing, etc.
[0128] After the modification is completed, then the second agent 22 will call the simulation tool configured therein to simulate and evaluate the modified initial NPU module, determine the PPA (power consumption, performance, and area) of the initial reference NPU module and compare it with the PPA required by the user. Furthermore, according to the comparison result, iterative optimization is carried out by adjusting parameters and improving module functions until the user requirements can be met. Among them, the PPA required by the user can be sent by the management agent 10 to the second agent 22. After that, the second agent will feedback the module document of the modified initial NPU module to the management agent 10, and the management agent 10 will review this module modification design according to a set of predefined data flow inspection criteria. During the review, as Figure 4b shown in the left figure in, the management agent 10 will, for example, pay attention to one or more of the following issues: unit redundancy, unit missing, potential data flow blockage or congestion, etc. If any problem is found, the management agent 10 will instruct the second agent 22 to make modifications. After the second agent 22 finally completes the modification design of this initial NPU module, the management agent 10 will flatten the entire modified initial NPU module for the final module evaluation. Among them, flattening a hardware core module (such as the modified reference NPU module here) usually means: converting a complex multi-level hardware core module structure into a single-layer form that is easier to understand and implement. For example, for the flattened modified initial NPU module mentioned above, the PE array, each processing unit in the Img2col unit, etc. are all at the top layer.
[0129] In summary, as Figure 4bAs shown in the right figure in , during the process of the second agent 22 modifying and designing a new hardware core module, an adaptive system based on feedback is formed between the management agent 10 and the second agent 22. The second agent 22 performs module design exploration according to predefined guiding principles (such as hierarchical (e.g., layered and top-down) and data flow-driven principles), while the management agent is responsible for supervision and verification evaluation. Among them, the interaction (feedback, communication) between the second agent 22 and the management agent 10 is carried out through structured module description texts, which include current module information and the summary context of the discussion.
[0130] After finally completing the design of a hardware core module, such as completing the modification of the above-mentioned initial NPU model, the second agent will determine that the designed hardware core module is the target hardware core module required as expected, and will write a structured module description document for this hardware core module to assist in subsequent RTL writing, etc. Among them, the content in the structured module description document includes but is not limited to module identification (such as module name, names of each unit in the module), function (such as key function description), structural hierarchy, interface signals, relationship with other external modules, relationship between each unit in the module, state machine description, etc. In addition, the second agent 22 can also generate and output a corresponding module design overview document for this hardware module, and the module design overview document includes an overview of the module design process of the hardware core module.
[0131] The above takes selecting a hardware core module from the hardware core library as the initial hardware core module as an example to detail the implementation of a new hardware core module design. Of course, in other embodiments, when the second agent cannot find a target hardware core module that meets the requirements in the hardware core library, other methods can also be used to determine the corresponding initial hardware core module. For example: based directly on the module information of the target hardware core module that is expected to be needed, a hierarchical determination and data flow-driven method can be used to automatically generate a corresponding initial hardware core module, and feedback to the management agent for evaluation for this initial hardware core module, and then modify this initial hardware core module according to the problems feedback by the management agent until the requirements are met.
[0132] Based on the above content, the second agent 22 is also used for: when a matching target hardware core module cannot be found in the hardware core library, generating a hardware core module as the target hardware core module based on the model information of the target hardware core module.
[0133] Among them, when the second agent 22 is used to generate a hardware core module as the target hardware core module based on the model information of the target hardware core module, it is specifically used for:
[0134] Determine an initial hardware core module, where the initial hardware core module is of the same type as the required target hardware core module;
[0135] Modify the reference hardware core module based on the module information of the target hardware core module;
[0136] Determine the modified reference module as the target hardware core module;
[0137] Among them, the initial hardware core module includes a hardware core module selected from the hardware core library; the initial hardware core module includes at least one level, and each level has at least one hardware unit. The modification includes performing at least one of the following modification operations on the hardware units on the level in a top-down manner: deleting redundant hardware units, adding hardware units, and modifying hardware units; when adding a hardware unit and / or modifying a hardware unit, the processing units to be included in the hardware unit are determined according to data flow driving.
[0138] Furthermore, the second agent 22 can also add the self-designed target hardware core module to the hardware core library to expand the hardware core library. Thus, the second agent 22 can also be used for:
[0139] Generate a module graph structure document and a module key function description according to the relevant information of the target hardware core module;
[0140] Add the target hardware core module to the appropriate module classification group in the hardware core library according to the module graph structure document and the module key function description;
[0141] Among them, the relevant information includes the code information (module code) of the target hardware core module obtained from the third agent. In addition, it can also include module documents, etc. The module documents can include technical documents, the content of which can describe in detail the design specifications, implementation, usage methods, performance characteristics, key functions, interfaces, etc. of this target hardware core module, and this can be generated by the second agent. The third agent is one of the multiple agents; a node in the module graph structure document represents a hardware unit in the target hardware core module and the attributes of the hardware unit, and the edge represents the relationship between different units in the target hardware core module.
[0142] The specific implementation process of adding the target hardware core code to the hardware core library is described here. For details, refer to the relevant process of establishing the hardware core library (OSH IP library) described in combination with 3 in other embodiments, and no specific elaboration will be made here. In addition, after the third agent and the fourth agent cooperate to generate the code of the target hardware core module and complete the test verification, the step of adding the target hardware core module to the hardware core library as described above can be triggered for execution.
[0143] 2.3. The third agent 23 (such as the "RTL writer" agent shown in Figure 2a etc.) and the fourth agent (such as the "RTL tester" agent shown in Figure 2a etc.).
[0144] After the above-mentioned second agent 22 generates a comprehensive module description document (such as the Arch document shown in Figure 2c etc.) and a module design overview document for the newly designed hardware core, these documents will be transferred to the code generation and verification stage for processing.
[0145] As shown in the example in Figure 2c referred to, for the code generation and verification stage, it can be divided into the following four steps: 1) IP code generation, 2) IP code verification, 3) SoC integration, 4) SoC verification. The above steps 1) and 3) are the responsibilities of the third agent 23 (i.e., the third agent is responsible for converting the corresponding description documents (such as module description documents) into RTL code, specifically Verilog code), and steps 2) and 4) are the responsibilities of the fourth agent 24. And the above-mentioned IP code can refer to RTL code, and the RTL code can come from the OSH IP library and / or be generated using LLM. The above-mentioned IP code generation and IP code verification refer to generating the corresponding IP code (as RTL code) for the hardware core module and verifying this IP code. Additionally, the RTL code generation follows a hierarchical-driven structure and a data-flow-driven orientation. Specifically, that is: the third agent 23 and the fourth agent 24 execute their respective tasks in a hierarchical-driven and data-flow-driven orientation based on the received module description documents, etc.
[0146] Specifically, for a self-designed hardware core module (such as the one combined with Figure 4aDescribe an NPU module designed and generated by the second agent 22. The third agent 23 can use the fine-tuned LLM model (such as the Llama3.170B model) therein to generate the RTL code of the hardware core module according to the received corresponding module description document and module design overview document. In addition, the module description document (or the module description document and module design overview document) will also be provided to the fourth agent 24 to facilitate the development of reference hardware core modules and test cases. After receiving the RTL code from the third agent 23, the fourth agent 24 will use the established UVM verification environment and tools such as VCS for test verification. Once all test cases pass, the self-designed hardware core module (hardware IP module) will be assembled and integrated with other hardware core modules obtained from the OSH IP library for SoC-level test verification. In this application, the processes for hardware IP and SoC verification are the same, except for the different test objects during verification. Also, for other hardware core modules obtained from the OSH IP library, their RTL code can be obtained from the OSH IP library or can also be generated using an LLM model, which is not specifically limited here.
[0147] It should be noted here that: In each figure (such as Figure 2c ), the hardware core modules shown, the hardware core modules with the "Self IP" identifier represent a hardware core module newly modified and designed by the second agent 22, and those with only the "IP" identifier are those previously stored in the hardware core library (such as the OSH IP library). Also, for the hardware core modules found in the hardware core library, in addition to obtaining their module codes from the hardware core library, in other embodiments, they can of course also be generated using an LLM. For example, the second agent 22 can send the module description information of this hardware core module obtained from the hardware core library to the third agent 23 and the fourth agent 24, and thus the third agent 23 and the fourth agent 24 cooperate to generate the module code of this hardware core module.
[0148] Exemplarily, refer to Figure 5Taking the example given within the dashed box a, for an NPU module included in a chip architecture, when the third agent 23 executes its own tasks in a hierarchical-driven (in a bottom-up manner) and data-flow-driven oriented way according to the module description document of this NPU module, it will first extract from the module description document the names, structures, and function descriptions, etc. of units such as PE units in the third level Level3. Subsequently, according to the information extracted, corresponding unit code files will be generated for the PE units. After the generation of the unit code files for each unit on the third level Level3 is completed, it will execute the extraction of the names, structures, and function descriptions, etc. of each unit in the second level Level2 from the module description document, so as to generate corresponding unit code for each unit thereon. For example, taking the PE array on the second level Level2 as an example, corresponding unit code files can be generated for the PE array according to the extracted information such as the name, structure, and function description of the PE array, its relationship with other external units, and the relationship between each PE unit therein, and in accordance with the transmission path of the data flow within it. According to the unit code files of each unit on the second level, in accordance with the data flow relationship of each unit component on the second level, the hierarchical code file of the second level Level2 can finally be generated, and this hierarchical file of the second level Level2 is the unit code file of the computing unit on the first level Level1. After the generation of the unit code files for each unit on the second level Level3 is completed, it will execute the extraction of the names, structures, and function descriptions, etc. of each unit in the first level Level1 from the module description document, so as to generate corresponding unit code for each unit thereon. For the generation of the unit code files of each unit such as the computing unit and storage unit on the first level Level1 here, reference can be made to the relevant content regarding the generation of the relevant unit code files described above for the first level Level1 and the second level Level2. Based on the generated unit codes above, the module code file of the NPU module will finally be generated.
[0149] It should be supplemented and explained here that: in this application, the format of the code files (such as unit code files, etc.) can be but is not limited to Verilog code files. Verilog code is a hardware description (HDL) code, specifically a kind of code at the RTL level. Also, for the third agent 23, in order to enhance the ability of the LLM model in Verilog code generation, this application fine-tuned the LLM model (such as the Llama 3.1 70B model) using an open-source dataset and integrated the fine-tuned LLM model as a core component into the third agent. That is, the code generation described in the above example is generated by the third agent 23 using the fine-tuned LLM model within it. Additionally, as can be seen Figure 2cRegarding the content shown in step 2), after the third intelligent agent 23 generates the module code file of a hardware core module, it will also call a code analysis tool (such as Linter, a static code analysis tool) to analyze and detect potential errors, style inconsistencies, and other issues in the module code, and output an inspection report. The inspection report (Linter report) usually contains a series of warning and error messages, indicating the areas that need improvement in the code and modification suggestions, etc. After that, the third intelligent agent 23 will also modify the generated module code file according to this inspection report until there are no errors in the module code file, and then trigger the execution to send the module code file to the fourth intelligent agent 24 for test verification.
[0150] Further, referring to Figure 5 Regarding the content shown in the dashed box b in
[0151] After the test verification passes, before performing step 3) SoC integration, that is, before assembling and integrating each hardware core module, as shown in Figure 5 Regarding the content shown in the dashed box c in Figure 2cThe DSSoC.V file shown, specifically a netlist file. After that, the third intelligent agent 23 will send the chip-level code file to the fourth intelligent agent 24, and the fourth intelligent agent 24 will conduct system-level test verification. Once the test verification passes, the chip-level code file will ultimately be transferred to the physical intelligent agent to enter the EDA physical implementation stage for subsequent physical implementation.
[0152] 2.4. The fifth intelligent agent (such as Figure 2a , 2d the "EDA expert" intelligent agent shown in etc.)
[0153] As shown in, for example, Figure 2d the fifth intelligent agent is equipped with relevant EDA scripts and tool manuals. In the EDA physical implementation stage, the physical intelligent agent develops and customizes scripts according to specific requirements and applies them to corresponding EDA tools (such as Genus and Innovus). Through iterative adjustment, the physical intelligent agent will ultimately generate and output a GDSII file (a physical layout file) that meets the specified design requirements.
[0154] Exemplarily, the generation process of the GDSSII file may include the following steps:
[0155] S1. Synthesis: Convert the code in the chip-level code file into a gate-level netlist.
[0156] S2. Floorplanning: Determine the overall layout of the chip.
[0157] S3. Power planning: Design the power distribution network.
[0158] S4. Placement: Place the logic cells on the chip.
[0159] S5. Clock tree synthesis (CTS): Generate a clock tree to ensure low jitter of the clock signal.
[0160] S6. Routing: Connect all logic cells and power networks.
[0161] S6. Static timing analysis (STA): Verify whether the timing meets the requirements.
[0162] Through the above, after completing the physical implementation steps such as synthesis, placement and routing, a physical view (chip layout) of the chip design will be formed, and based on this physical view, the design file of the chip can be generated. According to this design file, subsequent chip manufacturing can be carried out.
[0163] In summary of the content in 2.3 - 2.4, the execution tasks of the third intelligent agent 23, the fourth intelligent agent 24 and the fifth intelligent agent 25 can be briefly described as follows:
[0164] The third intelligent agent 23 is used to execute a code testing task to generate the module code of the hardware core module according to the module description document of the hardware core module;
[0165] The fourth intelligent agent 24 is used to execute a code testing task to generate test cases and a reference hardware core module according to the module description document of the hardware core module, and test the module code of the target hardware core module based on the test cases and the reference hardware core module;
[0166] The third intelligent agent 23 is further used to, after determining that the verification intelligent agent has completed the code verification task, perform logic synthesis processing on the module codes of the hardware core modules (including the newly designed hardware modules by the second intelligent agent and the hardware core modules found from the hardware core library) to obtain a chip-level code file;
[0167] The fourth intelligent agent 24 is further used to test the chip-level code file;
[0168] The fifth intelligent agent 25 is used to, after determining that the chip-level code file passes the test, trigger the execution of a physical implementation task to generate a chip layout according to the chip-level code file, and generate the chip design file according to the chip layout. The chip design file contains chip layout information and is used for chip manufacturing.
[0169] It should be supplemented here that: The chip-level code file described in this application can be understood as a netlist (such as the netlist shown in Figure 1 For example, a gate-level netlist). A netlist is a text representation form of circuit design. It is an important link in the chip design process. It converts high-level design descriptions into low-level logic gate-level representations, facilitating subsequent physical implementation and verification. That is, a netlist can be regarded as an important step in converting high-level design descriptions (such as Verilog or VHDL code) into physical implementation (layout design)
[0170] To verify this application, two chip design scenarios are listed below.
[0171] Using the solution of this application, two domain-specific system-on-chips (DSSoCs) are developed for the Internet of Things (IoT) and mobile application fields, with the design information listed in Table 2. Figure 1 The chip layouts generated for these two fields are shown. Among them, the IoT chip (whose layout is shown as Case A on the left in Figure 6 is designed to support algorithms such as MobileNet (a lightweight deep learning model designed specifically for mobile devices and embedded devices), ResNet (residual neural network), and DS-CNN (depthwise separable convolutional neural network), while the mobile system-on-chip (whose layout is shown as Figure 6The case B shown on the right in [Case] can run ViT, Llama2-7B, and 3D GS (3D Gaussian Scattering). For case A, the chip architecture includes a CPU, a neural network processing unit (NPU), and a digital signal processor (DSP). Among them, the NPU enhances sparse computing and img2col support. For case B, the chip architecture integrates a CPU, a GPU, and an NPU. The GPU is extended through a dedicated functional unit (SFU) and a tensor computing module. The NPU is further optimized for INT8 / FP16 mixed-precision computing and sparse computing.
[0172] During the code generation and SoC integration process, case A uses OSH CVA6 as the CPU, integrates multiple modules of OSH to build the DSP, and independently develops the NPU. For case B, a single-core OSH C910 CPU is used, and Nyuzi is selected as the GPU. The GPU is further enhanced by adding an SFU and a tensor computing unit. The third agent (the "RTL writer" agent) also independently develops the NPU code for case B. Finally, the physical agent (the "EDA expert" agent) implements case A through a flattening method, while case B adopts a modular method combined with a hierarchical physical design process. Different technology nodes (7nm and 22nm) are adopted according to different cost considerations. After completing the entire workflow design process, the final designs of the two SoCs are finally obtained, as Figure 6 shown in Table 2.
[0173] Table 2 Design Information
[0174]
[0175]
[0176] The above Table 2 shows the detailed design specifications of the two chips, and their operating frequencies are 500MHz and 1GHz respectively. The total area and power consumption of case A are 4.0 square millimeters and 419.6 milliwatts respectively, while those of case B are 30.5 square millimeters and 22.6 watts. Considering the manual debugging and tool usage in the design process, the design time of case A is about 2 weeks, while case B takes about 4 weeks. In contrast, an existing agile-designed SoC (system-on-chip) is reported to have a design time of 10 - 12 weeks. Obviously, the proposed solution in this application significantly shortens the design cycle, demonstrating its efficiency in rapid design.
[0177] Figure 7Shows the energy efficiency and performance improvements of Case A and Case B. By designing the DSA (Domain-Specific Architecture) architecture through the solution of this application, it is obvious that significant efficiency and performance enhancements have been obtained. For example, the MobileNet algorithm achieved a 23.45-fold increase in efficiency and a 26.72-fold increase in performance on the basic Case A architecture. During the architecture optimization process of this application, additional improvements of 1.47 times and 1.34 times were obtained. For Case B, the evaluation results show that the ViT algorithm achieved an impressive 36.64-fold increase in energy efficiency and a 43.44-fold increase in performance on the finally improved architecture.
[0178] In addition, this application also compares the energy efficiency of the existing Internet of Things terminal STM32MP25 SoC with that of Case A, and the energy efficiency of the mobile Jetson Nano with that of Case B. As Figure 8 shown, the energy efficiency of the DS-CNN running on Case A is 23.81 times that of the STM32MP25, while the energy efficiency of the Llama2-7B running on Case B is 32.43 times that of the Jetson Nano. Other algorithms also show different degrees of improvement.
[0179] To verify the impact of the multi-agent system in the theory of this application solution, the workflow performance under different numbers of agents was also evaluated. As Figure 9 shown, this application also shows the execution workflow of the neural processing unit (NPU) and the code generation process of the systolic array that changes with the number of agents in the workflow and the number of agents in the code generation. This application is standardized based on the highest energy efficiency, the smallest total area, and the shortest running time. For the NPU design task, when the number of agents in the workflow is too large, such as reaching 10, the design time reaches the maximum, which is about 7 times longer than that with 6 agents, and the energy efficiency drops to the lowest. Similarly, for the code generation of the systolic array, the code writing process involving 5 agents results in a design time that is almost 5 times the shortest time, and the total area expands to about 3 times the smallest area. This phenomenon is mainly attributed to the increased communication overhead between agents and the possible information distortion during the communication process, which ultimately leads to a decrease in efficiency.
[0180] Another embodiment of this application also provides a chip design method, which is implemented based on the system provided by the foregoing this application. Specifically, the method includes:
[0181] S1. Use multiple agents to cooperate to execute the following operation steps:
[0182] S11. Receive the design request of the user, and the design request includes chip design requirement information;
[0183] S12. Based on the chip design requirement information, perform chip design processing to generate a chip layout file.
[0184] For the specific implementation of each step in the above method embodiments, reference may be made to the relevant content in other embodiments. In addition, the above method embodiments not only include the steps shown above, but may also include other steps. For the other steps that may be included, reference may also be made to the relevant content in other embodiments, and no specific elaboration will be made here.
[0185] This application also provides an electronic device. The above system provided in other embodiments of this application is installed and deployed on the electronic device. The electronic device can be various terminal devices such as smart phones, personal computers (such as desktop computers, laptop computers), tablets, etc.; or it can also be a server device, such as a single server, a server cluster, a virtual server, etc.
[0186] Furthermore, the electronic device also includes other components, such as an interface module, a control panel, and other components.
[0187] Correspondingly, an embodiment of this application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the steps or functions in the method embodiments provided in this application.
[0188] In addition, an embodiment of this application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it can implement the steps or functions in the method embodiments provided in this application.
[0189] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM (Read Only Memory), RAM (Random Access Memory), magnetic disks, optical disks, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.
Claims
1. A multi-agent based chip design system, characterized in that: include: Multiple agents; The multiple agents collaborate to perform the following operations: Receiving a design request from a user, wherein the design request includes chip design requirement information; Based on the chip design requirement information, execute chip design processing to generate a chip design file; The chip design file contains chip layout information.
2. The system according to claim 1, characterized in that The chip design requirement information includes required function information, and the function information includes algorithm information; the multiple agents include at least one execution agent, and the at least one execution agent includes a first agent and a second agent; The first agent is used to generate a functional relationship graph based on the algorithm information; each node in the functional relationship graph represents a functional module, and each edge represents a data dependency relationship between two functional modules; each functional module is responsible for part of the algorithm tasks in the algorithm information; The second intelligent agent has a memory module, and the memory module is embedded with a hardware core library; the second intelligent agent is used to: assign a hardware core module to each functional module in the functional relationship diagram, and obtain module information of at least one required hardware core module; according to the module information of the hardware core module, search for a matching hardware core module from the hardware core library; wherein the module information includes configuration information of the hardware core module and the corresponding functional module.
3. The system according to claim 2, characterized in that The target hardware core module is one of the at least one required hardware core module; and The second agent is further used to: when no matching target hardware core module can be found in the hardware core library, generate a hardware core module as the target hardware core module based on the module information of the target hardware core module.
4. The system according to claim 3, characterized in that The second agent, when used to generate a hardware core module as the target hardware core module based on the model information of the target hardware core module, is specifically used to: Determine an initial hardware core module, the initial hardware core module is of the same type as the required target hardware core module; Modifying the initial hardware core module based on the module information of the target hardware core module; Determining the modified initial hardware core module as the target hardware core module; Among them, the initial hardware core module includes a hardware core module selected from a hardware core library; the initial hardware core module includes at least one level, each level has at least one hardware unit, and the modification includes performing at least one of the following modification operations on the hardware units on the level in a top-down manner: deleting redundant hardware units, adding hardware units, and modifying hardware units; when adding a hardware unit and / or modifying a hardware unit, the processing unit to be included in the hardware unit is determined according to the data flow driver.
5. The system according to claim 4, characterized in that The second agent is also used to: Generate a module graph structure document and a module key function description based on the relevant information of the target hardware core module; According to the module graph structure document and the module key function description, adding the target hardware core module to a corresponding module classification group in the hardware core library; Among them, the relevant information includes the module code of the target hardware core module obtained from the third agent, and the third agent is one of the at least one execution agent; a node in the module graph structure represents a hardware unit in the target hardware core module and the attributes of the hardware unit, and the edge represents the relationship between different units in the target hardware core module.
6. The system according to any one of claims 3 to 5, characterized in that The at least one execution agent also includes a third agent, a fourth agent, and a fifth agent; The second agent is further used to: generate a module description document of the target hardware core module; The third agent is used to generate the module code of the target hardware core module according to the module description document; The fourth agent is used to generate a test case and a reference hardware core module according to the module description document, and test the module code of the target hardware core module based on the test case and the reference hardware core module; The third agent is further used to perform logic synthesis processing on the module code of the target hardware core module and the module code of each hardware core module found from the hardware core library to obtain a chip-level code file after determining that the module code test of the target hardware core module passes; The fourth intelligent agent is further used to test the chip-level code file; The fifth intelligent agent is used to generate a chip layout according to the chip-level code file that has passed the test, and to generate the chip design file according to the chip layout.
7. The system according to claim 6, characterized in that The multiple agents also include: a management agent; The management agent is used to receive a design request from a user and perform chip design task decomposition based on chip design requirement information contained in the design request to allocate an adaptive processing task to the at least one execution agent.
8. A multi-agent based chip design method, characterized in that: include: Use multiple agents to collaborate and perform the following steps: Receiving a design request from a user, wherein the design request includes chip design requirement information; Based on the chip design requirement information, chip design processing is performed to generate a chip layout file.
9. An electronic device, characterized in that: include: The system of any one of claims 1 to 7 above.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method of claim 8 can be implemented.
Citation Information
Cited By
Stacked chip system cross-level error tracking and repairing method and system
CN121257419A
Simulation error correction method, device and system based on multi-agent cooperation
CN122152585A
RTL code generation method and system based on demand multilayer modeling
CN122219933A
A Method and System for RTL Code Generation Based on Multi-Level Requirements Modeling
CN122219933B
Chip full-link collaborative design method and device and storage medium
CN122363942A