Hardware simulator cluster and method for chip hardware simulation
Different parts of the ultra-large-scale chip design are simulated separately through the hardware emulator cluster, and the logical configuration match interface signal transmission function is used to solve the problem of insufficient simulation capacity and compilation time in the existing technology, and a more efficient simulation and compilation process is achieved.
Patent Information
- Application Number
- CN202311837922.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing hardware simulation solutions have insufficient simulation capacity and compilation time, which is difficult to meet the needs of ultra-large-scale chip design.
A hardware emulator cluster is used to simulate different parts of the circuit to be tested through multiple cabinet systems, and the logical configuration match interface signal transmission function is used to realize joint simulation of multiple parts.
It improves simulation capacity and compilation efficiency, reduces compilation time, and realizes synchronous simulation of multiple parts of the circuit to be tested.
Smart Images

Figure CN120215293A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure mainly relate to the field of chip hardware emulation. More specifically, embodiments of the present disclosure relate to a hardware emulator cluster, a method, a computer-readable storage medium, and a computer program product for chip hardware emulation. Background Art
[0002] Electronic design automation (EDA) tools are widely used in the processes of chip functional design, synthesis, verification, etc. Hardware emulation tools are an important part of EDA tools and can efficiently perform hardware emulation verification of chip designs. Generally, the code describing the chip design can be compiled into a compiled version that can be executed on the dedicated hardware circuit of a hardware emulator, and then the compiled version is executed by the hardware emulator to verify whether the chip design meets the expected goals.
[0003] With the development of semiconductor technology, the scale of chip design has reached tens of billions of gates, so a hardware emulation solution that supports ultra-large scale, such as a simulation capacity of more than tens of billions of gates, is required. Summary of the Invention
[0004] Some related hardware emulation solutions have problems of small emulation capacity and long compilation time. Embodiments of the present disclosure provide a solution for chip hardware emulation to at least partially solve the above problems.
[0005] In a first aspect of the present disclosure, a hardware emulator cluster for chip hardware emulation is provided. The hardware emulator cluster includes: a first cabinet system, which includes a plurality of interconnected first cabinets and first logic. The plurality of first cabinets are configured to run a first compiled version of a circuit under test to emulate a first part of the circuit under test and output a plurality of first signals through a plurality of first interfaces. The hardware emulator cluster further includes: a second cabinet system, which includes a plurality of interconnected second cabinets and second logic. The plurality of second cabinets are configured to run a second compiled version of the circuit under test to emulate a second part of the circuit under test and receive the plurality of first signals through a plurality of second interfaces. Wherein the first logic is connected to the second logic, and at least one of the first logic or the second logic is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
[0006] In this way, by using a hardware emulator cluster including multiple cabinet systems to emulate a circuit under test, a larger verification scale than a single cabinet system can be achieved. In addition, by separately compiling different parts of the circuit under test, the compilation efficiency can be improved and the compilation time can be reduced.
[0007] In some embodiments of the first aspect, the second logic is configured to: define the signal receiving functions of the plurality of second interfaces to match the plurality of first interfaces. In some embodiments of the first aspect, the first logic is configured to: define the signal transmitting functions of the plurality of first interfaces to match the plurality of second interfaces.
[0008] In some embodiments of the first aspect, the plurality of second cabinets are further configured to output a plurality of second signals through the plurality of second interfaces, and the plurality of first cabinets are further configured to receive the plurality of second signals through the plurality of first interfaces. In some embodiments of the first aspect, the second logic is configured to: define the signal transmitting functions of the plurality of second interfaces to match the plurality of first interfaces. In some embodiments of the first aspect, the first logic is configured to: define the signal receiving functions of the plurality of first interfaces to match the plurality of second interfaces.
[0009] In this way, by using one or more logics in each cabinet system to define the signal transmission functions of the interfaces, the same signals in different parts of the circuit under test simulated by each cabinet system can be interconnected, so as to realize the co-simulation of multiple parts of the circuit under test.
[0010] In some embodiments of the first aspect, when running the first compiled version of the circuit under test, the plurality of first cabinets are used to: control the simulation period of the first compiled version to be the same as the simulation period of the second compiled version. In some embodiments of the first aspect, when running the second compiled version of the circuit under test, the plurality of second cabinets are used to: control the simulation period of the second compiled version to be the same as the simulation period of the first compiled version. In this way, the transmission frequencies of the same signals in each cabinet system in the hardware emulator cluster can be matched, so as to better realize the co-simulation of multiple parts of the circuit under test.
[0011] In some embodiments of the first aspect, the hardware emulator cluster further includes: a management unit configured to manage at least one of the hardware simulation resources or clock resources of the first cabinet system and the second cabinet system.
[0012] In some embodiments of the first aspect, the first cabinet system further includes a first host configured to: initiate a snapshot saving operation to save the snapshot data of the hardware emulator cluster; and initiate a snapshot restoration operation to restore the hardware emulator cluster by using the snapshot data of the hardware emulator cluster.
[0013] In some embodiments of the first aspect, the first cabinet system further includes a first host system, and the first host system is configured to: initiate a waveform sampling operation based on a trigger condition to sample simulation waveforms from the hardware emulator cluster. In some embodiments of the first aspect, before initiating the waveform sampling operation, the first host system is configured to: initiate a simulation pause operation based on the trigger condition to pause the simulation process of the hardware emulator cluster.
[0014] In this way, trigger detection, signal sampling, and / or waveform analysis can be performed on the entire hardware emulator cluster, so that the simulation of the circuit under test running in the hardware emulator cluster can be conveniently debugged.
[0015] In a second aspect of the present disclosure, a method for chip hardware simulation is provided. The method includes: using a first cabinet system in a hardware emulator cluster to simulate a first part of a circuit under test, the first cabinet system including a plurality of interconnected first cabinets and first logic, the plurality of first cabinets being configured to run a first compiled version of the circuit under test and output a plurality of first signals through a plurality of first interfaces. The method further includes: using a second cabinet system in the hardware emulator cluster to simulate a second part of the circuit under test, the second cabinet system including a plurality of interconnected second cabinets and second logic, the plurality of second cabinets being configured to run a second compiled version of the circuit under test and receive the plurality of first signals through a plurality of second interfaces. The first logic is connected to the second logic, and at least one of the first logic or the second logic is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
[0016] In this way, by using a hardware emulator cluster including a plurality of cabinet systems to simulate a circuit under test, a larger verification scale than a single cabinet system can be achieved. In addition, by separately compiling different parts of the circuit under test, the compilation efficiency can be improved and the compilation time can be reduced.
[0017] In some embodiments of the second aspect, the second logic is configured to: define the signal receiving functions of the plurality of second interfaces to match the plurality of first interfaces. In some embodiments of the second aspect, the first logic is configured to: define the signal sending functions of the plurality of first interfaces to match the plurality of second interfaces.
[0018] In some embodiments of the second aspect, the plurality of second cabinets are further configured to output a plurality of second signals through the plurality of second interfaces, and the plurality of first cabinets are further configured to receive the plurality of second signals through the plurality of first interfaces.
[0019] In some embodiments of the second aspect, the second logic is configured to: define the signal transmission functions of the plurality of second interfaces to match the plurality of first interfaces. In some embodiments of the second aspect, the first logic is configured to: define the signal reception functions of the plurality of first interfaces to match the plurality of second interfaces.
[0020] In this way, by using one or more logics in each cabinet system to define the signal transmission functions of the interfaces, the same signals in different parts of the circuit under test emulated by each cabinet system can be interconnected, so as to realize the co-simulation of multiple parts of the circuit under test.
[0021] In some embodiments of the second aspect, running the first compiled version of the circuit under test includes: controlling the simulation period of the first compiled version of the circuit under test to be the same as that of the second compiled version of the circuit under test. In some embodiments of the second aspect, running the second compiled version of the circuit under test includes: controlling the simulation period of the first compiled version of the circuit under test to be the same as that of the second compiled version of the circuit under test. In this way, the transmission frequencies of the same signals in each cabinet system in the hardware emulator cluster can be matched, so as to better realize the co-simulation of multiple parts of the circuit under test.
[0022] In some embodiments of the second aspect, the method further includes: managing at least one of the hardware simulation resources or clock resources of the first cabinet system and the second cabinet system by a management unit in the hardware emulator cluster.
[0023] In some embodiments of the second aspect, the method further includes: the first host in the first cabinet system performs the following: initiating a snapshot save operation to save the snapshot data of the hardware emulator cluster; and initiating a snapshot restore operation to restore the hardware emulator cluster by using the snapshot data of the hardware emulator cluster.
[0024] In some embodiments of the second aspect, the method further includes: the first host in the first cabinet system performs the following: initiating a waveform sampling operation based on a trigger condition to sample the simulation waveform from the hardware emulator cluster. In some embodiments of the second aspect, the method further includes: before initiating the waveform sampling operation, the first host system initiates a simulation pause operation based on the trigger condition to pause the simulation process of the hardware emulator cluster.
[0025] In this way, trigger detection, signal sampling, and / or waveform analysis can be performed on the entire hardware emulator cluster, so that the simulation of the circuit under test running in the hardware emulator cluster can be conveniently debugged.
[0026] In a third aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method provided in the second aspect.
[0027] In a fourth aspect of the present disclosure, there is provided a computer program product including computer-executable instructions that, when executed by a processor, implement some or all of the steps of the method in the second aspect.
[0028] It can be understood that the computer storage medium in the third aspect or the computer program product in the fourth aspect provided above are both used to execute the method provided in the second aspect. Therefore, the explanations or descriptions regarding the second aspect also apply to the third aspect and the fourth aspect. In addition, the beneficial effects achievable by the third aspect and the fourth aspect can refer to the beneficial effects in the corresponding method, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0030] Figure 1 A flowchart showing the design and manufacturing process of a chip is shown;
[0031] Figure 2 A schematic diagram showing an example hardware emulator cluster for chip hardware emulation according to some embodiments of the present disclosure is shown;
[0032] Figure 3 A schematic diagram showing the data transmission of an example hardware emulator cluster according to some embodiments of the present disclosure is shown;
[0033] Figure 4 A schematic diagram showing an example hardware emulator cluster for clock synchronization according to some embodiments of the present disclosure is shown;
[0034] Figure 5 A schematic diagram showing an example hardware emulator cluster with clock homology according to some embodiments of the present disclosure is shown;
[0035] Figure 6 A flowchart showing an example method for chip hardware emulation according to some embodiments of the present disclosure; and
[0036] Figure 7 A block diagram showing a computing device capable of implementing multiple embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0037] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0038] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0039] As briefly mentioned above, there is a need for a hardware simulation solution that supports a very large simulation capacity. Currently, some related hardware simulation solutions have been proposed. In some solutions, the simulation capacity supported by a single chip in a hardware emulator can be increased by using more advanced chip processes. In other solutions, the simulation capacity supported by a single cabinet system can be increased by adding more cabinets (racks, also known as frames) in a single cabinet system.
[0040] However, the simulation capacity supported by these hardware simulation solutions is still difficult to meet the requirements of an increasingly large chip scale. The number of cabinets in a single cabinet system is limited by factors such as latency, so the total capacity supported by a single cabinet system is limited. In addition, as the chip scale increases, the difficulty of compiling a very large-scale chip design into an executable compilation version on a hardware emulator also increases significantly.
[0041] To at least partially address the above problems and other potential problems, various embodiments of the present disclosure provide a hardware emulator cluster for chip hardware emulation, as well as corresponding methods, media, and program products. The hardware emulator cluster includes: a first cabinet system, which includes a plurality of interconnected first cabinets (racks) and first logic. The plurality of first cabinets are configured to run a first compiled version of a circuit under test to emulate a first part of the circuit under test, and output a plurality of first signals through a plurality of first interfaces. The hardware emulator cluster further includes: a second cabinet system, which includes a plurality of interconnected second cabinets and second logic. The plurality of second cabinets are configured to run a second compiled version of the circuit under test to emulate a second part of the circuit under test, and receive the plurality of first signals through a plurality of second interfaces. Wherein the first logic is connected to the second logic, and at least one of the first logic or the second logic is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
[0042] In this way, by utilizing one or more logics for matching signal transmission functions in each cabinet system, joint emulation between multiple cabinet systems can be achieved, thereby enabling a larger emulation capacity than a single cabinet system. For example, assuming that the maximum emulation capacity supported by a single cabinet system is x hundred million gates, then by using the hardware emulator cluster of the present disclosure including n cabinet systems, a maximum emulation capacity of n×x hundred million gates can be supported, thus significantly enhancing the supported emulation capacity. In addition, by dividing the circuit under test into multiple parts and separately compiling the multiple parts to generate corresponding compiled versions, the compilation difficulty for each sub-circuit can be reduced and the size of a single compiled version can be decreased.
[0043] Various example embodiments of the present disclosure will be described below with reference to the accompanying drawings. Figure 1The flowchart of the chip design and manufacturing process 100 is shown. The design and manufacturing process 100 begins with specification development 110. In the stage of specification development 110, the requirements for the functions and performance that the integrated circuit needs to achieve are determined. In the stage of chip design 120, circuit design is carried out with the help of EDA tools to obtain, for example, a layout file for chip manufacturing. Based on the differences in circuits (such as digital circuits or analog circuits), the design 120 may include different design steps. In the stage of manufacturing 140, integrated circuits are formed on the wafer through processes such as lithography, etching, ion implantation, thin film deposition, and polishing. In the stage of packaging 150, the wafer is cut to obtain die, and the die are packaged through processes such as bonding, soldering, and molding to obtain chips. The obtained chips are tested in the stage of testing 160 to ensure that the performance of the finished chips meets the requirements determined in the specification development 110. The chips 170 that pass the test can be delivered to customers. It can be understood that the above process is only illustrative and does not limit the scope of the present disclosure. In some cases, the chip design and manufacturing process may be different. For example, tape-out may be performed before manufacturing 140. A small number of chips obtained from tape-out can be used for testing to verify whether the chip design meets the expectations. If the expectations are not met, this indicates that the tape-out fails, and it may be necessary to adjust the chip design or redesign the chip.
[0044] In some embodiments, the design 120 of a digital circuit may exemplarily include architecture design 121, register transfer level (RTL) design 123, functional simulation 125, synthesis 127, timing analysis 129, design for test (DFT) 131, verification check 133, placement and routing 135, design rule check (DRC) 137, and layout generation 139. The architecture design 121 includes, for example, designing the architecture of a chip. For example, EDA tools can be used to determine the types and quantities of components or sub-circuits included in a chip system, as well as the functions, connections, and interactions of each component or sub-circuit. In the stage of RTL design 123, the determined chip architecture can be described in code at the RTL level using a hardware programming language such as Verilog or VHDL. The functional simulation 125 is also referred to as RTL-level behavioral simulation or front-end simulation. The purpose of functional simulation is to analyze the correctness of the logic relationships of the designed circuit. Synthesis 127 can convert RTL into a gate-level netlist. Synthesis 127 can include, for example, translation, optimization, and mapping. In one embodiment, the EDA tool for synthesis can first convert the RTL code into a general Boolean equation and compile it. The netlist can be optimized according to constraints such as delay and area imposed by the designer, and then the RTL netlist can be mapped to a process library to generate a gate-level netlist.
[0045] Timing analysis 129 is usually static timing analysis, which mainly involves the timing calculation and prediction of digital circuits. By performing timing analysis on the paths in the digital circuit, it is determined whether timing convergence is achieved, thereby ensuring whether the timing of various circuits meets various timing requirements. This verification of digital circuits is usually done statically and does not require the simulation of digital logic. In the DFT 131 stage, various hardware logics for improving the testability of the chip (including controllability and observability) can be embedded in the design. By using this part of the logic, test vectors can be generated to achieve the purpose of testing large-scale digital circuits. DFT can include, for example, test methods based on scan chains or built-in self-test circuits (BIST). In the verification check 133 stage, formal verification and / or equivalence checking can be performed on the circuit. Formal verification can use mathematical methods to prove its correctness or incorrectness according to a certain or certain formal specifications or properties. Formal verification can include, for example, abstract interpretation, formal model checking (also known as property checking), and theorem proving. Equivalence checking can be used to verify the consistency between the register transfer level design and the gate-level netlist, and between the gate-level netlists.
[0046] In the placement and routing 135 stage, the chip circuit can be placed and routed. Placement can reasonably arrange the gate-level netlist generated by logic synthesis 127 in a rectangular area corresponding to the chip based on considerations such as area, critical path delay length, and power consumption. After that, the components or sub-circuits that have been placed can be routed to connect them. Routing generally expects the total wire length to be short, the wire delay to meet the timing requirements, and to comply with the wire rules in the process (such as wire density). Although placement and routing are described separately here, this is only illustrative and does not limit the scope of the present disclosure. In some cases, placement and routing can be performed simultaneously or alternately to achieve the optimization of placement and routing.
[0047] In the DRC 137 stage, it can be checked whether there are potential open circuits, short circuits, or adverse effects caused by violations of the design rules in the layout. After passing the DRC, a file representing the layout, such as a GDSII file, can be generated by EDA tools. It can be understood that the above links are only exemplary and do not limit the scope of the present disclosure. In the actual design process, the above links can be added, deleted, or modified according to the design needs. In addition, some of the above links can be implemented by different EDA tools or can be integrated in one or more EDA tools. The scope of the present disclosure is not limited here.
[0048] In one or more of the above stages, the hardware simulation solution of the present disclosure can be applied to verify the chip design. In some examples, after the RTL design 123, the RTL code can be compiled into a compiled version that can be executed on the dedicated circuits and logic of the hardware emulator cluster, so as to verify the chip design by performing RTL functional simulation executed by the hardware emulator cluster. In some examples, after the synthesis 127 stage, the netlist can be mapped into a compiled version that can be executed on the hardware emulator cluster, so as to verify the chip design by applying hardware simulation to the netlist. In some examples, after the placement and routing 135, factors such as component delay and routing delay can be considered to perform hardware simulation on the netlist to further verify the chip design.
[0049] The hardware simulation solution of the present disclosure can be applied to a variety of hardware simulation scenarios, and the scope of the present disclosure is not limited herein. In some embodiments, a test bench can be established to provide control over the input and environment of the circuit under test (also referred to as the design under test, DUT). For example, stimuli can be injected into the DUT. By comparing the output of the DUT with the expected target, the circuit under test can be verified. In some examples, the RTL code describing the circuit under test can be compiled into a compiled version that can be executed on the hardware emulator cluster, such as a bitstream file, through processes such as synthesis, partitioning, and / or placement and routing. The hardware emulator cluster can load the compiled version to simulate the operation of the circuit under test. The hardware emulator cluster can output signals for observing the internal state of the circuit under test. By sampling the signals of the hardware emulator cluster, performing trigger detection, and / or performing waveform analysis, the operation of the circuit under test can be observed and debugged.
[0050] Figure 2 FIG. shows a schematic diagram of an example hardware emulator cluster 200 for chip hardware simulation according to some embodiments of the present disclosure. As Figure 2 shown, the hardware emulator cluster 200 includes a first cabinet system 210 and a second cabinet system 220. The first cabinet system 210 includes a plurality of first cabinets 251 and a first logic 252 that are interconnected. The plurality of first cabinets 251 are configured to run a first compiled version of the circuit under test to simulate a first part of the circuit under test, and the plurality of first cabinets 251 output a plurality of first signals through a plurality of first interfaces.
[0051] The second cabinet system 220 includes a plurality of interconnected second cabinets 261 and a second logic 262. The plurality of second cabinets 261 are configured to run a second compiled version of the circuit under test for simulating a second part of the circuit under test, and the plurality of second cabinets 261 receive the plurality of first signals through a plurality of second interfaces. The first logic 252 is connected to the second logic 262, and at least one of the first logic 252 or the second logic 262 is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
[0052] In other words, when performing hardware simulation on a circuit under test using the hardware emulator cluster 200, the circuit under test can be divided into multiple parts, and the multiple parts can be respectively compiled into corresponding compiled versions. The corresponding compiled versions can be respectively loaded by the plurality of cabinet systems in the hardware emulator cluster 200 to simulate the corresponding parts of the circuit under test. In addition, in the joint simulation of the plurality of cabinet systems, the connection of the same signals between the multiple parts of the circuit under test is implemented by one or more logics in the plurality of cabinet systems. One or more logics are configured to match the signal transmission functions of the interfaces in each cabinet system so that the same signals can be interconnected when each cabinet system executes the corresponding compiled version. Specifically, in each cabinet system, the logic communicates with the corresponding plurality of cabinets and is configured to process the signals output by the corresponding plurality of cabinets through the interfaces and / or the signals received from the signals received through the interfaces, so that the signal transmission functions of the interfaces in each cabinet system match.
[0053] In some embodiments, the plurality of first cabinets 251 output a plurality of first signals from the plurality of first interfaces. The first logic 252 can receive the plurality of first signals and output the plurality of first signals. The second logic 262 can be configured to define the signal reception functions of the plurality of second interfaces in the second cabinet system 220 to match the plurality of first interfaces in the first cabinet system 210. For example, the second logic 262 can define the signal reception functions of the plurality of second interfaces by reordering the received signals.
[0054] For example, when the plurality of first cabinets 521 run the first compiled version of the circuit under test, the plurality of first cabinets 251 respectively output signals a, b, c, and d through interfaces A1, A2, A3, and A4. The second logic 262 can be configured to define the signal reception functions of the interfaces B1, B2, B3, and B4 of the plurality of second cabinets 261 so that the interfaces B1, B2, B3, and B4 respectively receive signals a, b, c, and d, so that the outputs corresponding to the interfaces A1, A2, A3, and A4 in the first part of the circuit under test are respectively connected to the inputs corresponding to the interfaces B1, B2, B3, and B4 in the second part of the circuit under test.
[0055] Additionally or alternatively, the first logic 252 may be configured to define the signal transmission functions of the plurality of first interfaces to match the plurality of second interfaces in the second cabinet system 220. For example, the first logic 252 may define the signal transmission functions of the plurality of first interfaces such that interfaces A1, A2, A3, and A4 output signals a, b, c, and d respectively. The second logic 262 may receive signals a, b, c, and d and input signals a, b, c, and d to interfaces B1, B2, B3, and B4 of the plurality of second cabinets 261 respectively, so that the outputs corresponding to interfaces A1, A2, A3, and A4 in the first part of the circuit under test are connected to the inputs corresponding to interfaces B1, B2, B3, and B4 in the second part of the circuit under test.
[0056] In some embodiments, the first logic 252 or the second logic 262 may be configured to define the signal transmission functions of the interfaces to achieve signal transmission matching between the plurality of first interfaces in the first cabinet system 210 and the plurality of second interfaces in the second cabinet system 220. In other embodiments, both the first logic 252 and the second logic 262 may be configured to define the signal transmission functions of the interfaces, thereby achieving signal transmission matching between the plurality of first interfaces in the first cabinet system 210 and the plurality of second interfaces in the second cabinet system 220.
[0057] In some embodiments, the plurality of first cabinets 251 may also receive a plurality of second signals through a plurality of third interfaces, and the plurality of second cabinets 261 may output a plurality of second signals through a plurality of fourth interfaces. The plurality of third interfaces in the first cabinet system 210 may be the same as or different from the plurality of first interfaces. The plurality of fourth interfaces in the second cabinet system 220 may be the same as or different from the plurality of second interfaces.
[0058] In some embodiments, the second logic 262 may receive a plurality of second signals from the plurality of second cabinets 261 and output the plurality of second signals. The first logic 252 may be configured to define the signal reception functions of the plurality of third interfaces of the plurality of first cabinets 251 to match the plurality of fourth interfaces in the second cabinet system 220.
[0059] For example, when multiple second cabinets 261 are running the second compiled version of the circuit under test, the multiple second cabinets 261 respectively output signal e, signal f, signal g, and signal h through interfaces B1, B2, B3, and B4. The first logic 252 can be configured to define the signal receiving functions of interfaces A1, A2, A3, and A4 of the multiple first cabinets 251, so that interfaces A1, A2, A3, and A4 respectively receive signal e, signal f, signal g, and signal h, thereby enabling the outputs corresponding to interfaces B1, B2, B3, and B4 in the second part of the circuit under test to be respectively connected to the inputs corresponding to interfaces A1, A2, A3, and A4 in the first part of the circuit under test.
[0060] Additionally or alternatively, the second logic 262 can be configured to define the signal sending functions of the multiple fourth interfaces to match the multiple third interfaces in the first cabinet system 210. For example, the second logic 262 can define the signal sending functions of the multiple fourth interfaces, so that the multiple second cabinets 261 respectively output signal e, signal f, signal g, and signal h through interfaces B1, B2, B3, and B4. The first logic 252 can receive signal e, signal f, signal g, and signal h and respectively input signal e, signal f, signal g, and signal h to interfaces A1, A2, A3, and A4 of the multiple first cabinets 251, thereby enabling the outputs corresponding to interfaces B1, B2, B3, and B4 in the second part of the circuit under test to be respectively connected to the inputs corresponding to interfaces A1, A2, A3, and A4 in the first part of the circuit under test.
[0061] In this way, by using one or more logics in each cabinet system, the same signals in different parts of the circuit under test respectively simulated by the first cabinet system 210 and the second cabinet system 220 can be interconnected. Therefore, by using the hardware emulator cluster of this solution, multiple compiled versions of the circuit under test can be respectively loaded in multiple cabinet systems to implement the joint simulation of multiple parts of the circuit under test.
[0062] Additionally, in some embodiments, the simulation cycle of the first compiled version running in the first cabinet system 210 is the same as the simulation cycle of the second compiled version running in the second cabinet system 220 to implement the synchronous simulation of multiple parts of the circuit under test. Herein, the term "simulation cycle of the compiled version" refers to the depth of the critical path in the circuit, that is, the number of levels of combinational logic between two registers in the critical path. The shorter the simulation cycle of the compiled version, the better the frequency performance of the compiled version.
[0063] In some embodiments, when the first compilation version of the circuit under test is running in the first cabinet system 210, multiple first cabinets 251 can make the simulation cycle of the first compilation version the same as that of the second compilation version running in the second cabinet system 220. Alternatively or additionally, when the second compilation version of the circuit under test is running in the second cabinet system 220, multiple second cabinets 261 can make the simulation cycle of the second compilation version the same as that of the first compilation version.
[0064] In other words, the simulation cycle of the corresponding compilation version can be controlled by one or more of the first cabinet system 210 and the second cabinet system 220, so that the simulation cycles of different compilation versions running in the first cabinet system 210 and the second cabinet system 220 are the same. In some embodiments, when loading the compilation version, the compilation version with higher frequency performance can be downclocked so that the simulation cycles of the two compilation versions running in the two cabinet systems are the same. For example, no-operation instructions can be added to the compilation version with higher frequency performance to reduce its frequency performance.
[0065] In this way, by making the simulation cycles of different compilation versions running in the first cabinet system 210 and the second cabinet system 220 the same, the transmission frequencies of the same signals in the first cabinet system 210 and the second cabinet system 220 can be matched, so as to better realize the co-simulation of multiple parts of the circuit under test.
[0066] It should be understood that Figure 2 the hardware emulator cluster 200 shown in
[0067] Figure 3 is only exemplary and does not limit the scope of the present disclosure. The number of cabinet systems in the hardware emulator cluster and the number of cabinets in a single cabinet system can be set according to specific application scenarios. The hardware emulator cluster can also include any other suitable components. Figure 3 FIG. shows a schematic diagram of data transmission of an example hardware emulator cluster 300 according to some embodiments of the present disclosure. As Figure 3In the illustrated example, the hardware emulator parallel machine A may further include a cluster physical top layer for managing the hardware emulation resources and / or clock resources of each cabinet system in the hardware emulator cluster 300.
[0068] In some embodiments, data transmission between the hardware emulator parallel machine A and the hardware emulator parallel machine B may be through optical fibers and driver boards. The driver board can enhance the driving ability of signal transmission between the parallel machines. Although not shown, the hardware emulator parallel machine A further includes Logic A, the hardware emulator parallel machine B further includes Logic B, and data transmission between the hardware emulator parallel machine A and the hardware emulator parallel machine B is via Logic A and Logic B.
[0069] As Figure 3 shown, in some embodiments, the hardware emulator cluster 300 may further include a management unit. The management unit may be connected to the hardware emulator parallel machines A and B. The management unit may be configured to manage at least one of the hardware emulation resources or clock resources of the hardware emulator parallel machines A and B. The cluster physical top layer in the hardware emulator parallel machine A may communicate with the management unit to manage each hardware emulator parallel machine in the hardware emulator cluster.
[0070] In some embodiments, the hardware emulator cluster 300 may further include a host system A connected to the hardware emulator parallel machine A and a host system B connected to the hardware emulator parallel machine B. The host system A (or B) is configured to communicate with the hardware emulator parallel machine A (or B) for instruction loading and read / write communication of snapshot data, trace data communication, etc. during the hardware emulation debug phase. In some embodiments, the host system may be a component in the cabinet system. The hardware emulator parallel machine A and the host system A may be included in the cabinet system A, and the hardware emulator parallel machine B and the host system B may be included in the cabinet system B. The cabinet system A may be Figure 2 an example of the first cabinet system 210 shown in Figure 2 and the cabinet system B may be an example of the second cabinet system 220 shown in. Alternatively, the host system may be an external component coupled to the cabinet system.
[0071] In some embodiments, one or more host systems in the hardware emulator cluster 300 may initiate a snapshot save operation to save the snapshot data of the hardware emulator cluster 300. In some embodiments, the snapshot data of the hardware emulator cluster 300 may include the snapshot data of each parallel machine and the N-shot data cached in Logic A and Logic B. Additionally, one or more host systems in the hardware emulator cluster 300 may initiate a snapshot recovery operation to recover the hardware emulator cluster 300 using the snapshot data of the hardware emulator cluster 300.
[0072] For example, the host system A can initiate a snapshot save operation to the hardware emulator parallel unit A to save the snapshot data of the hardware emulator cluster 300. In this case, the hardware emulator parallel unit A is the master node of the cluster, and the other parallel units in the cluster are slave nodes. The hardware emulator parallel unit A can send a request to the management unit to perform a snapshot save operation on the hardware emulator cluster 300. The management unit can send an instruction to the host system B to initiate a snapshot save operation on the hardware emulator parallel unit B based on this request. The management unit can maintain the cache data and its corresponding relationship with the hardware simulation resources of each parallel unit.
[0073] The host system A can initiate a snapshot restore operation to the hardware emulator parallel unit A to restore the hardware emulator cluster 300 using the snapshot data of the hardware emulator cluster 300. In this case, the hardware emulator parallel unit A is the master node of the cluster, and the other parallel units in the cluster are slave nodes. The hardware emulator parallel unit A can read the cache data and its corresponding relationship with the hardware simulation resources of each parallel unit from the management unit. The hardware emulator parallel unit A can send instructions to each parallel unit in the hardware emulator cluster 300 to perform a cluster snapshot restore. When restoring the hardware emulator cluster 300, not only the data in each parallel unit is restored, but also the N-shot data cached in Logic A and Logic B is restored.
[0074] In some embodiments, one or more host systems in the hardware emulator cluster 300 can initiate a waveform sampling operation based on a trigger condition to sample the simulation waveform from the hardware emulator cluster 300. In some embodiments, one or more host systems in the hardware emulator cluster 300 are also configured to initiate a simulation pause operation based on a trigger condition to pause the simulation process of the hardware emulator cluster 300. The trigger condition can be a failure or any other suitable condition. For example, the host system A can send a request to the management unit to perform a waveform sampling operation on the hardware emulator cluster 300 based on the trigger condition. The management unit can send instructions to each host system in the hardware emulator cluster 300 to perform waveform sampling based on this request.
[0075] In some embodiments, when the trigger condition is met, the parallel unit in the hardware emulator cluster 300 can only trigger the signals within the parallel unit. When a trigger event occurs in the simulation unit of the parallel unit, the parallel unit can report it to the synchronization unit on the top layer of the cluster physics. The synchronization unit on the top layer of the cluster physics can trigger the simulation pause of the entire cluster. The synchronization unit on the top layer of the cluster physics can also synchronously start the waveform sampling of the simulation unit.
[0076] Alternatively, when the trigger condition is satisfied, the parallel machines in the hardware emulator cluster 300 may initiate waveform sampling for the entire cluster. The cluster topology can be obtained from the management unit. This parallel machine is the business master node of the cluster, and the other parallel machines in the cluster are business slave nodes. The business master node of the cluster may issue an instruction to start waveform sampling of the cluster to each parallel machine in the cluster.
[0077] In this way, trigger detection, signal sampling, and / or waveform analysis can be performed on the entire hardware emulator cluster, so that the simulation of the circuit under test running in the hardware emulator cluster can be conveniently debugged. It should be understood that the solution of the present disclosure also supports performing trigger detection, signal sampling, and / or waveform analysis on a single cabinet system using conventional techniques. Therefore, by using the debugging function for the entire hardware emulator cluster or a single cabinet system, the flexibility of switching between single die and multi-die in the verification environment can be improved.
[0078] Figure 4 A schematic diagram of an exemplary hardware emulator cluster 400 for clock synchronization according to some embodiments of the present disclosure is shown. As Figure 4 shown, a cluster physical top node can be added to any one of the parallel machines, and the cluster physical top node distributes clocks to each parallel machine in the cluster to implement a hardware emulator cluster with clock synchronization. In the synchronous cluster scenario, the versions within each parallel machine can be compiled separately. Taking two parallel machines as an example, assume that the frequency performance of the compiled version corresponding to parallel machine A is perf1 (in MHz), and the frequency performance of the compiled version corresponding to parallel machine B is perf2 (in MHz), and perf1 > perf2. When loading the compiled version in the synchronous cluster scenario, the compiled versions within each parallel machine can be loaded separately. For the clock tree refresh process, a cluster load can be initiated on any one of the parallel machines, a request to obtain the cluster topology can be sent to the management unit, and the top of the clock tree of each parallel machine in the cluster can be refreshed to the cluster physical top layer. The clocks of each parallel machine in the cluster are distributed by the cluster physical top layer to achieve a synchronous cluster. In addition, version downscaling can be performed. For example, the frequency performance perf1 (MHz) of the compiled version corresponding to parallel machine A can be downscaled to perf2 (MHz). At this time, the frequency performance of the compiled versions loaded by each parallel machine in the cluster is perf2 (MHz).
[0079] In the process of performing hardware simulation using the hardware emulator cluster 400, the user of the hardware emulator cluster 400 can divide a very large-scale DUT into n large-scale DUTs according to die or subchip, denoted as DUT_0, DUT_1... DUT_n-1 respectively, where n is greater than or equal to 2. The compilation software can be used to compile DUT_0, DUT_1... DUT_n-1 respectively to obtain n compilation versions, denoted as db_0, db_1... db_n-1 respectively, and the frequency performance corresponding to each compilation version is perf_0 (MHz), perf_1 (MHz)... perf_n-1 (MHz). Assume that perf_0 > perf_1 >... > perf_n-1. Then, the corresponding compilation versions can be loaded on each parallel machine respectively, that is, load the compilation version db_0 corresponding to DUT_0 on parallel machine 0, load the compilation version db_1 corresponding to DUT_1 on parallel machine 1,..., load the compilation version db_n-1 corresponding to DUT_n-1 on parallel machine n-1, and at this time, the frequency performance of the compilation versions loaded on each parallel machine is perf_n-1. The clock synchronization unit at the physical top layer of the cluster can distribute clocks to each parallel machine in the cluster to achieve synchronous clustering. The user can start the hardware simulation executed in the synchronous cluster on any one of the parallel machines in the cluster. The user can start the debug executed in the synchronous cluster. The user can stop the debug executed in the synchronous cluster. The user can stop the hardware simulation executed on the synchronous cluster on the parallel machine where the simulation is started. After the simulation ends, each parallel machine can unload the compilation version in the parallel machine respectively.
[0080] Figure 5 FIG. shows a schematic diagram of an example hardware emulator cluster 500 with clock homology according to some embodiments of the present disclosure. As Figure 5As shown, a cluster physical top-level node can be added to any one of the parallel machines, and the cluster physical top-level node distributes the physical reference clock to each parallel machine within the cluster. Each parallel machine uses its own parallel machine physical top-level node to distribute the clock within the parallel machine, so as to implement a hardware emulator cluster with clock synchronization. In an asynchronous cluster scenario, the versions within each parallel machine can be compiled separately. Taking two parallel machines as an example, assume that the frequency performance of the compiled version corresponding to parallel machine A is perf1 (unit: MHz), and the frequency performance of the compiled version corresponding to parallel machine B is perf2 (unit: MHz), and perf1 > perf2. When loading the compiled version in an asynchronous cluster scenario, the compiled versions within each parallel machine can be loaded separately. The physical reference clocks of each parallel machine within the cluster are distributed by the cluster physical top-level, and other clocks are still distributed by the physical top-levels of each parallel machine. In addition, version downscaling can be performed. For example, the frequency performance perf1 (MHz) of the compiled version corresponding to parallel machine A is downscaled to perf2 (MHz). At this time, the frequency performance of the compiled versions loaded by each parallel machine within the cluster is all perf2 (MHz).
[0081] In the process of performing hardware emulation using the hardware emulator cluster 500, similar to the hardware emulation process described in the reference Figure 4 The processes such as DUT partitioning, compilation, and loading can be performed, and the specific details will not be elaborated here. The difference from the hardware emulation process described in the reference Figure 4 is that the physical top-levels of each parallel machine can distribute their respective clocks, and the cluster physical top-level only distributes the reference clocks of each parallel machine to achieve asynchronous cluster emulation with clock synchronization. The user can start the hardware emulation executed in the synchronous cluster on any one of the parallel machines within the cluster. The user can start the debug executed in the synchronous cluster. The user can stop the debug executed in the synchronous cluster. The user can stop the hardware emulation executed in the synchronous cluster on the parallel machine where the emulation is started. After the emulation ends, each parallel machine can unload the compiled version within the parallel machine separately.
[0082] Figure 6 FIG. shows a flowchart of an example method 600 for chip hardware emulation according to some embodiments of the present disclosure. Process 600 can be implemented by any suitable computing unit. For example, it can be executed by a computer or other electronic devices with computing or circuit design capabilities. Specifically, for example, the EDA tool can be implemented by a processor of a computer according to the data and / or instructions stored in the memory to execute process 600.
[0083] In block 610, the processor uses the first cabinet system in the hardware emulator cluster to emulate the first part of the circuit under test. The first cabinet system includes a plurality of interconnected first cabinets and first logic. The plurality of first cabinets are configured to run the first compiled version of the circuit under test and output a plurality of first signals through a plurality of first interfaces.
[0084] At block 620, the processor uses a second cabinet system in the hardware emulator cluster to emulate a second portion of the circuit under test, the second cabinet system including a plurality of interconnected second cabinets and second logic, the plurality of second cabinets being configured to run a second compiled version of the circuit under test and receive the plurality of first signals through a plurality of second interfaces, wherein the first logic is connected to the second logic, and at least one of the first logic or the second logic is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
[0085] In some embodiments, the second logic is configured to: define the signal reception functions of the plurality of second interfaces to match the plurality of first interfaces. In some embodiments, the first logic is configured to: define the signal transmission functions of the plurality of first interfaces to match the plurality of second interfaces.
[0086] In some embodiments, the plurality of second cabinets are further configured to output a plurality of second signals through the plurality of second interfaces, and the plurality of first cabinets are further configured to receive the plurality of second signals through the plurality of first interfaces. In some embodiments, the second logic is configured to: define the signal transmission functions of the plurality of second interfaces to match the plurality of first interfaces. In some embodiments, the first logic is configured to: define the signal reception functions of the plurality of first interfaces to match the plurality of second interfaces.
[0087] In some embodiments, running the first compiled version of the circuit under test includes: controlling the simulation period of the first compiled version of the circuit under test to be the same as the simulation period of the second compiled version of the circuit under test. In some embodiments, running the second compiled version of the circuit under test includes: controlling the simulation period of the first compiled version of the circuit under test to be the same as the simulation period of the second compiled version of the circuit under test.
[0088] In some embodiments, the method further includes: managing at least one of the hardware simulation resources or clock resources of the first cabinet system and the second cabinet system by a management unit in the hardware emulator cluster.
[0089] In some embodiments, the method further includes: a first host in the first cabinet system performing the following: initiating a snapshot save operation to save snapshot data of the hardware emulator cluster; and initiating a snapshot restore operation to restore the hardware emulator cluster using the snapshot data of the hardware emulator cluster.
[0090] In some embodiments, the method further includes: a first host in the first cabinet system performs the following: initiating a waveform sampling operation based on a trigger condition to sample simulation waveforms from the hardware emulator cluster. In some embodiments, the method further includes: before initiating the waveform sampling operation, a first host system initiates a simulation pause operation based on the trigger condition to pause the simulation process of the hardware emulator cluster.
[0091] Figure 7 FIG. shows a schematic block diagram of an example device 700 that may be used to implement embodiments of the present disclosure. The device 700 may be a computing device running EDA software, and a user may initiate a hardware simulation operation in the EDA software to perform hardware simulation using a hardware emulator cluster communicating with the computing device. As shown, the device 700 includes a computing unit 701 that may execute various appropriate actions and processes according to computer program instructions stored in a random access memory (RAM) 703 and / or a read-only memory (ROM) 702, or computer program instructions loaded from a storage unit 708 into the RAM 703 and / or the ROM 702. In the RAM 703 and / or the ROM 702, various programs and data required for the operation of the device 700 may also be stored. The computing unit 701 and the RAM 703 and / or the ROM 702 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0092] Multiple components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0093] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as process 600. For example, in some embodiments, process 600 may be implemented as a computer software program, specifically an EDA program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the RAM and / or ROM and / or the communication unit 709. When the computer program is loaded into the RAM and / or ROM and executed by the computing unit 701, one or more steps of the process 600 described above may be executed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute process 600 in any other suitable manner (e.g., by means of firmware).
[0094] In the above embodiments, the method flow can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a server or a terminal, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial optical cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by the server or the terminal, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, and a magnetic tape, etc.), an optical medium (such as a digital video disk (DVD), etc.), or a semiconductor medium (such as a solid-state drive, etc.).
[0095] In addition, although the operations are depicted in a particular order, this should be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.
[0096] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A hardware emulator cluster, characterized in that, Comprising: A first cabinet system, including a plurality of interconnected first cabinets (racks) and first logic, the plurality of first cabinets being configured to run a first compiled version of a circuit under test to simulate a first part of the circuit under test and output a plurality of first signals through a plurality of first interfaces; And A second cabinet system, including a plurality of interconnected second cabinets and second logic, the plurality of second cabinets being configured to run a second compiled version of the circuit under test to simulate a second part of the circuit under test and receive the plurality of first signals through a plurality of second interfaces; Wherein the first logic is connected to the second logic, and at least one of the first logic or the second logic is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
2. The hardware emulator cluster according to claim 1, characterized in that, The second logic is configured to: Define the signal reception functions of the plurality of second interfaces to match the plurality of first interfaces.
3. The hardware emulator cluster according to claim 1 or 2, characterized in that, The first logic is configured to: Define the signal transmission functions of the plurality of first interfaces to match the plurality of second interfaces.
4. The hardware emulator cluster according to any one of claims 1 to 3, characterized in that The plurality of second cabinets are further configured to output a plurality of second signals through the plurality of second interfaces, and the plurality of first cabinets are further configured to receive the plurality of second signals through the plurality of first interfaces.
5. The hardware emulator cluster according to claim 4, wherein The second logic is configured to: Define the signal transmission functions of the plurality of second interfaces to match the plurality of first interfaces.
6. The hardware emulator cluster according to claim 4 or 5, characterized in that The first logic is configured to: Define the signal reception functions of the plurality of first interfaces to match the plurality of second interfaces.
7. The hardware emulator cluster according to any one of claims 1 to 6, characterized in that, When running the first compiled version of the circuit under test, the plurality of first cabinets are used to: control the simulation period of the first compiled version to be the same as the simulation period of the second compiled version.
8. The hardware emulator cluster according to any one of claims 1 to 7, characterized in that, When running the second compiled version of the circuit under test, the plurality of second cabinets are used to: control the simulation period of the second compiled version to be the same as the simulation period of the first compiled version.
9. The hardware emulator cluster according to any one of claims 1 to 8, characterized in that Further comprising: A management unit, configured to manage at least one of the hardware simulation resources or clock resources of the first cabinet system and the second cabinet system.
10. The hardware emulator cluster according to any one of claims 1 to 9, characterized in that, The first cabinet system further includes a first host, the first host being configured to: Initiate a snapshot saving operation to save the snapshot data of the hardware emulator cluster; And Initiate a snapshot recovery operation to recover the hardware emulator cluster using the snapshot data of the hardware emulator cluster.
11. The hardware emulator cluster according to any one of claims 1 to 10, characterized in that, The first cabinet system further includes a first host system, the first host system being configured to: Initiate a waveform sampling operation based on a trigger condition to sample simulation waveforms from the hardware emulator cluster.
12. The hardware emulator cluster according to claim 11, wherein Before initiating the waveform sampling operation, the first host system is configured to: Initiate a simulation pause operation based on the trigger condition to pause the simulation process of the hardware emulator cluster.
13. A method for performing simulation by using a cluster of hardware emulators, characterized in that, Comprising: Using the first cabinet system in the hardware emulator cluster to simulate a first part of a circuit under test, the first cabinet system including a plurality of interconnected first cabinets and first logic, the plurality of first cabinets being configured to run the first compiled version of the circuit under test and output a plurality of first signals through a plurality of first interfaces; And Use the second cabinet system in the hardware emulator cluster to simulate the second part of the circuit under test. The second cabinet system includes a plurality of interconnected second cabinets and second logic. The plurality of second cabinets are configured to run the second compiled version of the circuit under test and receive the plurality of first signals through a plurality of second interfaces. Wherein the first logic is connected to the second logic, and at least one of the first logic or the second logic is configured to match the signal transmission functions of the plurality of first interfaces and the plurality of second interfaces.
14. The method according to claim 13, wherein The second logic is configured to: Define the signal reception functions of the plurality of second interfaces to match the plurality of first interfaces.
15. The method according to claim 13 or 14, characterized in that, The first logic is configured to: Define the signal transmission functions of the plurality of first interfaces to match the plurality of second interfaces.
16. The method according to any one of claims 13 to 15, characterized in that The plurality of second cabinets are further configured to output a plurality of second signals through the plurality of second interfaces, and the plurality of first cabinets are further configured to receive the plurality of second signals through the plurality of first interfaces.
17. The method according to claim 16, wherein The second logic is configured to: Define the signal transmission functions of the plurality of second interfaces to match the plurality of first interfaces.
18. The method according to claim 16 or 17, characterized in that The first logic is configured to: Define the signal reception functions of the plurality of first interfaces to match the plurality of second interfaces.
19. The method according to any one of claims 13 to 18, characterized in that, Running the first compiled version of the circuit under test includes: controlling the simulation period of the first compiled version of the circuit under test to be the same as the simulation period of the second compiled version of the circuit under test.
20. The method according to any one of claims 13 to 19, characterized in that, Running the second compiled version of the circuit under test includes: controlling the simulation period of the first compiled version of the circuit under test to be the same as the simulation period of the second compiled version of the circuit under test.
21. A computer-readable storage medium having a computer program stored thereon, the program, when executed by a processor, implementing the method according to any one of claims 13 to 20.
22. A computer program product including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 13 to 20.