Software and hardware combined simulation system and simulation method

By utilizing the high simulation speed of the software simulation domain and the high simulation accuracy of the hardware simulation domain through a software-hardware co-simulation system, the problems of low simulation accuracy and slow speed in chip design are solved, and efficient simulation-based collaborative design is achieved.

CN121809381APending Publication Date: 2026-04-07BEIJING AIJIE KEXIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, the separation of software simulation and hardware simulation results in low simulation accuracy and slow speed, making it difficult to perform chip design simulation efficiently and accurately.

Method used

A hardware-software co-simulation system is adopted. The software simulation domain is responsible for transaction scheduling and global time advancement, while the hardware simulation domain performs hardware simulation of peripheral devices, providing realistic latency feedback and bandwidth characteristics, and realizing the coordinated design of simulation accuracy and speed.

Benefits of technology

It achieves a synergistic improvement in simulation accuracy and speed, reduces communication overhead between the software simulation domain and the hardware simulation domain, and provides high-efficiency simulation accuracy and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809381A_ABST
    Figure CN121809381A_ABST
Patent Text Reader

Abstract

The invention provides a software and hardware combined simulation system and method, the simulation system comprises a software simulation domain and a hardware simulation domain, the software simulation domain comprises a plurality of processor cores and a software bridge, and the hardware simulation domain comprises a hardware bridge and peripheral equipment. The software simulation domain of the system is responsible for transaction scheduling and global time advancing, the advantage of high simulation speed of software simulation is fully utilized, meanwhile, hardware simulation is carried out on peripheral equipment through the hardware simulation domain of the system, real delay feedback and bandwidth characteristics can be provided, and the hardware simulation speed of the peripheral equipment is improved. The advantage of high simulation precision of hardware simulation is fully utilized, so that collaborative design of simulation precision and simulation speed is realized through combination of software and hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of simulation, and more specifically, to a hardware and software combined simulation system and simulation method. Background Technology

[0002] As chip design complexity increases dramatically, efficient and accurate simulation has become a key bottleneck in the chip design process. In related technologies, software simulation and hardware simulation are typically separated, each with its own significant limitations.

[0003] On the one hand, pure software simulation has advantages such as fast simulation speed, flexible debugging, and strong observability. However, the accuracy of pure software simulation is limited, as it typically makes highly abstract or idealized assumptions about the physical characteristics of the hardware, such as underlying timing, actual bandwidth, and peripheral interaction latency. Therefore, software simulation struggles to obtain accurate performance evaluation data, resulting in low simulation accuracy.

[0004] On the other hand, pure hardware simulation can provide extremely high timing accuracy and realistic hardware behavior, resulting in high simulation precision. However, due to the long compilation time of hardware simulation, it is generally very slow. In other words, hardware simulation has a low simulation speed. Summary of the Invention

[0005] This application provides at least one hardware-software co-simulation system and simulation method. The software simulation domain of the system is responsible for transaction scheduling and global time advancement, which fully utilizes the high simulation speed advantage of software simulation. At the same time, the hardware simulation domain of the system performs hardware simulation on peripheral devices, which can provide realistic latency feedback and bandwidth characteristics, fully utilizing the high simulation accuracy advantage of hardware simulation. Thus, the co-design of simulation accuracy and simulation speed is achieved through hardware-software co-simulation.

[0006] Firstly, this application provides a combined software and hardware simulation system, comprising a software simulation domain and a hardware simulation domain. The software simulation domain includes multiple processor cores and a software bridge, while the hardware simulation domain includes a hardware bridge and peripheral devices. Multiple processor cores are used to send multiple transaction requests to the software bridge; The software bridge is used to receive multiple transaction requests and send multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol. The hardware bridge is used to receive multiple hardware transaction requests via the interface protocol and send multiple hardware transaction requests to peripheral devices. The peripheral device is used to receive multiple hardware transaction requests in order to obtain the total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge; The hardware bridge is also used to receive overall feedback information; The software bridge is also used to read the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to multiple transaction requests, so as to send the feedback results and hardware simulation latency corresponding to multiple transaction requests to multiple processor cores. Multiple processor cores are also used to receive feedback results and hardware simulation latency corresponding to multiple transaction requests, so as to terminate multiple transaction requests and advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to multiple transaction requests.

[0007] Secondly, this application provides a hardware-software co-simulation method applied to a hardware-software co-simulation system. The simulation system includes a software simulation domain and a hardware simulation domain. The software simulation domain includes multiple processor cores and a software bridge, and the hardware simulation domain includes a hardware bridge and peripheral devices. The method includes: Multiple processor cores send multiple transaction requests to the software bridge; The software bridge receives multiple transaction requests and sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol. The hardware bridge receives multiple hardware transaction requests through the interface protocol and sends multiple hardware transaction requests to peripheral devices. The peripheral device receives multiple hardware transaction requests in order to obtain the total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge; The hardware bridge receives the overall feedback information; The software bridge reads the total feedback information and determines the feedback results and hardware simulation latency corresponding to each of the multiple transaction requests based on the total feedback information, so as to send the feedback results and hardware simulation latency corresponding to each of the multiple transaction requests to multiple processor cores. Multiple processor cores receive feedback results and hardware simulation latency corresponding to multiple transaction requests, so as to terminate multiple transaction requests and advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to each of the multiple transaction requests.

[0008] In summary, this application provides a hardware-software co-simulation system and method. The simulation system includes a software simulation domain and a hardware simulation domain. The software simulation domain includes multiple processor cores and a software bridge, while the hardware simulation domain includes a hardware bridge and peripheral devices. The multiple processor cores send multiple transaction requests to the software bridge. The software bridge receives the multiple transaction requests and sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge via an interface protocol. The hardware bridge receives the multiple hardware transaction requests via the interface protocol and sends them to the peripheral devices. The peripheral devices receive the multiple hardware transaction requests to obtain the total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge. The hardware bridge also receives the total feedback information. The software bridge reads the total feedback information and, based on the total feedback information, determines the feedback results and hardware simulation delays corresponding to the multiple transaction requests, so as to send the feedback results and hardware simulation delays corresponding to the multiple transaction requests to the multiple processor cores. The multiple processor cores also receive the feedback results and hardware simulation delays corresponding to the multiple transaction requests, so as to terminate the multiple transaction requests and advance the local simulation time based on the hardware simulation delays and transaction start times corresponding to the multiple transaction requests. The software simulation domain of the above system is responsible for transaction scheduling and global time advancement, which makes full use of the high simulation speed advantage of software simulation. At the same time, the hardware simulation domain of the above system performs hardware simulation of peripheral devices, which can provide realistic latency feedback and bandwidth characteristics, making full use of the high simulation accuracy advantage of hardware simulation. Thus, the collaborative design of simulation accuracy and simulation speed is achieved through the joint use of software and hardware.

[0009] Other advantages of this application will be explained in more detail in conjunction with the following description and figures.

[0010] It should be understood that the above description is merely an overview of the technical solution of this application, so as to enable a general understanding of the technical means of this application and to implement it in accordance with the contents of the specification. In order to make the above and other objects, features and advantages of this application more apparent and understandable, specific embodiments of this application are illustrated below. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. The accompanying drawings are incorporated in and constitute a part of this specification. These drawings illustrate embodiments conforming to this application and are used together with the specification to explain the technical solutions of this application. It should be understood that the drawings only illustrate certain embodiments of this application and should not be considered as a limitation on the scope of protection. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. Furthermore, the same reference numerals denote the same components throughout the drawings. In the drawings: Figure 1 A schematic diagram of a hardware and software combined simulation system provided for an embodiment of this application; Figure 2 A transaction state machine diagram of a hardware and software co-simulation system provided for embodiments of this application; Figure 3 A timing diagram of a hardware and software co-simulation system provided for an embodiment of this application; Figure 4 A transaction full-process latency diagram provided for embodiments of this application; Figure 5 A time hop diagram of multiple processor cores in a hardware-software co-simulation system provided in this application embodiment; Figure 6 A flowchart illustrating a hardware-software co-simulation method provided in this application embodiment. Detailed Implementation

[0012] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0013] In the description of embodiments of this application, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of the disclosed features, figures, steps, behaviors, components, portions or combinations thereof in this specification, and do not exclude the possibility of the presence of one or more other features, figures, steps, behaviors, components, portions or combinations thereof.

[0014] Unless otherwise stated, " / " means "or". For example, A / B can mean A or B. In this article, "and / or" is merely a way of describing the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A alone, A and B at the same time, and B alone.

[0015] The terms "first," "second," etc., are used only for ease of description to distinguish identical or similar technical features and should not be construed as indicating or implying the relative importance or number of these technical features. Therefore, a feature defined by "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of embodiments of this application, unless otherwise stated, the term "multiple" means two or more.

[0016] With the rapid increase in the complexity of integrated circuits, efficient and accurate simulation has become a key bottleneck in the integrated circuit design process. In related technologies, software simulation and hardware simulation are usually separated, each with its own obvious limitations.

[0017] On the one hand, pure software simulation has advantages such as fast simulation speed, flexible debugging, and strong observability. However, the accuracy of pure software simulation is limited, as it typically makes highly abstract or idealized assumptions about the physical characteristics of the hardware, such as underlying timing, actual bandwidth, and peripheral interaction latency. Therefore, software simulation struggles to obtain accurate performance evaluation data, resulting in low simulation accuracy.

[0018] On the other hand, pure hardware simulation can provide extremely high timing accuracy and realistic hardware behavior, resulting in high simulation precision. However, due to the long compilation time of hardware simulation, it is generally very slow. In other words, hardware simulation has a low simulation speed.

[0019] In view of this, this application provides a hardware and software co-simulation system and simulation method. The software simulation domain of the above system is responsible for transaction scheduling and global time advancement, which makes full use of the high simulation speed advantage of software simulation. At the same time, the hardware simulation domain of the above system performs hardware simulation of peripheral devices, which can provide realistic latency feedback and bandwidth characteristics, making full use of the high simulation accuracy advantage of hardware simulation. Thus, the co-design of simulation accuracy and simulation speed is achieved through hardware and software co-operation.

[0020] The following examples illustrate a hardware-software co-simulation system provided in this application. Figure 1 As shown, Figure 1 This is a schematic diagram of a hardware-software co-simulation system provided in an embodiment of this application. The hardware-software co-simulation system 100 includes a software simulation domain 110 and a hardware simulation domain 120. The software simulation domain 110 includes multiple processor cores and a software bridge, and the hardware simulation domain 120 includes a hardware bridge and peripheral devices. In practical applications, the software simulation domain can be implemented based on the SystemC language, and the hardware simulation domain can be implemented based on a Field-Programmable Gate Array (FPGA). Figure 2 , Figure 3 , Figure 4 and Figure 5 The following describes the hardware and software combined simulation system 100: Multiple processor cores are used to send multiple transaction requests to the software bridge.

[0021] Specifically, Transaction-Level Modeling (TLM) is a modeling method that uses "transactions" as the unit of communication and describes system behavior at a high level of abstraction.

[0022] Multiple processor cores can be used to send multiple transaction requests to the software bridge, such as Figure 2 As shown, processor cores 1 to N can act as transaction request initiators, and the software bridge can act as transaction request destinations, sending multiple transaction requests. Transaction requests can be TLM read / write requests.

[0023] The software bridge is used to receive multiple transaction requests and send multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol.

[0024] Specifically, such as Figure 2 As shown, the software bridge can be used to receive multiple transaction requests and send multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol.

[0025] In practical applications, the interface protocol can adopt the PCIe-hardware protocol. For example, the PCIe-AXI protocol can be used to realize high-speed interconnection between the software bridge and the hardware bridge. In this case, multiple hardware transaction requests can be Advanced eXtensible Interface (AXI) requests. That is, the software bridge can send multiple hardware transaction requests (AXI requests) corresponding to multiple transaction requests (TLM requests) to the hardware bridge through the interface protocol (PCIe-AXI).

[0026] The hardware bridge is used to receive multiple hardware transaction requests via an interface protocol and to send multiple hardware transaction requests to peripheral devices.

[0027] Specifically, such as Figure 2 As shown, the hardware bridge is used to receive multiple hardware transaction requests through the interface protocol and send multiple hardware transaction requests to peripheral devices.

[0028] In practical applications, peripheral devices can be Register Transfer Level (RTL) devices, and the hardware bridge can send multiple hardware transaction requests to peripheral devices in parallel.

[0029] The peripheral device is used to receive multiple hardware transaction requests in order to obtain the total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge.

[0030] Specifically, such as Figure 2 As shown, the peripheral device is used to receive multiple hardware transaction requests in order to obtain the total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge.

[0031] In practical applications, peripheral devices may include a memory subsystem consisting of on-chip networks, 3D dynamic random access memory, cache, etc.

[0032] The total feedback information is obtained through hardware simulation using peripheral devices. Hardware simulation can provide realistic delay feedback and bandwidth characteristics, thereby ensuring the simulation accuracy of the total feedback information.

[0033] The hardware bridge is also used to receive overall feedback information.

[0034] Specifically, such as Figure 2 As shown, the hardware bridge can also be used to receive total feedback information for the software bridge to read.

[0035] The software bridge is also used to read the total feedback information and, based on the total feedback information, determine the feedback results and hardware emulation latency corresponding to each of the multiple transaction requests, so as to send the feedback results and hardware emulation latency corresponding to each of the multiple transaction requests to multiple processor cores.

[0036] Specifically, the software bridge is also used to read the overall feedback information.

[0037] In practical applications, the total feedback information will include identification information, which may include a transaction identifier (Transaction ID). The transaction identifier is a unique number that identifies a transaction request and is used to distinguish different transaction requests. This allows the information in the total feedback information to be mapped one-to-one with multiple transaction requests, thereby determining the feedback results and hardware simulation latency corresponding to each of the multiple transaction requests.

[0038] Hardware emulation latency refers to the actual latency feedback obtained from emulating hardware transaction requests from peripheral devices. In other words, hardware emulation latency only corresponds to the peripheral device and does not include the latency of the software bridge and hardware bridge performing the corresponding processing. Figure 3 As shown, the time it takes for the hardware bridge to send each hardware transaction request to the peripheral device is the hardware transaction request start time, denoted as . The feedback information returned by the peripheral device to the hardware bridge for each hardware transaction request is the end time of the hardware transaction request, denoted as . The time difference between the start time and the end time of a hardware transaction request is the hardware simulation latency.

[0039] In practical applications, after hardware simulation, the total feedback information can include: Transaction ID: A unique identifier for a transaction request, used to distinguish different transaction requests; Cache hit: The hit status recorded by the multi-level cache controller; Cache miss: A cache miss status recorded by a multi-level cache controller; Location: The cell where the data is located (e.g., the nth L2 or the nth DDR); destination: The final point of data migration or replication; Read latency: Hardware simulation latency of the read path. Write latency: Hardware simulation latency of the write path.

[0040] Multiple processor cores are also used to receive feedback results and hardware simulation latency corresponding to multiple transaction requests, so as to terminate multiple transaction requests and advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to multiple transaction requests.

[0041] Specifically, such as Figure 2 and Figure 3 As shown, multiple processor cores are also used to receive feedback results and hardware simulation latency corresponding to multiple transaction requests.

[0042] Multiple processor cores can terminate multiple transaction requests after receiving the feedback results corresponding to each transaction request.

[0043] The transaction start time corresponding to multiple transaction requests refers to the time when multiple processor cores initiate their respective transaction requests, such as... Figure 3 As shown, it can be written as .

[0044] Multiple processor cores are used to advance their local simulation time based on the hardware simulation latency and the transaction start time corresponding to each of the multiple transaction requests. In other words, multiple processor cores can advance their local simulation time based on the sum of the hardware simulation latency and the transaction start time corresponding to the transaction request.

[0045] like Figure 4 As shown, the latency of the entire transaction process of transaction 1 is used as an example to illustrate the latency of the transaction request. Specifically, during the simulation, transaction 1 will sequentially go through the processor core - software bridge - hardware bridge - peripheral device - hardware bridge - software bridge from start to finish. Since the processing of transaction 1 by the software bridge and hardware bridge is for transaction transmission, the latency of the software bridge and hardware bridge in processing transaction 1 needs to be ignored. The final completion time of the transaction only needs to refer to the transaction start time of the corresponding transaction request initiated by the processor core. The start time of the hardware transaction request sent by the hardware bridge to the peripheral device. The end time of the hardware transaction request from the peripheral device to return feedback information corresponding to the hardware transaction request to the hardware bridge. Hardware transaction request start time Hardware transaction request end time The difference between the two is the hardware simulation latency, which is the final completion time of the transaction. ,like Figure 4 As shown in the purple time block.

[0046] like Figure 5 As shown in (a), for multiple transaction requests sent from multiple processor cores, the start times of the transactions corresponding to the multiple transaction requests may not be the same. The transaction start times are as follows: Figure 5 As shown by the red broken line in (a); Figure 5 As shown in (b), in the hardware emulation domain, the hardware emulation latency corresponding to multiple transaction requests determined by the peripheral device may not be the same. The hardware emulation latency is as follows: Figure 5 As shown by the black vertical line in (b); Figure 5 As shown in (c), based on the transaction start time and hardware simulation latency corresponding to the aforementioned multiple transaction requests, the local simulation time corresponding to the end of the multiple transaction requests can be determined. The local simulation time corresponding to the end of the multiple transaction requests is as follows: Figure 5 As shown by the red broken line at the top in (c). Furthermore, after advancing the local simulation time corresponding to multiple transaction requests, a leapfrog of the next batch of multiple transaction requests can be performed.

[0047] Therefore, by scheduling transactions through the software simulation domain of the aforementioned system, which offers the advantage of fast response times, and by advancing local simulation time based on the hardware simulation latency and start time of each transaction request, the system's software simulation domain effectively leverages the high simulation speed of software simulation. Simultaneously, the system's hardware simulation domain performs hardware simulation of peripheral devices, providing realistic latency feedback and bandwidth characteristics, thus fully utilizing the high simulation accuracy of hardware simulation. This combined software and hardware approach achieves a synergistic design of simulation accuracy and speed.

[0048] Furthermore, the simulation system of this application reduces communication overhead between the software simulation domain and the hardware simulation domain. A dual-bridge structure is designed between the two domains. The software bridge only transmits transaction control information to the hardware simulation domain, which autonomously performs memory access and performance statistics. The hardware bridge returns overall feedback information to the software bridge, which then returns the required data to each processor core and advances the local simulation time. In other words, only lightweight transaction control information is transmitted between the software and hardware simulation domains, without directly transmitting large-scale data. This significantly reduces bandwidth usage and communication overhead, achieving efficient software-hardware co-simulation and accurate performance evaluation.

[0049] In one possible implementation, the software bridge includes a mirrored storage hierarchy that simulates the storage hierarchy of peripheral devices. The software bridge is also used to send multiple transaction requests to the mirrored storage hierarchy after receiving multiple transaction requests. The mirrored storage hierarchy is also used to receive multiple transaction requests; The software bridge is also used to read the overall feedback information and, based on this information, determine the feedback results and hardware simulation latency corresponding to multiple transaction requests, including: The software bridge is also used to read the total feedback information and send it to the mirrored storage hierarchy; the mirrored storage hierarchy is also used to receive the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to multiple transaction requests.

[0050] Specifically, such as Figure 2 As shown, the software bridge includes a mirror memory hierarchy (MMH). The MMH is used to simulate the storage hierarchy of peripheral devices, which is equivalent to mirroring the peripheral devices on the software side, thereby ensuring data consistency between the software side and the hardware side.

[0051] The software bridge is also used to forward multiple transaction requests to the MMH after receiving them, so that the mirrored storage hierarchy can receive multiple transaction requests, such as... Figure 2 As shown, in addition to sending hardware transaction requests to the hardware bridge, the software bridge also sends multiple transaction requests to the MMH via the splitter.

[0052] When the MMH is configured in the software bridge, the software bridge is also used to receive and send total feedback information to the MMH; the MMH is also used to receive total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to multiple transaction requests respectively. Since the MMH has the same data as the hardware side, it can better determine the feedback results and hardware simulation latency corresponding to multiple transaction requests.

[0053] In one possible implementation, the hardware bridge is also used to send an interrupt signal to the software bridge after receiving the total feedback information. The software bridge is also used to read total feedback information, including: The software bridge is also used to read the total feedback information from the hardware bridge after receiving an interrupt signal.

[0054] Specifically, to ensure that the software bridge can accurately read the overall feedback information, such as Figure 2As shown, after the software bridge sends a hardware transaction request to the hardware bridge, it waits for a response from the hardware emulation domain. After receiving the total feedback information, the hardware bridge sends an interrupt signal to the software bridge so that the software bridge can read the total feedback information from the hardware bridge after receiving the interrupt signal.

[0055] In one possible implementation, the mirrored storage hierarchy is further used to receive total feedback information and, based on the total feedback information, determine the feedback results and hardware emulation latency corresponding to multiple transaction requests, including: The mirrored storage hierarchy is also used to receive overall feedback information and update the storage hierarchy based on the overall feedback information in order to determine the feedback results and hardware simulation latency corresponding to multiple transaction requests.

[0056] Specifically, to ensure data consistency between the MMH and the hardware side, such as Figure 3 As shown, after the MMH receives the total feedback information, it will update the storage hierarchy structure.

[0057] In one possible implementation, multiple processor cores are used to send multiple transaction requests, including: Multiple processor cores are used to send multiple transaction requests in parallel; The software bridge receives multiple transaction requests and sends multiple hardware transaction requests corresponding to these requests to the hardware bridge via an interface protocol, including: The software bridge is used to receive the multiple transaction requests, and when the multiple transaction requests meet the preset batch processing conditions, it sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge in batches through the interface protocol. Multiple processor cores are also used to receive feedback results and transaction completion times corresponding to multiple transaction requests, in order to terminate multiple transaction requests and advance the local simulation time according to the transaction completion times corresponding to the multiple transaction requests, including: Multiple processor cores are also used to receive feedback results and hardware simulation latency corresponding to multiple transaction requests, so as to asynchronously end multiple transaction requests and asynchronously advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to multiple transaction requests, and synchronize at the quantum time boundary after multiple transaction requests have ended.

[0058] Specifically, in this embodiment, multiple processor cores are used to send multiple transaction requests in parallel. That is, multiple processors can asynchronously advance multiple transactions within a quantum time interval without immediate global time synchronization. In practical applications, multiple processors can send multiple transaction requests to a transaction queue in parallel, and the transaction queue supports batch reception and caching of transaction requests.

[0059] To address this, the software bridge receives multiple transaction requests and, when these requests meet preset batch processing conditions, packages them into batches and sends them to the hardware bridge as multiple hardware transaction requests, such as... Figure 2 As shown, after the preset batch processing conditions are met, the software bridge sends a batch of hardware transaction requests through the splitter. The batch of hardware transaction requests may include request metadata, such as address, command, timestamp, etc.

[0060] Multiple processor cores are also used to receive feedback results and hardware simulation latency corresponding to multiple transaction requests, so as to asynchronously terminate multiple transaction requests and asynchronously advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to each of the multiple transaction requests. Figure 2 As shown, the local simulation time of each transaction request is asynchronously advanced in sequence according to the process modules of "read / write data" - "mirror storage hierarchy" - "feedback results" - "sorting transaction queue" - "transaction queue". After multiple transaction requests have been processed asynchronously, they are synchronized at the quantum time boundary, thereby ensuring the timing consistency of the software simulation domain.

[0061] In one possible implementation, the hardware emulation domain also includes a detector: The detector is used to collect feedback information from peripheral devices in response to multiple hardware transaction requests in order to obtain total feedback information, and then send the total feedback information to the hardware bridge.

[0062] Specifically, such as Figure 2 As shown, the hardware emulation domain also includes a monitor, which contains multiple probes. The monitor is used to collect feedback information from peripheral devices in response to multiple hardware transaction requests in order to obtain total feedback information, and then send the total feedback information to the hardware bridge.

[0063] In one possible implementation, the hardware bridge includes a dedicated register space corresponding to the interface protocol, which is used to store the total feedback information for the software bridge to read.

[0064] Specifically, such as Figure 2 As shown, the hardware bridge includes a dedicated register space (TraceBuffer) corresponding to the interface protocol. The dedicated register space is used to store the total feedback information for the software bridge to read. In practical applications, the dedicated register space stores the total feedback information in the form of a trace record.

[0065] Therefore, this application provides a hardware-software co-simulation system. The simulation system includes a software simulation domain and a hardware simulation domain. The software simulation domain includes multiple processor cores and a software bridge, while the hardware simulation domain includes a hardware bridge and peripheral devices. The multiple processor cores are used to send multiple transaction requests. The software bridge is used to receive multiple transaction requests and send multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge via an interface protocol. The hardware bridge is used to receive multiple hardware transaction requests via an interface protocol and send multiple hardware transaction requests to the peripheral devices. The peripheral devices are used to receive multiple hardware transaction requests to obtain total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge. The hardware bridge is also used to receive the total feedback information. The software bridge is also used to read the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation delays corresponding to the multiple transaction requests, so as to send the feedback results and hardware simulation delays corresponding to the multiple transaction requests to the multiple processor cores. The multiple processor cores are also used to receive the feedback results and hardware simulation delays corresponding to the multiple transaction requests, so as to terminate the multiple transaction requests and advance the local simulation time according to the hardware simulation delays and transaction start times corresponding to the multiple transaction requests. The software simulation domain of the above system is responsible for transaction scheduling and global time advancement, which makes full use of the high simulation speed advantage of software simulation. At the same time, the hardware simulation domain of the above system performs hardware simulation of peripheral devices, which can provide realistic latency feedback and bandwidth characteristics, making full use of the high simulation accuracy advantage of hardware simulation. Thus, the collaborative design of simulation accuracy and simulation speed is achieved through the joint use of software and hardware.

[0066] The following examples illustrate a hardware-software co-simulation method provided in this application. Figure 6 As shown, Figure 6 A flowchart of a hardware-software co-simulation method provided in this application embodiment is shown. This simulation method is applied to the aforementioned hardware-software co-simulation system. The simulation system includes a software simulation domain and a hardware simulation domain. The software simulation domain includes multiple processor cores and a software bridge, and the hardware simulation domain includes a hardware bridge and peripheral devices. The method includes: S601: Multiple processor cores send multiple transaction requests to the software bridge.

[0067] S602, the software bridge receives multiple transaction requests and sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol.

[0068] S603, the hardware bridge receives multiple hardware transaction requests through the interface protocol and sends multiple hardware transaction requests to peripheral devices.

[0069] S604. The peripheral device receives multiple hardware transaction requests in order to obtain the total feedback information corresponding to the multiple hardware transaction requests and send the total feedback information to the hardware bridge. S605, the hardware bridge receives the overall feedback information.

[0070] S606, the software bridge reads the total feedback information and determines the feedback results and hardware simulation latency corresponding to multiple transaction requests based on the total feedback information, so as to send the feedback results and hardware simulation latency corresponding to multiple transaction requests to multiple processor cores. S607: Multiple processor cores receive feedback results and hardware simulation delays corresponding to multiple transaction requests, so as to terminate multiple transaction requests and advance the local simulation time according to the hardware simulation delays and transaction start times corresponding to the multiple transaction requests.

[0071] In one possible implementation, the software bridge includes a mirrored storage hierarchy used to simulate the storage hierarchy of peripheral devices, and the method further includes: After receiving multiple transaction requests, the software bridge sends multiple transaction requests to the mirrored storage hierarchy. The mirrored storage hierarchy receives multiple transaction requests; The software bridge in S606 reads the overall feedback information and, based on this information, determines the feedback results and hardware simulation latency corresponding to multiple transaction requests, including: The software bridge reads the overall feedback information and sends it to the mirror storage hierarchy. The mirror storage hierarchy receives the overall feedback information and determines the feedback results and hardware simulation latency corresponding to the multiple transaction requests based on the overall feedback information.

[0072] In one possible implementation, the method further includes: After receiving the overall feedback information, the hardware bridge sends an interrupt signal to the software bridge. The software bridge in S606 reads the overall feedback information, including: After receiving the interrupt signal, the software bridge reads the overall feedback information from the hardware bridge.

[0073] In one possible implementation, the mirrored storage hierarchy receives the overall feedback information and, based on this information, determines the feedback results and transaction completion times for each of the multiple transaction requests, including: The mirrored storage hierarchy receives the overall feedback information and updates the storage hierarchy based on the overall feedback information in order to determine the feedback results and transaction completion times corresponding to multiple transaction requests.

[0074] In one possible implementation, multiple processor cores send multiple transaction requests, including: Multiple processor cores send multiple transaction requests in parallel; The software bridge receives multiple transaction requests and sends multiple hardware transaction requests corresponding to these requests to the hardware bridge via an interface protocol, including: The software bridge receives multiple transaction requests, and when the multiple transaction requests meet the preset batch processing conditions, it sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge in batches through the interface protocol. Multiple processor cores receive feedback results and hardware simulation latency corresponding to multiple transaction requests, in order to terminate the multiple transaction requests and advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to each of the multiple transaction requests, including: Multiple processor cores asynchronously receive feedback results and transaction completion times corresponding to multiple transaction requests, so as to asynchronously end multiple transaction requests and asynchronously advance the local simulation time according to the hardware simulation latency and transaction start time corresponding to multiple transaction requests, and synchronize at the quantum time boundary after all multiple transaction requests have ended.

[0075] In one possible implementation, multiple processor cores send multiple transaction requests in parallel, including: Multiple processor cores send multiple transaction requests in parallel, and the transaction start time corresponds to each of the multiple transaction requests; The software bridge receives the multiple transaction requests, and when the multiple transaction requests meet preset batch processing conditions, it sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge in batches via an interface protocol, including: The software bridge receives multiple transaction requests and the transaction start times corresponding to each transaction request. When multiple transaction requests meet preset batch processing conditions, the software bridge sorts the transaction requests based on the transaction start times corresponding to each transaction request. Then, it sends multiple sorted hardware transaction requests corresponding to the sorted multiple transaction requests to the hardware bridge in batches through the interface protocol.

[0076] Specifically, since multiple transaction requests need to be sent in batches, the software bridge can sort the transaction requests based on the transaction start time corresponding to each of the multiple transaction requests. For example, the transaction requests can be sorted from low to high based on the transaction start time corresponding to each of the multiple transaction requests, and the sorted multiple transaction requests can be packaged together so that the hardware bridge can also send hardware transaction requests in the same order.

[0077] like Figure 2 As shown, the software bridge can obtain a sorted transaction queue by sorting, and the hardware bridge can also obtain a sorted hardware transaction request queue according to the sorting of the software bridge.

[0078] In one possible implementation, the preset batch processing conditions include the number of pending transaction requests meeting a preset number threshold or the time difference between the start time of the latest transaction request and the start time of the earliest transaction request meeting a preset time threshold.

[0079] Specifically, in order to better handle multiple transaction requests, the preset batch processing conditions can be either that the number of pending transaction requests meets a preset number threshold, or that the time difference between the start time of the latest transaction request and the start time of the earliest transaction request meets a preset time threshold.

[0080] In one possible implementation, the hardware emulation domain also includes a detector, and the method further includes: The detector collects feedback information from peripheral devices in response to multiple hardware transaction requests to obtain total feedback information, and then sends the total feedback information to the hardware bridge.

[0081] It should be noted that the relevant descriptions in the method embodiments can be referred to the foregoing system embodiments, and will not be repeated here.

[0082] The following is based on Figure 2 and Figure 3 A more comprehensive and easier-to-understand explanation of the asynchronous operation in the hardware-software co-simulation system provided in this application is given below: Multiple processor cores can asynchronously advance multiple transactions within a quantum time interval without immediate global time synchronization. Multiple processor cores can send multiple transaction requests to a transaction queue in parallel. The transaction queue supports batch reception and caching of transaction requests and can be sorted based on the transaction start time of the requests to obtain a sorted transaction queue.

[0083] The software bridge is triggered only when the number of new transaction requests detected reaches a preset number threshold (batch size) or when the time difference between the start time of the latest transaction request and the start time of the earliest transaction request meets a preset time threshold (batch time).

[0084] During the "batch request sending" phase, when a batch of transaction requests arrives, the splitter divides the requests into two paths: one path is converted into a hardware transaction request via PCIe–AXI and sent to the hardware emulation domain; the other path is forwarded as a transaction request to the MMH. The first request is processed by the peripheral device, and the detector can collect the total feedback information of the peripheral device's processing of the hardware transaction request. After the first request and the collection of total feedback information are completed, the hardware bridge notifies the software bridge via an interrupt signal; the software bridge receives the interrupt signal, reads the total feedback information, and forwards it to the MMH. In other words, the MMH receives transaction requests and total feedback information at different times.

[0085] The software bridge then enters the "read / write data" phase. The MMH determines the received total feedback information based on the transaction identifier and updates the image hierarchy according to the content of the total feedback information. For example, it can update the cache / DDR hit and replacement status, and return the feedback results and hardware emulation latency corresponding to each of the multiple transaction requests in the batch. The software bridge then enters the "feedback result" phase, where multiple processor cores can asynchronously terminate multiple transaction requests and asynchronously advance the local emulation time according to the hardware emulation latency and transaction start time corresponding to each of the multiple transaction requests.

[0086] The following also combines Figure 6 Let's explain in detail the progression of simulation time in the simulation method: 1. Each processor core sends a transaction request and determines the local simulation time (local timeoffset) corresponding to each of the multiple transaction requests as the transaction start time. 2. When the software bridge receives a transaction request, it records the minimum transaction start time as the base synchronization point (bsp) and the maximum transaction start time as the top synchronization point (tsp) based on the transaction start time corresponding to the multiple transaction requests. 3. Within the batch time, if the number of transaction requests received in the software bridge reaches the preset threshold, the received transaction requests will be sorted from low to high according to the transaction start time and packaged in sorted order; or, if the time difference between the tsp and bsp meets the preset time threshold, the transaction requests in the software bridge will also be sorted from low to high according to the transaction start time and packaged in sorted order. 4. After receiving the packaged hardware transaction request, the hardware bridge then sends the hardware transaction requests in sequence within the hardware bridge. 5. If the hardware bridge has finished processing all hardware transaction requests, the hardware bridge state changes to Idle, and the FPGA state is frozen. 6. The software bridge reads the total feedback information from the hardware bridge; 7. Each processor core advances its local time offset based on the sum of the hardware simulation latency corresponding to the transaction request and the transaction start time corresponding to the transaction request. 8. Repeat the above steps iteratively.

[0087] In the description of this specification, references to terms such as "some possible implementations," "some implementations," "example," "specific example," or "some examples" indicate that a specific feature, structure, material, or characteristic described in connection with that implementation or example is included in at least one implementation or example of this application, and the aforementioned terms do not necessarily refer to the same implementation or example. Furthermore, the described specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more implementations or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different implementations or examples described in this specification, as well as the features of different implementations or examples.

[0088] While the spirit and principles of this application have been described above with reference to several specific embodiments, it should be understood that this application is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined. This application is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A hardware and software combined simulation system, characterized in that, The simulation system includes a software simulation domain and a hardware simulation domain. The software simulation domain includes multiple processor cores and a software bridge, and the hardware simulation domain includes a hardware bridge and peripheral devices. The plurality of processor cores are used to send a plurality of transaction requests to the software bridge; The software bridge is used to receive the multiple transaction requests and send multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol. The hardware bridge is used to receive the plurality of hardware transaction requests through an interface protocol and to send the plurality of hardware transaction requests to the peripheral device. The peripheral device is used to receive the plurality of hardware transaction requests in order to obtain the total feedback information corresponding to the plurality of hardware transaction requests and send the total feedback information to the hardware bridge; The hardware bridge is also used to receive the total feedback information; The software bridge is also used to read the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to the multiple transaction requests respectively, so as to send the feedback results and hardware simulation latency corresponding to the multiple transaction requests respectively to the multiple processor cores; The plurality of processor cores are also configured to receive feedback results and hardware simulation latency corresponding to the plurality of transaction requests respectively, so as to terminate the plurality of transaction requests and advance the local simulation time according to the hardware simulation latency corresponding to the plurality of transaction requests and the transaction start time corresponding to the plurality of transaction requests respectively.

2. The simulation system according to claim 1, characterized in that, The software bridge includes a mirrored storage hierarchy, which is used to simulate the storage hierarchy of the peripheral device. The software bridge is also used to send the multiple transaction requests to the mirrored storage hierarchy after receiving the multiple transaction requests; The mirrored storage hierarchy is also used to receive the multiple transaction requests; The software bridge is also used to read the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to the multiple transaction requests, including: The software bridge is also used to read the total feedback information and send the total feedback information to the mirror storage hierarchy; The mirror storage hierarchy is also used to receive the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to the multiple transaction requests respectively.

3. The simulation system according to claim 1, characterized in that, The hardware bridge is also used to send an interrupt signal to the software bridge after receiving the total feedback information; The software bridge is also used to read the total feedback information, including: The software bridge is also used to read the total feedback information from the hardware bridge after receiving the interrupt signal.

4. The simulation system according to claim 2, characterized in that, The mirrored storage hierarchy is also used to receive the total feedback information and, based on the total feedback information, determine the feedback results and hardware simulation latency corresponding to the multiple transaction requests, including: The mirrored storage hierarchy is also used to receive the total feedback information and update the storage hierarchy according to the total feedback information, so as to determine the feedback results and hardware simulation latency corresponding to the multiple transaction requests respectively.

5. The simulation system according to claim 1, characterized in that, The plurality of processor cores are used to send multiple transaction requests, including: The multiple processor cores are used to send multiple transaction requests in parallel; The software bridge is used to receive the multiple transaction requests and send multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through an interface protocol, including: The software bridge is used to receive the multiple transaction requests, and when the multiple transaction requests meet the preset batch processing conditions, it sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge in batches through the interface protocol. The plurality of processor cores are also configured to receive feedback results and transaction completion times corresponding to the plurality of transaction requests, so as to terminate the plurality of transaction requests and advance the local simulation time according to the transaction completion times corresponding to the plurality of transaction requests, including: The plurality of processor cores are also configured to receive feedback results and hardware simulation latency corresponding to the plurality of transaction requests respectively, so as to asynchronously end the plurality of transaction requests and asynchronously advance the local simulation time according to the hardware simulation latency and the transaction start time corresponding to the plurality of transaction requests respectively, and synchronize at the quantum time boundary after all the plurality of transaction requests have ended.

6. The simulation system according to claim 1, characterized in that, The hardware simulation domain also includes detectors: The detector is used to collect feedback information from the peripheral device in response to the multiple hardware transaction requests in order to obtain the total feedback information, and to send the total feedback information to the hardware bridge.

7. The simulation system according to claim 1, characterized in that, The hardware bridge includes a dedicated register space corresponding to the interface protocol, which is used to store the total feedback information so that the software bridge can read it.

8. A hardware-software co-simulation method, characterized in that, A simulation system applied to hardware and software integration, the simulation system comprising a software simulation domain and a hardware simulation domain, the software simulation domain comprising multiple processor cores and a software bridge, the hardware simulation domain comprising a hardware bridge and peripheral devices, the method comprising: The plurality of processor cores send a plurality of transaction requests to the software bridge; The software bridge receives the multiple transaction requests and sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge through the interface protocol. The hardware bridge receives the multiple hardware transaction requests through an interface protocol and sends the multiple hardware transaction requests to the peripheral device. The peripheral device receives the plurality of hardware transaction requests in order to obtain the total feedback information corresponding to the plurality of hardware transaction requests and send the total feedback information to the hardware bridge; The hardware bridge receives the total feedback information; The software bridge reads the total feedback information and determines the feedback results and hardware simulation latency corresponding to the multiple transaction requests based on the total feedback information, so as to send the feedback results and hardware simulation latency corresponding to the multiple transaction requests to the multiple processor cores. The multiple processor cores receive feedback results and hardware simulation latency corresponding to the multiple transaction requests, respectively, so as to terminate the multiple transaction requests and advance the local simulation time according to the hardware simulation latency and the transaction start time corresponding to the multiple transaction requests.

9. The simulation method according to claim 8, characterized in that, The software bridge includes a mirrored storage hierarchy, which is used to simulate the storage hierarchy of the peripheral device. The method further includes: After receiving the multiple transaction requests, the software bridge sends the multiple transaction requests to the mirror storage hierarchy; The mirrored storage hierarchy receives the multiple transaction requests; The software bridge reads the total feedback information and, based on the total feedback information, determines the feedback results and hardware simulation latency corresponding to the multiple transaction requests, including: The software bridge reads the total feedback information and sends the total feedback information to the mirror storage hierarchy; the mirror storage hierarchy receives the total feedback information and determines the feedback results and hardware simulation latency corresponding to the multiple transaction requests based on the total feedback information.

10. The simulation method according to claim 8, characterized in that, The method further includes: After receiving the total feedback information, the hardware bridge sends an interrupt signal to the software bridge. The software bridge reads the total feedback information, including: After receiving the interrupt signal, the software bridge reads the total feedback information from the hardware bridge.

11. The simulation method according to claim 9, characterized in that, The mirrored storage hierarchy receives the total feedback information and, based on the total feedback information, determines the feedback results and transaction completion times corresponding to the multiple transaction requests, including: The mirrored storage hierarchy receives the total feedback information and updates the storage hierarchy based on the total feedback information in order to determine the feedback results and transaction completion times corresponding to the multiple transaction requests.

12. The simulation method according to claim 8, characterized in that, The multiple processor cores send multiple transaction requests, including: The multiple processor cores send multiple transaction requests in parallel. The software bridge receives the multiple transaction requests and sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge via an interface protocol, including: The software bridge receives the multiple transaction requests, and when the multiple transaction requests meet the preset batch processing conditions, it sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge in batches through the interface protocol. The plurality of processor cores receive feedback results and hardware simulation latency corresponding to the plurality of transaction requests, respectively, in order to terminate the plurality of transaction requests and advance the local simulation time according to the hardware simulation latency and the transaction start time corresponding to the plurality of transaction requests, including: The multiple processor cores asynchronously receive the feedback results and transaction completion times corresponding to the multiple transaction requests, so as to asynchronously end the multiple transaction requests and asynchronously advance the local simulation time according to the hardware simulation latency and the transaction start time corresponding to the multiple transaction requests, and synchronize at the quantum time boundary after all the multiple transaction requests have ended.

13. The simulation method according to claim 12, characterized in that, The multiple processor cores send multiple transaction requests in parallel, including: The multiple processor cores send multiple transaction requests in parallel and the transaction start time corresponding to each of the multiple transaction requests; The software bridge receives the multiple transaction requests, and when the multiple transaction requests meet preset batch processing conditions, it sends multiple hardware transaction requests corresponding to the multiple transaction requests to the hardware bridge in batches via an interface protocol, including: The software bridge receives the multiple transaction requests and the transaction start times corresponding to the multiple transaction requests respectively. When the multiple transaction requests meet the preset batch processing conditions, it sorts the transaction requests based on the transaction start times corresponding to the multiple transaction requests respectively, and sends the sorted multiple hardware transaction requests corresponding to the sorted multiple transaction requests to the hardware bridge in batches through the interface protocol.

14. The simulation method according to claim 13, characterized in that, The preset batch processing conditions include the number of pending transaction requests meeting a preset number threshold or the time difference between the start time of the latest transaction request and the start time of the earliest transaction request meeting a preset time threshold.

15. The simulation method according to claim 8, characterized in that, The hardware simulation domain also includes a detector, and the method further includes: The detector collects feedback information from the peripheral devices in response to the multiple hardware transaction requests to obtain the total feedback information, and sends the total feedback information to the hardware bridge.