A SoC-oriented accelerator virtualization verification system
Patent Information
- Application Number
- CN202610159308.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-02-04
AI Technical Summary
在传统流程中,软件调试必须等待RTL设计达到较高成熟度,且在FPGA原型就绪后才能有效开展,导致软硬件开发周期严重串行化,整体项目周期长、成本高昂、效率低
[0019] The accelerator virtualization verification system and method for SoC provided in this application fundamentally changes the traditional serial mode of software and hardware development by constructing a virtual verification environment based on shared memory and transactional bus simulation. It enables the software (including drivers, operating systems, and applications) for dedicated accelerators to be fully developed and debugged in an environment that simulates the behavior of the target hardware with high precision before the hardware design is completed, realizing true parallel software and hardware development and significantly shortening the product development cycle.
Smart Images

Figure CN122086693B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual verification technology, and in particular to an accelerator virtualization verification system and method for SoC. Background Technology
[0002] With the rapid development of artificial intelligence, high-performance computing, and heterogeneous computing technologies, integrating various dedicated hardware accelerators (such as general-purpose matrix multiplication units (GEMMs) and video processing units (VPUs) into System-on-Chip (SoC) designs has become a key means to improve system energy efficiency and computing power. However, the increasing complexity of accelerator design has brought significant challenges to the development and verification of the accompanying software (including drivers, operating systems, and applications). Traditionally, hardware design and software development often follow a sequential process: the hardware design first completes the register-transfer-level (RTL) description, then performs functional verification through software simulation, and only after the design is relatively stable can it be ported to a field-programmable gate array (FPGA) prototype platform for system-level software integration and testing.
[0003] Currently, the industry mainly relies on two technical paths: one is the traditional verification process based on RTL simulation and FPGA prototyping; the other is virtual prototyping technology based on tools such as SystemC / TLM or QEMU. In the traditional process, software debugging must wait for the RTL design to reach a high level of maturity and can only be effectively carried out after the FPGA prototype is ready. This results in a serious serialization of the software and hardware development cycle, with long overall project cycles, high costs, and low efficiency.
[0004] Furthermore, RTL-level simulation is extremely slow, making it difficult to support large-scale software testing, and its limited observability of the internal hardware state hinders problem localization. While virtual prototyping technology can establish a runnable software model before the chip's physical implementation, its communication models are usually highly idealized. For example, it uses blocking communication with TLM 2.0 to approximate timing, making it difficult to accurately simulate the complex arbitration, queuing, bandwidth contention, and latency jitter effects in a real SoC bus. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the purpose of this application is to provide an accelerator virtualization verification system and method for SoC, which accurately simulates the real SoC bus process and improves the efficiency of accelerator verification for SoC.
[0006] To achieve the above objectives, this application provides an accelerator virtualization verification system for SoCs, comprising: A shared memory unit is set within the process address space of the general-purpose computing platform to construct a unified virtual memory space; The transaction processing module is communicatively connected to the shared memory unit and is configured to abstract and encapsulate the access operations of the master device represented by the upper-layer software to the hardware accelerator or memory into a transaction structure of a predetermined format, and to schedule the transaction structure. An accelerator function model unit, connected to the transaction processing module and the shared memory unit, is configured to execute the corresponding hardware acceleration function simulation in response to a transaction scheduled by the transaction processing module, and write the response result to the shared memory unit.
[0007] Furthermore, the transaction processing module includes: The transaction abstraction unit is configured to encapsulate access operations initiated by the master device represented by the upper-layer software into the transaction structure, wherein the transaction structure includes at least the following fields: target address, data length, transmission type, burst mode, transaction identifier, and priority identifier. A transaction queue is set in the shared memory unit to temporarily store transaction structures generated by the transaction abstraction unit in sequence; An arbitration scheduling unit, connected to the transaction queue, is configured to select transactions from the transaction queue for scheduling according to a preset scheduling strategy, and to simulate the timing behavior of the SoC bus during the scheduling process.
[0008] Furthermore, the preset scheduling strategy includes at least one of the following: a round-robin strategy, a fixed priority strategy, and a bandwidth reservation strategy.
[0009] Furthermore, it also includes: The Direct Memory Access Model (DMI) unit is communicatively connected to the transaction processing module and the shared memory unit. It is configured to run as an independent background thread to receive and execute DMA transfer transactions from the transaction processing module, asynchronously move data between different areas of the shared memory unit, and generate a completion event notification after the transfer is completed.
[0010] Furthermore, the completion event notification includes any of the following methods: Update the preset virtual interrupt status register in the shared memory unit; Call the callback function that was pre-registered by the upper-level software.
[0011] Furthermore, it also includes: The consistency maintenance unit is communicatively connected to the shared memory unit and provides a set of explicit software interfaces, including at least a cache clearing interface, a cache invalidation interface, and a memory barrier interface, for upper-layer software to call in order to simulate and maintain cache consistency and memory order semantics in the virtualization verification system.
[0012] Furthermore, the shared memory unit includes at least one of the following logical partitions: a register mapping area, a data buffer, and a direct memory access buffer.
[0013] Furthermore, it also includes: An executable file loading unit is configured to load an executable file in native format compiled for the actual target SoC hardware and establish a runtime environment for the executable file; The virtual memory output mapping unit is connected to the executable file loading unit and the shared memory unit, and is configured to redirect the memory-mapped input / output addresses defined in the executable file to the register mapping area in the shared memory unit when loading the executable file.
[0014] Furthermore, the executable file is an ELF format file; the executable file loading unit is configured to parse the program header and section information of the ELF format file to establish a process address space.
[0015] Furthermore, it also includes: An integrated debugging unit is communicatively connected to the shared memory unit, the transaction processing module, and the accelerator functional model unit, and is used to monitor and record the system status of the virtualization verification system; the system status includes the transaction structure flowing through the transaction queue, the internal variables of the accelerator functional model unit, and the memory content of the shared memory unit; the integrated debugging unit is also configured with breakpoint setting and single-step execution functions.
[0016] Furthermore, the accelerator functional model unit includes multiple accelerator model instances; the arbitration scheduling unit is also configured to simulate a multi-level or ring-shaped bus interconnect topology to handle concurrent access transactions from multiple master devices.
[0017] To achieve the above objectives, this application also provides a method for verifying accelerator virtualization for SoCs, running on a general-purpose computing platform, employing the accelerator virtualization verification system for SoCs as described above, the method comprising: A unified virtual memory space is constructed within the process address space of the general-purpose computing platform; The access of the main device, represented by the upper-layer software, to the hardware accelerator or memory is abstracted and encapsulated into a transaction structure of a predefined format. The transaction structure is scheduled, and the timing behavior of the SoC bus is simulated during the scheduling process; Based on the scheduling result, the corresponding accelerator function model is triggered to perform hardware acceleration function simulation, and the response result is written to the virtual memory space. The results in the virtual memory space are read to perform functional verification.
[0018] To achieve the above objectives, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.
[0019] The accelerator virtualization verification system and method for SoC provided in this application fundamentally changes the traditional serial mode of software and hardware development by constructing a virtual verification environment based on shared memory and transactional bus simulation. It enables the software (including drivers, operating systems, and applications) for dedicated accelerators to be fully developed and debugged in an environment that simulates the behavior of the target hardware with high precision before the hardware design is completed, realizing true parallel software and hardware development and significantly shortening the product development cycle.
[0020] The accelerator virtualization verification system and method for SoC provided in this application can accurately simulate the timing, bandwidth, latency, and asynchronous transmission behavior in a real SoC bus through a configurable transaction arbitration and scheduling mechanism and an independent asynchronous DMA model. This allows the software to expose hidden errors caused by timing races, resource conflicts, or improper handling of asynchronous events while running on the virtual platform, significantly shifting and mitigating the risks in the later hardware integration stage, and reducing the cost and risk of discovering serious software defects only after chip tape-out.
[0021] The accelerator virtualization verification system and method for SoC provided in this application offer explicit cache coherency interfaces and memory barrier support, forcing software developers to adhere to memory access and consistency maintenance standards consistent with real hardware in the virtual environment. This mechanism effectively avoids cache coherency problems masked by overly idealized virtual platform models, greatly improving the reliability and stability of the final software on real hardware.
[0022] The accelerator virtualization verification system and method for SoC provided in this application reduce the reliance on expensive hardware emulation equipment or FPGA prototype boards, and mainly run on general-purpose x86 server platforms, saving hardware costs and maintenance expenses. At the same time, its feature of supporting direct execution of native object code (ELF) avoids the extra workload of recompiling software for virtual environments, further improving the overall efficiency of the development and verification process.
[0023] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing this application. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the present application and form part of the specification. Together with the embodiments of the present application, they serve to explain the present application but do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the accelerator virtualization verification system for SoC according to Embodiment 1 of this application; Figure 2 A schematic diagram of transaction queues and arbitration scheduling in a virtualization verification system; Figure 3 A schematic diagram of the DMA model and interrupts for a virtualization verification system; Figure 4 A schematic diagram illustrating cache consistency and barriers in a virtualization verification system; Figure 5 A schematic diagram illustrating executable file loading and MMIO access in a virtualization verification system; Figure 6 This is a flowchart illustrating the accelerator virtualization verification method for SoC according to Embodiment 2 of this application; In the diagram: 100 - Shared memory unit, 200 - Transaction processing module, 201 - Transaction abstraction unit, 202 - Transaction queue, 203 - Arbitration scheduling unit, 300 - Accelerator functional model unit, 400 - Direct memory access model unit, 500 - Consistency maintenance unit, 600 - Executable file loading unit, 700 - Virtual memory output mapping unit, 800 - Upper-layer software, 900 - Integration debugging unit. Detailed Implementation
[0025] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0026] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] It should be noted that the terms "one" and "multiple" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless explicitly stated otherwise in the context, they should be understood as "one or more". "Multiple" should be understood as two or more.
[0029] SoC stands for System on a Chip, which is an integrated circuit that integrates a central processing unit (CPU), memory, peripherals, and dedicated processing units onto a single chip.
[0030] Hardware accelerators are modules in a SoC designed for specific computing tasks (such as matrix multiplication GEMM and video encoding / decoding VPU) to efficiently perform specific workloads and improve system energy efficiency.
[0031] A general-purpose computing platform refers to mainstream commercial computing devices used as host machines, such as servers or high-performance personal computers based on x86 or ARM architectures. It provides physical computing, memory, and operating system resources for virtual verification systems.
[0032] Upper-layer software refers to the target software that needs to be verified and runs on top of this virtual verification system. This includes device drivers written for the real target SoC hardware, operating system kernel components, and the final application. In this system, its behavior is considered the origin of hardware access.
[0033] A master device, derived from the concept of a bus architecture, refers to a device capable of actively initiating data transfer requests. In this virtual environment, when upper-layer software performs a memory read / write operation, this behavior is simulated as a bus transaction initiated by a master device (usually the CPU). The DMA controller is also a typical master device.
[0034] In this application, a transaction refers to a complete abstract description of a single hardware access operation. It transforms software-layer read / write requests into a standardized data packet containing all relevant information such as address, data, command, and timing requirements, facilitating scheduling, transmission, and processing within the virtual system.
[0035] DMA and asynchronous transfer: DMA stands for Direct Memory Access, a technology that allows external devices to transfer large blocks of data directly to memory without CPU intervention. Asynchronous transfer means that the execution of the transfer operation is not aligned with the execution of CPU instructions in time. DMA typically notifies the CPU of an interrupt signal upon completion.
[0036] Cache coherence and memory barriers: In modern multi-core processors, the cache is a small, high-speed memory used to accelerate data access. Cache coherence refers to ensuring that all processor cores and master devices like DMA see the same data at the same memory address consistently. A memory barrier is a synchronization instruction used to ensure that all memory access operations preceding the barrier are completed before subsequent accesses can be executed, thus controlling the global visibility of the execution order.
[0037] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0038] Example 1 One embodiment of this application provides an accelerator virtualization verification system for SoCs, running on general-purpose computing platforms such as x86. Figure 1 As shown, it includes: The shared memory unit 100 is located within the process address space of the general-purpose computing platform and is used to construct a unified virtual memory space. In this embodiment, the shared memory unit 100 is one or more contiguous areas in the host physical memory of a general-purpose computing platform that are accessed by all components of the virtual system, serving as the "physical memory" of the entire virtual SoC. Data exchange, state storage, and communication buffering all occur here, including a register mapping area, a data buffer, and a direct memory access (DMA) buffer. The register mapping area 101 maps each register of the virtual accelerator and DMA controller to a specific memory address. The "register read and write" of the upper-layer software 800 to the hardware is actually an access to the address corresponding to this area. The data buffer is used to store the computational data and results to be processed, and the DMA buffer is a source and destination data area dedicated to DMA transfer.
[0039] The transaction processing module 200 is communicatively connected to the shared memory unit and is configured to abstract and encapsulate the access operations of the master device represented by the upper-layer software to the hardware accelerator or memory into a transaction structure of a predetermined format, and to schedule the transaction structure. Transaction processing module 200 includes: The transaction abstraction unit 201 is configured to encapsulate access operations initiated by the master device represented by the upper-layer software into a transaction structure, wherein the transaction structure includes at least the following fields: target address, data length, transmission type, burst mode, transaction identifier, and priority identifier. For example, when upper-layer software (e.g., a driver for an AI application) needs to operate an accelerator, it executes a write-to-memory instruction with the target address being the address of a certain accelerator's "configuration register" (e.g., 0xF000_0000). The transaction abstraction unit 201 encapsulates the operation into a transaction structure.
[0040] A transaction queue 202 is set in the shared memory unit and is used to temporarily store transaction structures generated by the transaction abstraction unit 201 in sequence. It should be noted that the generated transaction structure is stored in the transaction queue 202 within the shared memory unit, and multiple queues with different priorities can be maintained.
[0041] Arbitration scheduling unit 203 is connected to the transaction queue 202 and is configured to select transactions from the transaction queue for scheduling according to a preset scheduling strategy, and to simulate the timing behavior of the SoC bus during the scheduling process.
[0042] In this embodiment, the arbitration scheduling unit 203 is an independent background thread that continuously polls the transaction queue 202. According to a preset scheduling strategy (such as priority-based polling), it retrieves a transaction from the queue and dispatches it. Before dispatching the transaction, it simulates the non-ideal characteristics of a real SoC bus, such as: Delay simulation: Depending on the configuration, the scheduling thread is made to "sleep" for several microseconds to simulate the propagation delay of signals on physical wires and the logic processing delay.
[0043] Bandwidth simulation: By using algorithms such as token bucket, the number of transactions dispatched per unit time is limited to simulate the upper limit of bus bandwidth.
[0044] Competition simulation: When multiple high-priority transactions arrive simultaneously, the arbitration logic is simulated to make some transactions wait in the queue.
[0045] In the embodiments of this application, the preset scheduling strategy includes at least one of the following: a round-robin strategy, a fixed priority strategy, and a bandwidth reservation strategy.
[0046] Understandably, the transaction processing module 200 converts the software's "instantaneous" access request into a hardware communication event with real timing characteristics.
[0047] For example, transaction queues and arbitration scheduling are as follows: Figure 2 As shown, Figure 2 This is a schematic diagram of the transaction queue and arbitration scheduling of a virtualization verification system.
[0048] Accelerator functional model unit 300 is connected to the transaction processing module and the shared memory unit, and is configured to execute the corresponding hardware acceleration function simulation in response to the transaction scheduled by the transaction processing module, and write the response result to the shared memory unit.
[0049] In the embodiments of this application, the accelerator functional model unit 300 is a module implemented in high-level languages such as C / C++ that simulates the computing functions of a specific accelerator (such as GEMM, VPU). It works by listening to or responding to dispatched transactions, performing calculations and writing the results back to the shared memory unit.
[0050] In other embodiments, the accelerator functional model unit further includes multiple accelerator model instances, and the arbitration scheduling unit is further configured to simulate a multi-level or ring-shaped bus interconnect topology to handle concurrent access transactions from multiple master devices; simulating complex timing scenarios when multiple master devices compete for the bus. This is used to verify advanced functions such as multi-core task scheduling and data stream parallelism.
[0051] Through the synergy of the above structures, a verification environment with basic bus communication timing simulation capabilities can be provided even without real hardware RTL code, supporting early software development.
[0052] In the embodiments of this application, it also includes: The Direct Memory Access Model (DMI) unit 400 is communicatively connected to the transaction processing module 200 and the shared memory unit 100. It is configured to run as an independent background thread to receive and execute DMA transfer transactions from the transaction processing module, asynchronously move data between different areas of the shared memory unit 100, and generate a completion event notification after the transfer is completed, so as to perform high-precision DMA behavior simulation.
[0053] It should be noted that the Direct Memory Access Model Unit 400 has its own channels, state machine, and register mapping. The interrupt process of the Direct Memory Access Model Unit 400 is as follows: Figure 3 As shown.
[0054] When the accelerator model or upper-layer software needs to initiate data transfer, a DMA transfer transaction is generated and placed in the transaction queue; the arbitration scheduling unit 203 dispatches the transaction to the direct memory access model unit 400. The direct memory access model unit 400 performs data transfer asynchronously (moving data between different regions of shared memory) and can simulate configurable transfer latency, bandwidth, and uncertainty in transfer completion time.
[0055] In this application embodiment, the direct memory access model unit 400 generates a completion event notification in any of the following ways: Update the preset virtual interrupt status register in the shared memory unit; Call the callback function that was pre-registered by the upper-level software.
[0056] This asynchronous completion event mechanism allows software to wait for DMA to complete, just like on real hardware, using interrupts or polling, thereby enabling the verification of software logic related to DMA timing.
[0057] In this embodiment of the application, the direct memory access model unit 400 is further provided with: The consistency maintenance unit 500 is communicatively connected to the shared memory unit 100 and provides a set of explicit software interfaces, including at least a cache clearing interface, a cache invalidation interface, and a memory barrier interface, for upper-layer software to call in order to simulate and maintain cache consistency and memory order semantics in the virtualization verification system.
[0058] like Figure 4 As shown, when the DMA model needs to read or write data that may be cached by the CPU, the relevant software driver must call the corresponding cache clearing or cache invalidation interface before and after the DMA operation. This forces developers to correctly handle consistency issues in the virtual environment, avoiding the discovery of such hidden errors on real hardware later. This unit can also simulate a weakly consistent memory model, allowing specific out-of-order execution phenomena under certain configurations to verify whether the software's protection of memory order is correct.
[0059] In the embodiments of this application, the following are also provided: The executable file loading unit 600 is configured to load an executable file in native format compiled for the actual target SoC hardware and establish a runtime environment for the executable file; The executable file is an ELF format file; the executable file loading unit is configured to parse the program header and section information of the ELF format file to establish the process address space.
[0060] Native ELF executables contain applications, drivers, operating system kernels, etc., and do not require recompilation for virtual platforms after loading.
[0061] The virtual memory output mapping unit 700 is connected to the executable file loading unit 600 and the shared memory unit 100, and is configured to redirect the memory-mapped input / output addresses defined in the executable file to the register mapping area in the shared memory unit when loading the executable file.
[0062] The Virtual Memory Output Mapping Unit 700 seamlessly redirects memory-mapped I / O (MMIO) addresses in the ELF file to the register mapping area in shared memory. This allows software access to hardware registers to be transparently translated into transactional bus access.
[0063] like Figure 5As shown, the system can directly load native ELF format executable files (containing applications, drivers, operating system kernels, etc.) compiled for real target SoC hardware without recompiling for the virtual platform; it seamlessly redirects memory-mapped I / O (MMIO) addresses in the ELF file to the register mapping area in shared memory. This allows software access to hardware registers to be transparently transformed into transactional bus access as in Example 1.
[0064] The integrated debugging unit 900 is communicatively connected to the shared memory unit 100, the transaction processing module 200, and the accelerator functional model unit 300, and is used to monitor and record the system status of the virtualization verification system; the system status includes the transaction structure flowing through the transaction queue, the internal variables of the accelerator functional model unit, and the memory content of the shared memory unit; the integrated debugging unit is also configured with breakpoint setting and single-step execution functions.
[0065] In the embodiments of this application, since all states are under software control, their observability far exceeds that of real hardware or FPGA prototypes.
[0066] Example 2 One embodiment of this application provides an accelerator virtualization verification method for SoCs, running on a general-purpose computing platform and applied to the aforementioned accelerator virtualization verification system for SoCs. The following will refer to... Figure 6 This application provides a detailed description of the SoC-oriented accelerator virtualization verification method, including: S101: Construct a unified virtual memory space in the process address space of the general-purpose computing platform; S102: Abstract and encapsulate the access of the master device, represented by the upper-layer software, to the hardware accelerator or memory into a transaction structure of a predefined format. S103: Schedule the transaction structure and simulate the timing behavior of the SoC bus during the scheduling process; S104: Based on the scheduling result, trigger the corresponding accelerator function model to perform hardware acceleration function simulation, and write the response result into the virtual memory space; S105: Read the results from the virtual memory space to perform functional verification.
[0067] Example 3 One embodiment of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the SoC-oriented accelerator virtualization verification method as described above.
[0068] The above description is merely a partial embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
[0069] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0070] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An accelerator virtualization verification system for SoC, characterized in that, Shared memory units are located within the process address space of a general-purpose computing platform and are used to construct a unified virtual memory space. The transaction processing module is communicatively connected to the shared memory unit and is configured to abstract and encapsulate the access operations of the master device represented by the upper-layer software to the hardware accelerator or memory into a transaction structure of a predetermined format, and to schedule the transaction structure. The transaction processing module includes: a transaction abstraction unit configured to encapsulate access operations initiated by the master device represented by the upper-layer software into the transaction structure, wherein the transaction structure includes at least the following fields: target address, data length, transmission type, burst mode, transaction identifier, and priority identifier; a transaction queue set in the shared memory unit for sequentially storing the transaction structures generated by the transaction abstraction unit; and an arbitration scheduling unit connected to the transaction queue, configured to select transactions from the transaction queue for scheduling according to a preset scheduling strategy, and simulate the timing behavior of the SoC bus during the scheduling process. An accelerator function model unit is connected to the transaction processing module and the shared memory unit, and is configured to execute the corresponding hardware acceleration function simulation in response to a transaction scheduled by the transaction processing module, and write the response result to the shared memory unit.
2. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, The preset scheduling strategy includes at least one of the following: a round-robin strategy, a fixed priority strategy, and a bandwidth reservation strategy.
3. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, Also includes: The Direct Memory Access Model (DMI) unit is communicatively connected to the transaction processing module and the shared memory unit. It is configured to run as an independent background thread to receive and execute DMA transfer transactions from the transaction processing module, asynchronously move data between different areas of the shared memory unit, and generate a completion event notification after the transfer is completed.
4. The accelerator virtualization verification system for SoC according to claim 3, characterized in that, The completion event notification includes any of the following methods: Update the preset virtual interrupt status register in the shared memory unit; Call the callback function that was pre-registered by the upper-level software.
5. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, Also includes: The consistency maintenance unit is communicatively connected to the shared memory unit and provides a set of explicit software interfaces, including at least a cache clearing interface, a cache invalidation interface, and a memory barrier interface, for upper-layer software to call in order to simulate and maintain cache consistency and memory order semantics in the virtualization verification system.
6. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, The shared memory unit includes at least one of the following logical partitions: register mapping area, data buffer, and direct memory access buffer.
7. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, Also includes: An executable file loading unit is configured to load an executable file in native format compiled for the actual target SoC hardware and establish a runtime environment for the executable file; The virtual memory output mapping unit is connected to the executable file loading unit and the shared memory unit, and is configured to redirect the memory-mapped input / output addresses defined in the executable file to the register mapping area in the shared memory unit when loading the executable file.
8. The accelerator virtualization verification system for SoC according to claim 7, characterized in that, The executable file is an ELF format file; the executable file loading unit is configured to parse the program header and section information of the ELF format file to establish the process address space.
9. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, Also includes: An integrated debugging unit is communicatively connected to the shared memory unit, the transaction processing module, and the accelerator functional model unit, and is used to monitor and record the system status of the virtualization verification system; the system status includes the transaction structure flowing through the transaction queue, the internal variables of the accelerator functional model unit, and the memory content of the shared memory unit; the integrated debugging unit is also configured with breakpoint setting and single-step execution functions.
10. The accelerator virtualization verification system for SoC according to claim 1, characterized in that, The accelerator functional model unit includes multiple accelerator model instances; the arbitration scheduling unit is also configured to simulate a multi-level or ring-shaped bus interconnect topology to handle concurrent access transactions from multiple master devices.
11. A method for verifying accelerator virtualization for SoCs, running on a general-purpose computing platform, employing the accelerator virtualization verification system for SoCs as described in any one of claims 1-10, the method comprising: A unified virtual memory space is constructed within the process address space of the general-purpose computing platform; The access of the main device, represented by the upper-layer software, to the hardware accelerator or memory is abstracted and encapsulated into a transaction structure of a predefined format. The transaction structure is scheduled, and the timing behavior of the SoC bus is simulated during the scheduling process; Based on the scheduling result, the corresponding accelerator function model is triggered to perform hardware acceleration function simulation, and the response result is written to the virtual memory space. The results in the virtual memory space are read to perform functional verification.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in claim 11.
Citation Information
Patent Citations
System-on-chip simulation method and system based on virtual machine
CN116401984A