DPI Multithreaded Simulation Acceleration Method, Storage Medium

The DPI multi-threaded simulation acceleration method addresses performance issues by using a thread pool and asynchronous execution to improve simulation efficiency in chip verification.

CN114721782BActive Publication Date: 2025-07-15YUNHE ZHIWANG (SHANGHAI) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210418234.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-07-15
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

In the chip verification process, when calling the C model method through the DPI interface, the simulation process is frequently interrupted and waited, resulting in low simulation efficiency.

Method used

The thread pool class management C model method is adopted, and the C++ std::future and std::queue class templates are used to realize multi-threaded parallel execution of the C model method and asynchronous result return, which are decomposed into two asynchronous processes predictor and evaluator for calculation and comparison.

Benefits of technology

By executing the C model method in parallel with multithreading, the interrupt waiting time of the simulation process is reduced and the simulation efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114721782B_ABST
    Figure CN114721782B_ABST
Patent Text Reader

Abstract

The present invention discloses a DPI multi-threaded simulation acceleration method and a storage medium. Among them, a DPI multi-threaded simulation acceleration method uses C++ advanced template class methods in combination with the UVM verification methodology. According to the needs of the project, it innovatively divides the original method of directly calling the C model method through the DPI interface into two asynchronous processes. That is, one process is that the predictor component calls the predict_call_by_predictor method to send the C model method to be called into the thread pool for parallel operation, and the other process is that the evaluator component calls the get_result_call_by_evaluator method to retrieve the results completed in the thread pool for comparison and analysis. The present invention avoids the problem of low simulation efficiency caused by interrupting the simulation and waiting every time the C model method is called by using multi-threaded parallel execution of the C model method, and finally realizes the acceleration of the simulation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip verification, and particularly relates to a DPI multi-threaded simulation acceleration method and a storage medium. Background Art

[0002] Generally, the verification work of a chip starts after the designer provides a mature, stable and self-tested RTL design code.

[0003] A typical architecture based on the UVM verification platform is as Figure 1 shown. As can be seen from the figure, generally, chip verification personnel will complete the connection between the RTL design (the green part DUT in the figure) provided by the designer and the verification environment on the left through the interface. Then, the verification personnel will extract the functional characteristics of the corresponding RTL design (i.e., DUT) according to the design document, and then write corresponding test cases according to the functional characteristic requirements and perform simulation verification.

[0004] Generally, we will use a scoreboard to check whether the behavior function of the DUT meets the expectations. It is one of the components of the above-mentioned UVM verification platform and is derived from the uvm_scoreboard class.

[0005] As Figure 2 shown, it consists of two parts:

[0006] (1) Predictor, that is, the reference model, which is used to complete the same function as the DUT.

[0007] By subscribing to the analysis_port of the monitor, the transaction data incentives sent to the DUT and the transaction data results output by the DUT are obtained, and then the incentives are applied to the reference model to generate the expected results, and then the actual output results of the DUT are analyzed and compared. Briefly speaking, here the reference model predictor and the DUT receive the same test incentives, and after the operation is completed, the respective output results are sent to the evaluator for analysis and comparison to help verify the correctness of the DUT function.

[0008] (2) Evaluator, that is, the part used to compare the expected value and the actual value and output the result.

[0009] The Predictor doesn't have a dedicated base class and is generally derived from uvm_component. Usually, the interface method for calculating the expected value can be written in C, C++, SystemVerilog, or SystemC. Generally, for algorithmic chips, C is used to write the reference model, and then the DPI (Direct Programming Interface) of SystemVerilog is used to call the interface method in the reference model written in C to calculate the expected value. Figure 3 It is the implementation structure of the existing solution.

[0010] The following are the specific implementation steps of the existing solution:

[0011] First step, write a C function method to implement the reference model (c_model) for the components on the SystemVerilog side in the verification environment, namely the Predictor, to call. Two points need to be noted during the process:

[0012] (1) The mapping relationship between data types in C and SystemVerilog. For example, in this example, the mapping conversion between the svBitVecVal data type on the C side and the bit[31:0] data type on the SystemVerilog side. In addition, the mapping of port data types also needs to be noted:

[0013] Input port: const svBitVecVal data → input bit[31:0] data

[0014] Output port: svBitVecVal* result → output bit[31:0] result

[0015] (2) Import the interface file svdpi.h for SystemVerilog and C.

[0016] Second step, in the Predictor, obtain the input stimuli monitored by the Monitor, then call the C function for calculating the expected value to calculate the expected result, and after the calculation is completed, send it to the Evaluator through the TLM communication port for analysis and comparison.

[0017] In the third step, obtain the actual output result monitored by the monitor at the output port of the DUT in the evaluator, and obtain the expected result completed by the c_model operation called by the predictor in the previous step, and then compare them to determine the correctness of the DUT function. The above existing solution is feasible, but when calling the interface method of the reference model written in C language, the simulation performance will decrease.

[0018] This is because after calling the C interface method of the reference model, the simulation tool will stop and wait for the result returned by the called interface method operation. Only after obtaining the returned operation result can the code in the verification environment be continued to execute to continue the subsequent simulation. Then, if many such C interface methods need to be called during the simulation process, then it is necessary to interrupt and wait for the return of the operation result multiple times during the entire simulation process. Therefore, the entire simulation process will become slower. Summary of the Invention

[0019] According to an embodiment of the present invention, a DPI multi-threaded simulation acceleration method is provided, including the following steps:

[0020] Create a thread pool class, instantiate the thread pool in the thread pool class declaration, and provide a get_instance interface method for components in the external verification platform to obtain;

[0021] Use the std::future and std::queue class templates and their methods provided by C++ to manage the C model methods to be called through the thread pool;

[0022] Perform operations and return on the asynchronous results of the C model methods through the mechanism provided by the std::future class template to access the asynchronous operation results;

[0023] Create a predictor, and the predictor obtains the monitored input stimulus;

[0024] Create a predict_call_by_predictor method, import it into the predictor through the DPI interface, and call the predict_call_by_predictor method to obtain the expected result;

[0025] Create an evaluator, and the evaluator obtains the actual output result monitored from the output port of the DUT;

[0026] Create a get_result_call_by_evaluator method, import it into the evaluator through the DPI interface, call the get_result_call_by_evaluator method, obtain the expected result and compare it with the actual output result to determine the correctness of the DUT function.

[0027] Furthermore, the thread pool class instantiates the thread pool using the singleton pattern.

[0028] Furthermore, using the std::future and std::queue class templates and their methods provided by C++, the C model methods to be called are managed through a thread pool, which includes the following sub-steps:

[0029] Add tasks that need to call C model method operations to the queue tasks declared by the std::queue class template, that is, add them to the thread pool to be executed in parallel within multiple threads;

[0030] Pass the C model method to be called as a parameter through the add_job method of the std::future class template.

[0031] Furthermore, the asynchronous results of the C model methods are calculated and returned through the mechanism provided by the std::future class template to access the asynchronous operation results, which includes the following sub-steps:

[0032] Use the std::queue class template to provide a queue;

[0033] Call the emplace method provided by the queue to write the asynchronous thread into the queue;

[0034] Retrieve the operation return result from the queue tasks through the queue method of the std::queue class template.

[0035] According to another embodiment of the present invention, a storage medium with non-volatile program code executable by a processor is provided, and the program code causes the processor to run the DPI multithreaded simulation acceleration method.

[0036] The DPI multi-threaded simulation acceleration method according to the embodiments of the present invention avoids the problem of low simulation efficiency caused by interrupting the simulation and waiting every time the C model method is called by using multi-threads to execute the C model method in parallel, and finally realizes the acceleration of the simulation process. Originally, each time the C model method was called, it was necessary to wait for the operation to complete and return the result. The improved solution can realize the multi-threaded parallel execution of the C model method for multiple calls and asynchronously return the operation result, thereby reducing the interrupt waiting time of the original solution and improving the simulation efficiency. The patent solution uses the C++ advanced template class method combined with the UVM verification methodology. According to the needs of the project, it innovatively divides the original direct call of the C model method through the DPI interface into two asynchronous processes. One process is that the predictor component calls the predict_call_by_predictor method to send the C model method to be called into the thread pool for parallel operation, and the other process is that the evaluator component calls the get_result_call_by_evaluator method to retrieve the result of the operation completed in the thread pool for comparison and analysis.

[0037] It is to be understood that both the foregoing general description and the following detailed description are exemplary and are intended to provide further explanation of the claimed technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 FIG. is a typical architecture diagram of a UVM-based verification platform.

[0039] Figure 2 FIG. is a composition structure diagram of a scoreboard.

[0040] Figure 3 FIG. is a composition structure diagram of the scoreboard of the existing solution.

[0041] Figure 4 FIG. is a flowchart of the DPI multi-threaded simulation acceleration method according to the embodiments of the present invention.

[0042] Figure 5 FIG. is a flowchart of the first sub-step of the DPI multi-threaded simulation acceleration method according to the embodiments of the present invention.

[0043] Figure 6 FIG. is a flowchart of the second sub-step of the DPI multi-threaded simulation acceleration method according to the embodiments of the present invention.

[0044] Figure 7 FIG. is a structure diagram of the DPI multi-threaded simulation acceleration method according to the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The following will further elaborate on the present invention by describing in detail the preferred embodiments of the present invention in conjunction with the accompanying drawings.

[0046] First, this will be combined with Figures 4 - 7 Describe the DPI multi-threaded simulation acceleration method according to an embodiment of the present invention, which is used for the verification of chips and has a wide range of application scenarios.

[0047] As Figures 4 - 7 shown, the DPI multi-threaded simulation acceleration method according to an embodiment of the present invention includes the following steps:

[0048] In S1, as Figure 4 shown, create a thread pool thread_pool class. The thread pool thread_pool class declares and instantiates the thread pool thread_pool, and provides a get_instance interface method for components in the external verification platform to obtain. In this embodiment, further, the thread pool thread_pool class is declared and instantiated using the singleton pattern singleton.

[0049] In S2, as Figure 4 shown, use the std::future and std::queue class templates and their methods provided by C++ to manage the C model methods to be called through the thread pool thread_pool. At this time, there is no need to interrupt the simulation and wait for its operation to complete every time a C model method is called, because at this time all C model methods will be executed in parallel within multiple threads, thereby improving the simulation efficiency. And the C model method is used to predict the expected result value for comparison with the actual result value of the design under test.

[0050] Further, using the std::future and std::queue class templates and their methods provided by C++ to manage the C model methods to be called through the thread pool thread_pool includes the following sub-steps:

[0051] In S21, as Figure 5 shown, add the tasks that need to call the C model method operations to the queue task c_tasks declared by the std::queue class template, that is, add them to the thread pool thread_pool to be executed in parallel within multiple threads.

[0052] In S22, as Figure 5 shown, pass the C model method to be called as a parameter through the add_job method of the std::future class template.

[0053] In S3, as Figure 4 shown, perform operations and return on the asynchronous results of the C model method through the mechanism provided by the std::future class template to access the asynchronous operation results.

[0054] Furthermore, the asynchronous result of the C model method is calculated and returned by using the mechanism provided by the std::future class template to access the asynchronous operation result, including the following sub-steps:

[0055] In S31, as Figure 6 shown, a queue is provided using the std::queue class template.

[0056] In S32, as Figure 6 shown, the emplace method provided by the queue is called to write the asynchronous thread into the queue.

[0057] In S33, as Figure 6 shown, the result of calculation and return is retrieved from the queue task c_tasks through the queue method of the std::queue class template.

[0058] In S4, as Figure 4 shown, a predictor is created, and the predictor obtains the monitored input excitation;

[0059] In S5, as Figure 4 shown, a predict_call_by_predictor method is created and imported into the predictor through the DPI interface, and the predict_call_by_predictor method is called to obtain the expected result;

[0060] In S6, as Figure 4 shown, an evaluator is created, and the evaluator obtains the actual output result monitored from the output port of the DUT;

[0061] In S7, as Figure 4 shown, a get_result_call_by_evaluator method is created and imported into the evaluator through the DPI interface, and the get_result_call_by_evaluator method is called to obtain the expected result and compare it with the actual output result to judge the correctness of the DUT function.

[0062] According to another embodiment of the present invention, a storage medium with non-volatile program code executable by a processor is provided, and the program code causes the processor to run the DPI multithreaded simulation acceleration method.

[0063] Above, with reference to Figures 4 - 7A DPI multi-threaded simulation acceleration method according to an embodiment of the present invention is described. By using multi-threading to execute the C model method in parallel, the problem of low simulation efficiency caused by interrupting the simulation and waiting every time the C model method is called is avoided, and finally the acceleration of the simulation process is achieved. Previously, each time the C model method was called, it was necessary to wait for the operation to complete and return the result. The improved solution can implement the multi-threaded parallel execution of multiple calls to the C model method and asynchronously return the operation result, thereby reducing the interrupt waiting time of the original solution and improving the simulation efficiency. The patent solution uses the C++ advanced template class method combined with the UVM verification methodology. According to the needs of the project, the original direct call to the C model method through the DPI interface is innovatively split into two asynchronous processes. One process is that the predictor component calls the predict_call_by_predictor method to send the C model method to be called into the thread pool for parallel operation, and the other process is that the evaluator component calls the get_result_call_by_evaluator method to retrieve the result of the operation completed in the thread pool for comparison and analysis.

[0064] It should be noted that in this specification, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0065] Although the content of the present invention has been described in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present invention. After those skilled in the art have read the above content, various modifications and alternatives to the present invention will be obvious. Therefore, the protection scope of the present invention should be defined by the appended claims.

Claims

1. A DPI multi-threaded simulation acceleration method, characterized in that, It includes the following steps: Create a thread pool class, which instantiates the thread pool and provides a get_instance interface method for components in the external verification platform to obtain; Use the std::future and std::queue class templates and their methods provided by C++ to manage the C model methods to be called through the thread pool; Operate and return the asynchronous results of the C model methods through the mechanism provided by the std::future class template to access the asynchronous operation results; Create a predictor, which obtains the monitored input stimuli; Create a predict_call_by_predictor method and import it into the predictor through the DPI interface, and call the predict_call_by_predictor method to obtain the expected result; Create an evaluator, which obtains the actual output results monitored from the output port of the DUT; Create a get_result_call_by_evaluator method and import it into the evaluator through the DPI interface, call the get_result_call_by_evaluator method, obtain the expected result and compare it with the actual output result to judge the correctness of the DUT function; The use of the std::future and std::queue class templates and their methods provided by C++ to manage the C model methods to be called through the thread pool includes the following sub-steps: Add the tasks that need to call the C model method operations to the queue tasks declared by the std::queue class template, that is, add them to the thread pool to be executed in parallel within multiple threads; Pass the C model method to be called as a parameter through the add_job method of the std::future class template; The operation and return of the asynchronous results of the C model methods through the mechanism provided by the std::future class template to access the asynchronous operation results includes the following sub-steps: Use the std::queue class template to provide a queue; Call the emplace method provided by the queue to write the asynchronous thread into the queue; Take out the operation return result from the queue tasks through the queue method of the std::queue class template.

2. The DPI multi-threaded simulation acceleration method according to claim 1, wherein The thread pool class instantiates the thread pool using the singleton pattern.

3. A storage medium having non-volatile program code executable by a processor, characterized in that, The program code causes the processor to run the DPI multi-threaded simulation acceleration method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Simulator multithread running method using PERL scripts

    CN104899369A

  • Verification method and device and related product

    CN110543395A