A Method and System for Fast and Accurate Co-Simulation of a Coprocessor

By quickly and accurately simulating the execution of the real CPU on the FPGA side, and combining the FPGA monitoring program on the host side, the contradiction between coprocessor algorithm design and CPU collaborative execution in simulation execution speed and accuracy is solved, and the requirement of rapid iterative design of coprocessor algorithm and CPU design is realized.

CN119249989BActive Publication Date: 2025-05-30WUXI CORE FIELD MICROELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411766646.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-30
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The existing technology is difficult to support the rapid iterative design of coprocessor algorithms and the rapid collaborative execution of CPU design while ensuring operational accuracy and speed, resulting in a contradiction between simulation execution speed and accuracy.

Method used

By designing the FPGA side to quickly and accurately simulate the execution of the real CPU, and running the FPGA monitoring program with the host side, it realizes the interface output of CPU instruction execution track, performance parameters and coprocessor execution instructions. When there is an instruction to use the coprocessor, an analog request is sent to the host side and the CPU is suspended at the appropriate time. After receiving the host simulation operation results, it continues to execute from the breakpoint.

Benefits of technology

It realizes the performance of different coprocessor implementation solutions in the early stage of design, thereby supporting the rapid iterative design of coprocessor algorithms and the rapid collaborative execution of CPU design to meet the needs of rapid modification and rapid execution evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119249989B_ABST
    Figure CN119249989B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for fast and accurate co-simulation of a coprocessor. The system includes: an FPGA side, which is used to place the target CPU hardware implementation, including an FPGA development board deployed with the target CPU hardware design, a CPU control module, an instruction execution trace acquisition module, a coprocessor request processing module, and a communication interface module; a host side, which is communicatively connected to the FPGA development board through a high-speed interface of the FPGA development board and is used to place an FPGA monitoring program, a target algorithm program, and an instruction execution trace acquisition program. Through the designs of the FPGA side and the host side, the performance of the target algorithm under different coprocessor implementation schemes can be quickly obtained, so that the advantages and disadvantages of different design schemes can be evaluated in the early stage of design, and rapid iteration can be carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of CPU design and coprocessor algorithms, and specifically to a method and system for fast and accurate co-simulation of coprocessors, mainly used to support the evaluation of CPU design, coprocessor algorithms, or dedicated algorithm instructions. Background Art

[0002] A CPU coprocessor is a hardware algorithm implementation circuit developed on the coprocessor extension interface provided by CPU design. It is closely combined with CPU design to jointly complete computing / processing tasks. The main purpose of designing a coprocessor is to improve the computing / processing ability of the CPU chip in certain specific fields or dedicated scenarios.

[0003] During the design process of the coprocessor, the algorithm needs to be coordinated with the CPU design to jointly execute the program. The general process is as follows:

[0004] The CPU sends instructions and operands to the coprocessor, waits for the coprocessor to complete the execution of the corresponding instructions, and returns the results and completion status to the CPU. For complex coprocessor algorithms, in the early stage of design, the algorithm may be a program described in a high-level language such as C or Python, without a mature circuit structure. However, it is necessary to evaluate the execution effect of the CPU program with the coprocessor in a timely manner. Therefore, the coprocessor algorithm design platform needs to support rapid iteration of algorithm design and be able to quickly execute to obtain feedback on the effect, so as to facilitate exploring the algorithm design space.

[0005] The existing methods for the coordinated operation of coprocessor algorithms and CPU design are mainly divided into two categories:

[0006] 1. Incorporate the coprocessor algorithm into software tools such as CPU instruction simulators and microarchitecture simulators, and analyze by running software simulation programs to obtain execution data;

[0007] 2. Combine with the actual CPU design (usually Verilog / SytemVerilog), connect the coprocessor design with the interface provided by the CPU design, and analyze by running programs in the CPU design verification environment to obtain data.

[0008] For the first method, since software tools such as CPU instruction simulators and microarchitecture simulators cannot be accurately matched with the actual CPU design, it is difficult to obtain the accurate CPU instruction running beats after adding the coprocessor / accelerator, thus affecting the accuracy of analysis and evaluation. And if an instruction simulator / microarchitecture simulator that matches the actual CPU design is to be obtained, it will take a lot of manpower and time for calibration, with a low input-output ratio. At the same time, as more design details are added to the instruction simulator / microarchitecture simulator, its running speed will decrease and slow down until it loses its practical significance.

[0009] For the second method, if the CPU test stimuli / programs are run in the EDA soft simulation environment, the speed is very slow and it is difficult to meet the requirements of rapid modification and rapid execution evaluation of the coprocessor algorithm. If the entire design is placed in the FPGA prototype to increase the running speed, on the one hand, there is no detailed circuit design in the early stage of the coprocessor algorithm design, making it difficult to generate the FPGA circuit. On the other hand, even if the algorithm structure is described using HDL (such as Verilog), the time cycle for iterative modification and generation of the FPGA circuit is also very long, and it is also difficult to meet the requirements of rapid modification and rapid execution evaluation of the coprocessor algorithm.

[0010] The main reason for the above problems is that the existing method ideas are difficult to solve the contradiction between the coprocessor algorithm design and the CPU co-execution in terms of simulation execution speed and accuracy.

[0011] The existing practices are either to reduce the CPU model accuracy in exchange for execution speed, or to require the coprocessor to complete a detailed design to be added to the FPGA for operation, at the cost of lengthening the design iteration cycle in exchange for high-precision operation.

[0012] Therefore, a method needs to be proposed to support the rapid iterative design of the coprocessor algorithm and the rapid co-execution with the CPU design on the premise of ensuring operation accuracy / accuracy and operation speed. Summary of the Invention

[0013] Object of the Invention: To provide a method and system for rapid and accurate co-simulation of a coprocessor to solve the above problems existing in the prior art.

[0014] Technical Solution: A method for rapid and accurate co-simulation of a coprocessor includes:

[0015] Step 1, lead out an instruction execution trace output interface, a performance statistics output interface, and a coprocessor execution interface for the target CPU, and design a CPU control module, an instruction execution trace acquisition module, a coprocessor request processing module, and a communication interface module;

[0016] Step 2, synthesize the target CPU hardware design, the CPU control module, the instruction execution trace acquisition module, the coprocessor request processing module, and the communication interface module and deploy them to the FPGA development board;

[0017] Step 3, run the FPGA monitoring program on the host side, load the main test program into the target CPU through the CPU control module, and start the execution;

[0018] Step 4, during the execution of the test program by the target CPU, the instruction execution trace acquisition module tracks the running situation of the CPU in real time and saves the instruction execution trace;

[0019] Step 5, the coprocessor request processing module records the target algorithm instructions received during the target CPU executing the test program, and initiates a simulation request to the host side through the communication interface module;

[0020] Step 6: After receiving the coprocessor call request, the host uses the pre-deployed target algorithm program to perform simulation execution, and returns the running result and the number of beats to be run to the coprocessor request processing module through the high-speed interface;

[0021] Step 7: After receiving the simulated execution result of the target algorithm instruction, the coprocessor request processing module feeds back the result to the target CPU according to the number of beats to be run. The target CPU writes the result back and continues execution from the breakpoint.

[0022] Step 8: After the FPGA test program is executed, the host reads the instruction execution trace information and performance statistics through the communication interface.

[0023] Designers can continue to optimize and iterate the target algorithm based on program execution and performance simulation;

[0024] By designing the FPGA end, the execution of the real CPU can be quickly and accurately simulated. At the same time, the CPU instruction execution trajectory, performance parameters and coprocessor execution instructions are brought out in the form of an interface. When there is an instruction using the coprocessor, a simulation request is sent to the host end and paused at the appropriate time. After the pause, the target CPU clock cycle will no longer increase, and the PC will no longer change. After receiving the host simulation operation result, it will continue to execute from the breakpoint.

[0025] The host-side FPGA monitoring program is designed to schedule the main test program to the target CPU on the FPGA side for execution, monitor and capture the coprocessor call requests issued by the target CPU, simulate the execution using the target algorithm program and send the calculation results back to the FPGA side, and collect the instruction execution trace and operation performance parameters of the target CPU.

[0026] The performance of the target algorithm under different coprocessor implementation schemes can then be quickly obtained, so that the pros and cons of different design schemes can be evaluated in the early stages of the design and rapid iteration can be performed.

[0027] In a further embodiment, the CPU control module is responsible for loading instructions to the target CPU, providing clock / reset signals, controlling the startup and debugging of the target CPU, and providing a read and write interface for the target CPU internal control status register.

[0028] In a further embodiment, the instruction execution trace collection module is used to track the operation of the CPU in real time and collect the instruction execution trace.

[0029] In a further embodiment, the co-processor request processing module is configured to intercept co-processor algorithm instructions emitted by the target CPU during the execution of a test program, send a target algorithm simulation request to the host side, receive the operation result of the target algorithm program on the host side at the same time, and feedback the operation result to the target CPU.

[0030] In a further embodiment, the communication interface module is configured to convert the simulation request sent by the co-processor request processing module into an interface signal according to the communication protocol and send it to the host side, and at the same time parse the read / write request of the control status register and the algorithm simulation result sent by the host side, and send the read / write request of the control status register and the algorithm simulation result to the CPU control module, the instruction execution trace acquisition module, and the co-processor request processing module for processing respectively.

[0031] In a further embodiment, step 4 further includes:

[0032] When the target CPU encounters a target algorithm instruction, information related to the target algorithm instruction is extracted and sent to the co-processor request processing module, and at the same time, the operation of the target CPU is paused at an appropriate time according to the number of cycles to be executed by the target algorithm instruction. The host side can view the operation status of the CPU at any time through the CPU control module.

[0033] A co-processor fast and accurate co-simulation system, comprising:

[0034] The FPGA side is used to place the hardware implementation of the target CPU, including an FPGA development board deployed with the hardware design of the target CPU, a CPU control module, an instruction execution trace acquisition module, a co-processor request processing module, and a communication interface module;

[0035] The host side is communicatively connected to the FPGA development board through the high-speed interface of the FPGA development board, and is used to place the FPGA monitoring program, the target algorithm program, and the instruction execution trace acquisition program.

[0036] Beneficial effects: The present invention discloses a method and system for fast and accurate co-simulation of a co-processor. By designing the FPGA side to quickly and accurately simulate the execution of a real CPU, and at the same time extracting the CPU instruction execution trace, performance parameters, and co-processor execution instructions in the form of an interface. When an instruction uses the co-processor, a simulation request is sent to the host side and paused at an appropriate time. After pausing, the target CPU clock cycle no longer increases, and the PC also no longer changes. When the host simulation operation result is received, it continues to execute from the breakpoint;

[0037] By designing an FPGA monitoring program to run on the host side, it is responsible for scheduling the main test program to the target CPU on the FPGA side, monitoring and capturing the coprocessor call requests sent by the target CPU, using the target algorithm program to simulate the execution and then sending the operation results back to the FPGA side, and at the same time collecting the instruction execution traces and running performance parameters of the target CPU;

[0038] Furthermore, the performance of the target algorithm under different coprocessor implementation schemes can be quickly obtained, so that the advantages and disadvantages of different design schemes can be evaluated in the early stage of design, and rapid iteration can be carried out. Brief Description of the Drawings

[0039] Figure 1 It is a schematic diagram of the system structure of the present invention.

[0040] Figure 2 It is a schematic diagram of the prior art structure of the present invention. Detailed Embodiments

[0041] This application relates to a method and system for fast and accurate co-simulation of a coprocessor. The existing methods for co-running a coprocessor algorithm and a CPU design are mainly divided into two categories:

[0042] 1. Incorporate the coprocessor algorithm into software tools such as CPU instruction simulators and microarchitecture simulators, and analyze the execution data by running software simulation programs;

[0043] 2. Combine with the actual CPU design (usually Verilog / SytemVerilog), dock the interface provided by the coprocessor design with the CPU design, and run programs in the CPU design verification environment to obtain data for analysis.

[0044] For the first method, since software tools such as CPU instruction simulators and microarchitecture simulators cannot accurately match the actual CPU design, it is difficult to obtain the accurate CPU instruction running cycle count after adding a coprocessor / accelerator, thus affecting the accuracy of analysis and evaluation. And if an instruction simulator / microarchitecture simulator that matches the actual CPU design is to be obtained, it will take a lot of manpower and time for calibration, with a low input-output ratio. At the same time, as more design details are added to the instruction simulator / microarchitecture simulator, its running speed will decrease and slow down until it loses its practical significance.

[0045] For the second method, if the CPU test stimuli / programs are run in an EDA soft simulation environment, the speed is very slow, making it difficult to meet the requirements of rapid modification and rapid execution evaluation of the coprocessor algorithm; if the entire design is placed in an FPGA prototype to increase the running speed, firstly, there is no detailed circuit design in the early stage of coprocessor algorithm design, making it difficult to generate FPGA circuits; secondly, even if the algorithm structure is described using HDL (such as Verilog), the time period for iterative modification and generation of FPGA circuits is also very long, and it is also difficult to meet the requirements of rapid modification and rapid execution evaluation of the coprocessor algorithm.

[0046] The main reason for the above problems is that the existing method ideas are difficult to solve the contradiction between the coprocessor algorithm design and CPU collaborative execution in terms of simulation execution speed and accuracy.

[0047] The existing practices are either to reduce the CPU model accuracy in exchange for execution speed, or to require the coprocessor to complete a detailed design to be added to the FPGA for operation, sacrificing the design iteration cycle in exchange for high-precision operation.

[0048] Therefore, a method needs to be proposed that supports rapid iterative design of the coprocessor algorithm and rapid collaborative execution with the CPU design on the premise of ensuring operation accuracy / accuracy and operation speed;

[0049] The present invention designs the FPGA side to quickly and accurately simulate the execution of a real CPU, and at the same time leads out the CPU instruction execution trace, performance parameters, and coprocessor execution instructions in the form of an interface. When an instruction uses the coprocessor, a simulation request is sent to the host side and paused at an appropriate time. After pausing, the target CPU clock cycle no longer increases, and the PC also no longer changes. When the host simulation operation result is received, it continues to execute from the breakpoint;

[0050] By designing the host side to run the FPGA monitoring program, which is responsible for scheduling the main test program to the target CPU on the FPGA side, monitoring and capturing the coprocessor call requests sent by the target CPU, simulating and executing using the target algorithm program and then sending the operation result back to the FPGA side, and at the same time collecting the instruction execution trace and operation performance parameters of the target CPU;

[0051] Furthermore, the performance of the target algorithm under different coprocessor implementation schemes can be quickly obtained, so that the advantages and disadvantages of different design schemes can be evaluated in the early stage of design for rapid iteration.

[0052] The following will be explained in detail through specific implementation manners.

[0053] A method for rapid and accurate co-simulation of a coprocessor, comprising:

[0054] Step 1: Draw out the instruction execution trace output interface, performance statistics output interface and coprocessor execution interface for the target CPU, and design the CPU control module, instruction execution trace acquisition module, coprocessor request processing module and communication interface module;

[0055] Step 2: Deploy the target CPU hardware design, CPU control module, instruction execution trace acquisition module, coprocessor request processing module and communication interface module to the FPGA development board after integration;

[0056] Step 3: The host runs the FPGA monitoring program, loads the main test program into the target CPU through the CPU control module, and starts execution;

[0057] Step 4: When the target CPU executes the test program, the instruction execution trajectory acquisition module tracks the CPU operation status in real time and saves the instruction execution trajectory;

[0058] Step 5, the coprocessor request processing module records the target algorithm instructions received during the target CPU executing the test program, and initiates a simulation request to the host side through the communication interface module;

[0059] Step 6: After receiving the coprocessor call request, the host uses the pre-deployed target algorithm program to perform simulation execution, and returns the running result and the number of beats to be run to the coprocessor request processing module through the high-speed interface;

[0060] Step 7: After receiving the simulated execution result of the target algorithm instruction, the coprocessor request processing module feeds back the result to the target CPU according to the number of beats to be run. The target CPU writes the result back and continues execution from the breakpoint.

[0061] Step 8: After the FPGA test program is executed, the host reads the instruction execution trace information and performance statistics through the communication interface.

[0062] Designers can continue to optimize and iterate the target algorithm based on program execution and performance simulation;

[0063] By designing the FPGA end, the execution of the real CPU can be quickly and accurately simulated. At the same time, the CPU instruction execution trajectory, performance parameters and coprocessor execution instructions are brought out in the form of an interface. When there is an instruction using the coprocessor, a simulation request is sent to the host end and paused at the appropriate time. After the pause, the target CPU clock cycle will no longer increase, and the PC will no longer change. After receiving the host simulation operation result, it will continue to execute from the breakpoint.

[0064] By designing an FPGA monitoring program to run on the host side, it is responsible for scheduling the main test program to the target CPU on the FPGA side, monitoring and capturing the coprocessor call requests sent by the target CPU, using the target algorithm program to simulate the execution and then sending the operation results back to the FPGA side, and at the same time collecting the instruction execution traces and running performance parameters of the target CPU;

[0065] Furthermore, the performance of the target algorithm under different coprocessor implementation schemes can be quickly obtained, so that the advantages and disadvantages of different design schemes can be evaluated in the early stage of design and rapid iteration can be carried out.

[0066] The CPU control module is responsible for loading instructions to the target CPU, providing clock / reset signals, controlling the startup and debugging of the target CPU, and providing read / write interfaces for the internal control status registers of the target CPU.

[0067] The instruction execution trace acquisition module is used to track the running situation of the CPU in real time and collect instruction execution traces.

[0068] The coprocessor request processing module is used to intercept the coprocessor algorithm instructions emitted by the target CPU during the execution of the test program, send a target algorithm simulation request to the host side, and at the same time receive the running results of the target algorithm program on the host side and feedback the running results to the target CPU.

[0069] The communication interface module is used to convert the simulation request sent by the coprocessor request processing module into an interface signal according to the communication protocol and send it to the host side, and at the same time parse the read / write requests for the control status register and the algorithm simulation results sent by the host side, and send the read / write requests for the control status register and the algorithm simulation results to the CPU control module, the instruction execution trace acquisition module and the coprocessor request processing module for processing respectively.

[0070] Step 4 further includes:

[0071] When the target CPU encounters a target algorithm instruction, it extracts the information related to the target algorithm instruction and sends it to the coprocessor request processing module, and at the same time pauses the operation of the target CPU at an appropriate time according to the number of cycles to be executed by the target algorithm instruction. The host side can view the running status of the CPU at any time through the CPU control module.

[0072] A system for fast and accurate co-simulation of a coprocessor includes:

[0073] The FPGA side is used to place the hardware implementation of the target CPU, including an FPGA development board deployed with the hardware design of the target CPU, a CPU control module, an instruction execution trace acquisition module, a coprocessor request processing module, and a communication interface module;

[0074] The host side is connected to the FPGA development board through the FPGA development board high-speed interface, and is used to place the FPGA monitoring program (cooperative control software), target algorithm program and instruction execution trajectory acquisition program.

[0075] Working principle description:

[0076] Derive the instruction execution trace output interface, performance statistics output interface and coprocessor execution interface for the target CPU, and design the CPU control module, instruction execution trace acquisition module, coprocessor request processing module and communication interface module;

[0077] The target CPU hardware design, CPU control module, instruction execution trace acquisition module, coprocessor request processing module and communication interface module are integrated and deployed to the FPGA development board;

[0078] The host runs the FPGA monitoring program, loads the main test program into the target CPU through the CPU control module, and starts execution;

[0079] When the target CPU executes the test program, the instruction execution trajectory acquisition module tracks the CPU operation status in real time and saves the execution trajectory;

[0080] When the target CPU encounters the target algorithm instruction, it sends the target algorithm instruction to the coprocessor request processing module, and at the same time, according to the number of cycles to be executed by the target algorithm instruction, the operation of the target CPU is suspended at an appropriate time. The host side checks the CPU operation status at any time through the CPU control module;

[0081] The coprocessor request processing module records the received target algorithm instructions and initiates a simulation request to the host through the communication interface module;

[0082] After receiving the coprocessor call request, the host uses the pre-deployed target algorithm program to perform simulation execution, and returns the running results and the number of beats to be run to the coprocessor request processing module through the high-speed interface;

[0083] After receiving the simulated execution result of the target algorithm instruction, the coprocessor request processing module feeds back the result to the target CPU according to the number of beats to be run. The target CPU writes the result back and continues execution from the breakpoint.

[0084] When the test program on the FPGA side is executed, the host side reads the instruction execution trajectory information and performance statistics through the communication interface.

[0085] The preferred specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above specific embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for fast and accurate co-simulation of a coprocessor, characterized in that: include: Step 1: Draw out the instruction execution trace output interface, performance statistics output interface and coprocessor execution interface for the target CPU, and design the CPU control module, instruction execution trace acquisition module, coprocessor request processing module and communication interface module; Step 2: Deploy the target CPU hardware design, CPU control module, instruction execution trace acquisition module, coprocessor request processing module and communication interface module to the FPGA development board after integration; Step 3: The host runs the FPGA monitoring program, loads the main test program into the target CPU through the CPU control module, and starts execution; Step 4: When the target CPU executes the test program, the instruction execution trajectory acquisition module tracks the CPU operation status in real time and saves the instruction execution trajectory; When the target CPU encounters the target algorithm instruction, the target algorithm instruction related information is extracted and sent to the coprocessor request processing module. At the same time, the operation of the target CPU is suspended at a predetermined time according to the number of cycles to be executed by the target algorithm instruction. The host side checks the CPU operation status at any time through the CPU control module. Step 5, the coprocessor request processing module records the target algorithm instructions received during the target CPU executing the test program, and initiates a simulation request to the host side through the communication interface module; Step 6: After receiving the coprocessor call request, the host uses the pre-deployed target algorithm program to perform simulation execution, and returns the running result and the number of beats to be run to the coprocessor request processing module through the high-speed interface; Step 7: After receiving the simulated execution result of the target algorithm instruction, the coprocessor request processing module feeds back the result to the target CPU according to the number of beats to be run. The target CPU writes the result back and continues execution from the breakpoint. Step 8: After the FPGA test program is executed, the host reads the instruction execution trace information and performance statistics through the communication interface.

2. The method for fast and accurate co-simulation of a coprocessor according to claim 1, characterized in that: The CPU control module is responsible for loading instructions to the target CPU, providing clock / reset signals, controlling the startup and debugging of the target CPU, and providing a read and write interface for the internal control status register of the target CPU.

3. The method for fast and accurate co-simulation of a coprocessor according to claim 1, characterized in that: The instruction execution trace collection module is used to track the operation status of the target CPU in real time and collect the instruction execution trace.

4. The method for fast and accurate co-simulation of a coprocessor according to claim 1, characterized in that: The coprocessor request processing module is used to intercept the target algorithm instruction and send the target algorithm simulation request to the host side, and at the same time receive the running result of the target algorithm program on the host side and feed back the running result to the target CPU.

5. The method for fast and accurate co-simulation of a coprocessor according to claim 1, characterized in that: The communication interface module is used to convert the simulation request issued by the coprocessor request processing module into an interface signal according to the communication protocol and send it to the host end, while parsing the control read and write requests and algorithm simulation results sent by the host end, and sending the control read and write requests and algorithm simulation results to the CPU control module, the instruction execution trajectory acquisition module and the coprocessor request processing module for processing respectively.

6. A system for fast and accurate co-simulation of coprocessors, used to implement the method according to any one of claims 1 to 5, characterized in that: include: The FPGA side is used to place the target CPU hardware implementation, including an FPGA development board with a target CPU hardware design, a CPU control module, an instruction execution trace acquisition module, a coprocessor request processing module, and a communication interface module; The host side is connected to the FPGA development board through the FPGA development board high-speed interface and is used to place the FPGA monitoring program, target algorithm program and instruction execution trajectory acquisition program.

Citation Information

Patent Citations

  • Computational fluid dynamics (CFD) coprocessor-enhanced system and method

    US20070219766A1

  • FPGA upgrade method based on pcie interface

    US20210294615A1