Chip simulation system construction method and device and medium
By deploying simulated CPU and hardware logic models in different processes and clock domains, and using inter-process communication, the problem that the operation efficiency of simulated CPU is slowed down by SystemC hardware logic is solved, and the overall efficiency of chip simulation system is achieved.
Patent Information
- Application Number
- CN202510237426.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-23
AI Technical Summary
In chip simulation systems, the operating efficiency of the simulation CPU is slowed down by the inefficiency of the SystemC simulation hardware logic, resulting in the slowdown in the efficiency of the entire simulation system.
By deploying the simulated central processing unit model in one process, deploying the simulated hardware logic model in another process, and making the clock domains of the two processes different, data interaction is achieved using inter-process communication mechanisms.
This method improves the overall simulation efficiency of the chip simulation system and avoids the low efficiency of SystemC simulation hardware logic slowing down the operation efficiency of the simulation CPU.
Smart Images

Figure CN120029841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip simulation technology, and in particular to a chip simulation system construction method, device and medium. Background Art
[0002] When simulating a System on Chip (SOC), it is necessary to simulate the central processing unit (CPU) and the hardware logic together. Currently, Gem5 (a modular platform for computer system architecture research) is commonly used to simulate the entire SOC chip. Figure 1 As shown in the figure, the simulation model based on the Gem5 platform mainly includes two parts: the simulation CPU and the hardware simulation logic, which communicate with each other based on the internal address bus. Among them, the hardware simulation logic is generally implemented through SystemC (a collection based on libraries and methods, which is common in simulating the hardware behavior of chips).
[0003] However, in the above simulation model, since the clocks of the hardware simulated by SystemC and the simulated CPU are the same, the simulation CPU model in Gem5 needs to wait for the SystemC execution time to advance. The operating efficiency of the hardware logic simulated by SystemC is slower than that of the simulated CPU. Therefore, the operating efficiency of the simulated CPU will be slowed down by SystemC, which in turn slows down the efficiency of the entire simulation system.
[0004] Therefore, technicians in this field are in urgent need of a chip simulation system construction method to solve the problem that when simulating a chip system, the operating efficiency of the simulated CPU will be slowed down by SystemC, resulting in a slowdown in the efficiency of the entire simulation system. Summary of the invention
[0005] The purpose of the present invention is to provide a chip simulation system construction method, device and medium to improve the overall simulation efficiency of the chip simulation system.
[0006] In order to solve the above technical problems, the present invention provides a chip simulation system construction method, comprising:
[0007] deploying the simulated CPU model in a first process;
[0008] Deploying the simulation hardware logic model in a second process; wherein the clock domains of the first process and the second process are different;
[0009] The simulated central processing unit model and the simulated hardware logic model realize data interaction through an inter-process communication mechanism.
[0010] In a possible embodiment, each time the simulated central processing unit model sends an access request to the simulated hardware logic model, the method further includes:
[0011] Determining a theoretical communication transmission delay according to the data size of the access request and the communication transmission rate between the simulated central processing unit model and the simulated hardware logic model;
[0012] Determining a time advancement deadline based on the theoretical communication transmission delay;
[0013] advancing the system time of the simulated CPU model;
[0014] If the simulation CPU model still has not received the target response when the system time reaches the time advancement deadline, the advancement of the system time is stopped until the simulation CPU model receives the target response;
[0015] The target response is the response returned by the simulation hardware logic model for the access request.
[0016] In a possible embodiment, before advancing the system time of the simulated CPU model, the method further includes:
[0017] Determining the time advancement step according to the theoretical communication transmission delay;
[0018] Then the system time of the simulated CPU model is advanced forward including:
[0019] Using the time advancement step as a system time advancement step, advancing the system time of the simulation CPU model forward;
[0020] The time advancement period is a positive integer multiple of the time advancement step.
[0021] In a possible embodiment, the access request and the target response carry time information;
[0022] After the simulation CPU model receives the target response, the method further includes:
[0023] Determine the target time according to the time information; wherein the target time is: after sending the access request, the system time of the simulated central processing unit model should be advanced by the time required for the simulated hardware logic model to process the access request;
[0024] The system time of the simulated CPU model is advanced to the target time.
[0025] In a possible embodiment, the time information carried in the access request includes: the time when the access request is sent;
[0026] The time information carried in the target response includes: the time when the access request is received and the time when the target response is sent;
[0027] Determining the target time according to the time information includes:
[0028] Determine a target duration according to the difference between the target response sending time and the access request receiving time;
[0029] The target time is determined according to the sum of the access request sending time and the target duration.
[0030] In a possible embodiment, the simulated central processing unit model and the simulated hardware logic model implement data interaction through an inter-process communication mechanism, including:
[0031] Converting the Gem5 address bus access signal sent by the simulated central processing unit model / the simulated hardware logic model into an external communication signal for inter-process communication, and sending it to the second process / the first process;
[0032] The external communication signal sent by the second process / the first process is converted into the Gem5 address bus access signal, and sent to the simulated central processing unit model / the simulated hardware logic model.
[0033] In a possible embodiment, the first process and the second process both include: a transaction-level modeling conversion module and an inter-process communication module;
[0034] Wherein, the transaction-level modeling conversion module is used to: convert the external communication signal sent by the inter-process communication module into a Gem5 address bus access signal and send it to the simulation central processing unit model / the simulation hardware logic model, or to convert the Gem5 address bus access signal sent by the simulation central processing unit model / the simulation hardware logic model into the external communication signal and send it to the inter-process communication module;
[0035] The inter-process communication module is used to implement the transmission of the external communication signal between the first process and the second process.
[0036] In order to solve the above technical problems, the present invention further provides a chip simulation system construction device, comprising:
[0037] A first deployment module, used to deploy the simulation CPU model in a first process;
[0038] A second deployment module, configured to deploy the simulation hardware logic model in a second process; wherein the clock domains of the first process and the second process are different;
[0039] The data communication module is used for the simulation central processing unit model and the simulation hardware logic model to realize data interaction through an inter-process communication mechanism.
[0040] In order to solve the above technical problems, the present invention further provides a chip simulation system construction device, comprising:
[0041] Memory for storing computer programs;
[0042] The processor is used to implement the steps of the chip simulation system construction method as described above when executing the computer program.
[0043] In order to solve the above technical problems, the present invention also provides a non-volatile storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the chip simulation system construction method as described above are implemented.
[0044] The present invention provides a chip simulation system construction method, which realizes the simulation CPU model realized in Gem5 and the simulation hardware logic model realized in SystemC through two processes, and the clock domains of the two processes are different. On the one hand, the chip simulation system can be run in two physical CPU cores to improve the CPU resources of each simulation model, thereby improving the simulation efficiency. On the other hand, the clock domains of the two processes are different, and the simulation CPU can run independently according to the respective system time when not accessing the hardware simulation logic, thereby solving the problem that the SystemC simulation hardware logic efficiency is poor and slows down the simulation CPU efficiency, and effectively improves the simulation efficiency of the entire chip simulation system.
[0045] The chip simulation system construction device and non-volatile storage medium provided by the present invention correspond to the above method and have the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 This is a structural diagram of a chip simulation system based on Gem5;
[0048] Figure 2 A flowchart of a chip simulation system construction method provided by an embodiment of the present invention;
[0049] Figure 3 A structural diagram of a chip simulation system provided by an embodiment of the present invention;
[0050] Figure 4 A schematic diagram of a simulation delay provided by an embodiment of the present invention;
[0051] Figure 5 A schematic diagram of a target time determination principle provided by an embodiment of the present invention;
[0052] Figure 6 A structural diagram of a chip simulation system provided by an embodiment of the present invention;
[0053] Figure 7 A structural diagram of a TLM conversion module provided by an embodiment of the present invention;
[0054] Figure 8 A schematic diagram of an inter-process communication module provided by an embodiment of the present invention;
[0055] Fig. 9 A structural diagram of a line signal conversion module provided by an embodiment of the present invention;
[0056] Fig.10 A structural diagram of a chip simulation system construction device provided by an embodiment of the present invention;
[0057] Fig.11 A structural diagram of another chip simulation system construction device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] The core of the present invention is to provide a chip simulation system construction method, device and medium.
[0060] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0061] A simulation system is a computer software or hardware system that simulates a system or process in the real world. For example, for an integrated circuit (chip), a chip system can be simulated, and the chip behavior can be simulated by building a chip simulation system, and then tested, debugged, and optimized.
[0062] In the chip simulation process, SystemC is currently generally used to simulate hardware behavior to evaluate hardware functions and performance data. SystemC is a collection of libraries and methods based on C++ (a computer language) that can be used to create models accurate to clock cycles and implement simulations at the software algorithm, hardware architecture, SOC interface, and system level. When the entire chip system needs to be simulated, Gem5 (a modular platform for computer system architecture research) is generally used.
[0063] like Figure 1 As shown, the chip simulation system based on Gem5 is implemented through a process. Gem5 can simulate the chip central processing unit (CPU) to obtain a simulated CPU model. Then SystemC is deployed in Gem5 to simulate the hardware logic of the chip to obtain a simulated hardware logic model. Based on the above chip simulation system, Gem5 can simulate the chip system in which the CPU and hardware logic run together.
[0064] However, if Gem5 is deployed in a process, the clocks of the simulated CPU model and the simulated hardware logic model in the Gem5 simulation system are the same. This means that the Gem5 simulated CPU model needs to wait for the execution time of the SystemC simulated hardware logic model to advance. At this time, the system time of the simulated CPU model will be slowed down due to the slow running efficiency of SystemC, which in turn slows down the efficiency of the entire chip simulation system.
[0065] To solve the above problems, the present invention provides a chip simulation system construction method, such as Figure 2 As shown, including:
[0066] S11: deploying the simulated CPU model in the first process.
[0067] S12: deploying the simulation hardware logic model in the second process.
[0068] The clock domains of the first process and the second process are different.
[0069] S13: The simulation central processing unit model and the simulation hardware logic model realize data interaction through an inter-process communication mechanism.
[0070] It should be noted that this embodiment only changes the deployment location of the simulated CPU model and the simulated hardware logic model, and does not limit how to implement the simulated CPU model and the simulated hardware logic model. The simulated CPU model can also be obtained through Gem5 simulation, and the simulated hardware logic model can also be obtained through SystemC simulation.
[0071] Based on the above steps S11 to S13, the chip simulation system constructed by the method provided in this embodiment is as follows: Figure 3 As shown: Gem5 and SystemC are deployed respectively through the two independent processes of the first process and the second process to simulate the simulated CPU model and the simulated hardware logic model. In order to simulate the interaction between the CPU and the hardware logic in the entire chip simulation system, the communication between the simulated CPU model and the simulated hardware logic model deployed in different processes is realized through inter-process communication to ensure the realization of the normal functions of the chip simulation system.
[0072] It is not difficult to see from the above that, because the simulation CPU model and the simulation hardware logic model are deployed in different processes, and each process can correspond to different actual physical CPU cores, that is, compared with the original chip simulation system, the chip simulation system constructed by this method is expanded from the original simulation by one physical CPU core to the simulation by two physical CPU cores, thereby improving the overall simulation efficiency.
[0073] In addition, since the first process and the second process are two independent processes, the clock domains of the first process and the second process may be different, and the simulated CPU model and the simulated hardware logic model are no longer of the same origin. When the simulated CPU model does not need to access the simulated hardware logic model, the simulated CPU model can continue to run based on the independent clock domain to advance the system time without being limited by the operating efficiency of the simulated hardware logic model, thereby improving the overall simulation efficiency of the chip simulation system from another level.
[0074] In summary, the chip simulation system construction method provided by the present invention realizes simulation by deploying the simulation CPU model and the simulation hardware logic model in different processes, thereby increasing the physical CPU resources available to the chip simulation system; and based on the deployment method of different processes, the simulation CPU model and the simulation hardware logic model can run in independent clock domains, and the operating efficiency of the simulation CPU model is no longer limited by the performance bottleneck of the SystemC simulation hardware logic; thereby improving the overall simulation efficiency of the chip simulation system from the above two aspects.
[0075] On the other hand, the above embodiment focuses on the chip simulation system deployed based on the present method. Since the simulation CPU model and the simulation hardware logic model can run in independent clock domains, when the simulation CPU model does not need to access the simulation hardware logic model, the simulation CPU model can continue to run based on the independent clock domain to advance the system time, and will not be limited by the operating efficiency of the simulation hardware logic model, so as to achieve the purpose of improving the overall efficiency of the chip simulation system. However, when the simulation CPU model needs to access the simulation hardware logic model, the simulation CPU model generally needs to wait for the response of the simulation hardware logic model. The present invention does not limit the specific way in which the simulation CPU model waits. A common response waiting scheme is that the simulation CPU model starts waiting for the response of the simulation hardware logic model after issuing an access request.
[0076] However, this embodiment also provides another implementation scheme which is different from the above-mentioned response waiting scheme. Whenever the simulation central processing unit model sends an access request to the simulation hardware logic model, the method further includes:
[0077] S21: Determine a theoretical communication transmission delay according to the data size of the access request and the communication transmission rate between the simulated central processing unit model and the simulated hardware logic model.
[0078] S22: Determine the time advancement deadline based on the theoretical communication transmission delay.
[0079] S23: Advance the system time of the simulated CPU model.
[0080] S24: If the simulation CPU model still has not received the target response when the system time reaches the time advancement deadline, the advancement of the system time is stopped until the simulation CPU model receives the target response.
[0081] The target response is the response returned by the simulation hardware logic model for this access request.
[0082] It should be noted that in the chip simulation system built based on this method, since the simulation CPU model and the simulation hardware logic model are deployed in different processes, the communication between them needs to rely on the inter-process communication mechanism. Therefore, based on the communication transmission rate of the inter-process communication and the data size of this access request, the transmission delay for completing the transmission of this access request in theory can be calculated, that is, the theoretical communication transmission delay.
[0083] Based on the theoretical communication transmission delay, it is possible to determine how long it will take for the simulated hardware logic model to receive the access request after the simulated CPU model issues the access request under ideal circumstances, and even further estimate the arrival time of the simulated hardware logic model's response to the access request. Based on the time determined above, a time advancement deadline can be obtained, which represents how long the simulated CPU model can continue to advance the system time after issuing the access request without affecting the subsequent processing of the access request.
[0084] Based on the obtained time advancement deadline, the simulation CPU model does not need to stop the system time to wait for the response of the simulation hardware logic model after issuing an access request, but can continue to advance the system time to process other tasks. In other words, based on the method provided by this embodiment, the simulation CPU model can continue to advance the system time up to the time advancement deadline to process other tasks, compared with the solution of directly stopping the system time to wait for the response after the access request, when accessing the simulation hardware logic model each time. Thereby alleviating the problem that the simulation CPU model's own operating efficiency is slowed down by the slow operating efficiency of the simulation hardware logic model when the simulation CPU model needs to access the simulation hardware logic model. At the same time, based on the setting of the time advancement deadline, the maximum limit of the system time that the simulation CPU model can advance in one access request is limited. Thereby avoiding the abnormal interaction between the simulation CPU model and the simulation hardware logic model due to excessive advancement of the system time, and ensuring the normal implementation of the entire chip system simulation.
[0085] It should also be noted that the above time advancement deadline is the maximum system time that the simulation CPU model can advance after issuing an access request. However, if the simulation CPU model receives a response returned by the simulation hardware logic model before the time advancement deadline is reached, steps S21 to S24 provided in this embodiment can be skipped, and the simulation CPU model can continue to advance the system time normally to complete the processing task, without stopping the advancement of the system time after reaching the time advancement deadline.
[0086] Based on this embodiment and in combination with the above embodiments, in a chip simulation system constructed by this method, the simulation CPU model can continue to run with its own independent clock domain when it does not need to access the simulation hardware logic model, and the operating efficiency will not be slowed down by the simulation hardware logic model. When the simulation CPU model needs to access the simulation hardware logic model, it can also continue to advance the system time up to the time advancement deadline after each access request, thereby greatly alleviating the problem of slowing down the operating efficiency by the simulated hardware logic model, and further improving the overall efficiency of the chip simulation system.
[0087] Furthermore, in a possible embodiment, assuming that the calculated theoretical communication transmission delay is ΔT, a time parameter k may be preset to quickly obtain a time advancement deadline = k*ΔT.
[0088] This embodiment does not limit the value of the time parameter k, and it can be freely selected according to different actual needs. In general, the larger the k value is, the better the effect of reducing the simulation hardware logic model from slowing down the running efficiency of the simulation CPU model. However, the k value cannot be infinite, otherwise it will cause abnormal interaction between the CPU and the hardware logic in the simulation chip system. In general, k is usually less than 1. In a possible implementation scheme, k can be 0.8.
[0089] Furthermore, this embodiment does not limit how to advance the system time of the simulated CPU model, and the system time can be advanced based on the original clock frequency or processing needs. However, this embodiment also provides another possible system time advancement solution. Before the above step S23, the method further includes:
[0090] S25: Determine the time advancement step according to the theoretical communication transmission delay.
[0091] Then step S23 specifically includes:
[0092] The system time of the simulation CPU model is advanced by taking the time advancement step as one system time advancement step.
[0093] Among them, the time advancement period is a positive integer multiple of the time advancement step.
[0094] It can be seen from the above embodiments that the time advancement deadline is the maximum value of the system time that can be advanced by the simulation CPU model after an access request is issued, and the system time advanced cannot exceed the time advancement deadline. Therefore, when advancing the system time, the step of advancing the system time can be set to a value that can be divided by the time advancement deadline, so that after each access request is issued, it can be reached without exceeding the time advancement deadline after a certain number of steps (i.e., the multiple between the time advancement deadline and the time advancement step). On the one hand, it can ensure the normal interaction between the simulation CPU model and the simulation hardware logic model, and on the other hand, it can also maximize the use of the time advancement deadline. That is, in the case where the response does not arrive in advance, it can be ensured that the system time advanced by the simulation CPU model each time can just reach the time advancement deadline, providing the maximum system time for the simulation CPU model to process other tasks, thereby improving the overall simulation efficiency.
[0095] On the other hand, a key purpose of simulating the entire chip system is to simulate the software and hardware interaction between the CPU and the hardware logic to obtain accurate performance data. However, since the simulated CPU model and the simulated hardware logic model run in different clock domains in this method, due to the existence of external communication delays, the simulated CPU model cannot know the exact time when the simulated hardware logic model responds to access requests.
[0096] For example, Figure 4 As shown in the figure, assuming that the simulated CPU model in Gem5 accesses the simulated hardware logic model in SystemC at time T0, and receives the response returned by the simulated hardware logic model at time T1, Gem5 will use T1-T0 as the time required for SystemC to respond to this access request. However, in fact, due to the existence of external communication delays in inter-process communication, the time difference of T1-T0 cannot accurately represent the actual time required for SystemC to respond to this access request, resulting in inaccurate performance data obtained by simulation.
[0097] To solve the above problems, this application provides a possible implementation scheme:
[0098] Access requests and target responses carry time information.
[0099] After the simulation CPU model receives the target response, the method further includes:
[0100] S31: Determine the target time according to the time information.
[0101] The target time is: after sending the access request, the system time of the simulated central processing unit model should be advanced by the time required for the simulated hardware logic model to process the access request.
[0102] S32: Advance the system time of the simulated CPU model to the target time.
[0103] Specifically, this embodiment carries time information in the access request sent by the simulated CPU model and the response returned by the simulated hardware logic model to obtain the time relationship between the request and the response, and then calculates the accurate time required for the simulated hardware logic model to process the access request. Combining this time with the time when the simulated CPU model sends the access request, the exact time when the simulated CPU model should receive the target response when ignoring the external communication delay can be obtained, thereby achieving synchronization between the simulated CPU model and the simulated hardware logic model, and ensuring that the performance data obtained by simulation is accurate.
[0104] Furthermore, this embodiment does not limit the specific scheme for determining the target time. Accordingly, the time information carried in the access request and the target response should also be set according to the specific scheme for determining the target time, and this embodiment does not limit this.
[0105] However, this embodiment also provides a possible solution for determining the target time:
[0106] The time information carried in the access request includes: the time when the access request is sent.
[0107] The time information carried in the target response includes: the time when the access request is received and the time when the target response is sent.
[0108] Then the above step S32 specifically includes:
[0109] S321: Determine the target duration according to the difference between the target response sending time and the access request receiving time.
[0110] S322: Determine the target time according to the sum of the access request sending time and the target duration.
[0111] For example, Figure 5 As shown, the time when the simulated CPU model in Gem5 sends the access request is T0 (corresponding to the time T0' on the SystemC side), and the time when the target response is received is T1 (corresponding to the time T3' on the SystemC side). And the time when the simulated hardware logic model in SystemC actually receives the access request is T1', and the time when the target response is actually issued is T2'. The above time points are all based on the actual physical time.
[0112] Depend on Figure 5 It can be seen that on the Gem5 side, the actual perceived time required for the simulated hardware logic model to process this access request is T1-T0. However, in fact, the time required for the simulated hardware logic model to process this access request is T2'-T1'. Therefore, this embodiment calculates the exact value of T2'-T1' and resets the system time when the simulated CPU model receives the target response to ensure the accuracy of the system time advancement of the simulated CPU model in Gem5.
[0113] Specifically, the time at which the access request is sent is equivalent to Figure 5 The access request receiving time mentioned above is equivalent to T0. Figure 5 T1' in the above target response sending time is equivalent to Figure 5 T2' in the above embodiment. Therefore, based on the above steps of this embodiment, we have:
[0114] Target duration = T2'-T1';
[0115] Target time = T0 + T2' - T1';
[0116] Based on the system time advancement solution provided in this embodiment, it can be ensured that when the external communication delay is uncertain and the simulation CPU model continues to advance the system time after the access is issued, the system time can still be accurately advanced to the correct moment, thereby realizing cross-platform (cross-process or cross-clock domain) clock synchronization between Gem5 and SystemC, and ensuring the accuracy of the performance data obtained by simulation.
[0117] On the other hand, regarding how to implement communication between the simulated CPU model and the simulated hardware logic model in the above step S13, this embodiment provides a possible implementation scheme, and step S13 specifically includes:
[0118] S131: Convert the Gem5 address bus access signal sent by the simulated central processing unit model / simulated hardware logic model into an external communication signal for inter-process communication, and send it to the second process / first process.
[0119] S132: Convert the external communication signal sent by the second process / the first process into a Gem5 address bus access signal, and send it to the simulated central processing unit model / the simulated hardware logic model.
[0120] It can be seen from the above embodiments that this embodiment does not limit the specific simulation implementation method of the simulation CPU model and the simulation hardware logic model, and can be implemented through Gem5 and SystemC as in the original solution. Figure 1 As shown in the figure, the simulation CPU model based on Gem5 and the simulation hardware logic model based on SystemC originally support the communication mode within the Gem5 platform, that is, communication is realized on the internal address bus within the Gem5 platform based on the Gem5 address bus access signal. In addition, there are currently many mature technical solutions for the inter-process communication mechanism, such as sockets, shared memory, file mapping, message queues and many other inter-process communication solutions.
[0121] Based on the above, in order to reduce the complexity of the communication mechanism, improvements are made directly on the original communication mechanism, adding conversions between the Gem5 address bus access signals used for internal communication in Gem5 and the external communication signals used for communication between external processes, so as to utilize the existing communication mechanism to simply and quickly realize the communication between the simulated CPU model and the simulated hardware logic model.
[0122] Furthermore, this embodiment also provides a possible implementation scheme for how to implement the conversion between the Gem5 address bus access signal and the external communication signal in the above embodiment, such as Figure 6 As shown:
[0123] Both the first process and the second process include: a transaction-level modeling conversion module and an inter-process communication module.
[0124] Among them, the transaction-level modeling conversion module is used to: convert the external communication signal sent by the inter-process communication module into a Gem5 address bus access signal and send it to the simulation central processing unit model / simulation hardware logic model, or to convert the Gem5 address bus access signal sent by the simulation central processing unit model / simulation hardware logic model into an external communication signal and send it to the inter-process communication module.
[0125] The inter-process communication module is used to realize the transmission of external communication signals between the first process and the second process.
[0126] TLM: Transaction-Level Modeling. It is a high-level abstraction standard used in SystemC. The TLM conversion module is a functional module supported by the existing SystemC that can convert external communication signals into Gem5 address bus access signals for internal communication.
[0127] In a possible embodiment, the TLM conversion module is as follows: Figure 7 As shown, the TLM conversion module is used to convert the bus communication interface between Gem5 and SystemC into external communication, and call the communication module interface for data transmission and reception. The data transmitted and received by the communication module is converted into TLM bus reading and writing. Specifically, the TLM interface is divided into Initiator_SOCket and target_SOCket. Among them, the interface Initiator_SOCket is used to initiate data reading and writing, and the interface target_SOCket is used to process the reading and writing of external data. Since there are many mature solutions for the implementation of the TLM conversion module, this embodiment will not be repeated here.
[0128] The implementation of the inter-process communication module may be implemented through various inter-process communication schemes such as sockets, shared memory, file mapping, message queues, etc. as mentioned in the above embodiments, and this embodiment does not impose any limitation on this.
[0129] In a possible embodiment, the inter-process communication module is as follows: Figure 8 As shown in the figure, this module is used for data exchange between Gem5 and SystemC simulation system, by encapsulating the interface of communication between operating system processes. It defines channel creation and data receiving and sending interface. It also defines data format and interaction mechanism.
[0130] The channel is created by encapsulating the operating system process communication interface to create a communication SOCket. A receiving thread is created to call the receiving processing module to receive external data.
[0131] The receiving thread periodically determines whether there is data from the outside. If there is, it processes the data. After receiving a complete packet of data, it determines whether the data is a response packet or a request packet. If it is a response data packet, it releases an event signal to trigger other modules for subsequent processing. If it is a request data packet, it calls back to process the data packet.
[0132] The data receiving interface calls the operating system interface to receive external data and analyzes the data. The data integrity is determined based on the data length in the data frame header. When a complete data packet is received, the buffer-flag (a flag) corresponding to the data is set for processing by other modules.
[0133] The data transmission interface encapsulates the data to be sent into the corresponding data format package and sets the buffer-flag. The buffer-flag is used to mark the association between the sent data and the received data. The sending interface is protected by a mutex lock to prevent the situation where the data is not sent and is preempted by other modules when multiple modules call the sending interface.
[0134] In addition, the data in the communication module can be specifically divided into request data and response data.
[0135] The data format includes command type, length, packet unique number (ID), request device ID, flag, and data payload. The flag is used to distinguish whether the data is request data or response data. The packet ID is used to match the request data with the response data. For the above command types that require time synchronization, the time information can be carried in the data payload.
[0136] Furthermore, in a possible embodiment, the first process and the second process mentioned above both further include: a line signal conversion module.
[0137] The line signal conversion module is used to realize the transmission of single-line signals, specifically to transmit the interrupt signal to the outside through the inter-process communication module, or to receive the external interrupt signal for processing, and is used to deal with the interrupt processing interconnection after Gem5 and SystemC are distributed. Its structure is as follows Fig. 9 shown.
[0138] It should be noted that the TLM conversion module can also realize the transmission of single-line signals. However, this embodiment separates the transmission and processing of single-line signals from the TLM conversion module by setting up separate line signal conversion modules, thereby improving the signal transmission efficiency and achieving the purpose of improving the overall efficiency of the chip simulation system.
[0139] It is particularly important to note that the TLM conversion module, communication module and line signal conversion module mentioned in the above embodiments are all functional modules that can be supported and implemented by the current SystemC. Therefore, for the simulated hardware logic model deployed in the second process and implemented by SystemC, the above functional modules can be deployed in SystemC to implement the corresponding functions. As for the simulated CPU model deployed in the first process and implemented by Gem5, it can be seen from the above embodiments that the Gem5 platform supports the deployment of SystemC, so SystemC can also be deployed in the Gem5 platform of the first process to implement the above functional modules.
[0140] In addition, steps S21 to S25, steps S31 to S32 and their respective sub-steps provided in the above embodiment can also be implemented through functional modules as in the above embodiment. Figure 6 As shown, the above steps S21-S25 and steps S31-S32 are implemented by setting a time control module in the first process. Specifically, SystemC is a collection of libraries and methods based on C++. Based on the method steps provided in the above embodiment, the time control module implementing the above steps can be obtained by compiling in C++ language, and this embodiment will not be repeated here.
[0141] In addition to the embodiment of a chip simulation system construction method provided in the above embodiment, the present invention also provides an embodiment corresponding to a computer program product. A computer program product includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the chip simulation system construction method described in any of the above embodiments can be implemented.
[0142] Since the embodiments of the computer program product part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the computer program product part, which will not be repeated here.
[0143] In the above embodiment, a chip simulation system construction method is described in detail, and the present invention also provides a corresponding embodiment of a chip simulation system construction device. It should be noted that the present invention describes the embodiment of the device part from two perspectives, one is based on the functional module perspective, and the other is based on the hardware perspective.
[0144] Based on the perspective of functional modules, such as Fig.10 As shown, this embodiment provides a chip simulation system construction device, including:
[0145] A first deployment module 11, used to deploy the simulation CPU model in a first process;
[0146] A second deployment module 12 is used to deploy the simulation hardware logic model in a second process; wherein the clock domains of the first process and the second process are different;
[0147] The data communication module 13 is used to realize data interaction between the simulation central processing unit model and the simulation hardware logic model through an inter-process communication mechanism.
[0148] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.
[0149] Fig.11 A structural diagram of a chip simulation system construction device provided by another embodiment of the present invention, such as Fig.11 As shown, a chip simulation system construction device includes: a memory 20 for storing a computer program;
[0150] The processor 21 is used to implement the steps of a chip simulation system construction method in the above embodiment when executing a computer program.
[0151] The chip simulation system construction device provided in this embodiment may include but is not limited to a mobile terminal, a personal computer, a workstation, etc.
[0152] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0153] The memory 20 may include one or more computer-readable storage media, which may be non-transitory. The memory 20 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201, wherein, after the computer program is loaded and executed by the processor 21, it can implement the relevant steps of a chip simulation system construction method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 20 may also include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. Data 203 may include, but is not limited to, a chip simulation system construction method, etc.
[0154] In some embodiments, a chip simulation system construction device may further include a display screen 22 , an input / output interface 23 , a communication interface 24 , a power supply 25 , and a communication bus 26 .
[0155] Those skilled in the art will understand that Fig.11 The structure shown in the figure does not constitute a limitation on a chip simulation system construction device, and may include more or fewer components than those shown in the figure.
[0156] An embodiment of the present invention provides a chip simulation system construction device, which includes a memory and a processor. When the processor executes a program stored in the memory, it can implement the following method: a chip simulation system construction method.
[0157] Finally, the present invention also provides an embodiment corresponding to a non-volatile storage medium. The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.
[0158] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes.
[0159] The above is a detailed introduction to a chip simulation system construction method, device and medium provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referenced to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
[0160] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
Claims
1. A chip simulation system construction method, characterized in that: include: deploying the simulated CPU model in a first process; Deploying the simulation hardware logic model in a second process; wherein the clock domains of the first process and the second process are different; The simulated central processing unit model and the simulated hardware logic model realize data interaction through an inter-process communication mechanism.
2. The chip simulation system construction method according to claim 1, characterized in that: Whenever the simulated CPU model sends an access request to the simulated hardware logic model, the method further comprises: Determining a theoretical communication transmission delay according to the data size of the access request and the communication transmission rate between the simulated central processing unit model and the simulated hardware logic model; Determining a time advancement deadline based on the theoretical communication transmission delay; advancing the system time of the simulated CPU model; If the simulation CPU model has not received the target response when the system time reaches the time advancement deadline, the advancement of the system time is stopped until the simulation CPU model receives the target response; The target response is the response returned by the simulation hardware logic model for the access request.
3. The chip simulation system construction method according to claim 2, characterized in that: Before advancing the system time of the simulated CPU model, the method further includes: Determining the time advancement step according to the theoretical communication transmission delay; Then the system time of the simulated CPU model is advanced forward including: Using the time advancement step as a system time advancement step, advancing the system time of the simulation CPU model forward; The time advancement period is a positive integer multiple of the time advancement step.
4. The chip simulation system construction method according to claim 2, characterized in that: The access request and the target response carry time information; After the simulation CPU model receives the target response, the method further includes: Determine the target time according to the time information; wherein the target time is: after sending the access request, the system time of the simulated central processing unit model should be advanced by the time required for the simulated hardware logic model to process the access request; The system time of the simulated CPU model is advanced to the target time.
5. The chip simulation system construction method according to claim 4, characterized in that: The time information carried in the access request includes: the time when the access request is sent; The time information carried in the target response includes: the time when the access request is received and the time when the target response is sent; Determining the target time according to the time information includes: Determine a target duration according to the difference between the target response sending time and the access request receiving time; The target time is determined according to the sum of the access request sending time and the target duration.
6. The chip simulation system construction method according to any one of claims 1 to 5, characterized in that: The simulation central processing unit model and the simulation hardware logic model realize data interaction through an inter-process communication mechanism, including: Converting the Gem5 address bus access signal sent by the simulated central processing unit model / the simulated hardware logic model into an external communication signal for inter-process communication, and sending it to the second process / the first process; The external communication signal sent by the second process / the first process is converted into the Gem5 address bus access signal, and sent to the simulated central processing unit model / the simulated hardware logic model.
7. The chip simulation system construction method according to claim 6, characterized in that: The first process and the second process both include: a transaction-level modeling conversion module and an inter-process communication module; Wherein, the transaction-level modeling conversion module is used to: convert the external communication signal sent by the inter-process communication module into a Gem5 address bus access signal and send it to the simulation central processing unit model / the simulation hardware logic model, or to convert the Gem5 address bus access signal sent by the simulation central processing unit model / the simulation hardware logic model into the external communication signal and send it to the inter-process communication module; The inter-process communication module is used to implement the transmission of the external communication signal between the first process and the second process.
8. A chip simulation system construction device, characterized in that: include: A first deployment module, used to deploy the simulation CPU model in a first process; A second deployment module, configured to deploy the simulation hardware logic model in a second process; wherein the clock domains of the first process and the second process are different; The data communication module is used for the simulation central processing unit model and the simulation hardware logic model to realize data interaction through an inter-process communication mechanism.
9. A chip simulation system construction device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the chip simulation system construction method as described in any one of claims 1 to 7 when executing the computer program.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the chip simulation system construction method according to any one of claims 1 to 7 are implemented.