Interpreted instruction set simulator based on LRU cache and instruction execution method thereof

By introducing an optimization method based on LRU cache in the interpreted instruction set simulator, the problem of slow speed in the instruction decoding process of traditional simulators is solved, and more efficient simulator operation and more significant performance improvement is achieved.

CN119512563BActive Publication Date: 2025-05-09LIANZIXIN INTELLIGENT TECHNOLOGY (HANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510073503.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-09
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Traditional interpreted instruction set simulators are slow in the instruction decoding process, resulting in inefficient verification and optimization processes. Existing optimization methods such as interleaved code and binary translation have limitations.

Method used

The interpreted instruction set simulator based on LRU cache is adopted to efficiently cache the decoding results of jump instruction blocks, reduce the resource consumption of repeated decoding, avoid the redundant decoding process, and accelerate the decoding speed.

Benefits of technology

It significantly improves the running speed of the simulator, improves overall accuracy and flexibility, solves the performance bottleneck problem in the decoding stage, and shows more significant performance improvements in complex loop programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119512563B_ABST
    Figure CN119512563B_ABST
Patent Text Reader

Abstract

The present invention discloses an interpreted instruction set simulator based on LRU cache and an instruction execution method thereof. The simulator comprises an executable file loading module, an instruction set module, a register module, a memory module and a disassembly module. By loading the entry address of the program in the ELF file, the instruction simulation work is completed according to the instruction execution method, and the disassembly result is displayed. The instruction execution method stores the address of the jump instruction through a sub-cache, and when the jump instruction appears for the second time, the instruction block decoding result between the two jump instructions is stored in the LRU cache, thereby reducing unnecessary cache operations and maximizing the use of cache resources. Before subsequent decoding, the decoding result is directly taken out from the LRU cache to avoid multiple and repeated decoding of an instruction block, thereby improving the operation efficiency of the simulator. The cache space is cleaned up based on the least recently used strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer technology, and relates to an optimized design of an interpreted instruction set simulator, and in particular to an interpreted instruction set simulator based on an LRU cache and an instruction execution method thereof. Background Art

[0002] With the continuous improvement of chip integration and the widespread application of deep submicron technology, designers are facing unprecedented challenges of complex hardware and software interaction. This complexity makes existing design and verification methods stretched. Traditional design methods rely on a cyclical "design-implementation-evaluation-improvement" process. However, as the number of transistors on a chip increases to hundreds of millions, this process becomes inefficient. In the early stages of design, if potential performance bottlenecks cannot be accurately identified and resolved, subsequent development and production will face great challenges and high costs. Instruction set simulators can more effectively analyze and handle the performance challenges brought by advanced technologies such as cache management, branch prediction, and out-of-order execution by providing detailed and accurate architecture simulations, filling the gaps in traditional theoretical models in system performance analysis.

[0003] Traditional interpreted simulators process instructions one by one by fetching, decoding, and executing them. Although accuracy is ensured, the time-consuming translation of each instruction significantly reduces the running speed of the simulator. The speed of the simulator is crucial to the verification and optimization process, because it requires rapid evaluation of design solutions and accurate simulation of performance in different scenarios in real application scenarios. In particular, when instruction blocks need to be executed repeatedly, slow instruction decoding often becomes a performance bottleneck. Therefore, improving simulation speed is one of the main challenges facing current instruction set simulators. Existing optimization methods, such as thread code and binary translation, have certain limitations in improving speed. Thread code technology relies on high instruction reuse rate, but in actual applications, instruction reuse is often concentrated in specific areas, affecting cache hit rate and making it difficult to significantly improve performance. Although binary translation can quickly execute compiled native code, it cannot support key functions of interpreted simulators, such as single-step debugging and fine-grained state tracking. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention proposes an interpreted instruction set simulator based on LRU (least recently used) cache and its instruction execution method. The method reduces the resource consumption of repeated decoding and avoids redundant decoding processes by efficiently caching the decoding results of jump instruction blocks, thereby significantly accelerating the decoding speed. This not only solves the performance bottleneck problem of the decoding stage in the prior art, but also improves the overall accuracy and flexibility of the simulator, providing strong support for design verification and early decision-making.

[0005] An interpreted instruction set simulator based on LRU cache includes an executable file loading module, an instruction set module, a register module, a memory module and a disassembly module.

[0006] The executable file loading module is used to parse and load the ELF (Executable and Linkable Format) file, extract the entry address of the program, and accurately map each segment in the file to the corresponding position of the memory module to ensure the effective execution of the program.

[0007] The instruction set module is used to perform decoding and execution operations on instructions. The decoding operation loads instructions from the program entry address of the memory module, extracts relevant operands and offsets and stores them in registers, prepares for subsequent instruction execution, and cooperates with LRU cache and sub-cache to achieve decoding acceleration. The execution operation is responsible for executing the corresponding simulation function according to the result of the decoding operation, and writing the result back to the register, thereby completing the behavior of the simulation instruction.

[0008] The register module is used to define and manage the names, quantities and data types of all registers. The register is responsible for storing the address of the next instruction to be executed and various temporary data.

[0009] The memory module is used to store all instructions and data, and provides specific memory read and write functions to perform memory access operations to meet various computing requirements.

[0010] The disassembly module accesses the data obtained by decoding the instruction set module, uses the corresponding disassembly function to generate and display the corresponding disassembly result, and supports the disassembly display of custom instructions.

[0011] The instruction execution method of the interpreted instruction set simulator based on LRU cache has the following specific steps:

[0012] Step 1: Initialize the simulator. Use a data structure that combines a doubly linked list and a hash table to implement the LRU cache. Build a jump address hash table as a sub-cache. Load the ELF file and load the program entry address into the program counter (PC). Execute step 2.

[0013] Step 2: Fetch the instruction pointed to by the program counter from the memory and decode it. If it is a jump instruction, store the jump address in the sub-cache. Then execute the instruction according to the decoding result. Execute step 3.

[0014] Step 3: Fetch the next instruction to be executed from the memory. Go to step 4.

[0015] Step 4: Decode the currently fetched instruction. If the instruction is not a jump instruction, execute it immediately according to the decoding result, and repeat step 3. If it is a jump instruction, check whether the jump address already exists in the sub-cache. If not, store the jump address in the sub-cache and execute the instruction, and repeat step 3; if it exists, execute the instruction and go to step 5.

[0016] Step 5: Extract and decode the next instruction, and cache the decoding result of the current instruction in the LRU cache. When the LRU cache reaches the upper limit of capacity, remove the instruction that has not been used for the longest time. If the current instruction is a jump instruction, execute step 6; if not, continue to execute step 5.

[0017] Step 6: Extract the next instruction to be executed from the memory and check whether the instruction exists in the LRU cache. If not, execute step 4; if so, extract the decoding result from the LRU cache and execute the instruction, and then repeat step 6.

[0018] The present invention has the following beneficial effects:

[0019] 1. This method introduces a cache search step before decoding. Unlike the low hit rate of traditional instruction cache, this method only stores the instruction block that is accessed for the second time into the cache. This not only reduces unnecessary cache operations, but also maximizes the utilization efficiency of cache resources and significantly improves the timeliness of decoding. To ensure the effect, the cache size is dynamically adjusted according to the principle of locality of the program to ensure that the most critical decoding information can be stored in the limited cache space.

[0020] 2. Divide the executed instructions into different instruction blocks according to jump instructions, and cache the decoding results of these instruction blocks, thereby effectively reducing the number of repeated decoding. Test results show that this optimization strategy can improve the operating efficiency of the simulator.

[0021] 3. The simulator integrates an executable file loading module to simplify the program loading process. In addition, the newly added customizable instruction disassembly module allows users to customize the disassembly format and adjust the disassembly output according to needs, further enhancing the applicability and flexibility of the simulator.

[0022] 4. The functional correctness and consistency of the simulator were verified through simulation experiments. Compared with the existing simulators, the average execution speed (MIPS) was improved by about 6.52%. Especially in complex loop programs, the performance improvement is more significant. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of an interpreted instruction set simulator based on LRU cache;

[0024] Figure 2A flow chart of an instruction execution method of an interpreted instruction set simulator based on an LRU cache;

[0025] Figure 3 This is a schematic diagram of the LRU cache structure in the embodiment;

[0026] Figure 4 It is a sub-cache structure in the embodiment;

[0027] Figure 5 This is a comparison of instruction execution speed before and after simulator optimization in the embodiment. DETAILED DESCRIPTION

[0028] The present invention will be further explained below with reference to the accompanying drawings;

[0029] An interpreted instruction set simulator based on LRU cache, such as Figure 1 As shown, it includes an executable file loading module, an instruction set module, a register module, a memory module and a disassembly module.

[0030] This embodiment is simulated based on an interpreted instruction set simulator based on the RISC-V 32-bit architecture, aiming to support 172 widely used instructions in the RISC-V instruction set: 42 instructions of the basic integer instruction set (I-set), 13 instructions of the multiplication and division extension (M-set), 30 instructions of the single-precision floating-point extension (F-set), 34 instructions of the double-precision floating-point extension (D-set), 38 instructions of the compressed instruction set (C-set) and 15 instructions of the atomic operation extension (A-set).

[0031] The executable file loading module is used to implement the support function of the instruction set simulator for ELF (Executable and Linkable Format) files. The executable file loading module first parses the ELF header of the ELF file to accurately obtain the program entry address, thereby pointing to the first instruction to be executed in the program, ensuring that the program is correctly started in the simulation environment. Then the program header table is analyzed, and the location and size of each segment that needs to be loaded into the memory module are determined based on the detailed information provided in the program header table on how each segment (Segments) in the file is mapped to the memory. Thereby reconstructing the operation logic and memory layout of the ELF file in the virtual computing environment. Finally, the initial program counter (PC) is set for the simulator to point to the previously determined entry point address. In this way, the simulator can simulate execution from the correct program starting point, and gradually parse and execute the instructions loaded into the memory.

[0032] The instruction set module is used to decode and execute instructions. By importing the compiled binary file from the memory into the instruction set module and performing decoding operations, relevant operands and offsets are extracted and stored in registers to prepare for subsequent instruction execution. The execution operation in the instruction set module is responsible for using this information, executing the corresponding simulation function, and writing the result back to the register, thereby completely simulating the behavior of the instruction.

[0033] The instruction set module adopts the Load / Store model for memory access, which emphasizes the principle that all memory operations except loading executable files need to be performed through explicit load and store instructions. In the translation, the lowest two bits of the instruction are first checked: if the lowest two bits are not 11, the instruction is regarded as a compressed instruction. After that, the simulator compares the instruction with the predefined instruction template through binary bit-level pattern matching. Once the match is successful, the corresponding operation function is immediately called to execute the instruction.

[0034] In order to simulate the operating state of the processor, a flexible and independent data structure is used in the register module to define the name, quantity and data type of registers, including general registers, floating-point registers and special registers.

[0035] In the RISC-V architecture, 32 general-purpose registers are labeled x0 to x31. The x0 register is a hard-coded register that always maintains a value of 0. This feature simplifies the design of the instruction set and improves efficiency during execution. The remaining registers (x1 to x31) are fully readable and writable to support a variety of general-purpose operations and data storage needs. The 32 independent floating-point registers are labeled f0~f31 and are independent of the general-purpose registers and are specifically used for floating-point operations. Unlike the fixed value characteristics of the x0 register, all floating-point registers can be flexibly used to store single-precision or double-precision data. In addition, for processors with both RV32F and RV32D functions, single-precision data only occupies the lower 32 bits of the floating-point register. This design strategy effectively expands the bandwidth and capacity of the registers, thereby improving the overall performance of the processor.

[0036] The special registers mainly include the program counter (PC) and the status control register. In order to refine the behavior of the program counter, three variables are used for simulation: pc is used for the current instruction address, snpc is responsible for recording the instruction address after the jump, and dnpc points to the next instruction address during sequential execution. When processing various jump and branch instructions, the program counter is accurately updated, reducing the probability of synchronization errors.

[0037] The memory module adopts an address mapping strategy to store and manage program code and data. The physical memory of the target machine is simulated through the virtual memory system of the host machine to effectively simulate the application level main memory. When the test program is loaded, it will first be placed in the virtual space. When the application running on the target machine accesses the memory, its physical address must be converted to the virtual address of the host machine to ensure accurate memory addressing and security protection in the host machine environment. This method can not only effectively reproduce the memory space of the target machine on the host machine's architecture, but also make full use of the host machine's memory management characteristics to improve the efficiency and flexibility of the simulator.

[0038] The disassembly module is used to convert binary machine language instructions into human-readable assembly language, which improves the intuitiveness and efficiency of instruction analysis and error location. Users can customize the disassembly format by modifying the source code, or add support for unofficial instruction sets, which improves the usability of the simulator and provides developers who focus on specific application requirements with a wider operating space and optimization possibilities. In addition, users can select the level of detail of the assembly information in the settings. In detailed mode, the corresponding information in the ELF file will be displayed before the auipc instruction, and jump instructions such as c.jal will additionally display their jump addresses in detailed mode. Table 1 shows the disassembly results of the auipc instruction and the c.jal instruction in abbreviated mode and detailed mode:

[0039] Table 1

[0040]

[0041] Traditional interpreted simulators run instructions in a loop mode of fetching, decoding, and executing. Although this ensures accuracy, the huge decoding time caused by translating each instruction one by one affects the running speed. The interpreted instruction set simulator based on LRU cache adopts an optimized instruction execution method to improve the instruction decoding process. For decoded instructions, it will decide whether to store them in the cache based on conditions. In the subsequent instruction fetch process, the simulator first checks the cache to determine whether the corresponding decoding result already exists. If a cached result exists, the decoding stage is skipped and the cache information is used directly. This improvement significantly shortens the instruction translation time and increases the simulation speed. The specific steps are as follows: Figure 2 As shown, the details are as follows:

[0042] Step 1: Initialize the simulator. Use a data structure that combines a bidirectional linked list and a hash table to implement LRU cache, such as Figure 3 At the same time, a jump address hash table is constructed as a sub-cache, as shown in Figure 4As shown. Load the ELF file through the executable file loading module and load the program entry address into the program counter (PC). Execute step 2.

[0043] The LRU cache is implemented by a structure combining a hash table and a bidirectional linked list, wherein the hash table is used to quickly locate data in the cache, and the bidirectional linked list is used to maintain the access order of the data. The virtual head node of the linked list is used to store the most recently used data, and the virtual tail node is used to delete the data that has not been used for the longest time. The status flag has two states: 0 means that the LRU cache is closed, and there is no need to cache the decoding result; 1 means that the LRU cache is turned on and the current decoding result needs to be cached.

[0044] In the sub-cache, each hash bucket includes two parts: an index and a flag, wherein the index is used to store the jump address, and the flag has two states. When the flag is 0, it indicates that the jump address is the first jump; when the flag is 1, it indicates that the jump address is a repeated jump.

[0045] Step 2: Fetch the instruction pointed to by the program counter (PC) from the memory and decode it. If it is a jump instruction, store the jump address in the sub-cache. Then execute the instruction according to the decoding result. Execute step 3.

[0046] Step 3: Fetch the next instruction to be executed from the memory. Go to step 4.

[0047] Step 4: Decode the currently fetched instruction. If the instruction is not a jump instruction, execute it immediately according to the decoding result, and repeat step 3. If it is a jump instruction, check whether the jump address already exists in the sub-cache. If not, store the jump address in the sub-cache and execute the instruction, and repeat step 3; if it exists, execute the instruction and go to step 5.

[0048] Step 5: Extract and decode the next instruction. Determine whether the LRU cache has reached its capacity limit. If so, remove the longest unused instruction. Cache the decoding result of the current instruction in the LRU cache. If the current instruction is a jump instruction, execute step 6; if not, continue to execute step 5.

[0049] Step 6: Extract the next instruction to be executed from the memory and check whether the instruction exists in the LRU cache. If not, execute step 4; if so, extract the decoding result from the LRU cache and execute the instruction, and then repeat step 6.

[0050] Through the LRU cache mechanism, the simulator can store and reuse the decoding results of jump instruction blocks during execution, thereby significantly reducing repeated pattern matching and instruction decoding operations. Unlike the low hit rate of traditional instruction caches, this method stores the decoding results of jump instructions in the cache only when they are accessed for the second time, and does not stop the caching process until the jump instruction is encountered again. This not only reduces unnecessary cache operations, but also maximizes the use of cache resources and improves timeliness. At the same time, in order to ensure effectiveness, the cache size is dynamically adjusted based on the principle of program locality to ensure that the most important decoding information can be stored in the limited cache space.

[0051] In order to further demonstrate the effectiveness of the application, the simulator was tested for function and performance using the method. Functional testing is mainly used to verify the correctness and completeness of the simulator's functions, while performance testing is intended to evaluate the simulator's execution speed and the actual effect of the proposed optimization technology.

[0052] In the functional testing phase, the RISC-V cross-compilation tool chain is used to compile a series of C language codes for specific instructions into an ELF format executable file. Taking the addition operation in the data processing instruction as an example, the following C language code sample double_imafdc is used:

[0053] double mul(double a, double b) {

[0054] return a + b;

[0055] }

[0056] double main() {

[0057] double a = 6.66;

[0058] double b = 3.14;

[0059] double res = mul(a, b);

[0060] return res;

[0061] }

[0062] After loading the compiled double_imafdc.elf file into the simulator, the console shows that the address of the first instruction is 0x100a0, confirming that the program is loaded correctly. The simulator then single-steps through the first five instructions, and the register changes tracked and recorded during this period are shown in Table 2:

[0063] Table 2

[0064]

[0065] This result is completely consistent with the tracking result obtained by performing the same operation on the Spike simulator, verifying the accuracy of the simulator.

[0066] Correspondingly, the disassembly results of the five instructions are shown in Table 3, which are consistent with the disassembly information generated by the readelf tool:

[0067] Table 3

[0068]

[0069] After executing the remaining instructions of the program, the content displayed by the simulator in the floating point register fa0 is converted into a binary system and is consistent with the expected result of the C language test program. This shows that the simulator proposed by this method can accurately simulate the execution of various instructions.

[0070] In performance testing, the size of the LRU cache is a key factor. The choice of cache capacity not only affects the hit rate, but also directly affects the efficiency of memory use. In order to determine the optimal cache size that can balance high hit rate and memory efficiency, performance tests were conducted on three test cases a, b, and c with increasing instruction numbers, respectively, to fully evaluate the impact of cache capacity on performance. The results are shown in Table 4:

[0071] Table 4

[0072]

[0073] As can be seen from Table 3, as the cache capacity increases, the execution time of all programs decreases, especially the test case c with the largest number of instructions, which has the most significant performance improvement. When the cache reaches 1024 bytes, the execution time of test case c tends to be stable, and further increasing the capacity to 2048 bytes does not bring significant improvement. This shows that a 1024-byte cache is sufficient to meet the program requirements, and for test cases a and b, 1024 bytes is also sufficient. Therefore, in subsequent performance tests, a 1024-byte LRU cache is used uniformly.

[0074] Finally, we verify the execution speed of the simulator after introducing the disassembly module and recording the registers, CPU status and other information, as well as the improvement of the simulator performance when executing instructions through this method. The results are as follows: Figure 5As shown in the figure, we can see that after introducing the disassembly module and recording registers, CPU status and other information, the average execution speed of the simulator is about 1.8 MIPS (million instructions per second), while the simulator that executes instructions through this method has an average execution speed of 1.9 MIPS, a performance improvement of about 6.52%, especially in the loop program. The performance is significantly improved. It proves that this method can effectively improve the operating efficiency of the simulator in most program scenarios, especially in the case of highly repetitive instructions, the improvement effect is particularly significant.

Claims

1. An instruction execution method of an interpreted instruction set simulator based on an LRU cache, characterized in that: The specific steps are as follows: Step 1, initialize the simulator; use a data structure combining a bidirectional linked list and a hash table to implement the LRU cache, where the hash table is used to locate the data in the cache, the bidirectional linked list is used to maintain the access order of the data, the virtual head node of the linked list is used to store the most recently used data, and the virtual tail node is used to delete the longest unused data, and the status flag has two states: 0 means that the LRU cache is closed, and there is no need to cache the decoding result; 1 means that the LRU cache is turned on and the current decoding result needs to be cached; build a jump address hash table as a sub-cache, in the jump address hash table, each hash bucket includes an index and a flag, where the index is used to store the jump address, and the flag has two states; when the flag is 0, it indicates that the jump address is the first jump; when the flag is 1, it indicates that the jump address is a repeated jump; load the ELF file and load the program entry address into the program counter; execute step 2; Step 2: Extract the instruction pointed to by the program counter and decode it; If it is a jump instruction, the jump address is stored in the sub-cache; then the instruction is executed according to the decoding result; and step 3 is executed; Step 3, extract the next instruction to be executed from the memory; execute step 4; Step 4: Decode the currently extracted instruction; if the instruction is not a jump instruction, execute it immediately according to the decoding result, and repeat step 3; If it is a jump instruction, check whether the jump address already exists in the sub-cache; If it does not exist, store the jump address in the sub-cache, execute the instruction, and repeat step 3; If it exists, execute the instruction and go to step 5; Step 5: extract and decode the next instruction, and cache the decoding result of the current instruction into the LRU cache; If the current instruction is a jump instruction, execute step 6; if not, continue to execute step 5; Step 6: Fetch the next instruction to be executed from the memory and check whether the instruction exists in the LRU cache; If it does not exist, execute step 4; if it does exist, extract the decoding result from the LRU cache and execute the instruction, and then repeat step 6.

2. The instruction execution method of the interpreted instruction set simulator based on LRU cache as claimed in claim 1, characterized in that: When the LRU cache reaches capacity, the least recently used instructions are removed.

3. An interpreted instruction set simulator based on LRU cache, characterized by: Including instruction set module, register module and memory module; The instruction set module is used to decode and execute instructions according to the method described in claim 1 or 2; The decoding operation loads instructions from the program entry address of the memory module, extracts the relevant operands and offsets and stores them in the register; the execution operation is responsible for executing the corresponding simulation function according to the result of the decoding operation and writing the result back to the register; The register module is used to define and manage the names, quantities and data types of all registers; Registers are responsible for storing the address of the next instruction to be executed and various temporary data; The memory module is used to store all instructions and data, and provides specific memory read and write functions to perform memory access operations.

4. The interpreted instruction set simulator based on LRU cache as claimed in claim 3, characterized in that: It also includes an executable file loading module; the executable file loading module is used to parse and load the ELF file, extract the entry address of the program, and map each segment in the file to the corresponding position of the memory module.

5. The interpreted instruction set simulator based on LRU cache as claimed in claim 3, characterized in that: It also includes a disassembly module; the disassembly module generates and displays corresponding disassembly results by accessing the data obtained by decoding the instruction set module using the corresponding disassembly function, and supports the disassembly display of custom instructions.

Citation Information

Patent Citations

  • Timing method for dynamic binary translation instruction set simulator

    CN103955357A

  • Optimization of instruction groups across group boundaries

    CN105593807A