A processor
By introducing a dual instruction cache structure into the processor, the processor's performance problem is solved when compatible with different instruction systems, the efficient compatibility effect of hardware translation is achieved, and the binary translation performance of cross-instruction systems is improved.
Patent Information
- Application Number
- CN202210578451.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-05-25
AI Technical Summary
When existing processors are compatible with different instruction systems, there are problems with insufficient software translation performance, especially due to dynamic relocation of indirect transfer of target addresses, switching and saving of contextual environments in and out of translation modules, translation basic block code trace connection and address space overlap.
A dual instruction cache processor is designed, including the host-level instruction cache and the client-level instruction cache. By selecting the circuit to switch the use of the cache in different states, it ensures that the translated code has an independent operating environment and address space, and uses hardware translation technology.
It improves the binary translation performance of cross-instruction systems, realizes efficient compatibility of processors among different instruction systems, reduces hardware costs and improves the performance of software translation.
Smart Images

Figure CN114968576B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of processor design. Specifically, it relates to the field related to processor instruction sets. More specifically, it relates to a processor that supports cross-instruction sets, that is, a processor applied to a dual-instruction cache. Background Art
[0002] As the arithmetic and control core of a computer system, a processor is the final execution unit for information processing and program execution. As Figure 1 shown, existing processors all have only one level-1 instruction cache (L1 IC) and one level-1 data cache (L1 DC). The instruction fetch component of the processor reads instructions from the instruction cache and hands them over to the execution component for execution. When the execution component executes instructions, it reads / writes the data cache as needed. The design of processors is all for efficiently implementing a specific instruction set, such as the X86 instruction set, ARM instruction set, MIPS instruction set, RISCV instruction set, and so on. Existing processor design methods can efficiently execute the code of the specific instruction set specified in the design. However, if it is also necessary to execute the code of other instruction sets at the same time, only the software binary translation technology can be adopted. The translation itself from the code of one instruction set to the code of another instruction set does not cost too much, which is equivalent to replacing the hardware decoding with software decoding. For example, the Java virtual machine adopts the binary translation technology, and the performance impact is not significant. However, compared with the Java virtual machine, the binary translation of code between different instruction sets has a series of additional costs that are difficult to solve by software. This is mainly because in software translation, sharing the same address space and running environment between the translation module and the execution module will cause dynamic relocation of the target address of indirect jumps, saving the context switching when entering and exiting the translation module, connecting the code traces of translated basic blocks, and overlapping the address spaces of the translation module and the translated code, etc., making the performance of binary translation of code between different instruction sets relatively poor. Summary of the Invention
[0003] Therefore, the object of the present invention is to overcome the defects of the above-mentioned prior art and provide a processor that can support software translation and avoid problems such as dynamic relocation of the target address of indirect jumps, saving the context switching when entering and exiting the translation module, connecting the code traces of translated basic blocks, and overlapping the address spaces of the translation module and the translated code.
[0004] According to a first aspect of the present invention, there is provided a processor which can operate in a host state or a virtual guest state according to application requirements. The host state and the virtual guest state respectively correspond to different instruction sets. The processor includes: a host first-level instruction cache for caching instructions executable by the processor operating in the host state corresponding to the instruction set of the host state; a guest first-level instruction cache for caching instructions executable by the processor operating in the virtual guest state corresponding to the instruction set of the virtual guest state; and a selection circuit for selecting to connect the host first-level instruction cache or the guest first-level instruction cache to the corresponding components of the processor according to the operating state of the processor, so that when the processor operates in the host state, it fetches instructions from the host first-level instruction cache, and when the processor operates in the virtual guest state, it fetches instructions from the guest first-level instruction cache.
[0005] Preferably, in the host state, there runs: a binary translation program which is executable by the processor operating in the host state and is used to translate the source program corresponding to the instruction set of the virtual guest state into code executable by the processor operating in the virtual guest state and store it in the guest first-level instruction cache. Wherein, the guest first-level instruction cache is configured to support the read and write of the execution components of the processor when the processor operates in the host state and support the access read and write outside the processor.
[0006] In some embodiments of the present invention, the guest first-level instruction cache is configured as a multi-way set-associative cache structure, and each cache line includes: a tag field for indicating the value of the source program counter before the source program corresponding to the cache line where it is located is translated; a continuation line field for indicating whether there is a continuation line in the translated virtual guest code; a code length field for storing the code length of the cache line where it is located; and a translated instruction code field.
[0007] Preferably, when the processor operates in the virtual guest state, it is configured to fetch instructions from the guest first-level instruction cache in the following manner: use the value of the source program counter in the virtual guest state to look up the cache line in the guest first-level instruction cache, return the instructions of the cache line whose source program counter value matches the corresponding tag field, and calculate the source program counter value for the next instruction fetch according to the source program counter value indicated by the tag field and the code length field in the cache line where the instruction fetch is finally completed. In some embodiments of the present invention, the processor is further configured to determine whether to continue looking for a continuation line in the adjacent next set of cache lines according to the indication of the continuation line field.
[0008] Preferably, the source program counter value for the next instruction fetch is the sum of the source program counter value indicated by the tag field of the cache line where the instruction fetch is completed this time and the value of the code length field
[0009] In some embodiments of the present invention, the processor is configured to: when an interrupt is generated due to a cache miss in the client-level instruction cache, run in the host state in response to the interrupt, and call a binary translator to complete the subsequent unfinished translation of the source program corresponding to the instruction system of the virtual client state and write the translated code into the client-level instruction cache.
[0010] According to a second aspect of the present invention, there is provided a method for running a dual instruction system based on the processor described in the first aspect of the present invention. The method includes: in response to an application requirement, running the processor in a state corresponding to the application requirement, where the processor can run in the host state or the virtual client state, and the host state and the virtual client state respectively correspond to different instruction systems; based on the state corresponding to the application requirement, selecting to connect the host-level instruction cache or the client-level instruction cache to the corresponding components of the processor, so that when the processor runs in the host state, it fetches instructions from the host-level instruction cache, and when the processor runs in the virtual client state, it fetches instructions from the client-level instruction cache.
[0011] Compared with the prior art, the advantages of the present invention are as follows: The present invention addresses the problem of insufficient performance when an existing processor uses binary translation technology to be compatible with programs of other instruction systems, and proposes a processor design method with dual instruction caches, which can greatly improve the binary translation performance of being compatible with other instruction systems. As can be seen from the above embodiments, the present invention enables the processor to efficiently execute the code of other instruction systems by adding a small amount of hardware support, and solves the problem of efficient software compatibility across instruction systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The following further describes the embodiments of the present invention with reference to the accompanying drawings, where:
[0013] Figure 1 It is a schematic diagram of the brief structure of a processor under the prior art according to an embodiment of the present invention;
[0014] Figure 2 It is a schematic diagram of the brief structure of a dual-instruction-cache processor according to an embodiment of the present invention;
[0015] Figure 3 It is a schematic diagram of the cache line data structure in the client-level instruction cache of a dual-instruction-cache processor according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] During the process of researching and designing processors, the inventors found that processors with complex instruction sets (such as X86 processors) are actually binary translation systems for different instruction sets. However, this binary translation system is implemented in hardware, which translates complex X86 instructions into microcode instructions similar to those of the reduced instruction set internally, while the performance of the X86 processor itself is not affected at all. The biggest difference between hardware translation and software translation is precisely that the translation component and the execution component have independent operating environments and address spaces. Therefore, problems such as dynamic relocation of indirect transfer target addresses that cannot be overcome by software translation do not exist in hardware translation. However, the cost of using hardware translation far exceeds that of software translation. It can be seen from this that if software binary translation can have an independent operating environment and address space during operation, similar to hardware translation, through processor design improvements, software binary translation will have performance comparable to that of hardware translation, which can greatly reduce hardware costs and improve the performance of software translation.
[0018] The technical idea of the present invention will be briefly introduced below.
[0019] Modern high-performance processors all support CPU virtualization technology. That is, on the basis of the original operating state of the processor, a guest operating state is added to run virtual guest programs, and the original operating state is used to run local programs, that is, host programs. Based on this, the inventor proposes a processor design scheme: the original operating state of the processor is called the host state, and the binary translation program runs in the host state, and the translated code runs in the guest state, so that the binary translation program and the translated code can have independent operating environments and address spaces. Among them, from the perspective of practical applications, the hardware instruction system actually executed by the virtual guest state is the same as that of the host state. Through binary translation technology, the virtual guest state appears to be running an instruction system different from that of the host state. However, since the size of the translated code has changed compared with the original source code, usually it has become larger, which is called code expansion. In order to keep the running address space of the translated code consistent with the original source code, the organization and storage method of the translated code in the address space should be different from that of other programs on the host. And the existing processors only have a single-level instruction cache, which is shared by the host program and the translated code of the guest. This leads to the translated code of the guest breaking the independence of its own address space in order to use this shared single-level instruction cache. Therefore, the inventor proposes to add a dedicated guest single-level instruction cache in the processor that supports CPU virtualization technology, which is specifically used when the processor runs virtual guest programs, so that the above problems caused by sharing the single-level instruction cache can be completely solved, thus achieving similar performance for software binary translation and hardware translation.
[0020] According to an embodiment of the present invention, as Figure 2As shown in the figure, the present invention provides a processor design scheme with a dual-instruction cache. The processor includes two level-1 instruction caches. One is a traditional level-1 instruction cache (L1 IC) for the host state, and the other is a level-1 instruction cache (L1 GIC) for the virtual guest state. The two instruction caches are connected to the instruction fetch component of the processor through a multiplexer circuit, and a selection signal sel is used to select from which instruction cache the processor fetches instructions. When the processor runs in the virtual guest state, the sel selection signal controls the multiplexer circuit to select the path from the instruction fetch component to the L1 GIC, so that the instruction fetch component fetches instructions from the guest level-1 instruction cache. When the processor runs in the host state, the sel selection signal controls the multiplexer circuit to select the path from the instruction fetch component to the L1 IC, so that the instruction fetch component fetches instructions from the host level-1 instruction cache. According to an embodiment of the present invention, the L1 GIC can be read / written by the execution component of the processor and can also be read / written by accesses from outside the processor. Usually, the execution component is only allowed to read / write the L1 GIC when the processor works in the host state. Since the binary translation program runs in the host state and the translated code runs in the guest state, allowing the execution component to read / write the L1 GIC can implement storing the translated code into it, so that the processor allows fetching instructions from the L1 GIC when running in the virtual guest state.
[0021] According to an embodiment of the present invention, the organization of the L1 GIC is the same as that of the traditional L1 IC, adopting a multi-way set-associative cache structure. Therefore, this structure will not be elaborated here. However, there are two differences in the specific cache line structure of the L1 GIC compared with the traditional L1 IC. Its structure is as Figure 3 shown. Among them, PC is the source instruction code program counter value (tag field), C is the continuation line flag field, CLEN is the source instruction code length field, and Inst is the translated instruction code. Specifically, the differences between the L1 GIC and the L1 IC are as follows:
[0022] The first difference: In the cache line of the L1 GIC, the tag field is the source program counter value (PC) before translation of the guest, while the tag field in the traditional L1 IC is the physical memory address of this line of code.
[0023] The second difference: The cache line of the L1 GIC adds two new fields: the continuation line field (C) and the source program code length field (CLEN). Among them, the continuation line field indicates whether there is a continuation line in the translated code. If the continuation line field is 1, it means there is a continuation line. If the continuation line field is 0, it means there is no continuation line. If there is a continuation line, it is necessary to continue fetching instructions in the next cache line; the source program code length field (CLEN) is used to calculate the source program counter value before translation of the subsequent code (the next fetched instruction).
[0024] In addition, the present invention also designs the processor as follows: If a failure occurs in the client-level instruction cache L1 GIC, a client high-speed instruction cache miss interrupt / exception will be generated, and this interrupt / exception can be sent to this processor for processing or to an external processor for processing. If it is sent to this processor, the processor will return from the virtual client state to the host state to respond to this interrupt / exception. For example, the subsequent translation of the client instruction code can be completed by calling a translator in the host state and writing it into the client instruction cache L1 GIC to respond to this interrupt / exception.
[0025] To better understand the working principle of the processor of the present invention, the following takes a single processor supporting CPU virtualization technology as an example for illustration:
[0026] The processor initially runs in the host state. When it is necessary to execute a client program across instruction systems, it enters the virtual client state by responding to the instruction to enter the client state.
[0027] When the processor enters the client state, the selection signal sel of the two-way selection circuit is set to make the instruction fetch component of the processor switch to fetch instructions from the client-level instruction cache. For example, the sel signal when entering the client state is 1, and the sel signal when returning to the host state is 0. Different instruction caches are selected for instruction fetching by setting different values of 0 / 1.
[0028] The processor uses the source program counter value of the client program to search for a cache line in the client-specific instruction cache. If the value in the PC field of the cache line matches the source program counter value of the client program to be searched, it means a hit. If there is a hit, the instruction of the hit cache line is returned. At the same time, if the continuation line field of the cache line is 1, the continuation line is searched according to a preset continuation line search strategy (for example, continue to search for the continuation line in the adjacent next set of cache lines, but the continuation line search strategy is not limited to this strategy and can be set to any other convenient search strategy according to requirements). It should be noted that when the processor runs in the client state, it sequentially executes the instructions of the code line that hits the instruction fetch in L1 GIC. After executing one line of instructions, it is necessary to calculate the code line that needs to be executed after the execution of this code line, that is, it is necessary to calculate the source program counter value for the next instruction fetch. If the continuation line field is 0, the source program counter value for the next instruction fetch is calculated according to the program counter value and the code length field of this cache line, that is, the source program counter value for the next instruction fetch = the PC value of this cache line + the CLEN value of this cache line; if the continuation line field is 1, the source program counter value for the next instruction fetch is calculated in the same way after completing the continuation line fetch, that is, the source program counter value for the next instruction fetch = the PC value of the cache line where the fetch is completed + the CLEN value of the cache line where the fetch is completed.
[0029] If the client-level instruction cache lookup fails, a client instruction cache miss exception is generated. The processor returns to the host state, executes the binary translator, updates the client-level instruction cache through the processor execution unit. After the update is completed, it re-enters the virtual client state to continue executing the client program across instruction systems.
[0030] In view of the performance deficiency problem of existing processors when compatible with programs of other instruction systems through binary translation technology, the present invention proposes a processor design method with a dual instruction cache, which can greatly improve the binary translation performance of compatible programs of other instruction systems. It can be seen from the above embodiments that the present invention enables the processor to efficiently execute the code of other instruction systems by adding a small amount of hardware support, thus solving the problem of efficient software compatibility across instruction systems.
[0031] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order, as long as the required functions can be achieved.
[0032] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0033] A computer-readable storage medium may be a tangible device that retains and stores instructions for use by an instruction execution device. A computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.
[0034] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A processor that can operate in a host state or a virtual guest state according to application requirements, where the host state and the virtual guest state respectively correspond to different instruction sets, and is characterized in that, The processor includes: A host first-level instruction cache for caching instructions of the instruction system corresponding to the host state that can be executed by the processor running in the host state; A guest first-level instruction cache for caching instructions of the instruction system corresponding to the virtual guest state that can be executed by the processor running in the virtual guest state; and A selection circuit for selecting to connect either the host first-level instruction cache or the guest first-level instruction cache to the corresponding components of the processor according to the running state of the processor, so that when the processor runs in the host state, it fetches instructions from the host first-level instruction cache, and when the processor runs in the virtual guest state, it fetches instructions from the guest first-level instruction cache.
2. The processor according to claim 1, wherein When running in the host state: A binary translation program that can be executed by the processor running in the host state to translate the source program of the instruction system corresponding to the virtual guest state into code executable by the processor running in the virtual guest state and store it in the guest first-level instruction cache.
3. The processor according to claim 2, wherein The guest first-level instruction cache is configured to support read and write operations of the processor execution components when the processor runs in the host state and support read and write accesses external to the processor.
4. The processor according to claim 3, characterized in that, The guest first-level instruction cache is configured as a multi-way set-associative cache structure, where each cache line includes: A tag field for indicating the source program counter value of the source program corresponding to the cache line before translation of the virtual guest; A continuation field for indicating whether there is a continuation in the translated virtual guest code; A code length field for storing the code length of the cache line where it is located; And a translated instruction code field.
5. The processor according to claim 4, wherein When the processor runs in the virtual guest state, it is configured to fetch instructions from the guest first-level instruction cache in the following manner: Use the source program counter value of the virtual guest state to look up a cache line in the guest first-level instruction cache, return the instructions of the cache line where the source program counter value matches the corresponding tag field, and calculate the source program counter value for the next instruction fetch according to the source program counter value indicated by the tag field and the code length field in the cache line where the instruction fetch is finally completed.
6. The processor according to claim 5, wherein The processor is further configured to determine whether to continue looking for a continuation in the adjacent next set of cache lines according to the indication of the continuation field.
7. The processor according to claim 5, characterized in that, The source program counter value for the next instruction fetch is the sum of the source program counter value indicated by the tag field of the cache line where the current instruction fetch is completed and the value of the code length field.
8. The processor according to claim 2, wherein The processor is configured to: when an interrupt is generated due to a miss in the guest first-level instruction cache, run in the host state to respond to the interrupt, and call the binary translation program to complete the subsequent unfinished translation of the source program of the instruction system corresponding to the virtual guest state and write the translated code into the guest first-level instruction cache.
9. A method for running a dual instruction system based on the processor according to any one of claims 1-8, characterized in that, The method includes: In response to an application requirement, running the processor in a state corresponding to the application requirement, where the processor can run in the host state or the virtual guest state, and the host state and the virtual guest state respectively correspond to different instruction systems; Based on the state corresponding to the application requirement, select to connect the host-level instruction cache or the guest-level instruction cache to the corresponding components of the processor, so that when the processor runs in the host state, it fetches instructions from the host-level instruction cache, and when the processor runs in the virtual guest state, it fetches instructions from the guest-level instruction cache.
10. An electronic device, characterized in that, The device includes: A storage device; One or more processors as described in any one of claims 1-8.
Citation Information
Patent Citations
Processor and instruction fetching method thereof
CN117270970A