A system and method for high-speed simulation of non-target instruction sets
By introducing API/ABI automatic scanning module, kernel processing module, instruction translation module, ABI conversion module and exception handling module in simulation technology, high-speed simulation of non-target instruction sets is achieved, and the problems of large program overhead, high memory usage and slow operation efficiency in existing simulation technologies are solved, which significantly improves program operation efficiency and practicality.
Patent Information
- Application Number
- CN202210789593.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-06
AI Technical Summary
When existing simulation technologies realize cross-instruction set operation, they need to simulate a complete system in full disk, resulting in large program overhead, high memory usage and slow operation efficiency.
Through the API/ABI automatic scanning module, kernel processing module, instruction translation module, ABI conversion module and exception processing module, high-speed simulation of non-target instruction sets is realized, and instructions are converted and simulated only for instruction set programs that need to be converted.
Improve program operation efficiency and reduce memory usage. The automatically scanned API/ABI information automatically contains all the information required during the ABI conversion of different instruction sets, which significantly improves practicality and progress.
Smart Images

Figure HDA0003733031040000011 
Figure HDA0003733031040000021 
Figure HDA0003733031040000031
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a system and method for high-speed simulation of non-target instruction sets. Background Art
[0002] With the continuous development of computer technology, people have higher and higher requirements for the operating efficiency of computers and their internal CPUs as productivity tools. Major technology companies have developed CPUs with different instruction sets to meet this demand. However, this development trend has led to a problem, that is, programs compiled for one instruction set can only run on the CPU of that instruction set alone, and cannot run across instruction sets. For example, a program developed for a mobile phone ARM CPU cannot run on a PC based on an INTEL CPU. In order to solve this problem, simulation technology came into being. Simulation technology achieves the purpose of running across instruction sets by translating one instruction set into another instruction set.
[0003] However, this technology has many flaws. First, in order to allow programs with another instruction set to run, the technology simulates an entire system, including the kernel, framework, and application. As a result, the program overhead is very large, the memory usage is high, and the running efficiency is slow, which needs to be improved. Summary of the invention
[0004] The object of the present invention is to provide a system and method for high-speed simulation of non-target instruction sets to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A system for high-speed simulation of a non-target instruction set, the system comprising:
[0007] API / ABI automatic scanning module, used to automatically scan the system SDK and prepare ABI conversion information for all the ABIs of the APIs that may be called;
[0008] The kernel processing module is used to modify the program type of the instruction set that the kernel can accept, so that the simulation program set can be successfully loaded into the memory;
[0009] An instruction translation module, used for translating the simulation instruction set into the target instruction set;
[0010] The ABI conversion module is responsible for converting the ABI formats of different instruction sets when the simulation program needs to interact with native instructions;
[0011] The exception handling module is responsible for ensuring that the program can smoothly switch between different instruction sets when an exception occurs and Unwind is performed.
[0012] Furthermore, the API / ABI automatic scanning module is based on LLVM open source code.
[0013] Furthermore, the instruction translation module is based on QEMU open source code.
[0014] A method for high-speed simulation of a non-target instruction set, the method for high-speed simulation of a non-target instruction set comprising the following steps:
[0015] S1, first run the API / ABI automatic scanning module to obtain and save the ABI records of all APIs that can be called by the program;
[0016] S2, then loads the kernel processing module, so that the kernel allows loading of programs simulating the instruction set;
[0017] S3, running the program simulating the instruction set. At this time, since the kernel processing module has been loaded, the instruction translation module, ABI conversion module and exception handling module will be loaded for the program simulating the instruction set;
[0018] S4, after the simulated instruction set program runs, the instruction translation module obtains the running right first, and starts the instruction translation work at this time, and starts to execute the translated instructions after completion;
[0019] S5, when the translation instruction runs to the API that needs to call the target instruction set, the ABI conversion module reads the ABI conversion information generated in step S1, and performs ABI conversion according to the ABI conversion information, so that the API of the target instruction set can be correctly called;
[0020] S6, when the API of the target instruction set needs to call back the simulated instruction set, since all callbacks are theoretically registered through the API, it is only necessary to record the API with callbacks in S5 and record the ABI information for the callback address, and the ABI conversion process is also performed by the ABI conversion module;
[0021] S7, when an exception occurs in the simulation program and unwinding is required, the exception handling module handles it and performs exception unwinding according to the mixed stack of the target instruction set and the simulation instruction set until the exception is handled.
[0022] Furthermore, in step S1, obtaining the ABI records of all APIs callable by the program includes the following steps:
[0023] Step 1: Read the SDK header file and obtain the API function prototype;
[0024] Step 2: Record the function name, number of function parameters, function parameter types, and function return type;
[0025] Step 3: If there are function type parameters, continue to record the parameters as the number of function parameters, function parameter type, and function return type.
[0026] Furthermore, in step S4, the instruction translation process is completed by the open source software qemu.
[0027] Further, in steps S5 and S6, the ABI conversion includes the following steps:
[0028] Step 1: read the ABI record created in step S1;
[0029] Step 2: Get the name or ABI information of the required function according to the target address to be called;
[0030] Step 3: Combine the ABI information read in step 1 and query the number of function parameters, function parameter types and function return type according to the function name;
[0031] Step 4: Based on the information in step 3, the caller's parameter information is read according to the ABI information. If the information is in a different form under the ABI of a different instruction set, it is adjusted at this time. After all the conversion work is completed, the converted information is written into the parameter ABI position of the instruction set of the callee.
[0032] Step 5. When the callee returns, according to the information in step 3, start reading the return information of the callee according to the ABI information. If the form is different under the ABI of different instruction sets, adjust it at this time. After all conversion work is completed, write the converted information into the return value ABI position of the caller's instruction set.
[0033] Furthermore, in step S7, the abnormal Unwind specifically includes the following steps:
[0034] Step 1: When an exception occurs, search for registered exception handling functions layer by layer on the network according to the stack content;
[0035] Step 2: If the instruction set to which the current stack belongs is the target instruction set, the exception is directly handed over to the exception handling function of the current layer to handle the exception;
[0036] Step 3: If the instruction set to which the current stack belongs is a simulated instruction set, the exception handling function is simulated and executed by an instruction simulation program;
[0037] Step 4: Repeat steps 2 and 3 until an exception handling function is found that is willing to handle the exception.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] Compared with traditional simulation systems and methods, this system and method for high-speed simulation of non-target instruction sets saves the overhead of simulating an entire system and improves program running efficiency because only the non-target instruction set program itself needs to perform instruction conversion and simulation. In addition, the automatically scanned API / ABI information automatically includes all the information required during the ABI conversion of different instruction sets, further ensuring program running efficiency. It has strong practicality and significant progress.
[0040] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0041] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 A structural schematic diagram of a system for high-speed simulation of non-target instruction sets is proposed for the present invention;
[0044] Figure 2 A flow chart of a method for high-speed simulation of non-target instruction sets is proposed for the present invention;
[0045] Figure 3 The present invention proposes an ABI record algorithm flow chart of a method for high-speed simulation of non-target instruction sets;
[0046] Figure 4 The present invention proposes a flow chart of an ABI conversion algorithm for a method of high-speed simulation of a non-target instruction set;
[0047] Figure 5 The present invention proposes a flow chart of an ABI conversion algorithm for a method of high-speed simulation of a non-target instruction set;
[0048] Figure 6 The present invention proposes an abnormal Unwind algorithm flow chart of a method for high-speed simulation of non-target instruction sets. DETAILED DESCRIPTION
[0049] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0050] Embodiment 1
[0051] In this embodiment, it includes:
[0052] System SDK, version is iOS 13.4.1;
[0053] Non-target instruction program, whose instruction set is arm64 instruction;
[0054] The system framework and operating system kernel are both based on OSX Catalina 10.15.1.
[0055] See also Figure 1 , a system for high-speed simulation of non-target instruction sets, the system for high-speed simulation of non-target instruction sets comprising:
[0056] API / ABI automatic scanning module, based on LLVM open source code, is used to automatically scan the system SDK and prepare ABI conversion information for all APIs that may be called, while using clang's syntax analyzer;
[0057] The kernel processing module is used to modify the program type of the instruction set that the kernel can accept, so that the simulation program set can be successfully loaded into the memory;
[0058] Instruction translation module, based on QEMU open source code, used to translate the simulation instruction set into the target instruction set;
[0059] The ABI conversion module is responsible for converting the ABI formats of different instruction sets when the simulation program needs to interact with native instructions;
[0060] The exception handling module is responsible for ensuring that the program can smoothly switch between different instruction sets when an exception occurs and Unwind is performed.
[0061] Embodiment 2
[0062] See also Figure 2-6 , a method for high-speed simulation of non-target instruction sets, the method for high-speed simulation of non-target instruction sets comprising:
[0063] S1, first run the API / ABI automatic scanning module to obtain and save the ABI records of all APIs that can be called by the program;
[0064] Specifically, obtaining the ABI records of all APIs that can be called by the program includes the following steps:
[0065] Step 1: Read the SDK header file and obtain the API function prototype;
[0066] Step 2: Record the function name, number of function parameters, function parameter types, and function return type;
[0067] The specific steps include:
[0068] I. Save the return type information of the function under AArch64;
[0069] II. Save the return type information of the function under x86_64;
[0070] III. Save the number of function parameters;
[0071] IV. Save the type information of each function parameter under AArch64;
[0072] V. Save the type information of each function parameter under x86_64;
[0073] VI. Save the total size of all parameters of the function under AArch64;
[0074] VII. Save the total size of all parameters of the function under x86_64;
[0075] VIII. Save the number of SSE registers that the function will use under x86_64;
[0076] IX. Save whether the function is a C++ constructor;
[0077] X. Save whether the function is a variable-length parameter function.
[0078] Step 3: If there is a function type parameter, continue to record the parameter as the function parameter number, function parameter type, and function return type;
[0079] S2, then loads the kernel processing module, so that the kernel allows loading of programs simulating the instruction set;
[0080] S3, running the program simulating the instruction set. At this time, since the kernel processing module has been loaded, the instruction translation module, ABI conversion module and exception handling module will be loaded for the program simulating the instruction set;
[0081] S4, after the simulated instruction set program runs, the instruction translation module obtains the running right first, and starts the instruction translation work at this time. After completion, the translated instructions are executed. The instruction translation process is completed by the open source software qemu;
[0082] S5, when the translation instruction runs to the API that needs to call the target instruction set, the ABI conversion module reads the ABI conversion information generated in step S1, and performs ABI conversion according to the ABI conversion information, so that the API of the target instruction set can be correctly called;
[0083] The specific process of the above step S5 is as follows:
[0084] When calling the entry point of one instruction set to another:
[0085] I. If the return value is Indirect and in memory on x86_64, or Direct and in a register on AArch64, reserve space on the stack for the return value.
[0086] II. If the return value is Indirect in both x86_64 and AArch64, then load the register contents that are the structure return value in AArch64 as the first argument in x86_64;
[0087] III. Start converting parameters;
[0088] IV. If the argument is Direct and in a register in both x86_64 and AArch64, jump to V, otherwise jump to VI;
[0089] V. If the number of registers used by the parameter in x86_64 and AArch64 is the same, then fetch it from the corresponding source register and store it in the corresponding target register. Otherwise, concatenate the contents stored in more registers into fewer registers, and then jump to XI.
[0090] VI. If the parameter is Direct in both x86_64 and AArch64, the former is in a register and the latter is in memory, then transfer it from memory to the register and then jump to XI;
[0091] VII. If the parameter is Direct in both x86_64 and AArch64, the former is in memory and the latter is in register, then transfer it from register to memory and then jump to XI;
[0092] VIII. If the parameter is Direct and exists in memory under both x86_64 and AArch64, fetch it from the corresponding source memory and store it in the corresponding target memory location, then jump to XI;
[0093] IX. If the argument is Direct and in memory on x86_64, or Indirect and in a register on AArch64, get the argument size and copy it from memory, transfer it to a register, and then jump to XI.
[0094] X. If the argument is Direct and in memory on x86_64 or Indirect and in memory on AArch64, get the argument size and copy it from the source memory location to the target memory location.
[0095] XI. If the current parameter is not the last parameter, return to III, otherwise end.
[0096] When returning from one instruction set to another instruction set's exit:
[0097] I. When AArch64 returns to x86_64, if this function is a C++ constructor, the this pointer is used as the return value;
[0098] II. If the return value is Direct and in a register in both x86_64 and AArch64, jump to III, otherwise jump to IV;
[0099] III. If the return value uses the same number of registers under x86_64 and AArch64, fetch it from the corresponding source register and store it in the corresponding target register. Otherwise, concatenate the contents stored in more registers into fewer registers and then end.
[0100] If the return value is Indirect in x86_64 and Direct in AArch64, the value is taken from the AArch64 return register and stored in the memory address of x86_64.
[0101] S6, when the API of the target instruction set needs to call back the simulated instruction set, since all callbacks are theoretically registered through the API, it is only necessary to record the API with callbacks in S5 and record the ABI information for the callback address, and the ABI conversion process is also performed by the ABI conversion module;
[0102] Specifically, ABI conversion includes the following steps:
[0103] Step 1: read the ABI record created in step S1;
[0104] Step 2: Get the name or ABI information of the required function according to the target address to be called;
[0105] Step 3: Combine the ABI information read in step 1 and query the number of function parameters, function parameter types and function return type according to the function name;
[0106] Step 4: Based on the information in step 3, the caller's parameter information is read according to the ABI information. If the information is in a different form under the ABI of a different instruction set, it is adjusted at this time. After all the conversion work is completed, the converted information is written into the parameter ABI position of the instruction set of the callee.
[0107] Step 5: When the callee returns, according to the information in step 3, the return information of the callee is read according to the ABI information. If the form is different under the ABI of different instruction sets, it is adjusted at this time. After all conversion work is completed, the converted information is written into the return value ABI position of the caller's instruction set;
[0108] S7, when an exception occurs in the simulation program and unwinding is required, the exception handling module handles it and performs the exception unwinding according to the mixed stack of the target instruction set and the simulation instruction set until the exception is handled;
[0109] Specifically, abnormal Unwind includes the following steps:
[0110] Step 1: When an exception occurs, search for registered exception handling functions layer by layer on the network according to the stack content;
[0111] Step 2: If the instruction set to which the current stack belongs is the target instruction set, the exception is directly handed over to the exception handling function of the current layer to handle the exception;
[0112] Step 3: If the instruction set to which the current stack belongs is a simulated instruction set, the exception handling function is simulated and executed by an instruction simulation program;
[0113] Step 4: Repeat steps 2 and 3 until an exception handling function is found that is willing to handle the exception.
[0114] It can be seen from the above embodiments that, compared with traditional simulation systems and methods, since only the non-target instruction set program itself needs to perform instruction conversion and simulation, the overhead of simulating an entire system is eliminated, the program running efficiency is improved, and the automatically scanned API / ABI information automatically includes all the information required during the ABI conversion of different instruction sets, further ensuring the program running efficiency, with strong practicality and significant progress.
[0115] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A system for high-speed simulation of non-target instruction sets, characterized in that: The system for high-speed simulation of non-target instruction sets comprises: API / ABI automatic scanning module, used to automatically scan the system SDK and prepare ABI conversion information for all the ABIs of the APIs that may be called; The kernel processing module is used to modify the program type of the instruction set that the kernel can accept, so that the simulation program set can be successfully loaded into the memory; An instruction translation module, used for translating the simulation instruction set into the target instruction set; The ABI conversion module is responsible for converting the ABI formats of different instruction sets when the simulation program needs to interact with native instructions; The exception handling module is responsible for ensuring that the program can smoothly switch between different instruction sets when an exception occurs and Unwind is performed.
2. The system for high-speed simulation of non-target instruction sets according to claim 1, characterized in that: The API / ABI automatic scanning module is based on LLVM open source code.
3. The system for high-speed simulation of non-target instruction sets according to claim 1, characterized in that: The instruction translation module is based on QEMU open source code.
4. A method for high-speed simulation of non-target instruction sets, characterized in that: The method for high-speed simulation of non-target instruction sets comprises the following steps: S1, first run the API / ABI automatic scanning module to obtain and save the ABI records of all APIs that can be called by the program; S2, then loads the kernel processing module, so that the kernel allows loading of programs simulating the instruction set; S3, running the program simulating the instruction set. At this time, since the kernel processing module has been loaded, the instruction translation module, ABI conversion module and exception handling module will be loaded for the program simulating the instruction set; S4, after the simulated instruction set program runs, the instruction translation module obtains the running right first, and starts the instruction translation work at this time, and starts to execute the translated instructions after completion; S5, when the translation instruction runs to the API that needs to call the target instruction set, the ABI conversion module reads the ABI conversion information generated in step S1, and performs ABI conversion according to the ABI conversion information, so that the API of the target instruction set can be correctly called; S6, when the API of the target instruction set needs to call back the simulated instruction set, since all callbacks are theoretically registered through the API, it is only necessary to record the API with callbacks in S5 and record the ABI information for the callback address, and the ABI conversion process is also performed by the ABI conversion module; S7, when an exception occurs in the simulation program and unwinding is required, the exception handling module handles it and performs the exception unwinding according to the mixed stack of the target instruction set and the simulation instruction set until the exception is handled; In steps S5 and S6, the ABI conversion includes the following steps: Step 1: read the ABI record created in step S1; Step 2: Get the name or ABI information of the required function according to the target address to be called; Step 3: Combine the ABI information read in step 1 and query the number of function parameters, function parameter types and function return type according to the function name; Step 4: Based on the information in step 3, the caller's parameter information is read according to the ABI information. If the information is in a different form under the ABI of a different instruction set, it is adjusted at this time. After all the conversion work is completed, the converted information is written into the parameter ABI position of the instruction set of the callee. Step 5. When the callee returns, according to the information in step 3, start reading the return information of the callee according to the ABI information. If the form is different under the ABI of different instruction sets, adjust it at this time. After all conversion work is completed, write the converted information into the return value ABI position of the caller's instruction set.
5. The method for high-speed simulation of non-target instruction sets according to claim 4, characterized in that: In step S1, obtaining the ABI records of all APIs callable by the program includes the following steps: Step 1: Read the SDK header file and obtain the API function prototype; Step 2: Record the function name, number of function parameters, function parameter types, and function return type; Step 3: If there are function type parameters, continue to record the parameters as the number of function parameters, function parameter type, and function return type.
6. The method for high-speed simulation of non-target instruction sets according to claim 4, characterized in that: In step S4, the instruction translation process is completed by the open source software qemu.
7. The method for high-speed simulation of non-target instruction sets according to claim 4, characterized in that: In step S7, the abnormal Unwind specifically includes the following steps: Step 1: When an exception occurs, search for registered exception handling functions layer by layer on the network according to the stack content; Step 2: If the instruction set to which the current stack belongs is the target instruction set, the exception is directly handed over to the exception handling function of the current layer to handle the exception; Step 3: If the instruction set to which the current stack belongs is a simulated instruction set, the exception handling function is simulated and executed by an instruction simulation program; Step 4: Repeat steps 2 and 3 until an exception handling function is found that is willing to handle the exception.
Citation Information
Patent Citations
System and method for execution of application code compiled according to two instruction set architectures
CN107077337A
Dynamic allocation of executable code for multi-architecture heterogeneous computing
US11113059B1