Operating system emulator across instruction sets
Patent Information
- Application Number
- CN202610920268.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明提供一种跨指令集的操作系统模拟器,用以解决现有技术中全系统模拟方法翻译开销大、执行效率低,而用户态兼容方法通用性差、无法利用宿主原生动态链接库加速、难以适配复杂应用的缺陷,整个模拟器通过动态翻译层在翻译执行闭源客户机可执行文件时,如需调用动态链接库,在有宿主机动态链接库可用时直接调用其中对应的宿主机函数接口,避免冗余翻译,仅在无宿主机动态链接库可用时才将客户机动态链接库的指令翻译为宿主机指令并执行,从而在保证通用性的前提下,最大限度减少动态翻译范围,降低翻译开销,并利用宿主机动态链接库实现执行加速,使复杂应用能够高效运行,从而有效突破现有模拟方法在性能和兼容性方面的局限,为在x86_64硬件平台上高效运行鸿蒙等新兴移动操作系统及其应用生态提供可行技术路径
[0016]本发明提供的跨指令集的操作系统模拟器,包括动态翻译层,所述动态翻译层包括动态翻译引擎和跨指令集应用二进制接口ABI调用层;其中,所述动态翻译引擎,用于在运行第一闭源客户机可执行文件的情况下,将客户机指令翻译为宿主机指令,并执行;所述跨指令集ABI调用层,用于在运行第二闭源客户机可执行文件,且所述第二闭源客户机可执行文件需调用动态链接库的情况下,若操作系统中存在宿主机动态链接库,则根据所述第二闭源客户机可执行文件中的客户机指令,调用所述宿主机动态链接库中与客户机函数接口对应的宿主机函数接口,所述客户机指令与所述客户机函数接口对应;所述动态翻译引擎,还用于在运行所述第二闭源客户机可执行文件,且所述第二闭源客户机可执行文件需调用动态链接库的情况下,若所述操作系统中仅存在客户机动态链接库,则翻译执行所述客户机动态链接库。整个模拟器通过动态翻译层在翻译执行闭源客户机可执行文件时,如需调用动态链接库,在有宿主机动态链接库可用时直接调用其中对应的宿主机函数接口,避免冗余翻译,仅在无宿主机动态链接库可用时才将客户机动态链接库的指令翻译为宿主机指令并执行,从而在保证通用性的前提下,最大限度减少动态翻译范围,降低翻译开销,并利用宿主机动态链接库实现执行加速,使复杂应用能够高效运行,从而有效突破现有模拟方法在性能和兼容性方面的局限,为在x86_64硬件平台上高效运行鸿蒙等新兴移动操作系统及其应用生态提供可行技术路径。
Smart Images

Figure CN122837982A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer architecture and operating system virtualization technology, and in particular to a cross-instruction set operating system simulator. Background Technology
[0002] With the rapid development of operating system (such as HarmonyOS and Android) ecosystems, mobile terminal software systems based on the ARM64 instruction set have formed a relatively complete application and service system. Meanwhile, given the abundance of x86_64 hardware platform resources, mature development toolchains, and the high penetration rate of servers and personal computers, how to efficiently run ARM64 architecture mobile operating systems on the x86_64 hardware platform, thus facilitating mobile software developers to efficiently develop and debug ARM64 mobile terminal software in an x86_64 desktop computer environment, has become an important issue.
[0003] Existing cross-architecture runtime technologies can generally be divided into two categories: those based on full-system simulation and those based on user-space compatibility layers. However, the former suffers from drawbacks such as high translation overhead and low execution efficiency, while the latter has disadvantages such as poor versatility, inability to utilize native dynamic link libraries of the host system for acceleration, and difficulty in adapting to complex applications.
[0004] Therefore, there is an urgent need to propose a new cross-instruction set operating system simulator to effectively overcome the limitations of existing simulation methods in terms of performance and compatibility. Summary of the Invention
[0005] This invention provides a cross-instruction set operating system simulator to address the shortcomings of existing full-system simulation methods, such as high translation overhead and low execution efficiency, as well as user-mode compatible methods, which suffer from poor versatility, inability to utilize native host dynamic link libraries for acceleration, and difficulty in adapting to complex applications. The simulator, through a dynamic translation layer, translates and executes closed-source client executables. When dynamic link libraries are needed, it directly calls the corresponding host function interfaces if available, avoiding redundant translation. Only when no host dynamic link libraries are available are the client dynamic link library instructions translated into host instructions and executed. This minimizes the scope of dynamic translation and reduces translation overhead while ensuring versatility, and utilizes host dynamic link libraries for execution acceleration, enabling complex applications to run efficiently. This effectively overcomes the performance and compatibility limitations of existing simulation methods, providing a feasible technical path for efficiently running emerging mobile operating systems such as HarmonyOS and their application ecosystems on x86_64 hardware platforms.
[0006] This invention provides a cross-instruction set operating system simulator, including a dynamic translation layer. The dynamic translation layer comprises a dynamic translation engine and a cross-instruction set application binary interface (ABI) calling layer. Specifically, the dynamic translation engine, when running a first closed-source client executable, translates client instructions into host instructions and executes them. The cross-instruction set ABI calling layer, when running a second closed-source client executable that requires calling dynamic link libraries, if a host dynamic link library exists in the operating system, calls the host function interface corresponding to the client function interface in the host dynamic link library based on the client instructions in the second closed-source client executable. The dynamic translation engine is further configured to, when running the second closed-source client executable that requires calling dynamic link libraries, if only a client dynamic link library exists in the operating system, translate and execute the client dynamic link library.
[0007] According to the present invention, a cross-instruction set operating system simulator includes a dynamic translation engine comprising an engine front-end, an optimizer, and a translation back-end. The dynamic translation engine is used to translate guest instructions into host instructions, comprising: the engine front-end, for each basic block in a second closed-source guest executable file, translating the guest instructions within the basic blocks into an intermediate representation of a first underlying virtual machine (LLVM IR); the optimizer, for performing register allocation optimization on the first LLVM IR to obtain a second LLVM IR; and performing further optimization on the second LLVM IR to obtain a third LLVM IR; and the translation back-end, for compiling the third LLVM IR into the host instructions.
[0008] According to the present invention, a cross-instruction set operating system simulator includes an optimizer configured to optimize register allocation in a first LLVM IR to obtain a second LLVM IR. Specifically, the optimizer is configured to: obtain the register access sequence in the first LLVM IR; uniformly represent different access widths of the same register in the register access sequence; convert the uniform access widths into a consistent register value propagation form and output a fourth LLVM IR with unified register access semantics; reduce redundant load / store operations on the same register in the fourth LLVM IR within each basic block to obtain a fifth LLVM IR; and delete redundant load / store operations in the fifth LLVM IR across basic blocks according to the control flow relationship between the basic blocks to obtain the second LLVM IR.
[0009] According to the present invention, a cross-instruction set operating system simulator is provided, wherein the second closed-source guest executable file is a guest binary file; the cross-instruction set ABI calling layer is used to call the host machine function interface corresponding to the guest function interface in the host machine dynamic link library according to the guest instructions in the second closed-source guest executable file, including: the cross-instruction set ABI calling layer is specifically used to perform parameter conversion on the guest function interface according to the function signature when it is recognized that the target address to be jumped to in the guest binary file points to the host machine dynamic link library, so as to obtain the host machine function interface.
[0010] According to the present invention, a cross-instruction set operating system simulator further includes a native execution layer; wherein, the cross-instruction set ABI calling layer is further configured to, when the native execution layer executes a native function on the host instruction corresponding to the host function interface and obtains a function return value, convert the function return value from the register corresponding to the host instruction back to the register corresponding to the guest instruction.
[0011] According to the present invention, an operating system simulator across instruction sets is provided, wherein the optimizer is used to perform optimization cycles on the second LLVM IR to obtain a third LLVM IR, comprising: the optimizer specifically performing loop optimization, dead code elimination, and vectorization on the second LLVM IR to obtain the third LLVM IR.
[0012] According to the present invention, a cross-instruction set operating system simulator further includes: a native execution layer; wherein the native execution layer is used to run the host dynamic link library when running the second closed-source client executable file and the second closed-source client executable file needs to call dynamic link libraries.
[0013] According to the present invention, a cross-instruction set operating system emulator further includes: a hardware virtualization layer; wherein the hardware virtualization layer is used to provide a standard virtualization environment by utilizing the KVM hardware-accelerated virtualization technology based on the fast emulator QEMU, the standard virtualization environment being used to run a target operating system based on the host machine instruction set.
[0014] According to the present invention, a cross-instruction set operating system emulator includes a host dynamic link library comprising the system library corresponding to the operating system and the host instruction set natively compiled version of the third-party software vendor library.
[0015] According to the present invention, a cross-instruction set operating system simulator is provided, wherein the system library includes system services, a user interface (UI) framework, and a graphics library.
[0016] The present invention provides a cross-instruction set operating system simulator, comprising a dynamic translation layer, which includes a dynamic translation engine and a cross-instruction set application binary interface (ABI) calling layer. The dynamic translation engine is used to translate client instructions into host instructions and execute them when running a first closed-source client executable file. The cross-instruction set ABI calling layer is used to, when running a second closed-source client executable file that requires calling a dynamic link library, if a host dynamic link library exists in the operating system, call the host function interface corresponding to the client function interface in the host dynamic link library according to the client instructions in the second closed-source client executable file. The dynamic translation engine is also used to, when running the second closed-source client executable file that requires calling a dynamic link library, if only a client dynamic link library exists in the operating system, translate and execute the client dynamic link library. The entire emulator, through a dynamic translation layer, translates and executes closed-source client executables. When dynamic link libraries are needed, if they are available, the corresponding host function interfaces are directly called, avoiding redundant translation. Only when no host dynamic link libraries are available are the instructions from the client dynamic link libraries translated into host instructions and executed. This minimizes the scope of dynamic translation and reduces translation overhead while ensuring versatility. It also utilizes host dynamic link libraries to accelerate execution, enabling complex applications to run efficiently. This effectively overcomes the performance and compatibility limitations of existing emulation methods and provides a feasible technical path for efficiently running emerging mobile operating systems such as HarmonyOS and their application ecosystems on x86_64 hardware platforms. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is one of the structural diagrams of the cross-instruction set operating system simulator provided by the present invention.
[0019] Figure 2 This is a schematic diagram of the structure of the dynamic translation engine provided by the present invention.
[0020] Figure 3 This is a schematic diagram illustrating a scenario where register allocation optimization is performed on the first LLVM IR, as provided by the present invention.
[0021] Figure 4 This is the second schematic diagram of the cross-instruction set operating system simulator provided by the present invention.
[0022] Figure 5 This is a schematic diagram of two function parameter passing and conversion methods provided by the present invention.
[0023] Figure 6 This is the third schematic diagram of the structure of the cross-instruction set operating system simulator provided by the present invention.
[0024] Figure 7 This is the fourth schematic diagram of the cross-instruction set operating system simulator provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0026] To better understand this invention, existing cross-architecture operation technologies will first be described in detail: Existing cross-architecture execution technologies can generally be divided into two categories: those based on full-system simulation and those based on user-space compatibility layers. Among these, the Quick Emulator (QEMU) full-system binary translator is a typical full-system simulation method. This method performs a comprehensive virtual model of the target architecture's processor, memory management unit, peripheral bus, interrupt controller, and boot chain, and then runs the target operating system as a whole through dynamic binary translation. In this method, the target operating system can run in a relatively complete virtual hardware environment, thus possessing good system behavior consistency and debug controllability. However, this method typically requires extensive dynamic translation of the target architecture's machine instructions. The system kernel, system services, graphics services, and applications are all in a unified simulated execution path, resulting in a significant consumption of computational resources on basic instruction translation and state maintenance. This introduces an additional overhead of approximately tens of times, making it unsuitable for daily development and debugging.
[0027] While compatibility execution methods targeting only user-space programs can avoid complete simulation of the entire kernel and underlying hardware platform, thus reducing translation costs to some extent, these solutions typically require the host operating system to directly support the target program's execution model, thus lacking universality. For example, Company A's Rosetta 2 is deeply coupled with macOS and cannot be ported to other operating systems; Android Houdini is a closed-source project from Company B, targeting only the Android operating system ecosystem and not applicable to other operating systems like HarmonyOS; while the QEMU user-space binary translator is an open-source project with the potential to be ported to operating systems like HarmonyOS, it cannot utilize the x86_64 native dynamic link library to accelerate applications, requiring a complete translation of the ARM64 binary library used by the application, resulting in poor performance and difficulty in adapting to complex translation needs such as graphical interfaces.
[0028] In other words, the full-system simulation method has disadvantages such as high translation overhead and low execution efficiency, while the user-space compatibility layer method has disadvantages such as poor versatility, inability to utilize the host's native dynamic link library for acceleration, and difficulty in adapting to complex applications.
[0029] Therefore, there is an urgent need to propose a new cross-instruction set operating system simulator to effectively overcome the limitations of existing simulation methods in terms of performance and compatibility.
[0030] To address the aforementioned technical problems, embodiments of the present invention provide a cross-instruction set operating system simulator comprising a dynamic translation layer, which includes a dynamic translation engine and a cross-instruction set application binary interface (ABI) calling layer. The dynamic translation engine, when running a first closed-source client executable file, translates client instructions into host instructions and executes them. The cross-instruction set ABI calling layer, when running a second closed-source client executable file that requires calling dynamic link libraries, if a host dynamic link library exists in the operating system, calls the host function interface corresponding to the client function interface in the host dynamic link library based on the client instructions in the second closed-source client executable file. The dynamic translation engine is further configured to translate and execute the client dynamic link library when running a second closed-source client executable file that requires calling dynamic link libraries, if only a client dynamic link library exists in the operating system. The entire emulator, through a dynamic translation layer, translates and executes closed-source client executables. When dynamic link libraries are needed, if they are available, the corresponding host function interfaces are directly called, avoiding redundant translation. Only when no host dynamic link libraries are available are the instructions from the client dynamic link libraries translated into host instructions and executed. This minimizes the scope of dynamic translation and reduces translation overhead while ensuring versatility. It also utilizes host dynamic link libraries to accelerate execution, enabling complex applications to run efficiently. This effectively overcomes the performance and compatibility limitations of existing emulation methods and provides a feasible technical path for efficiently running emerging mobile operating systems such as HarmonyOS and their application ecosystems on x86_64 hardware platforms.
[0031] The following is combined with Figures 1 to 7 The cross-instruction set operating system simulator provided in this embodiment of the invention will be described in detail below: Figure 1 This is one of the structural diagrams of the cross-instruction set operating system simulator provided by the present invention, such as... Figure 1 As shown, the operating system simulator includes a dynamic translation layer 101, which includes a dynamic translation engine 1011 and an application binary interface (ABI) call layer 1012. Among them, the dynamic translation engine 1011 is used to translate client instructions into host instructions and execute them when running the first closed-source client executable file; The cross-instruction set ABI calling layer 1012 is used when running a second closed-source client executable file, and the second closed-source client executable file needs to call a dynamic link library. If a host dynamic link library exists in the operating system, the host function interface corresponding to the client function interface in the host dynamic link library is called according to the client instruction in the second closed-source client executable file. The client instruction corresponds to the client function interface. The dynamic translation engine 1011 is also used to translate and execute the client dynamic link library when running a second closed-source client executable file that needs to call dynamic link libraries, provided that only the client dynamic link library exists in the operating system.
[0032] The dynamic translation layer 101, also known as the dynamic binary translation layer, is a dynamic binary translator based on the low-level virtual machine (LLVM). The dynamic translation engine 1011 is a dynamic binary translation engine.
[0033] A closed-source client executable file refers to a dynamic link library file that uses a client instruction set encoding (such as the ARM64 instruction set) and does not provide source code, such as an ARM64 closed-source dynamic link library (.so) file. Optionally, a closed-source client executable file includes closed-source client executable programs and closed-source client dynamic link libraries.
[0034] It should be noted that the first closed-source client executable does not need to call dynamic link libraries at runtime, while the second closed-source client executable does need to call dynamic link libraries at runtime. Optionally, the dynamic link libraries include host dynamic link libraries and client dynamic link libraries.
[0035] Host dynamic link libraries (such as x86_64 dynamic link libraries) refer to dynamic link library files encoded using the host instruction set, which can be directly executed natively by the central processing unit on the target platform. Optionally, these host dynamic link libraries include system libraries corresponding to the operating system and natively compiled versions of third-party software vendor libraries using the host instruction set. x86_64 indicates a 64-bit extension of the x86 architecture.
[0036] Client dynamic link libraries (such as ARM64 dynamic link libraries) are dynamic link library files encoded using the client instruction set. They cannot be executed natively on the host machine (such as the x86_64 platform) and need to be translated and executed by the dynamic translation engine 1011.
[0037] It should be noted that the system libraries corresponding to the aforementioned operating systems are dynamic link libraries pre-installed with or provided with the operating system. Optionally, these system libraries include, but are not limited to, system services, user interface (UI) frameworks, and graphics libraries.
[0038] For example, the aforementioned host dynamic link library consists of host native compiled versions of system services, UI frameworks, and graphics libraries, namely x86_64 native compiled versions.
[0039] In this embodiment of the invention, when the main program process needs to run a first closed-source client executable file, and the first closed-source client executable file does not need to call dynamic link libraries, it indicates that the first closed-source client executable file is in an independent running state. At this time, it can be started within the address space of the main program process, and then the client instructions in the first closed-source client executable file are translated into host machine instructions and executed, thereby enabling direct cross-instruction set execution of the closed-source client program. Here, the client instructions are the translated instructions, and the host machine instructions are the translated instructions.
[0040] When the cross-instruction set ABI call layer 1012 runs a second closed-source client executable file in the dynamic translation layer 101, and this second closed-source client executable file needs to call a dynamic link library, it indicates that the second closed-source client executable file has external dependencies. At this time, it is necessary to determine whether the host dynamic link library exists in the operating system. If it exists, it means that there is no need to translate the code of the host dynamic link library, and the host dynamic link library can be reused directly. In this case, the host function interface corresponding to the client function interface in the host dynamic link library can be called according to the client instructions in the second closed-source client executable file. For example, when translating the code segment of the ARM64 closed-source dynamic link library to the x86_64 dynamic link library, the x86_64 dynamic link library can be directly translated and executed, thereby avoiding redundant translation processes and realizing efficient cross-instruction set inter-call.
[0041] In the dynamic translation layer 101, when the second closed-source guest executable file is running, and this second closed-source guest executable file needs to call dynamic link libraries, the dynamic translation engine 1011 needs to determine whether the host dynamic link library exists in the operating system. If it does not exist, it means that only the guest dynamic link library exists in the operating system. In this case, the guest dynamic link library can be directly translated and executed, thereby ensuring that the main program process can correctly execute all functions. The entire process together achieves the versatility of the main program process, enabling it to seamlessly load and execute dynamic link library files with different instruction set architectures without modifying the source code.
[0042] It should be noted that the aforementioned operating system emulator is a simulator that efficiently runs ARM64 architecture mobile operating systems and their application ecosystems on the x86_64 hardware platform. It can run complete mobile operating systems such as HarmonyOS on the x86_64 hardware platform with near-native performance, and loads and translates closed-source ARM64 third-party dynamic link libraries within the mobile operating system. In other words, when translating and executing closed-source guest executable files through the dynamic translation layer 101, if dynamic link libraries are needed, the emulator directly calls the corresponding host function interfaces when host dynamic link libraries are available, avoiding redundant translation. Only when no host dynamic link libraries are available are the guest dynamic link library instructions translated into host instructions and executed. This minimizes the scope of dynamic translation and reduces translation overhead while ensuring versatility, and utilizes host dynamic link libraries to accelerate execution, enabling complex applications to run efficiently. This effectively overcomes the performance and compatibility limitations of existing simulation methods, providing a feasible technical path for efficiently running emerging mobile operating systems such as HarmonyOS and their application ecosystems on the x86_64 hardware platform.
[0043] In some embodiments, Figure 2 This is a schematic diagram of the structure of the dynamic translation engine provided by the present invention, as shown below. Figure 2 As shown, the dynamic translation engine 1011 includes an engine front-end 10111, an optimizer 10112, and a translation back-end 10113; the dynamic translation engine 1011 is used to translate client instructions into host machine instructions, including: Engine front-end 10111 is used to translate the client instructions in the basic blocks of the second closed-source client executable into the intermediate representation (LLVMIR) of the first underlying virtual machine. Optimizer 10112 is used to optimize register allocation for the first LLVM IR to obtain the second LLVM IR; and to perform an optimization pass on the second LLVM IR to obtain the third LLVM IR; Translation backend 10113 is used to compile the third LLVM IR into host machine instructions.
[0044] In this embodiment of the invention, the engine front-end 10111 uses an ARM64 translator to translate the ARM64 binary file corresponding to the client instructions in the second closed-source client executable file into a first LLVM IR, based on each basic block, to generate an initial intermediate representation. Next, the optimizer 10112 performs two levels of optimization on the first LLVM IR: first, register allocation optimization is performed to obtain a second LLVM IR, reducing redundant memory access operations and lowering the complexity of backend register allocation; then, the second LLVM IR is optimized again to obtain a third LLVM IR, improving the quality of the translated code. Finally, the translation back-end 10113 uses a Just-In-Time (JIT) compiler in LLVM to compile the third LLVM IR into host instructions, improving the overall execution efficiency of the operating system simulator. Furthermore, the code segment corresponding to the host instructions is cached for direct execution when accessing the same basic block later, avoiding redundant translation overhead.
[0045] Among them, client instructions can also be called ARM64 machine code, and host instructions can also be called x86_64 machine code.
[0046] In some embodiments, optimizer 10112 is used to perform register allocation optimization on the first LLVM IR to obtain a second LLVM IR, including: optimizer 10112 is specifically used to obtain the register access sequence in the first LLVM IR; uniformly represent the different access widths of the same register in the register access sequence; convert the unified access width into a consistent register value propagation form, and output a fourth LLVM IR with unified register access semantics; For the fourth LLVM IR within each basic block, reduce redundant load / store operations on the same register in the fourth LLVM IR to obtain the fifth LLVM IR; For the fifth LLVM IR that spans multiple basic blocks, redundant load / store operations in the fifth LLVM IR are removed based on the control flow relationships between the basic blocks to obtain the second LLVM IR.
[0047] In this embodiment of the invention, to improve the translation performance of the LLVM-based dynamic translation engine 1011, the optimizer 10112 can optimize register allocation for the first LLVM IR. This addresses the problem that traditional methods easily introduce a large number of cross-block register propagation issues in large translation regions, leading to high register pressure and high compilation overhead. While maintaining LLVM compatibility and the Static Single Assignment (SSA) format, it performs phased optimization of memory access related to ARM64 client registers to reduce redundant memory accesses and lower the complexity of backend register allocation. Specifically, as follows... Figure 3 As shown, from Figure 3 As can be seen, optimizer 10112 takes the first LLVM IR as input and first unifies the read and write (e.g., load / store) instructions related to the ARM64 client registers in the first LLVM IR. Specifically, it takes the register access sequence in the first LLVM IR as input, unifies the different access widths of the same register in the register access sequence, and obtains the unified access width, providing a unified width benchmark for subsequent register value propagation. Then, it converts the unified access width into a consistent register value propagation form to eliminate the semantic fragmentation between different bit widths of memory access and outputs a fourth LLVM IR with unified register access semantics, thus providing a consistent input for subsequent register value propagation. Subsequently, optimizer 10112 reduces redundant load / store operations on the same register in the fourth LLVM IR within each basic block, resulting in a fifth LLVM IR. The IR (Register Registry) is used to perform local register promotion optimization within the ARM64 basic block scope. This reduces redundant load / store operations and enables intra-block reuse of ARM64 client register values while avoiding large-scale cross-block register propagation. Finally, the optimizer 10112 performs controlled propagation optimization on ARM64 client register values across LLVM basic blocks. Specifically, for the fifth LLVM IR across basic blocks, it selectively propagates cross-block register values based on the control flow relationship between basic blocks and eliminates redundant write-back operations that can be overridden in the control flow path. That is, it removes redundant load / store operations in the fifth LLVM IR to obtain the second LLVM IR, thereby reducing unnecessary memory access overhead while controlling register pressure and compilation complexity.
[0048] In some embodiments, optimizer 10112 is used to perform optimization passes on the second LLVM IR to obtain a third LLVM IR, including: optimizer 10112 is specifically used to perform loop optimization, dead code elimination and vectorization on the second LLVM IR to obtain a third LLVM IR.
[0049] Optionally, loop optimization includes classic optimizations such as loop-invariant code hoisting, loop unrolling, and loop merging.
[0050] In this embodiment of the invention, during the optimization process of the second LLVM IR, the optimizer 10112 can perform loop optimization on the second LLVM IR to obtain a first sub-LLVM IR, thereby optimizing the loop structure, reducing redundant calculations in the loop iteration, and improving loop execution efficiency; then, dead code elimination is performed on the first sub-LLVM IR to obtain a second sub-LLVM IR, thereby deleting invalid code that cannot be executed or has no impact on the program result, which can reduce the size of the translated code and improve instruction cache efficiency; then, the second sub-LLVM IR is vectorized to obtain a third LLVM IR, and data-level parallelism is achieved by using Single Instruction Multiple Data (SIMD) instructions to improve the execution efficiency of the translated code and achieve data-level parallel acceleration.
[0051] In other words, optimizer 10112 can utilize the functional modules in LLVM that are specifically responsible for looping, deleting useless code, and vectorization (processing in each pass) to process the second LLVM IR sequentially, thereby improving the quality of the final generated third LLVM IR (running faster and smaller).
[0052] It should be noted that the aforementioned cross-instruction set ABI call layer 1012 can support cross-instruction set ABI inter-calls. This cross-instruction set ABI inter-call only needs to consider the inter-call between the closed-source client (ARM64) binary file and the host (x86_64) dynamic link library.
[0053] In some embodiments, the second closed-source client executable file is a client binary file; the cross-instruction set ABI calling layer 1012 is used to call the host machine function interface corresponding to the client function interface in the host machine dynamic link library according to the client instructions in the second closed-source client executable file, including: the cross-instruction set ABI calling layer 1012 is specifically used to perform parameter conversion on the client function interface according to the function signature when it is recognized that the target address to be jumped to in the client binary file points to the host machine dynamic link library, so as to obtain the host machine function interface.
[0054] The function signature refers to the type identification information of a function under the constraints of the calling convention, including the function's return type, parameter types, number and order of parameters, and the location of each parameter in the register or stack. This function signature is used to determine how parameter mapping and conversion are performed between client instructions and host instructions.
[0055] In this embodiment of the invention, to avoid redundant translation of x86_64 dynamic link libraries, an instruction set ABI inter-call method can be adopted. Specifically, when the cross-instruction set ABI calling layer 1012 recognizes that the target address to be jumped to in the client binary file points to the host function interface in the host dynamic link library, it can perform ABI-level conversion (including parameter type, quantity, order, and register / stack mapping) on the parameters and encoding of the client instruction according to the function signature to obtain the host instruction, so as to achieve seamless mapping from the client instruction set calling convention to the host instruction set calling convention.
[0056] In some embodiments, combined with Figure 1 , Figure 4 This is the second schematic diagram of the structure of the cross-instruction set operating system simulator provided by the present invention, as shown below. Figure 4 As shown, the operating system simulator also includes: native execution layer 102; Among them, the cross-instruction set ABI call layer 1012 is also used to execute the native function corresponding to the host machine instruction of the host machine function interface in the native execution layer 102, and when the function return value is obtained, convert the function return value from the register corresponding to the host machine instruction back to the register corresponding to the guest machine instruction.
[0057] Among them, the native execution layer 102 can also be called the native x86_64 execution layer.
[0058] In this embodiment of the invention, since the cross-instruction set ABI call layer 1012 adopts the instruction set ABI inter-call method, the native execution layer 102 can execute the native function (such as the x86_64 native function) after obtaining the aforementioned host machine instruction, and obtain the function return value. This function return value is used to characterize the execution result of the native function call. After execution, the cross-instruction set ABI call layer 1012 can convert the function return value from the register corresponding to the host machine instruction (such as the x86_64 Accumulator Register (RAX)) back to the register corresponding to the guest machine instruction (such as the ARM64 X0 register), thus completing the inter-call process. The entire process does not require translation of the host machine dynamic link library and is executed directly in a native manner, thereby significantly reducing the overhead of cross-instruction set calls.
[0059] It should be noted that, to ensure the performance and compatibility of inter-calls within the cross-instruction set ABI call layer 1012, the cross-instruction set ABI call layer 1012 can employ two methods for function parameter passing and conversion, such as... Figure 5As shown, firstly, for x86_64 dynamic link libraries with source code, a custom calling convention is added by customizing the LLVM compiler backend, so that x86_64 functions adopt the ARM64-compatible parameter passing layout at compile time, making the x86_64 functions directly compatible with ARM64 calls and avoiding additional overhead; secondly, for x86_64 dynamic link libraries without source code, parameter register remapping is performed through automatically generated stub functions, and function parameter conversion is performed at runtime.
[0060] In some embodiments, combined with Figure 4 The native execution layer 102 is used to run the host dynamic link library when the host dynamic link library second closed-source client executable file needs to call the dynamic link library.
[0061] In this embodiment of the invention, the native execution layer 102 runs the main program that supports cross-platform operation, as well as the host machine's natively compiled version of the system libraries corresponding to the operating system, namely the x86_64 natively compiled version. The code run by the native execution layer 102 accounts for more than 90% of the total code, and all of it runs the x86_64 native compilation without translation, thereby avoiding translation overhead.
[0062] In some embodiments, combined with Figure 4 , Figure 6 This is the third schematic diagram of the structure of the cross-instruction set operating system simulator provided by this invention, as shown below. Figure 6 As shown, the operating system simulator also includes a hardware virtualization layer 103; wherein, the hardware virtualization layer 103 is used to provide a standard virtualization environment by utilizing the kernel-based virtual machine (KVM) hardware-accelerated virtualization technology based on the fast simulator QEMU, and the standard virtualization environment is used to run the target operating system based on the host machine instruction set.
[0063] Among them, the QEMU-based KVM hardware-accelerated virtualization technology refers to the direct use of the hardware virtualization extension of the physical central processing unit (CPU) through the KVM kernel module, enabling LLVM to directly execute most host machine instructions (such as x86_64 instructions). At the same time, it uses QEMU as a user-space emulator to provide device simulation and virtual machine management capabilities, thereby avoiding full system simulation and achieving native instruction execution with performance close to that of a physical machine.
[0064] A standard virtualization environment refers to a virtual machine environment created by the aforementioned hardware-accelerated virtualization technology that can run native instructions from the host machine. This virtual machine environment does not rely on software emulation of specific guest instructions (such as ARM64 instructions), but enables the guest operating system to execute host machine instructions directly in a native manner through virtualized CPU, memory, and input / output (I / O) devices.
[0065] In this embodiment of the invention, the hardware virtualization layer 103 utilizes QEMU-based KVM hardware accelerated virtualization technology to provide a standard virtualization environment, enabling the operating system kernel (such as the Linux / OHOS kernel) to run in this standard virtualization environment based on the host machine instruction set (such as the x86_64 native instruction set), thereby avoiding the high overhead introduced by full system emulation.
[0066] To better understand the embodiments of the present invention, the cross-instruction set operating system simulator provided by the present invention will be further elaborated below: For example, Figure 7 This is the fourth schematic diagram of the structure of the cross-instruction set operating system simulator provided by the present invention, as shown below. Figure 7 As shown, this operating system emulator is a cross-instruction set operating system emulator based on process-level hybrid execution, mainly consisting of three parts: a hardware virtualization layer, a native x86_64 execution layer, and a dynamic binary translation layer.
[0067] The hardware virtualization layer is used to provide a standard virtualization environment by leveraging QEMU-based KVM hardware-accelerated virtualization technology.
[0068] The native x86_64 execution layer is used to run the x86_64 native compiled version of the main program that supports cross-platform operation and the corresponding system libraries of the operating system.
[0069] The dynamic translation engine in the dynamic binary translation layer is used to translate ARM64 instructions into x86_64 instructions for execution when the dynamic binary translation layer needs to run the ARM64 closed-source dynamic link library; In the case where the cross-instruction set ABI calling layer in the dynamic binary translation layer needs to run the ARM64 closed-source dynamic link library and the ARM64 closed-source dynamic link library needs to call dynamic link libraries, if the operating system has an x86_64 dynamic link library, then according to the ARM64 instructions in the ARM64 closed-source dynamic link library, the x86_64 instructions in the x86_64 dynamic link library corresponding to the ARM64 instructions are called. The aforementioned dynamic translation engine is also used in situations where the dynamic binary translation layer needs to run an ARM64 closed-source dynamic link library, and this ARM64 closed-source dynamic link library needs to call a dynamic link library. If the operating system does not have an x86_64 dynamic link library, but only an ARM64 dynamic link library, then the translation will execute the ARM64 dynamic link library.
[0070] In other words, the core idea of the aforementioned operating system emulator is that the mobile operating system itself executes in a standard virtualization environment using native x86_64 instructions, and only dynamically translates the ARM64 closed-source dynamic link libraries that need to be compatible within the same process address space.
[0071] In combination with the above Figures 1 to 7 The aforementioned operating system emulator has the following beneficial effects: 1. The operating system simulator provided by this invention executes in the x86_64 native compiled version at startup. Dynamic translation is only activated as needed when the application actually loads the ARM64 closed-source dynamic link library, thereby avoiding the overhead introduced by fully translating the operating system.
[0072] 2. The dynamic translation engine is based on LLVM and introduces a lightweight register allocation optimization method. By eliminating the problems of high compilation overhead and high register pressure in the register allocation process of LLVM, the translation speed of the dynamic translation engine is improved.
[0073] 3. The dynamic translation engine introduces a cross-instruction set ABI inter-call method, which realizes cross-instruction set conversion and passing of function parameter calls through a custom compiler and automatic generation of Thunk functions.
[0074] The following section uses CoreMark and 7z benchmark tests to simulate and evaluate the cross-instruction set operating system simulator provided by this invention.
[0075] To verify the binary translation performance of the aforementioned operating system emulator, a simulation evaluation experiment was conducted based on the CoreMark benchmark test. The computationally intensive CoreMark benchmark does not rely on third-party dynamic link libraries; all code must be translated and executed, and acceleration using native x86_64 dynamic link libraries is not possible. This test reflects the binary translation performance of the aforementioned operating system emulator under worst-case conditions. The experimental results are shown in Table 1.
[0076] Table 1: Test metrics QEMU User-Space Solution for Fast Emulator Technical solution of the present invention Performance Comparison (Comparison with QEMU) CoreMark 200K execution time 47.9 seconds 23.6 seconds 203% CoreMark steady-state speed 0.237 milliseconds / iteration 0.114 milliseconds / iteration 208% As can be seen from Table 1, the performance of the operating system simulator provided by this invention is more than twice that of the binary translator in the existing QEMU user-space scheme.
[0077] To verify the performance of the aforementioned operating system simulator in real-world application scenarios, a simulation evaluation experiment was conducted based on the 7z benchmark. The experimental results are shown in Table 2.
[0078] Table 2:
[0079] The unit of the test metric is Million Instructions Per Second (MIPS).
[0080] As can be seen from Table 2, since the operating system simulator provided by this invention can directly utilize the native dynamic link library of x86_64 for acceleration, the simulation execution performance of this operating system simulator can reach 96.5% of that of native execution on x86_64 host, with a performance loss of only 3.5%, which is more than 4 times the performance of binary translator in the existing QEMU user-space scheme.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-instruction set operating system emulator, characterized in that, include: The dynamic translation layer includes a dynamic translation engine and a cross-instruction set application binary interface (ABI) call layer; wherein, The dynamic translation engine is used to translate client instructions into host instructions and execute them when running a first closed-source client executable file; The cross-instruction set ABI calling layer is used to call the host machine function interface corresponding to the guest function interface in the host machine dynamic link library according to the guest instruction in the second closed-source guest executable file when the second closed-source guest executable file needs to call the dynamic link library. The guest instruction corresponds to the guest function interface. The dynamic translation engine is also used to translate and execute the client dynamic link library when running the second closed-source client executable file, and the second closed-source client executable file needs to call dynamic link libraries, if only the client dynamic link library exists in the operating system.
2. The cross-instruction set operating system emulator according to claim 1, characterized in that, The dynamic translation engine includes an engine front-end, an optimizer, and a translation back-end; The dynamic translation engine is used to translate client instructions into host machine instructions, including: The engine front-end is used to translate the client instructions in each basic block of the second closed-source client executable into an intermediate representation of the first underlying virtual machine, LLVM IR. The optimizer is used to perform register allocation optimization on the first LLVM IR to obtain a second LLVM IR; and to perform optimization on the second LLVM IR to obtain a third LLVM IR; The translation backend is used to compile the third LLVM IR into the host machine instructions.
3. The cross-instruction set operating system emulator according to claim 2, characterized in that, The optimizer, used to perform register allocation optimization on the first LLVM IR to obtain a second LLVM IR, includes: The optimizer is specifically used to obtain the register access sequence in the first LLVM IR; to uniformly represent the different access widths of the same register in the register access sequence; to convert the uniform access width into a consistent register value propagation form; and to output a fourth LLVM IR with uniform register access semantics. For the fourth LLVM IR within each basic block, reduce redundant load / store operations on the same register in the fourth LLVM IR to obtain the fifth LLVM IR; For the fifth LLVM IR that spans multiple basic blocks, redundant load / store operations in the fifth LLVM IR are removed based on the control flow relationships between the basic blocks to obtain the second LLVM IR.
4. The cross-instruction set operating system emulator according to any one of claims 1-3, characterized in that, The second closed-source client executable is a client binary file; The cross-instruction set ABI calling layer is used to call the host machine function interface corresponding to the guest machine function interface in the host machine dynamic link library according to the guest machine instruction in the second closed-source guest machine executable file, including: The cross-instruction set ABI calling layer is specifically used to perform parameter conversion on the client function interface based on the function signature when it is recognized that the target address to be jumped to in the client binary file points to the host dynamic link library, so as to obtain the host function interface.
5. The cross-instruction set operating system emulator according to claim 4, characterized in that, It also includes the native execution layer; among which, The cross-instruction set ABI calling layer is also used to convert the function return value from the register corresponding to the host instruction back to the register corresponding to the guest instruction when the host instruction executes the native function corresponding to the host instruction in the native execution layer and obtains the function return value.
6. The cross-instruction set operating system emulator according to claim 2 or 3, characterized in that, The optimizer, used to perform an optimization pass on the second LLVM IR to obtain a third LLVM IR, includes: The optimizer is specifically used to perform loop optimization, dead code elimination, and vectorization on the second LLVM IR to obtain the third LLVM IR.
7. The cross-instruction set operating system emulator according to any one of claims 1-3, characterized in that, Also includes: The native execution layer; among which, The native execution layer is used to run the host machine dynamic link library when the second closed-source client executable file is running and the second closed-source client executable file needs to call dynamic link libraries.
8. The cross-instruction set operating system emulator according to any one of claims 1-3, characterized in that, Also includes: Hardware virtualization layer; among which, The hardware virtualization layer is used to provide a standard virtualization environment by utilizing KVM hardware-accelerated virtualization technology based on the fast emulator QEMU. The standard virtualization environment is used to run a target operating system based on the host machine instruction set.
9. The cross-instruction set operating system emulator according to any one of claims 1-3, characterized in that, The host dynamic link library includes the system library corresponding to the operating system and the native compiled version of the host instruction set of the third-party software vendor library.
10. The cross-instruction set operating system emulator according to claim 9, characterized in that, The system library includes system services, user interface (UI) framework, and graphics library.