Binary translation system and method applied to X86 program

By introducing a backward data flow analysis module in the binary translation system, analyzing and eliminating redundant register operations in the X86 program, the problem that translation performance in the prior art is affected by redundant operations is solved, and a more efficient translation process is achieved.

CN120122994APending Publication Date: 2025-06-10INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510179804.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When the existing binary translation system translates source programs that follow X86 semantics, it retains the high-bit clearing or high-bit reservation operations of redundant general registers, which in turn affects translation performance.

Method used

Design a binary translation system for X86 programs, including data acquisition module, disassembly module, backward data flow analysis module and translation module. The backward data flow analysis module analyzes the register status of each instruction in the order from back to front, and eliminates the redundant instructions generated in translation through interpolation, difference and concurrent operations.

Benefits of technology

By eliminating redundant high-bit clearing or retention operations, the performance of the binary translation system is improved, ensuring the accuracy and efficiency of translation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122994A_ABST
    Figure CN120122994A_ABST
Patent Text Reader

Abstract

The invention provides a binary translation system applied to an X86 program, which is used for translating a source program following X86 semantics into a target program following other semantics, and comprises a data acquisition module used for acquiring the source program; the disassembling module is used for dividing a source program into a plurality of basic blocks and analyzing subsequent basic blocks corresponding to each basic block; the backward data flow analysis module is used for sequentially analyzing each instruction in each basic block from back to front so as to obtain a target definition set and a target subsequent use set corresponding to each instruction in the source program; and the translation module is used for eliminating redundant instructions of high-order zero clearing or high-order retention of the general register generated in translation. According to the technical scheme, the register state of the general register corresponding to each instruction of the source program is analyzed through the backward data flow analysis module so as to eliminate redundant instructions generated in the translation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of code optimization, specifically to binary translation technology in the field of code optimization, and more specifically, to a binary translation system and a translation method applied to X86 programs. Background Art

[0002] With the increasing popularity of virtual machines and the diversification of instruction set architectures (ISAs), the importance of binary translation technology has become increasingly prominent. Binary translation technology enables applications built for the source ISA to run directly on computers with the target ISA, and has a wide range of applications in many fields such as legacy code migration, binary instrumentation, and hot path analysis and optimization.

[0003] X86 series machines are the current mainstream models, with a rich collection of application software and legacy code. Therefore, a binary translation system with X86 as the source ISA is crucial. X86 uses a complex instruction set. To implement the function of one X86 instruction, the target ISA usually requires multiple instructions, which leads to code bloat and performance degradation. The segmented use of general registers in X86 is a key reason for code bloat.

[0004] In the X86 architecture, the operand widths of general registers include 64 bits, 32 bits, 16 bits, high 8 bits, and low 8 bits. When the destination register of an X86 instruction is 32 bits, the high 32 bits of the corresponding 64-bit register will be cleared; when the destination register is 16 bits or 8 bits, the remaining bits of the 64-bit register will be preserved. Therefore, the target ISA not only needs to simulate the arithmetic operations of X86 instructions but also needs to emit additional instructions to simulate the high-bit clearing or preservation operations. For example, when the target ISA simulates a 32-bit operation of X86, it needs to emit at least one instruction for the corresponding calculation and two shift instructions to clear the high 32 bits of the destination register; when simulating a 16-bit or 8-bit operation of X86, more shift instructions are needed to preserve the remaining bits of the destination register.

[0005] If the binary translation system strictly follows the semantics of X86, simulating the semantics of one X86 instruction requires executing multiple instructions on the target ISA, which will seriously affect the translation performance. However, not every additional operation of high-bit clearing or remaining-bit preservation in X86 will affect the correctness of program execution. For example, if an X86 program no longer uses the %rax register (the 64-bit register corresponding to the %eax register) after assigning a value to the %eax register, then the high 32-bit clearing operation on %rax is redundant.

[0006] In summary, when the existing binary translation system translates a source program that follows X86 semantics, it will retain the operations of clearing the high bits or retaining the high bits of redundant general-purpose registers, thereby affecting the translation performance.

[0007] It should be noted that: This background technology is only used to introduce relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily prior art. Without evidence indicating that the relevant information was publicly available before the filing date of the present invention, the relevant information should not be regarded as prior art. Summary of the Invention

[0008] Therefore, the object of the present invention is to overcome the above-mentioned defects of the prior art and provide a binary translation system for X86 programs and a binary translation method for X86 programs.

[0009] The object of the present invention is achieved by the following technical solutions.

[0010] According to a first aspect of the present invention, there is provided a binary translation system for X86 programs, which is used to translate a source program that follows X86 semantics into a target program that follows other semantics. The system includes: a data acquisition module for acquiring the source program, where the source program includes multiple instructions related to the use operation and def operation of general-purpose registers; a disassembly module for dividing the source program into multiple basic blocks and analyzing the successor basic blocks corresponding to each basic block; where a basic block represents a continuous instruction sequence in the source program, and the successor basic block of a basic block represents other basic blocks that can be directly reached through the control flow from this basic block; a backward data flow analysis module for sequentially analyzing each instruction in each basic block in a reverse order to obtain the target definition set and the target subsequent use set corresponding to each instruction in the source program; where the target definition set includes all general-purpose registers actually defined by the corresponding instruction and the final register state of each general-purpose register, and the target subsequent use set includes all general-purpose registers actually used by the corresponding instruction and subsequent instructions and the final register state of each general-purpose register, and the register state indicates that the operation bit width of the general-purpose register is 64 bits, 32 bits, 16 bits, high 8 bits, low 8 bits or 0 bits; a translation module for translating the source program into a target program with other semantics and eliminating redundant instructions of clearing the high bits or retaining the high bits of general-purpose registers generated during the translation based on the target definition set and the target subsequent use set corresponding to each instruction in the source program.

[0011] In some embodiments of the present invention, the backward data flow analysis module is configured with a register state lattice for performing intersection operations, difference operations, and union operations. The register state lattice is composed of a register state set and a register state partial order relationship, where: the register state set includes , , , , , , where indicates that the operation bit width is 64 bits, indicates that the operation bit width is 32 bits, indicates that the operation bit width is 16 bits, indicates that the operation bit width is the high 8 bits, indicates that the operation bit width is the low 8 bits, indicates that the operation bit width is 0 bits; the register status partial order relationship includes , , , , , , where indicates that 0 bits are part of the low 8 bits, indicates that 0 bits are part of the high 8 bits, indicates that the low 8 bits are part of 16 bits, indicates that the high 8 bits are part of 16 bits, indicates that 16 bits are part of 32 bits, indicates that 32 bits are part of 64 bits.

[0012] In some embodiments of the present invention, the backward data flow analysis module is configured to perform the intersection operation in the following manner: determine the operation bit widths corresponding to the register statuses of the two general registers for which the intersection operation is performed, and calculate the greatest lower bound between the operation bit widths to obtain the intersection operation result; wherein, the greatest lower bound between the 0-bit operation bit width and other bit operation bit widths is the 0-bit; the greatest lower bound between the low 8-bit operation bit width and the high 8-bit operation bit width is the 0-bit, and the greatest lower bound between the low 8-bit operation bit width and the 16-bit, 32-bit, and 64-bit operation bit widths is the low 8-bit; the greatest lower bound between the high 8-bit operation bit width and the 16-bit, 32-bit, and 64-bit operation bit widths is the high 8-bit; the greatest lower bound between the 16-bit operation bit width and the 32-bit and 64-bit operation bit widths is the 16-bit; the greatest lower bound between the 32-bit operation bit width and the 64-bit operation bit width is the 32-bit.

[0013] In some embodiments of the present invention, the backward data flow analysis module is configured to perform the difference operation in the following manner: determine the operation bit widths corresponding to the register statuses of the two general registers for which the difference operation is performed, and calculate the difference set between the operation bit widths to obtain the difference operation result; wherein, the difference set between the 16-bit operation bit width and the high 8-bit operation bit width is the low 8-bit, the difference set between the 16-bit operation bit width and the low 8-bit operation bit width is the high 8-bit, and the difference sets between other operation bit widths are all 0 bits.

[0014] In some embodiments of the present invention, the backward data flow analysis module is configured to perform an AND operation in the following manner: determine the operation bit widths corresponding to the register states of two general-purpose registers for performing the AND operation, and calculate the least upper bound between the operation bit widths to obtain the result of the AND operation; wherein, the least upper bound between a 64-bit operation bit width and other bit operation bit widths is 64 bits, and the least upper bound between a 32-bit operation bit width and 0-bit, low 8-bit, high 8-bit, and 16-bit operation bit widths is 32 bits; the least upper bound between a 16-bit operation bit width and 0-bit, low 8-bit, and high 8-bit operation bit widths is 16 bits; the least upper bound between a high 8-bit operation bit width and a low 8-bit operation bit width is 16 bits, and the least upper bound between a high 8-bit operation bit width and a 0-bit operation bit width is the high 8-bit; the least upper bound between a low 8-bit and a 0-bit operation bit width is the low 8-bit.

[0015] In some embodiments of the present invention, the backward data flow analysis module is configured to analyze each instruction in each basic block in a backward order as follows: obtain the initial use set, initial definition set, and initial subsequent use set corresponding to the current instruction; wherein, the initial use set includes all general-purpose registers directly used by the corresponding instruction semantics itself and the initial register states of each general-purpose register, the initial definition set includes all general-purpose registers directly defined by the corresponding instruction semantics itself and the initial register states of each general-purpose register, and the initial subsequent use set is the target subsequent use set of the next instruction corresponding to the current instruction; perform an intersection operation on the register states of each general-purpose register in the initial definition set corresponding to the current instruction and the register states of the corresponding general-purpose register in the initial subsequent use set corresponding to the current instruction to obtain the target definition set corresponding to the current instruction; perform a difference operation on the register states of each general-purpose register in the initial subsequent use set corresponding to the current instruction and the register states of the corresponding general-purpose register in the target definition set corresponding to the current instruction to obtain the intermediate subsequent use set corresponding to the current instruction; perform an OR operation on the register states of each general-purpose register in the intermediate subsequent use set corresponding to the current instruction and the register states of the corresponding general-purpose register in the initial use set corresponding to the current instruction to obtain the target subsequent use set corresponding to the current instruction; wherein, if there is no successor basic block for the basic block, when analyzing the last instruction in the basic block, the initial subsequent use set corresponding to the last instruction is a plurality of pre-set general-purpose registers and the fixed register states of each general-purpose register; if there is a successor basic block for the basic block, when analyzing the last instruction in the basic block, the initial subsequent use set corresponding to the last instruction is obtained by performing an OR operation on the target subsequent use sets of all successor basic blocks of the basic block, and the target subsequent use set of the successor basic block is the target subsequent use set corresponding to the starting instruction in the successor basic block.

[0016] In some embodiments of the present invention, the backward data flow analysis module is configured to obtain the initial use set and the initial definition set corresponding to the current instruction in the following manner: determine all general-purpose registers involved in the use operation and the def operation in the current instruction, and initialize the operation bit width corresponding to the register state of each general-purpose register to 0 bits; wherein, all general-purpose registers involved in the use operation constitute the source use set corresponding to the current instruction; all general-purpose registers involved in the def operation constitute the source definition set corresponding to the current instruction; for each general-purpose register involved in the use operation, perform a union operation on the register state corresponding to the general-purpose register in the source use set and the register state of the general-purpose register directly corresponding to the semantics of the current instruction to obtain the updated source use set of the current instruction; for each general-purpose register involved in the def operation, when the operation bit width corresponding to the register state of the general-purpose register in the semantics of the current instruction is 32 bits or 64 bits, update the operation bit width corresponding to the register state of the general-purpose register in the source definition set to 64 bits; for each general-purpose register involved in the def operation, when the operation bit width corresponding to the register state of the general-purpose register in the semantics of the current instruction is not 32 bits or 64 bits, update the register state of the general-purpose register in the updated source use set of the current instruction to an operation bit width of 64 bits; wherein, the source use set and the source definition set after all updates are completed are used as the initial use set and the initial definition set corresponding to the current instruction.

[0017] According to a second aspect of the present invention, there is provided a binary translation method for an X86 program, which is used to translate a source program conforming to X86 semantics into a target program conforming to other semantics. The method includes: Step S1, obtain the source program; Step S2, use the system as described in the first aspect of the present invention to translate the source program to obtain a target program conforming to other semantics.

[0018] Compared with the prior art, the advantages of the present invention are as follows: (1) A backward data flow analysis module is provided in the binary translation system to analyze the register states of general-purpose registers corresponding to each instruction in the source code in the reverse analysis order, so that the translation module can eliminate redundant instructions such as high-bit clearing or high-bit retention of general-purpose registers that may occur during the translation process based on the register states of general-purpose registers corresponding to each analyzed instruction; (2) When analyzing each instruction, the backward data flow analysis module makes the analysis result of the register state of the general-purpose register corresponding to each instruction more accurate through data flow analysis across basic blocks, and further enables the translation module to eliminate more redundant instructions during the translation process; (3) The register state lattice is configured in the backward data flow analysis module to perform intersection operation, difference operation and union operation, so as to analyze the register states of general-purpose registers corresponding to each instruction more accurately, and further enables the translation module to eliminate more redundant instructions during the translation process. Description of the Drawings

[0019] The embodiments of the present invention will be further described below with reference to the accompanying drawings, where:

[0020] Figure 1 It is a schematic diagram of the composition of a binary translation system according to an embodiment of the present invention;

[0021] Figure 2 It is a register Hasse diagram according to an embodiment of the present invention. Specific embodiments

[0022] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0023] As mentioned in the background art section, when the existing binary translation system translates a source program that follows X86 semantics, it will retain the operations of clearing the high bits or retaining the high bits of redundant general-purpose registers, thereby affecting the translation performance.

[0024] To solve the above problems, the inventor proposes a new binary translation system, which sets up a backward data flow analysis module to analyze the register status of the general-purpose register corresponding to each instruction in the source code in the backward analysis order, so that the translation module can eliminate the redundant instructions of clearing the high bits or retaining the high bits of the general-purpose register that may be generated during the translation process based on the register status of the general-purpose register corresponding to each analyzed instruction.

[0025] Generally speaking, as Figure 1As shown in the figure, the present invention provides a binary translation system for X86 programs, which is used to translate a source program that follows X86 semantics into a target program that follows other semantics. The system includes: a data acquisition module, which is used to acquire the source program. Among them, the source program includes multiple instructions related to the use operation and def operation of general registers; a disassembly module, which is used to divide the source program into multiple basic blocks and analyze the successor basic blocks corresponding to each basic block. Among them, a basic block represents a continuous instruction sequence in the source program, and the successor basic block of a basic block represents other basic blocks that can be directly reached through the control flow starting from this basic block; a backward data flow analysis module, which is used to analyze each instruction in each basic block in turn in the reverse order to obtain the target definition set and the target subsequent use set corresponding to each instruction in the source program. Among them, the target definition set includes all general registers actually defined by the corresponding instruction and the final register state of each general register, and the target subsequent use set includes all general registers actually used by the corresponding instruction and subsequent instructions and the final register state of each general register. The register state indicates that the operation bit width of the general register is 64 bits, 32 bits, 16 bits, high 8 bits, low 8 bits or 0 bits; a translation module, which is used to translate the source program into a target program with other semantics, and eliminate the redundant instructions of high-bit clearing or high-bit retention of general registers generated during translation based on the target definition set and the target subsequent use set corresponding to each instruction in the source program.

[0026] To better understand the present invention, the following will specifically describe each component module in the binary translation system with reference to specific embodiments.

[0027] I. Data Acquisition Module

[0028] The data acquisition module is used to acquire the source program. Among them, the source program includes multiple instructions related to the use operation and def operation of general registers.

[0029] To better understand X86 programs, the following will take the X86 assembly instruction sequence shown in Table 1 as an example for illustration.

[0030] As can be seen from Table 1, the movl instruction indicates that the value of %ebx is assigned to %eax. Therefore, in this instruction, the general-purpose register %ebx performs a use operation, and the general-purpose register %ebx is the source register (for simplicity of reference, the general-purpose register involved in the use operation is referred to as the source register in the following text); the general-purpose register %eax performs a def operation, and the general-purpose register %eax is the destination register (for simplicity of reference, the general-purpose register involved in the def operation is referred to as the destination register in the following text). The cdqe instruction indicates that the general-purpose register %eax is sign-extended to %rax. Therefore, in this instruction, the general-purpose register %eax performs a use operation, and the general-purpose register %eax is the source register; the general-purpose register %rax performs a def operation, and %rax is the destination register.

[0031] Table 1

[0032]

[0033] II. Disassembly Module

[0034] The disassembly module is used to divide the source program into multiple basic blocks and analyze the successor basic blocks corresponding to each basic block; among them, a basic block represents a continuous instruction sequence in the source program, and the successor basic block of a basic block represents other basic blocks that can be directly reached through the control flow starting from this basic block.

[0035] It should be noted that the reason for determining the successor basic blocks corresponding to each basic block is that the number of instructions in a single basic block is usually small, and the accuracy of data flow analysis for the instructions in a single basic block is limited, and the high-order zero-clearing or redundant instructions that can be eliminated are few. Therefore, when the backward data flow analysis module performs data flow analysis, it will call the disassembly module to try to explore more basic blocks until the successor cannot be determined statically and then stop discovering basic blocks. By expanding the number of basic blocks in the data flow analysis, the accuracy of data flow analysis for a single basic block can be improved, so that more redundant instructions will be eliminated.

[0036] III. Backward Data Flow Analysis Module

[0037] The backward data flow analysis module is used to analyze each instruction in each basic block in sequence from back to front to obtain the target definition set and the target subsequent use set corresponding to each instruction in the source program; wherein, the target definition set includes all general registers actually defined by the corresponding instruction and the final register state of each general register, and the target subsequent use set includes all general registers actually used by the corresponding instruction and subsequent instructions and the final register state of each general register, and the register state indicates that the operation bit width of the general register is 64 bits, 32 bits, 16 bits, high 8 bits, low 8 bits or 0 bits. It should be noted that when the operation bit width of the general register is 0 bits, it means that the general register is not used.

[0038] According to an embodiment of the present invention, the backward data flow analysis module is configured with a register status lattice for performing intersection operation, difference operation and union operation. The register status lattice is composed of a register status set and a register status partial order relationship, wherein: the register status set includes , , , , , , wherein, indicates that the operation bit width is 64 bits, indicates that the operation bit width is 32 bits, indicates that the operation bit width is 16 bits, indicates that the operation bit width is high 8 bits, indicates that the operation bit width is low 8 bits, indicates that the operation bit width is 0 bits; the register status partial order relationship includes , , , , , , wherein, indicates that 0 bits is a part of low 8 bits, indicates that 0 bits is a part of high 8 bits, indicates that low 8 bits is a part of 16 bits, indicates that high 8 bits is a part of 16 bits, indicates that 16 bits is a part of 32 bits, indicates that 32 bits is a part of 64 bits. Among them, in order to better understand the register status partial order relationship, the register status partial order relationship can be converted into a register lattice Hasse diagram as shown in Figure 2 . It can be seen from Figure 2 that the register lattice Hasse diagram has multiple nodes and multiple directed edges connecting two nodes. Among them, the directed edge indicates the inclusion relationship between the two nodes it connects. For example, node and The directed edge between nodes indicates that the 0 bit is part of the high 8 bits. For another example, The node and The directed edge between nodes indicates that the high 8 bits are part of the 16 bits.

[0039] According to an embodiment of the present invention, the backward data flow analysis module is configured to perform the intersection operation in the following manner: determine the operation bit widths corresponding to the register states of the two general registers for which the intersection operation is performed, and calculate the greatest lower bound between the operation bit widths to obtain the intersection operation result; wherein, the greatest lower bound between the 0-bit operation bit width and the other bit operation bit widths is the 0 bit; the greatest lower bound between the low 8-bit operation bit width and the high 8-bit operation bit width is the 0 bit, and the greatest lower bounds between the low 8-bit operation bit width and the 16-bit, 32-bit, and 64-bit operation bit widths are all the low 8 bits; the greatest lower bounds between the high 8-bit operation bit width and the 16-bit, 32-bit, and 64-bit operation bit widths are all the high 8 bits; the greatest lower bounds between the 16-bit operation bit width and the 32-bit and 64-bit operation bit widths are all the 16 bits; the greatest lower bound between the 32-bit operation bit width and the 64-bit operation bit width is the 32 bits.

[0040] It should be noted that when performing the intersection operation, according to the Figure 2 Register lattice Hasse diagram as shown to determine the greatest lower bound between the operation bit widths. The specific process is as follows: when performing the intersection operation, first delete the Or Nodes in the register lattice Hasse diagram to degenerate the register lattice Hasse diagram into a tree structure (at this time, from bottom to top are: Node, Or Node, Node, Node and Node); then use the lowest common ancestor algorithm on the tree to start traversing downward from the two corresponding nodes on the register lattice Hasse diagram until Node; mark the nodes passed by the two nodes (the nodes themselves are also included in the traversal process), and use the first common marked point encountered during the downward traversal of the two nodes as the intersection operation result. Among them, when performing the intersection operation of Node and Node, it is regarded as a special case, and there is no need to delete the Node or Node in the register lattice Hasse diagram, and directly use Node and Node to traverse downward until the first common marked point Node encountered.

[0041] According to an embodiment of the present invention, the backward data flow analysis module is configured to perform a difference operation in the following manner: determine the operation bit widths corresponding to the register states of two general-purpose registers for performing the difference operation, and calculate the difference set between the operation bit widths to obtain the difference operation result; wherein, the difference set between the 16-bit operation bit width and the high 8-bit operation bit width is the low 8-bit, the difference set between the 16-bit operation bit width and the low 8-bit operation bit width is the high 8-bit, and the difference sets between other operation bit widths are all 0 bits.

[0042] According to an embodiment of the present invention, the backward data flow analysis module is configured to perform a union operation in the following manner: determine the operation bit widths corresponding to the register states of two general-purpose registers for performing the union operation, and calculate the least upper bound between the operation bit widths to obtain the union operation result; wherein, the least upper bound between the 64-bit operation bit width and other bit operation bit widths is all 64 bits, the least upper bound between the 32-bit operation bit width and the 0-bit, low 8-bit, high 8-bit, and 16-bit operation bit widths is all 32 bits; the least upper bound between the 16-bit operation bit width and the 0-bit, low 8-bit, and high 8-bit operation bit widths is all 16 bits; the least upper bound between the high 8-bit operation bit width and the low 8-bit operation bit width is 16 bits, and the least upper bound between the high 8-bit operation bit width and the 0-bit operation bit width is the high 8-bit; the least upper bound between the low 8-bit and the 0-bit operation bit width is the low 8-bit.

[0043] It should be noted that when performing the union operation, the least upper bound between the operation bit widths is determined according to the register lattice Hasse diagram as shown in Figure 2 . The specific process is as follows: when performing the union operation, first delete the or nodes in the register lattice Hasse diagram to degenerate the register lattice Hasse diagram into a tree structure (at this time, from bottom to top are: node, or node, node, node and node); then adopt the nearest common ancestor algorithm on the tree, start traversing upward from the two corresponding nodes on the register lattice Hasse diagram until node; mark the nodes passed (the nodes themselves are also included in the traversal process), and use the first common marked point encountered during the upward traversal of the two nodes as the union operation result. Among them, when performing the union and intersection operations of and nodes, it is regarded as a special case, and there is no need to delete the or nodes in the register lattice Hasse diagram. Directly start traversing upward from the and nodes until the first common marked point node encountered.

[0044] According to an embodiment of the present invention, the backward data flow analysis module is configured to analyze each instruction in each basic block in a reverse order as follows: obtain the initial use set, the initial definition set, and the initial subsequent use set corresponding to the current instruction; wherein, the initial use set includes all general-purpose registers directly used by the instruction semantics itself and the initial register state of each general-purpose register, the initial definition set includes all general-purpose registers directly defined by the instruction semantics itself and the initial register state of each general-purpose register, and the initial subsequent use set is the target subsequent use set of the next instruction corresponding to the current instruction; perform an intersection operation on the register states of each general-purpose register in the initial definition set corresponding to the current instruction and the register states of the corresponding general-purpose registers in the initial subsequent use set corresponding to the current instruction to obtain the target definition set corresponding to the current instruction; perform a difference operation on the register states of each general-purpose register in the initial subsequent use set corresponding to the current instruction and the register states of the corresponding general-purpose registers in the target definition set corresponding to the current instruction to obtain the intermediate subsequent use set corresponding to the current instruction; perform a union operation on the register states of each general-purpose register in the intermediate subsequent use set corresponding to the current instruction and the register states of the corresponding general-purpose registers in the initial use set corresponding to the current instruction to obtain the target subsequent use set corresponding to the current instruction; wherein, if the basic block has no successor basic block, when analyzing the last instruction in this basic block, the initial subsequent use set corresponding to the last instruction is a plurality of pre-set general-purpose registers and the fixed register state of each general-purpose register; if the basic block has a successor basic block, when analyzing the last instruction in this basic block, the initial subsequent use set corresponding to the last instruction is obtained by performing a union operation on the target subsequent use sets of all successor basic blocks of this basic block, and the target subsequent use set of the successor basic block is the target subsequent use set corresponding to the starting instruction in this successor basic block.

[0045] Among them, when the basic block has no successor basic block, the initial subsequent use set corresponding to the last instruction of the basic block can be a plurality of pre-set general-purpose registers and the fixed register state of each general-purpose register. At this time, the plurality of pre-set general-purpose registers are all general-purpose registers, and the register state of each general-purpose register is , that is, the operation bit width of each general-purpose register is 64 bits. The purpose of such setting is to conservatively consider that all 64 bits of all general-purpose registers will be used after this basic block.

[0046] When the basic block has a successor basic block, the initial subsequent use set corresponding to the last instruction of this basic block is obtained by performing a union operation on the target subsequent use sets of all successor basic blocks of this basic block. For example, a basic block has two successor basic blocks, and the target subsequent use sets of these two successor basic blocks are {<%rax, >} and {<%rax, >}, then the initial subsequent use set corresponding to the last instruction in this basic block is the union of these two target subsequent use sets {<%rax, >}, that is, this basic block will subsequently use 32 bits of the general-purpose register %rax and will not use other general-purpose registers. It should also be noted that through the data flow analysis across basic blocks, the analysis of the initial subsequent use set of instructions can be made more accurate, thereby enabling the translation module to eliminate more redundant instructions during the translation process.

[0047] According to an embodiment of the present invention, the backward data flow analysis module is configured to obtain the initial use set and the initial definition set corresponding to the current instruction in the following manner: Determine all general-purpose registers involved in the use operation and the def operation in the current instruction, and initialize the operation bit width corresponding to the register state of each general-purpose register to 0 bits; among them, all general-purpose registers involved in the use operation constitute the source use set corresponding to the current instruction; all general-purpose registers involved in the def operation constitute the source definition set corresponding to the current instruction; for each general-purpose register involved in the use operation, perform a union operation on the register state corresponding to the general-purpose register in the source use set and the register state of the general-purpose register directly corresponding to the semantics of the current instruction to obtain the updated source use set of the current instruction; for each general-purpose register involved in the def operation, when the operation bit width corresponding to the register state of the general-purpose register in the semantics of the current instruction is 32 bits or 64 bits, update the operation bit width corresponding to the register state of the general-purpose register in the source definition set to 64 bits; for each general-purpose register involved in the def operation, when the operation bit width corresponding to the register state of the general-purpose register in the semantics of the current instruction is not 32 bits or 64 bits, update the register state of the general-purpose register in the updated source use set of the current instruction to an operation bit width of 64 bits; among them, the source use set and the source definition set that have completed all updates are used as the initial use set and the initial definition set corresponding to the current instruction.

[0048] IV. Translation Module

[0049] The one for translating the source program into a target program with other semantics, and eliminating redundant instructions of high-bit clearing or high-bit retention of general-purpose registers generated during translation based on the target definition set and the target subsequent use set corresponding to each instruction in the source program.

[0050] To better understand the binary translation system introduced in the foregoing embodiments, still taking the instruction sequence of X86 assembly shown in Table 1 as an example, it is illustrated how to use the binary translation system to translate it into a target program that meets the riscv architecture.

[0051] First, the data acquisition module acquires the instruction sequence of X86 assembly shown in Table 1 and passes it to the disassembly module for processing.

[0052] Then, the disassembly module analyzes the instruction sequence of the X86 assembly, determines that there is only one basic block in the instruction sequence of the X86 assembly, and passes the analysis result to the backward data flow analysis module.

[0053] Then, the backward data flow analysis module analyzes each instruction in reverse order to obtain the target definition set and the target subsequent use set corresponding to each instruction in the instruction sequence of the X86 assembly. The analysis process for each instruction is as follows.

[0054] First, analyze the cdqe instruction. The specific analysis steps are as follows.

[0055] Step 1: Determine the initial use set, the initial definition set, and the initial subsequent use set of the cdqe instruction. Among them, the initial subsequent use set of the cdqe instruction is {<%rax, >} (to simplify the analysis process, other general registers are ignored during the analysis, and ignoring other general registers does not affect the analysis result); for the initial use set and the initial definition set of the cdqe instruction, they are determined through analysis. The analysis process is as follows: The destination register of the cdqe instruction is %rax, and the source register of the cdqe is %eax. Initialize the operation bit width corresponding to the register states of the destination register and the source register of the cdqe to 0 bits, that is, the source definition set is {<%rax, >}, and the source use set is {<%eax, >}; for the source register %eax, perform a union operation with the register state corresponding to this general register in the source use set and the register state directly corresponding to the semantics of the cdqe instruction ({<%eax, >}), to obtain the updated source use set {<%eax, >}; for the destination register %rax, since the operation bit width corresponding to the register state of this general register in the semantics of the cdqe instruction is 64 bits, update the operation bit width corresponding to the register state of this general register in the source definition set of the cdqe instruction to 64 bits, to obtain the updated source definition set {<%rax, >}. Based on the above analysis, the initial use set of the cdqe instruction is {<%rax, >}, (since the general register %eax is the 32-bit of the general register %rax, for convenience of calculation, rewrite the initial use set as {<%rax, >}); the initial definition set of the cdqe instruction is {<%rax, >}, and the initial subsequent use set of the cdqe instruction is {<%rax, >}.

[0056] Step 2: Perform an intersection operation on the register states of the general-purpose registers in the initial definition set of the cdqe instruction ({<%rax, >}) and the register states of the corresponding general-purpose registers in the initial subsequent use set corresponding to the cdqe instruction ({<%rax, >}) to obtain the target definition set corresponding to the cdqe instruction {<%rax, >}. At this time, through the intersection operation, it can be determined that the cdqe instruction actually defines the general-purpose register %rax, and the operation bit width of the general-purpose register %rax is 64 bits.

[0057] Step 3: Perform a difference operation on the register states of the general-purpose registers in the initial subsequent use set corresponding to the cdqe instruction ({<%rax, >}) and the register states of the corresponding general-purpose registers in the target definition set corresponding to the cdqe instruction ({<%rax, >}) to obtain the intermediate subsequent use set corresponding to the cdqe instruction {<%rax, >}.

[0058] Step 4: Perform a union operation on the register states of the general-purpose registers in the intermediate subsequent use set corresponding to the cdqe instruction ({<%rax, >}) and the register states of the corresponding general-purpose registers in the initial use set corresponding to the cdqe instruction ({<%rax, >}) to obtain the target subsequent use set corresponding to the cdqe instruction {<%rax, >}. At this time, through the union operation, it can be determined that starting from the cdqe instruction, 32 bits of the general-purpose register %rax will be used.

[0059] Analyze the movl instruction next. The specific analysis steps are as follows.

[0060] Step 1: Determine the initial use set, initial definition set, and initial subsequent use set of the movl instruction. Among them, the initial subsequent use set of the movl instruction is {<%rax, >} (the target subsequent use set corresponding to the cdqe instruction); for the initial use set and initial definition set of the movl instruction, they are determined through analysis. The analysis process is as follows: the destination register of the movl instruction is %eax, the source register of the movl is %ebx, and the operation bit widths corresponding to the register states of the initialized destination register and source register of the movl are 0 bits, that is, the source definition set is {<%eax, >}, and the source use set is {<%ebx, >}; For the source register %ebx, perform a union operation on the register state corresponding to this register in the source use set and the register state of this general-purpose register directly corresponding to the semantics of the movl instruction ({%ebx, }), to obtain the updated source use set {%ebx, }; For the destination register %eax, since the register state corresponding to this general-purpose register in the semantics of the movl instruction has an operation bit width of 32 bits, update the register state of this general-purpose register in the source definition set to an operation bit width of 64 bits, to obtain the updated source definition set {<%eax, >}. Based on the above analysis, the initial use set of the movl instruction is {<%ebx, >}; the initial definition set of the movl instruction is {<%rax, >} (since the general-purpose register %eax is 32 bits of the general-purpose register %rax, for the convenience of calculation, rewrite the initial definition set as {<%rax, >}), and the initial subsequent use set of the movl instruction is {<%rax, >}.

[0061] Step 2: Perform an intersection operation on the register state of the general-purpose register in the initial definition set of the movl instruction ({<%rax, >}) and the register state of the corresponding general-purpose register in the initial subsequent use set corresponding to the movl instruction ({<%rax, >}) to obtain the target definition set {<%rax, >} corresponding to the movl instruction. At this time, through the intersection operation, it can be determined that the movl instruction actually defines 32 bits of the general-purpose register %rax. Therefore, when translating the movl instruction, there is no need to clear the register corresponding to the general-purpose register %rax in the target ISA (in riscv, these are two shift instructions).

[0062] Step 3: Perform a difference operation on the register state of the general-purpose register in the initial subsequent use set corresponding to the movl instruction ({<%rax, >}) and the register state of the corresponding general-purpose register in the target definition set corresponding to the movl instruction ({<%rax, >}) to obtain the intermediate subsequent use set {<%rax, >} corresponding to the movl instruction.

[0063] Step 4: Perform a difference operation on the register state of the general-purpose register in the intermediate subsequent use set corresponding to the movl instruction ({<%rax, >}) and the register state of the corresponding general-purpose register in the initial use set corresponding to the movl instruction {<%rax, >} (The general register %rax is not included in the initial usage set, indicating that the movl instruction does not use the general register %rax) performs a union operation to obtain {<%rax, >}; The register status of the general register <%ebx, > in the intermediate subsequent usage set corresponding to the movl instruction (the general register %ebx is not included in the intermediate subsequent usage set, so the status of the general register %ebx is set to ) is combined with the register status {<%ebx, >} of the corresponding general register in the initial usage set corresponding to the movl instruction to obtain {<%ebx, >}. At this time, the target subsequent usage set corresponding to the movl instruction is {<%rax, >,<%ebx, >}, so it can be determined that starting from the movl instruction, the general register %rax will not be used, and the 32 bits of the general register %ebx will be used.

[0064] Finally, the translation module translates the instruction sequence of the X86 assembly shown in Table 1 into add x1, x2, x0, where x0 is the zero register. Assume that %eax is mapped to x1 and %ebx is mapped to x2. It should be noted that if the analysis operation of the backward data flow analysis module is not performed, the instruction sequence of the X86 assembly shown in Table 1 will be translated into Instruction 1: add x1, x2, x0; Instruction 2: slli x1, x1, 32; Instruction 3: srli x1, x1, 32, where Instruction 2 and Instruction 3 are redundant shift instructions used to clear the upper 32 bits of the register corresponding to the general register %rax.

[0065] Based on the foregoing embodiments, the present invention also provides a binary translation method for an X86 program, which is used to translate a source program that follows X86 semantics into a target program that follows other semantics. The method includes: Step S1, obtaining the source program; Step S2, using the binary translation system described in the foregoing embodiments to translate the source program to obtain a target program that follows other semantics.

[0066] The beneficial effects of the present invention are as follows: (1) A backward data flow analysis module is set in the binary translation system to analyze the register states of the general registers corresponding to each instruction in the source code in the reverse analysis order from the back to the front, so that the translation module can eliminate the redundant instructions of high-order clearing or high-order retention of the general registers that may occur during the translation process based on the register states of the general registers corresponding to each analyzed instruction; (2) When analyzing each instruction, the backward data flow analysis module makes the analysis results of the register states of the general registers corresponding to each instruction more accurate through cross-basic block data flow analysis, and further enables the translation module to eliminate more redundant instructions during the translation process; (3) A register status lattice is configured in the backward data flow analysis module to perform intersection operation, difference operation and union operation, so as to more accurately analyze the register states of the general registers corresponding to each instruction, and further enable the translation module to eliminate more redundant instructions during the translation process.

[0067] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even the order can be changed, as long as the required functions can be achieved.

[0068] The present invention can be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0069] The computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing.

[0070] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A binary translation system for X86 programs, used to translate source programs that follow X86 semantics into target programs that follow other semantics, characterized in that: The system comprises: A data acquisition module, used for acquiring a source program, wherein the source program includes a plurality of instructions involving a general register use operation and a def operation; The disassembly module is used to divide the source program into multiple basic blocks and analyze the successor basic blocks corresponding to each basic block; wherein a basic block represents a continuous sequence of instructions in the source program, and the successor basic blocks of a basic block represent other basic blocks that can be directly reached from the basic block through the control flow; A backward data flow analysis module is used to analyze each instruction in each basic block in sequence from back to front to obtain a target definition set and a target subsequent use set corresponding to each instruction in the source program; wherein the target definition set includes all general registers actually defined by the corresponding instruction and the final register state of each general register, and the target subsequent use set includes all general registers actually used by the corresponding instruction and subsequent instructions and the final register state of each general register, and the register state indicates that the operation bit width of the general register is 64 bits, 32 bits, 16 bits, high 8 bits, low 8 bits or 0 bits; The translation module is used to translate the source program into a target program with other semantics, and to eliminate redundant instructions of clearing or retaining the high bits of general registers generated in the translation based on the target definition set and the target subsequent use set corresponding to each instruction in the source program.

2. The system according to claim 1, characterized in that The backward data flow analysis module is configured with a register state grid for performing intersection operations, difference operations and union operations, and the register state grid is composed of a register state set and a register state partial order relationship, wherein: The register state set includes , , , , , ,in, Indicates that the operation bit width is 64 bits. Indicates that the operation bit width is 32 bits. Indicates that the operation bit width is 16 bits. Indicates that the operation bit width is high 8 bits, Indicates that the operation bit width is the lower 8 bits. Indicates that the operation bit width is 0 bit; The register state partial order relationship includes , , , , , ,in, Indicates that bit 0 is part of the lower 8 bits, Indicates that bit 0 is part of the high 8 bits, Indicates that the lower 8 bits are part of the 16 bits, Indicates that the upper 8 bits are part of the 16 bits. Indicates that 16 bits are part of 32 bits, Indicates that 32 bits are part of 64 bits.

3. The system according to claim 2, characterized in that The backward data flow analysis module is configured to perform the intersection operation as follows: Determine the operation bit width corresponding to the register state of two general registers performing the intersection operation, and calculate the maximum infimum between the operation bit widths to obtain the intersection operation result; Among them, the maximum lower bounds of the 0-bit operation bit width and the other bit operation bit widths are both 0 bits; the maximum lower bounds of the low 8-bit operation bit width and the high 8-bit operation bit width are 0 bits, and the maximum lower bounds of the low 8-bit operation bit width and the 16-bit, 32-bit, and 64-bit operation bit widths are all low 8 bits; the maximum lower bounds of the high 8-bit operation bit width and the 16-bit, 32-bit, and 64-bit operation bit widths are all high 8 bits; the maximum lower bounds of the 16-bit operation bit width and the 32-bit and 64-bit operation bit widths are all 16 bits; the maximum lower bounds of the 32-bit operation bit width and the 64-bit operation bit width are 32 bits.

4. The system according to claim 3, characterized in that The backward data flow analysis module is configured to perform the difference operation as follows: Determine the operation bit width corresponding to the register states of the two general registers performing the difference operation, and calculate the difference set between the operation bit widths to obtain the difference operation result; The difference between the 16-bit operation bit width and the high 8-bit operation bit width is the low 8 bits, the difference between the 16-bit operation bit width and the low 8-bit operation bit width is the high 8 bits, and the differences between other operation bit widths are all 0 bits.

5. The system according to claim 4, characterized in that The backward data flow analysis module is configured to perform the following operations: Determine the operation bit width corresponding to the register states of the two general registers performing the union operation, and calculate the minimum supremum between the operation bit widths to obtain the union operation result; Among them, the minimum supremum of 64-bit operation bit width and other bit operation bit widths is 64 bits, the minimum supremum of 32-bit operation bit width and 0 bit, low 8 bits, high 8 bits and 16-bit operation bit width is 32 bits; the minimum supremum of 16-bit operation bit width and 0 bit, low 8 bits and high 8 bits operation bit width is 16 bits; the minimum supremum of high 8-bit operation bit width and low 8-bit operation bit width is 16 bits, the minimum supremum of high 8-bit operation bit width and 0 bit operation bit width is high 8 bits; the minimum supremum of low 8 bits and 0 bit operation bit width is low 8 bits.

6. The system according to claim 5, characterized in that The backward data flow analysis module is configured to analyze each instruction in each basic block in sequence from back to front in the following manner: Obtaining an initial use set, an initial definition set, and an initial subsequent use set corresponding to the current instruction; wherein the initial use set includes all general registers directly used by the corresponding instruction semantics itself and the initial register state of each general register, the initial definition set includes all general registers directly defined by the corresponding instruction semantics itself and the initial register state of each general register, and the initial subsequent use set is the target subsequent use set of the next instruction of the corresponding instruction; Perform an intersection operation on the register state of each general register in the initial definition set corresponding to the current instruction and the register state of the corresponding general register in the initial subsequent use set corresponding to the current instruction, so as to obtain a target definition set corresponding to the current instruction; Performing a difference operation on the register state of each general register in the initial subsequent use set corresponding to the current instruction and the register state of the corresponding general register in the target definition set corresponding to the current instruction, so as to obtain an intermediate subsequent use set corresponding to the current instruction; Perform a AND operation on the register state of each general register in the intermediate subsequent use set corresponding to the current instruction and the register state of the corresponding general register in the initial use set corresponding to the current instruction, so as to obtain a target subsequent use set corresponding to the current instruction; Among them, if the basic block has no successor basic block, when analyzing the last instruction in the basic block, the initial subsequent use set corresponding to the last instruction is a plurality of pre-set general registers and a fixed register state of each general register; if the basic block has a successor basic block, when analyzing the last instruction in the basic block, the initial subsequent use set corresponding to the last instruction is obtained by performing a AND operation on the target subsequent use sets of all the successor basic blocks of the basic block, and the target subsequent use set of the successor basic block is the target subsequent use set corresponding to the starting instruction in the successor basic block.

7. The system according to claim 6, characterized in that The backward data flow analysis module is configured to obtain the initial usage set and the initial definition set corresponding to the current instruction in the following manner: Determine all general registers involved in use operations and def operations in the current instruction, and initialize the operation bit width corresponding to the register state of each general register to 0 bits; wherein all general registers involved in use operations constitute the source use set corresponding to the current instruction; and all general registers involved in def operations constitute the source definition set corresponding to the current instruction; For each general register involved in the use operation, a register state corresponding to the general register in the source use set is combined with the register state of the general register directly corresponding to the semantics of the current instruction to obtain an updated source use set of the current instruction; For each general register involved in the def operation, when the operation bit width corresponding to the register state of the general register in the current instruction semantics is 32 bits or 64 bits, the operation bit width corresponding to the register state of the general register in the source definition set is updated to 64 bits; For each general register involved in the def operation, when the register state of the general register in the current instruction semantics corresponds to an operation bit width of not 32 bits or 64 bits, the register state of the general register in the updated source usage set of the current instruction is updated to an operation bit width of 64 bits; Among them, the source usage set and source definition set that have completed all updates are used as the initial usage set and initial definition set corresponding to the current instruction.

8. A binary translation method for an X86 program, for translating a source program conforming to X86 semantics into a target program conforming to other semantics, characterized in that: The method comprises: Step S1, obtaining the source program; Step S2: using the system as described in any one of claims 1 to 7 to translate the source program to obtain a target program that complies with other semantics.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to claim 8.

10. An electronic device, characterized in that: include: one or more processors, and memory; Wherein, each processor is configured with a system as described in any one of claims 1 to 7, which is used to translate a source program following X86 semantics into a target program following other semantics.

Citation Information

Cited By

  • Binary translation optimization method, binary translator and electronic equipment

    CN122018920A

  • Binary translation optimization method, binary translator and electronic device

    CN122018920B