Optimization method and optimizer applied to X86 vector instruction translation

Through mask optimization and vsetvli optimization steps, redundant instructions during the translation of X86 vector instructions into RISC-V vector instructions are eliminated, solving the problem of inefficient code execution after translation and achieving more efficient instruction conversion.

CN120335867AActive Publication Date: 2025-07-18INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510813877.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing open source dynamic binary translator generates redundant setting instructions during the process of translating X86 vector instructions into RISC-V vector instructions, affecting the execution efficiency of the translated code.

Method used

Through the mask optimization steps and vsetvli optimization steps, the extra mask register setting instructions and vtype register setting instructions during the X86 vector instruction translation process were deleted, including technical means such as pseudo instruction replacement, instruction scheduling and data flow analysis.

Benefits of technology

Improve the execution efficiency of X86 vector instructions translated into RISC-V vector instructions, eliminate redundant instructions, and improve the running performance of the code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335867A_ABST
    Figure CN120335867A_ABST
Patent Text Reader

Abstract

The invention provides an optimization method applied to X86 vector instruction translation, which is used for eliminating redundant instructions generated in an X86 vector instruction translation process, and comprises the following steps: obtaining a to-be-optimized code which is obtained after translation processing and comprises a plurality of instructions, in the mask optimization step, redundant mask register setting instructions in to-be-optimized codes are deleted according to a preset mask optimization rule so as to obtain mask optimization codes; and a vsetvli optimization step: deleting all csrr instructions and redundant vsetvli instructions in the mask optimization code according to a preset instruction optimization rule so as to obtain a target optimization code. According to the technical scheme, through mask optimization and vsetvli optimization, the problem that redundant instructions are generated in the X86 vector instruction translation process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of binary translation, specifically to instruction optimization techniques in the field of binary translation, and more specifically, to an optimization method and optimizer applied to X86 vector instruction translation. Background Art

[0002] With the development of computer architectures, there have emerged some different instruction set architectures. The vitality of an instruction set architecture lies in the perfection of its ecosystem. Currently, the two major mainstream ecosystems are the Wintel ecosystem with Windows and Intel as an alliance, and the AA ecosystem with Android and Arm as an alliance. It is worth noting that neither the X86 nor the Arm instruction set architectures are open-source instruction set architectures, which poses a technical barrier to the independent research and development of domestic processors. Against this background, the RISC-V instruction set architecture, which combines flexibility, openness, and simplicity, has become the best choice for the research and development of domestic processors. However, the key challenge faced by the RISC-V architecture is the lack of maturity of its ecosystem, and binary translation technology is needed to break through the ecological barriers.

[0003] Vector instructions generally refer to instructions in a computer architecture that can perform the same operation on multiple data elements simultaneously, thereby significantly improving the performance of the processor, reducing the number of instructions, and taking advantage of data parallelism. In the X86 architecture, vector instructions are mainly implemented through SIMD (Single Instruction Multiple Data) technology. SIMD instructions sequentially perform the same operation on multiple data, which are usually organized in a vector or an array. X86 vector instructions have gone through four main stages: MMX, SSE, AVX, and AVX512 over time. As of today, the X86 vector architecture has 32 512-bit vector registers, and the vector instructions and vector architecture in X86 provide powerful data parallel processing capabilities. In contrast, the RISC-V architecture adopts a modular vector extension scheme, supporting configurable vector lengths of 128 / 256 / 512 bits, and the element bit width is dynamically set through the sew field of the vtype register, which is significantly different from the design in X86 that solidifies the element bit width in the instruction opcode.

[0004] To break through the short board of the RISC-V ecosystem, dynamic binary translation technology can be used to convert X86 vector instructions into RISC-V instructions, so as to achieve the compatible operation of existing X86 software on RISC-V hardware. In the field of binary translation, the translation strategy of vector instructions has a significant impact on the execution efficiency of the translated code. At present, there are two main methods for translating vector instructions. One is to translate the vector instructions of the source architecture into scalar instructions of the target architecture. This method is simple to implement, does not require in-depth understanding of the vector instruction set of the target architecture, and can run on a target architecture that does not support vector instructions. However, this will reduce the execution efficiency of the translated code. The other method is to translate the vector instructions of the source architecture into vector instructions of the target architecture. This method requires in-depth understanding of the vector instruction sets of both architectures. This method is not applicable to a target architecture that does not support vector instructions, but it can make full use of the resources of the target architecture and significantly improve the execution efficiency of the translated code.

[0005] Although the existing two translation methods can both achieve the translation from the X86 architecture to the RISC-V architecture, considering the execution efficiency of the translated code, an open-source dynamic binary translator is usually used to translate the source X86 vector instructions into RISC-V vector instructions. However, some redundant setup instructions will be generated during the process of translating the source X86 vector instructions into RISC-V vector instructions by the open-source dynamic binary translator, and these redundant setup instructions will affect the execution efficiency of the translated code.

[0006] It should be noted that: This background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily prior art. Without evidence showing that the relevant information was publicly available before the filing date of the present invention, the relevant information should not be regarded as prior art. Summary of the Invention

[0007] Therefore, the purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide an optimization method and an optimizer applied to the translation of X86 vector instructions.

[0008] The purpose of the present invention is achieved by the following technical solutions.

[0009] According to a first aspect of the present invention, there is provided an optimization method applied to X86 vector instruction translation for eliminating redundant instructions generated during the X86 vector instruction translation process. The method includes obtaining the to-be-optimized code including multiple instructions obtained after translation processing, and performing the following steps: Mask optimization step: deleting redundant mask register setting instructions in the to-be-optimized code according to a preset mask optimization rule to obtain mask-optimized code; vsetvli optimization step: deleting all csrr instructions and redundant vsetvli instructions in the mask-optimized code according to a preset instruction optimization rule to obtain target-optimized code.

[0010] In some embodiments of the present invention, the preset mask optimization rule is: replacing all mask register setting instructions in the to-be-optimized code with pseudo-instructions according to a known instruction manual, and setting a high-bit identifier and a low-bit identifier in the pseudo-instruction to mark the usage status of the mask register when executing the corresponding pseudo-instruction; wherein, the to-be-optimized code includes multiple sequentially arranged code blocks, and some code blocks include instructions with labels, and the instructions with labels indicate that the instruction is a jump execution instruction corresponding to other instructions; setting a low-bit variable identifier and a high-bit variable identifier to mark the immediate usage status of the mask register, and analyzing each instruction in the to-be-optimized code one by one according to a preset pseudo-instruction elimination rule based on the immediate usage status of the mask register to delete redundant pseudo-instructions; selecting an instruction sequence corresponding to the pseudo-instruction from a preset instruction sequence list to replace each remaining pseudo-instruction in the to-be-optimized code to obtain mask-optimized code.

[0011] In some embodiments of the present invention, the preset pseudo-instruction elimination rule analyzes each instruction in the to-be-optimized code one by one in the following manner: determining whether the current instruction is a pseudo-instruction, if it is a pseudo-instruction, then analyzing whether the low-bit variable identifier and the high-bit variable identifier of the immediate usage status of the mask register are the same as the high-bit identifier and the low-bit identifier of the pseudo-instruction, if they are the same, deleting the pseudo-instruction, if they are not the same, updating the low-bit variable identifier and the high-bit variable identifier of the immediate usage status of the mask register to be the same as the high-bit identifier and the low-bit identifier of the pseudo-instruction, and ending the analysis of the current instruction; if the current instruction is not a pseudo-instruction, then analyzing whether the current instruction has a label, if the current instruction has a label, then analyzing whether the usage statuses of the mask registers corresponding to all other instructions that can jump to the current instruction are the same, if they are the same, updating the low-bit variable identifier and the high-bit variable identifier of the immediate usage status of the mask register to be the same as the usage status of the mask register corresponding to the other instructions, if they are not the same, initializing the low-bit variable identifier and the high-bit variable identifier and ending the analysis of the current instruction; if the current instruction does not have a label, then ending the analysis of the current instruction; wherein, initializing the low-bit variable identifier and the high-bit variable identifier means setting the immediate usage status of the mask register to unused.

[0012] In some embodiments of the present invention, the preset instruction optimization rules are as follows: all csrr instructions and redundant vsetvl instructions in the masked optimization code are deleted according to the preset conversion rules, and the remaining vsetvl instructions in the masked optimization code are replaced with vsetvli instructions to obtain intermediate optimized code; invalid vsetvli instructions in the intermediate optimized code are deleted according to the preset first-stage elimination rules to obtain first-stage optimized code; instruction scheduling is performed on the first-stage optimized code according to the preset instruction scheduling rules to adjust the arrangement order of each code block in the first-stage optimized code; and duplicate vsetvli instructions in the scheduled first-stage optimized code are deleted according to the preset second-stage elimination rules to obtain target optimized code.

[0013] In some embodiments of the present invention, the preset conversion rules are as follows: each vsetvl instruction in the masked optimization code and each crss instruction corresponding to each vsetvl are determined; each vsetvl instruction in the masked optimization code is analyzed; among them, the analysis process of each vsetvl instruction is as follows: starting from the current vsetvl instruction, traversing downward until the next vsetvl instruction or vsetvli instruction after the current vsetvl instruction is found, and analyzing whether other instructions between the current vsetvl instruction and the next vsetvl instruction or vsetvli instruction use a specific field in the vector configuration register; if not, the current vsetvl instruction and the corresponding csrr instruction are deleted, and the analysis of the current vsetvl instruction ends; if so, starting from the csrr instruction corresponding to the current vsetvl instruction, traversing upward until the previous vsetvli instruction of the csrr instruction is found, and replacing the current vsetvl instruction with the previous vsetvli instruction, while deleting the csrr instruction corresponding to the current vsetvl instruction and ending the analysis of the current vsetvl instruction; among them, the vector configuration register is the vtype register, and the specific field in the vector configuration register is the sew field in the vtype register.

[0014] In some embodiments of the present invention, the preset first-stage elimination rule is: perform first-stage elimination processing on each vsetvli instruction in the intermediate code to delete invalid vsetvli instructions; wherein, the first-stage elimination processing process for each vsetvli instruction is: starting from the current vsetvli instruction, traverse downward until the next vsetvli instruction after the current vsetvli instruction is found, and analyze whether other instructions between the current vsetvli instruction and the next vsetvli instruction use specific fields in the vector configuration register; if other instructions between the current vsetvli instruction and the next vsetvli instruction do not use specific fields in the vector configuration register, then delete the current vsetvli instruction and end the processing of the current vsetvli instruction; otherwise, do not delete the current vsetvli instruction and end the processing of the current vsetvli instruction.

[0015] In some embodiments of the present invention, the preset instruction scheduling rule is: perform data flow analysis on each code block in the first-stage optimized code to determine whether there is a data flow dependency relationship between the code blocks; wherein, the data flow dependency relationship means that when executing an instruction in a code block, it is necessary to rely on the execution results of other instructions in other code blocks that are sorted earlier; perform instruction scheduling on any two code blocks in the first-stage optimized code that do not have a data flow dependency relationship until the instruction scheduling of each code block is completed, wherein, perform instruction scheduling on any two code blocks in the following manner: determine the usage status of specific fields in the vector configuration register indicated by the last vsetvli instruction in the code block sorted earlier, and the usage status of specific fields in the vector configuration register indicated by the first vsetvli instruction in the code block sorted later, and analyze whether the usage status of specific fields in the vector configuration registers indicated by these two instructions is the same. If they are the same, then schedule the code block sorted later after the code block sorted earlier without affecting the execution results of other code blocks.

[0016] In some embodiments of the present invention, the preset two-stage elimination rule is as follows: A variable parameter identifier is set to mark the immediate use status of a specific field in the vector configuration register, and based on the immediate use status of the specific field in the vector configuration register marked by the variable parameter identifier, each instruction in the one-stage optimized code after instruction scheduling is subjected to two-stage elimination processing one by one; wherein, the two-stage elimination processing process for each instruction is as follows: Determine whether the current instruction is a vsetvli instruction. If it is a vsetvli instruction, analyze whether the immediate use status of the specific field in the vector configuration register marked by the variable parameter identifier is consistent with the use status of the specific field in the vector configuration register corresponding to the current instruction. If they are consistent, delete the current instruction and end the processing of the current instruction; if they are inconsistent, update the variable parameter identifier to be consistent with the use status of the specific field in the vector configuration register corresponding to the current instruction and end the processing of the current instruction; if the current instruction is not a vsetvli instruction, analyze whether the current instruction is set with a label; if the current instruction is set with a label, analyze whether the use status of the specific field in the vector configuration registers corresponding to all other instructions that can jump to the current instruction is consistent. If they are consistent, update the variable parameter identifier to be consistent with the use status of the specific field in the vector configuration registers corresponding to the other instructions and end the processing of the current instruction. If they are inconsistent, initialize the variable parameter identifier and end the processing of the current instruction; if the current instruction is not set with a label, end the processing of the current instruction; wherein, initializing the variable parameter identifier means setting the immediate use status of the specific field in the vector configuration register to unused.

[0017] According to a second aspect of the present invention, an optimizer is provided for eliminating redundant instructions generated during the translation of X86 vector instructions. The optimizer includes: a data acquisition module for acquiring the to-be-optimized code containing multiple instructions obtained after translation processing; a mask optimization module for deleting redundant mask register setting instructions in the to-be-optimized code according to a preset mask optimization rule to obtain mask-optimized code; a vsetvli optimization module for deleting all csrr instructions and redundant vsetvli instructions in the mask-optimized code according to a preset instruction optimization rule to obtain target optimized code.

[0018] Compared with the prior art, the advantages of the present invention are as follows: Redundant mask register setting instructions and vtype register setting instructions generated during the translation of X86 vector instructions into RISC-V vector instructions are deleted, improving the execution efficiency. Description of the Drawings

[0019] The following further illustrates embodiments of the present invention with reference to the accompanying drawings, wherein:

[0020] Figure 1Schematic diagram of the optimization method flow according to an embodiment of the present invention;

[0021] Figure 2 Schematic diagram of the preset pseudo-instruction elimination rule flow according to an embodiment of the present invention;

[0022] Figure 3 Schematic diagram of the preset conversion rule flow according to an embodiment of the present invention;

[0023] Figure 4 Schematic diagram of the two-stage elimination rule flow according to an embodiment of the present invention;

[0024] Figure 5 Schematic diagram of the optimizer composition according to an embodiment of the present invention. Detailed implementation manners

[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific embodiments with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0026] As mentioned in the background art section, some redundant setting instructions will be generated during the process of the open-source dynamic binary translator translating the source X86 vector instructions into RISC-V vector instructions, and these redundant setting instructions will affect the execution efficiency of the translated code.

[0027] To solve the above problems, the inventors analyzed the translation process of X86 vector instructions and found that when translating a single X86 instruction, due to the lack of context information between instructions, vtype registers and mask registers will be set separately for each instruction, resulting in a large number of redundant setting instructions being generated. Based on this, the inventors proposed an instruction optimization method for eliminating the redundant instructions generated during the translation process of X86 vector instructions. In this method, a mask optimization step and a vsetvli optimization step are included. Among them, in the mask optimization step, the redundant mask register setting instructions are deleted according to the preset mask optimization rules; in the vsetvli optimization step, the redundant vtype register setting instructions are deleted according to the preset instruction optimization rules.

[0028] Generally speaking, as Figure 1As shown in the figure, an optimization method for X86 vector instruction translation is used to eliminate redundant instructions generated during the X86 vector instruction translation process. The method includes obtaining the to-be-optimized code containing multiple instructions after translation processing, and performing the following steps: Mask optimization step: Delete the redundant mask register setting instructions in the to-be-optimized code according to the preset mask optimization rules to obtain the mask-optimized code; vsetvli optimization step: Delete all csrr instructions and redundant vsetvli instructions in the mask-optimized code according to the preset instruction optimization rules to obtain the target optimized code.

[0029] To better understand the present invention, the mask optimization step and the vsetvli optimization step will be described in detail below with specific embodiments.

[0030] I. Mask optimization step

[0031] In the mask optimization step, the redundant mask register setting instructions in the to-be-optimized code are deleted according to the preset mask optimization rules to obtain the mask-optimized code.

[0032] According to an embodiment of the present invention, the preset mask optimization rules are as follows: Replace all mask register setting instructions in the to-be-optimized code with pseudo-instructions according to the known instruction manual, and set high-order identifiers and low-order identifiers in the pseudo-instructions to mark the usage status of the mask register when executing the corresponding pseudo-instructions; wherein, the to-be-optimized code includes multiple sequentially arranged code blocks, and some code blocks contain instructions with labels, and the instructions with labels indicate that the instructions are jump execution instructions corresponding to other instructions; Set low-order variable identifiers and high-order variable identifiers to mark the immediate usage status of the mask register, and analyze each instruction in the to-be-optimized code one by one based on the immediate usage status of the mask register according to the preset pseudo-instruction elimination rules to delete redundant pseudo-instructions; Select the instruction sequence corresponding to the pseudo-instruction from the preset instruction sequence list to replace each remaining pseudo-instruction in the to-be-optimized code to obtain the mask-optimized code.

[0033] It should be noted that the reason for setting pseudo-instructions to replace the mask register setting instructions is that the mask register setting instructions involve multiple instructions, and the specific value of the mask register cannot be inferred from multiple instructions. If pseudo-instructions are not introduced for replacement, it is difficult to effectively analyze redundant mask register setting instructions during analysis.

[0034] According to an embodiment of the present invention, the preset pseudo-instruction elimination rule analyzes each instruction in the code to be optimized one by one in the following manner: Determine whether the current instruction is a pseudo-instruction. If it is a pseudo-instruction, analyze whether the low variable identifier and high variable identifier of the immediate use status of the mask register are the same as the high identifier and low identifier of the pseudo-instruction. If they are the same, delete the pseudo-instruction. If they are not the same, update the low variable identifier and high variable identifier of the immediate use status of the mask register to be consistent with the high identifier and low identifier of the pseudo-instruction, and end the analysis of the current instruction. If the current instruction is not a pseudo-instruction, analyze whether the current instruction is set with a label. If the current instruction is set with a label, analyze whether the use statuses of the mask registers corresponding to all other instructions that can jump to the current instruction are the same. If they are the same, update the low variable identifier and high variable identifier of the immediate use status of the mask register to be consistent with the use status of the mask register corresponding to the other instructions. If they are not the same, initialize the low variable identifier and high variable identifier and end the analysis of the current instruction. If the current instruction is not set with a label, end the analysis of the current instruction. Among them, the initialization of the low variable identifier and high variable identifier means setting the immediate use status of the mask register to unused.

[0035] To better understand the mask optimization step, the following will be described in detail with reference to the code example shown in Table 1 (the code example given in Table 1 contains 9 instructions and no instructions with labels) and Figure 2 the schematic diagram of the preset pseudo-instruction elimination rule process shown. Among them, the pseudo-instruction is represented by set_mask lo,hi, where lo represents the lower 64 bits (low identifier) of the mask register. When lo = 1, it means using the lower 64 bits of the mask register. When lo = 0, it means not using the lower 64 bits of the mask register. hi represents the upper 64 bits of the mask register. When hi = 1, it means using the upper 64 bits of the mask register. When hi = 0, it means not using the upper 64 bits of the mask register. For example, the pseudo-instruction set_mask 0x0,0x1 means not using the lower 64 bits of the mask register and using the upper 64 bits of the mask register. pre_lo represents the low variable identifier, and pre_hi represents the high variable identifier.

[0036] From Figure 2It can be seen that before starting mask optimization, the low variable identifier and the high variable identifier are initialized, that is, pre_lo = -1 and pre_hi = -1 are set, which means that the immediate usage status of the mask register is that the lower 64 bits and the upper 64 bits of the mask register are not used, and cur_ins = hd is initialized, where cur_ins represents the currently pointed instruction, and cur_ins = hd means that the currently pointed instruction is the first instruction in the code shown in Table 1; then, each instruction is analyzed one by one starting from the first instruction in the code example according to the preset pseudo-instruction elimination rules, and the code example after elimination processing as shown in Table 2 is obtained (Table 2 contains 7 instructions); finally, the instruction sequence corresponding to the pseudo-instruction is selected from the preset instruction sequence table to replace each remaining pseudo-instruction in Table 2 to obtain the mask optimization code.

[0037] Among them, when analyzing the first instruction set_mask 0x0, 0x1, cur_ins = set_mask 0x0, 0x1, pre_lo = -1, pre_hi = -1. The specific analysis process is as follows: first, it is judged whether the currently pointed instruction is not empty (cur_ins!= null). At this time, cur_ins points to the first instruction set_mask 0x0, 0x1, which satisfies cur_ins!= null; then, it is judged whether cur_ins points to a pseudo-instruction. At this time, cur_ins points to the pseudo-instruction set_mask 0x0, 0x1, and then it is judged whether the low variable identifier and the high variable identifier are the same as the low identifier and the high identifier in the pseudo-instruction. Since pre_lo = -1, pre_hi = -1, and in the pseudo-instruction represented by the first instruction, lo = 0, hi = 1, so pre_lo ≠ lo, pre_hi ≠ hi; the low variable identifier and the high variable identifier are modified to be the same as the low identifier and the high identifier in the pseudo-instruction pointed to by the current instruction, that is, pre_lo = 0, pre_hi = 1, and the analysis of the first instruction is ended, and cur_ins points to the next instruction.

[0038] When analyzing the second instruction vsetvli x0,x0,e64, cur_ins = vsetvli x0,x0,e64, pre_lo = 0, and pre_hi = 1. The specific analysis process is as follows: First, it is judged whether the currently pointed instruction is not empty (cur_ins!= null). At this time, cur_ins points to the second instruction vsetvli x0,x0,e64, satisfying cur_ins!= null. Then, it is judged whether cur_ins points to a pseudo-instruction. At this time, cur_ins points to a non-pseudo-instruction vsetvli x0,x0,e64. When cur_ins does not point to a pseudo-instruction, it is analyzed whether the instruction pointed to by cur_ins has a label. Since the instruction vsetvli x0,x0,e64 pointed to by cur_ins has no label, the analysis of the second instruction ends, and cur_ins points to the next instruction.

[0039] When analyzing the third instruction vfadd.vv v6,v6,v8,v0, cur_ins = vfadd.vv v6,v6,v8,v0, pre_lo = 0, and pre_hi = 1. The specific analysis process is as follows: First, it is judged whether the currently pointed instruction is not empty (cur_ins!= null). At this time, cur_ins points to the third instruction vfadd.vv v6,v6,v8,v0, satisfying cur_ins!= null. Then, it is judged whether cur_ins points to a pseudo-instruction. At this time, cur_ins points to a non-pseudo-instruction vfadd.vv v6,v6,v8,v0. When cur_ins does not point to a pseudo-instruction, it is analyzed whether the instruction pointed to by cur_ins has a label. Since the instruction vfadd.vv v6,v6,v8,v0 pointed to by cur_ins has no label, the analysis of the third instruction ends, and cur_ins points to the next instruction.

[0040] When analyzing the 4th instruction set_mask 0x0, 0x1, cur_ins = set_mask 0x0, 0x1, pre_lo = 0, and pre_hi = 1. The specific analysis process is as follows: First, it is judged whether the currently pointed instruction is not empty (cur_ins!= null). At this time, cur_ins points to the 4th instruction set_mask 0x0, 0x1, which satisfies cur_ins!= null. Then, it is judged whether cur_ins points to a pseudo-instruction. At this time, cur_ins points to the pseudo-instruction set_mask 0x0, 0x1. Then, it is judged whether the low variable flag and the high variable flag are consistent with the low flag and the high flag in the pseudo-instruction. Since pre_lo = 0, pre_hi = 1, and in the pseudo-instruction represented by the 4th instruction, lo = 0 and hi = 1, so pre_lo = lo and pre_hi = hi. The 4th instruction set_mask 0x0, 0x1 can be deleted, the analysis of the 4th instruction is ended, and cur_ins points to the next instruction.

[0041] The analysis process of the subsequent 5th to 9th instructions is the same as that of the previous instructions, and will not be elaborated here one by one.

[0042] Combined with Table 1 and Table 2, it can be seen that after analyzing and processing the code example shown in Table 1 according to the preset pseudo-instruction elimination rules, the redundant 4th and 7th corresponding pseudo-instructions in the code example can be deleted.

[0043] Table 1

[0044]

[0045]

[0046] Based on the foregoing embodiments, it can be seen that in the mask optimization step, pseudo-instructions are introduced to identify the usage status of the mask register, and the low variable flag and the high variable flag are also set to mark the immediate usage status of the mask register. By analyzing whether the usage status of the mask register corresponding to each instruction is the same as the immediate usage status of the mask register, redundant pseudo-instructions are deleted, and the remaining pseudo-instructions are replaced with a normal instruction sequence, realizing the optimization of the mask register setting instruction, and thus improving the code execution efficiency.

[0047] II. vsetvli Optimization Step

[0048] In the vsetvli optimization step, all csrr instructions and redundant vsetvli instructions in the mask optimization code are deleted according to the preset instruction optimization rules to obtain the target optimization code.

[0049] According to an embodiment of the present invention, the preset instruction optimization rule is as follows: all csrr instructions and redundant vsetvl instructions in the masked optimization code are deleted according to the preset conversion rule, and the remaining vsetvl instructions in the masked optimization code are replaced with vsetvli instructions to obtain intermediate optimization code; invalid vsetvli instructions in the intermediate optimization code are deleted according to the preset first-stage elimination rule to obtain first-stage optimization code; instruction scheduling is performed on the first-stage optimization code according to the preset instruction scheduling rule to adjust the arrangement order of each code block in the first-stage optimization code; and duplicate vsetvli instructions in the scheduled first-stage optimization code are deleted according to the preset second-stage elimination rule to obtain target optimization code.

[0050] According to an embodiment of the present invention, the preset conversion rule is as follows: each vsetvl instruction in the masked optimization code and each crss instruction corresponding to each vsetvl are determined; each vsetvl instruction in the masked optimization code is analyzed; wherein, the analysis process of each vsetvl instruction is as follows: starting from the current vsetvl instruction, traversing downward until the next vsetvl instruction or vsetvli instruction after the current vsetvl instruction is found, and analyzing whether other instructions between the current vsetvl instruction and the next vsetvl instruction or vsetvli instruction use a specific field in the vector configuration register; if not, the current vsetvl instruction and the corresponding csrr instruction are deleted, and the analysis of the current vsetvl instruction ends; if so, starting from the csrr instruction corresponding to the current vsetvl instruction, traversing upward until the previous vsetvli instruction of the csrr instruction is found, and replacing the current vsetvl instruction with the previous vsetvli instruction, while deleting the csrr instruction corresponding to the current vsetvl instruction and ending the analysis of the current vsetvl instruction; wherein, the vector configuration register is a vtype register, and the specific field in the vector configuration register is the sew field in the vtype register.

[0051] It should be noted that if an instruction related to the sew field in the vtype register is found between the current vsetvl instruction or vsetvli instruction and the next vsetvl instruction or vsetvli instruction, it is considered that the sew field in the vtype register is used; if no instruction related to the sew field in the vtype register is found between the current vsetvl instruction and the next vsetvl instruction or vsetvli instruction, it is considered that the sew field in the vtype register is not used. Among them, it can be judged whether an instruction is related to the sew field according to whether the value of the sew field affects the execution result of the instruction. If the value of the sew field affects the execution result of the instruction, the instruction is an instruction related to the sew field; if the value of the sew field does not affect the execution result of the instruction, the instruction is an instruction unrelated to the sew field. For example, scalar instructions, bitwise operation vector instructions, and shift vector instructions in RISC-V instructions are all instructions unrelated to the sew field; vector operation instructions are instructions related to the sew field.

[0052] To better understand the preset conversion rules, the following takes the code shown in Table 3 as an example (Table 3 contains 13 instructions and there are no instructions with labels), and is described in combination with Figure 3 the schematic diagram of the preset conversion rule process shown.

[0053] For the code example given in Table 3: First, determine each vsetvl instruction in the example and the corresponding csrr instruction for each vsetvl instruction. In the example, only the 6th instruction is a vsetvl instruction, and the 6th instruction corresponds to the 4th csrr instruction (the vsetvl instruction and the csrr instruction appear in pairs. The csrr instruction is used to save the state of the vtype register, and vsetvl is used to restore the state of the vtype register. The next one saved must be the one to restore); starting from the 6th vsetvl instruction, traverse the masked optimized code downward until the 8th vsetvli instruction is found. At this time, the 7th instruction vfadd.vvv6,v6,v8,v0 is an instruction related to the sew field, which means that the 7th instruction between the 6th instruction and the 8th instruction uses the sew field in the vector configuration register; starting from the 4th csrr instruction corresponding to the 6th vsetvl instruction, traverse upward until the previous vsetvli instruction (vsetvli x0,x0,e32) of the csrr instruction is found, and replace the 6th vsetvl instruction with the vsetvli x0,x0,e32 instruction, and at the same time delete the 4th csrr instruction corresponding to the 6th instruction, to obtain the converted code example shown in Table 4 (Table 4 contains 12 instructions).

[0054] Table 3

[0055]

[0056]

[0057] According to an embodiment of the present invention, the preset first-stage elimination rule is: perform first-stage elimination processing on each vsetvli instruction in the intermediate code to delete invalid vsetvli instructions; wherein, the first-stage elimination processing process for each vsetvli instruction is: starting from the current vsetvli instruction, traverse downward until the next vsetvli instruction after the current vsetvli instruction is found, and analyze whether other instructions between the current vsetvli instruction and the next vsetvli instruction use specific fields in the vector configuration register; if other instructions between the current vsetvli instruction and the next vsetvli instruction do not use specific fields in the vector configuration register, then delete the current vsetvli instruction and end the processing of the current vsetvli instruction; otherwise, do not delete the current vsetvli instruction and end the processing of the current vsetvli instruction.

[0058] To better understand the preset first-stage elimination rule, the following uses the code example given in Table 4 for illustration.

[0059] For the code example shown in Table 4, perform first-stage elimination processing starting from the first vsetvli instruction.

[0060] For the first vsetvli instruction: vsetvli x0,x0,e32, there are no other instructions between the first vsetvli instruction and the next vsetvli instruction (the second instruction), so the current vsetvli instruction is not deleted.

[0061] For the second vsetvli instruction: vsetvli x0,x0,e64, since the other instructions between the second vsetvli instruction and the next vsetvli instruction (the fifth instruction) are all instructions not related to the sew field, the second vsetvli instruction can be deleted.

[0062] For the third vsetvli instruction: vsetvli x0,x0,e32, since the other instructions between the third vsetvli instruction and the next vsetvli instruction (the seventh instruction) are instructions related to the sew field, the third vsetvli instruction cannot be deleted; and so on, to obtain the code example after the first-stage elimination processing as shown in Table 5 (Table 5 contains 10 instructions).

[0063]

[0064] According to an embodiment of the present invention, the preset instruction scheduling rule is as follows: perform data flow analysis on each code block in the first-stage optimized code to determine whether there is a data flow dependency relationship between the code blocks; wherein, the data flow dependency relationship means that when executing a certain instruction in a code block, it is necessary to rely on the execution results of other instructions in other code blocks that are sorted earlier; perform instruction scheduling on any two code blocks in the first-stage optimized code that do not have a data flow dependency relationship until the instruction scheduling of each code block is completed, wherein, the instruction scheduling of any two code blocks is performed in the following manner: determine the usage status of a specific field in the vector configuration register indicated by the last vsetvli instruction in the code block sorted earlier, and the usage status of a specific field in the vector configuration register indicated by the first vsetvli instruction in the code block sorted later, and analyze whether the usage status of the specific fields in the vector configuration registers indicated by these two instructions is the same. If they are the same, then schedule the code block sorted later after the code block sorted earlier without affecting the execution results of other code blocks.

[0065] It should be noted that the reason for performing the instruction scheduling process is that in this way, without affecting the execution results of each instruction in the code block, the vsetvli instructions repeatedly set in the code can be deleted to the greatest extent.

[0066] According to an embodiment of the present invention, the preset two-stage elimination rule is: set a variable parameter identifier to mark the immediate use status of a specific field in the vector configuration register, and based on the immediate use status of the specific field in the vector configuration register marked by the variable parameter identifier, perform two-stage elimination processing on each instruction in the one-stage optimized code after instruction scheduling; wherein, the two-stage elimination processing process for each instruction is: determine whether the current instruction is a vsetvli instruction. If it is a vsetvli instruction, analyze whether the immediate use status of the specific field in the vector configuration register marked by the variable parameter identifier is consistent with the use status of the specific field in the vector configuration register corresponding to the current instruction. If they are consistent, delete the current instruction and end the processing of the current instruction; if they are inconsistent, update the variable parameter identifier to be consistent with the use status of the specific field in the vector configuration register corresponding to the current instruction and end the processing of the current instruction; if the current instruction is not a vsetvli instruction, analyze whether the current instruction is set with a label; if the current instruction is set with a label, analyze whether the use status of the specific field in the vector configuration registers corresponding to all other instructions that can jump to the current instruction is consistent. If they are consistent, update the variable parameter identifier to be consistent with the use status of the specific field in the vector configuration register corresponding to the other instructions and end the processing of the current instruction. If they are inconsistent, initialize the variable parameter identifier and end the processing of the current instruction; if the current instruction is not set with a label, end the processing of the current instruction; wherein, initializing the variable parameter identifier means setting the immediate use status of the specific field in the vector configuration register to unused.

[0067] To better understand the preset two-stage elimination rule, the following takes the code example shown in Table 5 and combines it with Figure 4 the schematic diagram of the two-stage elimination rule process shown in

[0068] As Figure 4 known, before the two-stage elimination processing, first initialize the variable parameter identifier, that is, set cur_sew = -1, indicating that the sew field in the vtype register is not used, and initialize cur_ins = hd, where cur_ins represents the currently pointed instruction, and cur_ins = hd means that the currently pointed instruction is the first instruction in the code shown in Table 5; then, start processing each instruction one by one from the first instruction in the code example according to the preset two-stage elimination rule, and obtain the code example after two-stage elimination processing as shown in Table 6 (Table 6 contains 8 instructions).

[0069] Among them, when processing the first instruction vsetvli x0,x0,e64, cur_sew = -1, cur_ins = vsetvli x0,x0,e64. It is judged whether the first instruction is a vsetvli instruction. At this time, the first instruction is a vsetvli instruction. Then, it is analyzed whether the value of cur_sew is consistent with the value of the sew field in the first instruction. At this time, cur_sew = -1, sew = e64 in the first instruction, and cur_sew ≠ sew. Since cur_sew ≠ sew, the value of cur_sew is updated to e64 and then the processing of the first instruction ends.

[0070] When processing the second instruction li s6,1, cur_sew = e64, cur_ins = li s6,1. Since the second instruction is not a vsetvli instruction and has no label, the processing of the second instruction ends.

[0071] When processing the third instruction vmv.sx v0,s6, cur_sew = e64, cur_ins = vmv.sx v0,s6. Since the third instruction is not a vsetvli instruction and has no label, the processing of the third instruction ends.

[0072] When processing the fourth instruction vsetvli x0,x0,e32, cur_sew = e64, cur_ins = vsetvli x0,x0,e32. It is judged whether the fourth instruction is a vsetvli instruction. At this time, the fourth instruction is a vsetvli instruction. Then, it is analyzed whether the value of cur_sew is consistent with the value of the sew field in the fourth instruction. At this time, cur_sew = e64, sew = e32 in the fourth instruction, and cur_sew ≠ sew. Since cur_sew ≠ sew, the value of cur_sew is updated to e32 and then the processing of the fourth instruction ends.

[0073] When processing the fifth instruction vfadd.vv v6,v6,v8,v0, cur_sew = e32, cur_ins = vfadd.vv v6,v6,v8,v0. Since the fifth instruction is not a vsetvli instruction and has no label, the processing of the fifth instruction ends.

[0074] When processing the sixth instruction vor.vv v10,v10,v12, cur_sew = e32, cur_ins = vor.vv v10,v10,v12. Since the sixth instruction is not a vsetvli instruction and has no label, the processing of the sixth instruction ends.

[0075] When processing the 7th instruction vsetvli x0,x0,e32, cur_sew = e32, cur_ins = vsetvli x0,x0,e32. It is judged whether the 7th instruction is a vsetvli instruction. At this time, the 7th instruction is a vsetvli instruction. Then, it is analyzed whether the value of cur_sew is consistent with the value of sew in the 7th instruction. At this time, cur_sew = e32, sew = e32 in the 7th instruction, and cur_sew = sew. Since cur_sew = sew, the 7th instruction is deleted.

[0076] The processing procedures of the subsequent 8th to 10th instructions are the same as those of the foregoing instructions, and will not be elaborated one by one here.

[0077]

[0078] It should be noted that the optimization method proposed by the present invention is applicable to the code obtained by any open-source dynamic binary translator. Among them, some dynamic binary translators only support translating the vector instructions of the source X86 architecture into scalar instructions of the target RISC-V architecture. At this time, after converting the scalar instructions into vector instructions, the optimization method proposed by the present invention can be applied for optimization processing.

[0079] Based on the optimization method described in the foregoing embodiments, as Figure 5 shown, the present invention also proposes an optimizer for eliminating redundant instructions generated during the translation of X86 vector instructions. The optimizer includes: a data acquisition module for acquiring the code to be optimized containing multiple instructions obtained after translation processing; a mask optimization module for deleting redundant mask register setting instructions in the code to be optimized according to preset mask optimization rules to obtain mask-optimized code; a vsetvli optimization module for deleting all csrr instructions and redundant vsetvli instructions in the mask-optimized code according to preset instruction optimization rules to obtain target optimized code.

[0080] The beneficial effects of the present invention are as follows: redundant mask register setting instructions and vtype register setting instructions generated during the translation of X86 vector instructions into RISC-V vector instructions are deleted, improving the execution efficiency.

[0081] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even the order can be changed, as long as the required functions can be achieved.

[0082] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present invention.

[0083] A computer-readable storage medium may be a tangible device that retains and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing.

[0084] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.

Claims

1. An optimization method applied to X86 vector instruction translation, which is used to eliminate redundant instructions generated during the X86 vector instruction translation process, and is characterized in that, The method includes obtaining the to-be-optimized code containing multiple instructions after translation processing, and performing the following steps: Mask optimization step: Delete the redundant mask register setting instructions in the to-be-optimized code according to the preset mask optimization rules to obtain the mask-optimized code; vsetvli optimization step: Delete all csrr instructions and redundant vsetvli instructions in the mask-optimized code according to the preset instruction optimization rules to obtain the target-optimized code.

2. The method according to claim 1, wherein The preset mask optimization rules are as follows: Replace all mask register setting instructions in the to-be-optimized code with pseudo-instructions according to the known instruction manual, and set a high-bit identifier and a low-bit identifier in the pseudo-instructions to mark the usage status of the mask register when executing the corresponding pseudo-instruction; wherein, the to-be-optimized code includes multiple sequentially arranged code blocks, and some code blocks contain instructions with labels, and the instructions with labels indicate that the instruction is a jump execution instruction corresponding to other instructions; Set a low-bit variable identifier and a high-bit variable identifier to mark the immediate usage status of the mask register, and analyze each instruction in the to-be-optimized code one by one according to the preset pseudo-instruction elimination rules based on the immediate usage status of the mask register to delete redundant pseudo-instructions; Select the instruction sequence corresponding to the pseudo-instruction from the preset instruction sequence list to replace each remaining pseudo-instruction in the to-be-optimized code to obtain the mask-optimized code.

3. The method according to claim 2, wherein The preset pseudo-instruction elimination rules are to analyze each instruction in the to-be-optimized code one by one in the following manner: Judge whether the current instruction is a pseudo-instruction. If it is a pseudo-instruction, analyze whether the low-bit variable identifier and the high-bit variable identifier of the immediate usage status of the mask register are the same as the high-bit identifier and the low-bit identifier of the pseudo-instruction. If they are the same, delete the pseudo-instruction. If they are not the same, update the low-bit variable identifier and the high-bit variable identifier of the immediate usage status of the mask register to be consistent with the high-bit identifier and the low-bit identifier of the pseudo-instruction, and end the analysis of the current instruction; If the current instruction is not a pseudo-instruction, analyze whether the current instruction has a label. If the current instruction has a label, analyze whether the usage statuses of the mask registers corresponding to all other instructions that can jump to the current instruction are the same. If they are the same, update the low-bit variable identifier and the high-bit variable identifier of the immediate usage status of the mask register to be consistent with the usage status of the mask register corresponding to the other instruction. If they are not the same, initialize the low-bit variable identifier and the high-bit variable identifier and end the analysis of the current instruction; if the current instruction does not have a label, end the analysis of the current instruction; wherein, initializing the low-bit variable identifier and the high-bit variable identifier means setting the immediate usage status of the mask register to unused.

4. The method according to claim 3, wherein The preset instruction optimization rules are as follows: Delete all csrr instructions and redundant vsetvl instructions in the mask-optimized code according to the preset conversion rules, and replace the remaining vsetvl instructions in the mask-optimized code with vsetvli instructions to obtain the intermediate-optimized code; Delete the invalid vsetvli instructions in the intermediate-optimized code according to the preset one-stage elimination rules to obtain the one-stage optimized code; Schedule the instructions of the first-stage optimized code according to the preset instruction scheduling rules to adjust the arrangement order of each code block in the first-stage optimized code; And delete the duplicate vsetvli instructions in the scheduled first-stage optimized code according to the preset second-stage elimination rules to obtain the target optimized code.

5. The method according to claim 4, wherein The preset conversion rules are as follows: Determine each vsetvl instruction in the mask optimized code, and each crss instruction corresponding to each vsetvl; Analyze each vsetvl instruction in the mask optimized code; among them, the analysis process of each vsetvl instruction is as follows: Starting from the current vsetvl instruction, traverse downward until the next vsetvl instruction or vsetvli instruction after the current vsetvl instruction is found, and analyze whether the other instructions between the current vsetvl instruction and the next vsetvl instruction or vsetvli instruction use a specific field in the vector configuration register; If not used, delete the current vsetvl instruction and the corresponding csrr instruction, and end the analysis of the current vsetvl instruction; If it has been used, traverse upward starting from the csrr instruction corresponding to the current vsetvl instruction until the previous vsetvli instruction of the csrr instruction is found, and replace the current vsetvl instruction with the previous vsetvli instruction, and at the same time delete the csrr instruction corresponding to the current vsetvl instruction and end the analysis of the current vsetvl instruction; Among them, the vector configuration register is the vtype register, and the specific field in the vector configuration register is the sew field in the vtype register.

6. The method according to claim 5, wherein The preset first-stage elimination rules are as follows: Perform first-stage elimination processing on each vsetvli instruction in the intermediate code to delete invalid vsetvli instructions; among them, the first-stage elimination processing process of each vsetvli instruction is as follows: Starting from the current vsetvli instruction, traverse downward until the next vsetvli instruction after the current vsetvli instruction is found, and analyze whether the other instructions between the current vsetvli instruction and the next vsetvli instruction use a specific field in the vector configuration register; If the other instructions between the current vsetvli instruction and the next vsetvli instruction do not use the specific field in the vector configuration register, delete the current vsetvli instruction and end the processing of the current vsetvli instruction; otherwise, do not delete the current vsetvli instruction and end the processing of the current vsetvli instruction.

7. The method according to claim 6, wherein The preset instruction scheduling rules are as follows: Perform data flow analysis on each code block in the first-stage optimized code to determine whether there is a data flow dependency relationship between each code block; among them, the data flow dependency relationship means that when executing an instruction in a code block, it is necessary to rely on the execution results of other instructions in other code blocks sorted earlier; For any two code blocks without data flow dependencies in the first-stage optimized code, perform instruction scheduling until the instruction scheduling for each code block is completed. Among them, the instruction scheduling for any two code blocks is performed in the following manner: Determine the usage status of specific fields in the vector configuration register indicated by the last vsetvli instruction in the code block with a higher sort order, and the usage status of specific fields in the vector configuration register indicated by the first vsetvli instruction in the code block with a lower sort order, and analyze whether the usage status of specific fields in the vector configuration registers indicated by these two instructions is the same. If they are the same, schedule the code block with a lower sort order after the code block with a higher sort order without affecting the execution results of other code blocks.

8. The method according to claim 7, wherein The preset second-stage elimination rule is as follows: Set a variable parameter flag to mark the immediate usage status of specific fields in the vector configuration register, and based on the immediate usage status of specific fields in the vector configuration register marked by the variable parameter flag, perform second-stage elimination processing on each instruction in the first-stage optimized code after instruction scheduling; among them, the second-stage elimination processing process for each instruction is as follows: Judge whether the current instruction is a vsetvli instruction. If it is a vsetvli instruction, analyze whether the immediate usage status of specific fields in the vector configuration register marked by the variable parameter flag is consistent with the usage status of specific fields in the vector configuration register corresponding to the current instruction. If they are consistent, delete the current instruction and end the processing of the current instruction; if they are inconsistent, update the variable parameter flag to be consistent with the usage status of specific fields in the vector configuration register corresponding to the current instruction and end the processing of the current instruction. If the current instruction is not a vsetvli instruction, analyze whether the current instruction is set with a label; if the current instruction is set with a label, analyze whether the usage status of specific fields in the vector configuration registers corresponding to all other instructions that can jump to the current instruction is consistent. If they are consistent, update the variable parameter flag to be consistent with the usage status of specific fields in the vector configuration registers corresponding to other instructions and end the processing of the current instruction. If they are inconsistent, initialize the variable parameter flag and end the processing of the current instruction; if the current instruction is not set with a label, end the processing of the current instruction. Among them, initializing the variable parameter flag means setting the immediate usage status of specific fields in the vector configuration register to unused.

9. An optimizer based on the method according to any one of claims 1-8, for eliminating redundant instructions generated during the translation of X86 vector instructions, characterized in that, The optimizer includes: A data acquisition module for acquiring the code to be optimized containing multiple instructions obtained after translation processing; A mask optimization module for deleting redundant mask register setting instructions in the code to be optimized according to the preset mask optimization rule to obtain mask-optimized code; A vsetvli optimization module for deleting all csrr instructions and redundant vsetvli instructions in the mask-optimized code according to the preset instruction optimization rule to obtain the target optimized code.

10. A computer-readable storage medium, characterized in that, Stored thereon is a computer program, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1-8.

11. An electronic device, characterized in that, Including: One or more processors, and a memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1-8 by executing the executable instructions.

Citation Information

Patent Citations

  • Optimization method for RISC-V basic C library

    CN116860256A

  • Intermediate language support for change resilience

    US20110258615A1

  • Instruction transmitting unit, instruction execution unit, and related apparatus and method

    US20220147351A1

  • Instruction processing method and device, and storage medium and electronic device

    WO2025025991A1