Instruction optimization method and device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202511020769.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-07-23
AI Technical Summary
[0003]相关技术中,部分指令由于其操作对象或执行逻辑的特殊性,如涉及跨寄存器域的数据传输、复杂的控制状态变更等,往往导致其指令周期显著延长,执行延迟增大,进而对程序的执行效率产生不利影响,降低系统的整体性能
[0015] In this embodiment of the invention, an inefficient instruction set in the assembly instruction sequence is identified, and the target instruction function corresponding to the inefficient instruction set is obtained. The inefficient instruction set is used to perform floating-point exception judgment operations that match the target instruction function. Based on the instruction optimization rules that match the target instruction function, a target optimized instruction that is equivalent to the target instruction function and is used for floating-point exception judgment is generated. Based on the target optimized instruction, the inefficient instruction set in the assembly instruction sequence is replaced. In this way, by identifying the inefficient instruction set in the assembly instruction sequence and replacing it with a target optimized instruction that is equivalent to the target instruction function, the many internal operations or additional overhead that may exist during the execution of the inefficient instruction set are effectively avoided. While maintaining the integrity of the floating-point exception judgment function and ensuring code correctness, it can effectively reduce the execution efficiency and response speed during instruction execution, thereby improving the overall performance of the system.
Smart Images

Figure CN121008836B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an instruction optimization method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] With the continuous evolution of modern computer architecture, the improvement of processor performance has become one of the key factors driving the development of information technology. Against this backdrop, assembly language programs, as low-level programming languages that directly interact with hardware, have seen their execution efficiency become an important indicator for measuring the overall performance of a system. Different instructions, due to differences in functional complexity, hardware implementation difficulty, and microarchitectural design, may exhibit drastically different time consumption and latency characteristics during execution. These factors collectively constitute the key bottlenecks affecting the overall performance of a program.
[0003] In related technologies, some instructions, due to the special nature of their operation objects or execution logic, such as data transmission across register domains or complex control state changes, often have significantly longer instruction cycles and increased execution latency, which in turn adversely affect the execution efficiency of the program and reduce the overall performance of the system. Summary of the Invention
[0004] To overcome the problems existing in related technologies, the present invention provides an instruction optimization method, apparatus, electronic device, and readable storage medium.
[0005] In a first aspect, the present invention provides an instruction optimization method, the method comprising:
[0006] Identify inefficient instruction sets in the assembly instruction sequence and obtain the target instruction function corresponding to the inefficient instruction set; the inefficient instruction set is used to perform floating-point exception judgment operations that match the target instruction function;
[0007] Based on the instruction optimization rules that match the function of the target instruction, a target optimized instruction that is equivalent to the function of the target instruction and is used for floating-point exception detection is generated.
[0008] Based on the target optimized instructions, the inefficient instruction set in the assembly instruction sequence is replaced.
[0009] In a second aspect, the present invention provides an instruction optimization apparatus, the apparatus comprising:
[0010] The first identification module is used to identify inefficient instruction sets in the assembly instruction sequence and obtain the target instruction function corresponding to the inefficient instruction set; the inefficient instruction set is used to perform floating-point exception judgment operations that match the target instruction function.
[0011] The first generation module is used to generate a target optimized instruction that is functionally equivalent to the target instruction and is used for floating-point exception detection, based on instruction optimization rules that match the function of the target instruction.
[0012] The first substitution module is used to replace the inefficient instruction set in the assembly instruction sequence based on the target optimized instructions.
[0013] Thirdly, the present invention provides an electronic device comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the instruction optimization method described in any one of the first aspects above.
[0014] Fourthly, the present invention provides a readable storage medium that, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform steps in the instruction optimization method as described in any of the embodiments of the first aspect above.
[0015] In this embodiment of the invention, an inefficient instruction set in the assembly instruction sequence is identified, and the target instruction function corresponding to the inefficient instruction set is obtained. The inefficient instruction set is used to perform floating-point exception judgment operations that match the target instruction function. Based on the instruction optimization rules that match the target instruction function, a target optimized instruction that is equivalent to the target instruction function and is used for floating-point exception judgment is generated. Based on the target optimized instruction, the inefficient instruction set in the assembly instruction sequence is replaced. In this way, by identifying the inefficient instruction set in the assembly instruction sequence and replacing it with a target optimized instruction that is equivalent to the target instruction function, the many internal operations or additional overhead that may exist during the execution of the inefficient instruction set are effectively avoided. While maintaining the integrity of the floating-point exception judgment function and ensuring code correctness, it can effectively reduce the execution efficiency and response speed during instruction execution, thereby improving the overall performance of the system. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the steps of an instruction optimization method provided in an embodiment of the present invention;
[0018] Figure 2 This is a structural diagram of an instruction optimization device provided in an embodiment of the present invention;
[0019] Figure 3This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Figure 1 This is a flowchart illustrating the steps of an instruction optimization method provided in an embodiment of the present invention, as follows: Figure 1 As shown, the method may include:
[0022] Step 101: Identify the inefficient instruction set in the assembly instruction sequence and obtain the target instruction function corresponding to the inefficient instruction set; the inefficient instruction set is used to perform floating-point exception judgment operation matching the target instruction function.
[0023] In this embodiment of the invention, the instruction set, as the interface between computer hardware and software, directly determines the ease and efficiency of instruction execution due to its design rationality. Different instructions, due to differences in functional complexity and hardware implementation, will result in varying execution times and latency. Some instructions, involving complex operational logic, data transmission paths, or requiring additional hardware resources, often lead to longer execution times and greater latency. This latency not only reduces the execution speed of individual instructions but also accumulates during program execution, thus affecting the overall performance of the program. Therefore, inefficient instructions with long execution times and significant latency in the assembly instruction sequence can be identified, optimized, and their equivalent functions indirectly achieved using a set of simpler instructions.
[0024] Identify inefficient instruction sets within assembly instruction sequences. These inefficient instruction sets can be used for floating-point exception handling operations that match the functionality of the target instruction. During instruction-based floating-point exception handling, deep interaction between the General-Purpose Register (GR) and the Floating-Point Control and Status Register (FCSR) is unavoidable. This includes, but is not limited to, complex operations such as data transfer between registers, updating status flags, and potential exception handling mechanisms, often leading to extended execution cycles and increased instruction latency. Therefore, these inefficient instruction sets can be identified and optimized. Specifically, inefficient instruction sets within an assembly instruction sequence can be identified based on characteristics such as instruction execution time, instruction name, and instruction type.
[0025] Floating-point arithmetic is performed by the hardware floating-point unit (FPU). The hardware monitors for potential exceptions in real time based on the mathematical rules and logic of floating-point arithmetic to determine if a floating-point exception has occurred during the operation. A "floating-point exception" (or floating-point error) refers to an error that occurs during processor execution when the value of a floating-point number is outside the expected range. This includes, but is not limited to, overflow (i.e., the calculation result exceeds the floating-point representation range), underflow, and division by zero. To enhance program robustness, specific instructions are used in program design to perform floating-point exception checks on the results of floating-point instruction operations. Floating-point exceptions can be triggered by operations within floating-point instructions or vector instructions: floating-point instructions that directly involve floating-point operations, such as addition, subtraction, multiplication, division, and square root operations, may trigger floating-point exceptions.
[0026] However, not all floating-point operations will produce floating-point exceptions. Certain floating-point instructions, however, may trigger exceptions during execution due to various factors such as division by zero, overflow, and invalid operations, thus affecting the correct execution and stability of the program. For example, floating-point arithmetic and conversion instructions may encounter inaccuracy exceptions, overflow exceptions, and underflow exceptions. Therefore, inaccuracy, overflow, and underflow detection can be performed using exception detection instructions. Similarly, floating-point division instructions may encounter division by zero exceptions, so division by zero detection can be performed using exception detection instructions. Likewise, floating-point comparison instructions may encounter illegal operation exceptions, so illegal operation detection can be performed using exception detection instructions.
[0027] After identifying inefficient instruction sets in the assembly instruction sequence, the target instruction function corresponding to the inefficient instruction set is obtained. The target instruction function can include inaccurate detection, underflow detection, overflow detection, division by zero detection, and illegal operation detection. Specifically, based on the instruction semantics or instruction characteristics corresponding to the inefficient instruction set, the floating-point exception judgment type executed by the inefficient instruction set can be determined as the target instruction function. It is understood that due to the different instruction characteristics of different architectures, the method for determining the target instruction function of the inefficient instruction set can vary, depending on the instruction set architecture. For example, in the case of a LoongArch instruction set architecture, the target instruction function corresponding to the inefficient instruction set can be determined based on the bstrpick instruction in the inefficient instruction set. The bstrpick instruction is used to retrieve a specified number of bits from the FCSR register, and based on the value of the specified number of bits, it determines whether a floating-point exception has occurred and the type of floating-point exception. According to the instruction manual, bit 28 of the FCSR register indicates an illegal operation, bit 27 indicates division by zero, bit 26 indicates overflow, bit 25 indicates underflow, and bit 24 indicates inaccuracy. For example, the bstrpick instruction fetches bits 26-25 from the FCSR register, indicating overflow and underflow detections, respectively. Therefore, by retrieving bits from the FCSR register using the bstrpick instruction in an inefficient instruction set, the type of floating-point exception handling performed by that inefficient instruction set can be determined, and the function of the target instruction can be identified.
[0028] It's worth noting that this embodiment only uses LoongArch as an example. Mainstream instruction set architectures (x86, ARM, RISC-V) all provide methods for reading floating-point exception status. Although the names and methods differ slightly, their functionality is consistent with LoongArch's bstrpick+FCSR combination; they all determine the exception type by reading the status register and checking specific bits. This type of functionality is a common mechanism in floating-point instruction sets. For example, x86 / x64 instruction sets use the FSTSW instruction in conjunction with a mask to obtain the exception bit in the floating-point status register; these will not be elaborated upon here.
[0029] Step 102: Based on the instruction optimization rules that match the function of the target instruction, generate a target optimized instruction that is equivalent to the function of the target instruction and is used for floating-point exception detection.
[0030] In this embodiment of the invention, for different target instruction functions, corresponding instruction optimization rules are pre-defined. The principle of defining the instruction optimization rules is to use a set of simple instructions that can achieve the same semantics as the inefficient instructions in the inefficient instruction set to replace the inefficient instruction set. The instruction optimization rules are used to instruct the generation of a set of simple instructions equivalent to the inefficient instruction set. Here, simple instructions can refer to instructions with relatively basic and simple functions and relatively straightforward operation logic, mainly performing basic data processing such as logical operations, data loading, data shifting, and conditional judgments. Simple instructions can include load instructions, store instructions, arithmetic instructions, logical instructions, jump instructions, conditional instructions, shift instructions, and input / output instructions, etc. Based on the logical semantics of different instruction functions, instruction optimization rules are pre-defined for each instruction function. The instruction optimization rules can include a first optimization rule corresponding to the division-by-zero detection function or illegal operation detection function, a second optimization rule corresponding to the overflow detection function or underflow detection function, and a third optimization rule corresponding to the inaccurate detection function.
[0031] After determining the target instruction function, a target optimized instruction can be generated based on instruction optimization rules that match the target instruction function. This target optimized instruction is used for floating-point exception detection. The instruction function of the target optimized instruction is equivalent to the target instruction function of the inefficient instruction set, and the execution time and latency of the target optimized instruction are both less than those of the inefficient instruction set.
[0032] Step 103: Based on the target optimized instructions, replace the inefficient instruction set in the assembly instruction sequence.
[0033] In this embodiment of the invention, after generating the target optimized instructions, the optimized instructions can first undergo functional verification to ensure that they can correctly detect the same anomalies as the inefficient instruction set. Simultaneously, a performance evaluation is performed, comparing the execution efficiency, response speed, and resource consumption before and after optimization, ensuring that the optimized instruction set achieves the goal of reducing execution efficiency and response speed while maintaining correct functionality.
[0034] The inefficient instruction set is removed from the assembly instruction sequence, and the target optimized instructions that have undergone functional verification and performance evaluation are inserted into the position of the inefficient instruction set, thus completing the replacement of the inefficient instruction set.
[0035] In summary, in this embodiment of the invention, by identifying an inefficient instruction set in the assembly instruction sequence and obtaining the target instruction function corresponding to the inefficient instruction set, the inefficient instruction set is used to perform floating-point exception judgment operations that match the target instruction function. Based on the instruction optimization rules that match the target instruction function, a target optimized instruction equivalent to the target instruction function and used for floating-point exception judgment is generated. Based on the target optimized instruction, the inefficient instruction set in the assembly instruction sequence is replaced. Thus, by identifying an inefficient instruction set in the assembly instruction sequence and replacing it with a target optimized instruction equivalent to the target instruction function, the numerous internal operations or additional overhead that may exist during the execution of the inefficient instruction set are effectively avoided. This effectively reduces the execution efficiency and response speed during instruction execution while maintaining the integrity of the floating-point exception judgment function and ensuring code correctness, thereby improving the overall performance of the system.
[0036] Optionally, step 101, "identifying the inefficient instruction set in the assembly instruction sequence," may include the following steps:
[0037] Step 201: Obtain the instruction execution time corresponding to each instruction in the assembly instruction sequence.
[0038] In this embodiment of the invention, the execution time of each instruction in the assembly instruction sequence is obtained. An instruction manual matching the instruction set architecture can be obtained to determine the instruction latency information of each instruction. The instruction latency information may include the number of execution cycles (i.e., time consumption) for each instruction under a specific processor configuration. The assembly instruction sequence is run, and performance monitoring is performed using a performance analysis tool (such as linux-perf). Specifically, the linux-perf tool can be started via command line, specifying the program to be monitored and its execution parameters. After the program finishes running, the performance data generated by the linux-perf tool is analyzed to obtain the execution time of each instruction.
[0039] Step 202: Obtain the instruction type corresponding to the instruction whose execution time is greater than a preset threshold; the instruction type includes floating-point type and potential anomaly detection type.
[0040] In this embodiment of the invention, a preset threshold is pre-set to constrain the execution time of non-inefficient instructions. Based on the preset threshold, instructions whose execution time exceeds the preset threshold are filtered out. Furthermore, based on the instruction name, the instruction type of each instruction whose execution time exceeds the preset threshold is determined. The instruction type may include memory access type, transfer type, floating-point type, and potential anomaly detection type, etc.
[0041] Different instructions have unique encoding patterns at the machine code level. By analyzing the binary encoding of assembly instructions and combining it with the known instruction encoding rules corresponding to the instruction set architecture, the instruction type can be determined more accurately. For example, the encoding of certain floating-point instructions or exception detection instructions has specific prefixes or opcode ranges. By detecting these characteristics, floating-point instructions and potential exception detection instructions can be quickly identified. Alternatively, based on the instruction naming rules corresponding to the instruction set architecture, the instruction type can be determined according to the instruction name of the floating-point instruction and the instruction name of the exception detection instruction.
[0042] For example, for floating-point instructions and potential exception detection instructions, the instruction type can be determined by matching the instruction name of the floating-point instruction and the instruction names of related exception detection instructions involving the floating-point control status register with the instruction names of instructions whose execution time exceeds a preset threshold. For instance, if the instruction name `movfcsr2gr` or `movgr2fcsr` is matched, the instruction type can be determined as a potential exception detection type. If the instruction name `fadd` or `fsub` is matched, the instruction type can be determined as a floating-point type.
[0043] Step 203: Determine the inefficient instruction set based on the relative positional relationship, data flow relationship, and control flow relationship between the instructions of each potential anomaly detection type and the instructions of each floating-point type in the assembly instruction sequence.
[0044] In this embodiment of the invention, after determining the instruction type, it is equivalent to locating floating-point instructions and potential exception detection instructions in the assembly instruction sequence. Therefore, in order to locate the set of instructions representing exception detection logic for floating-point instructions in the assembly instruction sequence, the inefficient instruction set can be divided based on the relative positional relationship, data flow relationship, and control flow relationship between instructions of each potential exception detection type and instructions of each floating-point type in the assembly instruction sequence. The relative positional relationship between potential exception detection instructions and floating-point instructions in the instruction sequence is analyzed to filter out potential exception detection instructions that may be associated with floating-point instructions; further, the data flow and control flow relationships between instructions are examined to determine whether potential exception detection instructions access registers or memory locations involved in floating-point instruction operations, and whether there are conditional jump instructions based on the results of exception detection instructions.
[0045] For example, this paper analyzes the logical relationships and relative positions of potential exception detection instructions and floating-point instructions that have been categorized. In the assembly instruction sequence, exception detection instructions are usually located near floating-point instructions and are used to detect exceptions that may occur after the execution of floating-point instructions. Therefore, by determining the relative positions of instructions in the assembly instruction sequence, potential exception detection instructions that are associated with floating-point instructions can be preliminarily screened. For example, if a movfcsr2gr instruction immediately follows an FDIV (floating-point division) instruction, then this movfcsr2gr instruction is likely an exception detection instruction used to detect whether a floating-point exception occurs after the execution of the FDIV instruction.
[0046] Further consider the data flow and control flow relationships between instructions. Analyze whether potential exception detection instructions access registers or memory locations involved in floating-point instruction operations, and whether there are conditional jump instructions that branch based on the result of the exception detection instruction. For example, if a potential exception detection instruction reads the value of the floating-point control status register, and there is a subsequent instruction that performs a conditional jump based on that register value, then this potential exception detection instruction is likely an exception detection instruction used for floating-point exception detection.
[0047] After identifying the exception detection instructions, the instruction sequence is traced backwards from the identified instructions. Based on the exception judgment logic rules defined by the instruction set architecture, a logical flow for complete exception judgment of floating-point instruction results is derived. This logical flow clarifies the complete path from the generation of floating-point operation results by floating-point instructions to the judgment of whether an exception has occurred through a series of instructions, resulting in a set of inefficient instructions containing exception detection instructions. For example, during the tracing process, if a flag bit in the floating-point status register read by a potential exception detection instruction is found to be used to judge division by zero exceptions, the floating-point division instruction that sets this flag bit, and subsequent instructions that perform conditional jumps based on this flag bit to handle division by zero exceptions, are further searched, resulting in a set of inefficient instructions.
[0048] It is understood that the method for determining the above-mentioned inefficient instruction set can be selected according to the instruction set architecture, and the embodiments of the present invention do not limit this.
[0049] In this embodiment of the invention, based on instruction execution time and instruction type, the truly time-consuming instructions in the program can be located more accurately. Furthermore, by comprehensively considering the relative positions, data flow relationships, and control flow relationships of potentially exception-detecting instructions and floating-point instructions in the assembly instruction sequence, not only can instructions with inherently long execution times be identified, but also instruction combinations that cause performance degradation due to inter-instruction interactions can be uncovered. This avoids overlooking inefficient instructions and improves the comprehensiveness of performance optimization.
[0050] Optionally, step 101, "obtaining the target instruction function corresponding to the inefficient instruction set," may include the following steps:
[0051] Step 301: Determine the target instruction function corresponding to the inefficient instruction set based on the instruction type of the target floating-point instruction corresponding to the inefficient instruction set.
[0052] In this embodiment of the invention, the target instruction function corresponding to the inefficient instruction set can be determined based on the instruction type of the target floating-point instruction corresponding to the inefficient instruction set, according to the correspondence between the instruction types of floating-point instructions and the instruction functions of exception detection instructions. The target floating-point instruction corresponding to the inefficient instruction set can be a floating-point instruction that has data flow and control flow relationships with the inefficient instruction set. The inefficient instruction set is used to perform floating-point exception judgment on the floating-point operation results of the target floating-point instruction, matching the target instruction function. Different types of floating-point exception detection functions may involve different combinations and frequencies of floating-point operation instructions. Therefore, the correspondence between the instruction types of floating-point instructions and the instruction functions of exception detection instructions can be determined based on the instruction set architecture. For example, if the target floating-point instruction is a floating-point operation instruction and a floating-point conversion instruction, the target instruction function corresponding to the inefficient instruction set can be an inaccuracy detection function, an overflow detection function, and an underflow detection function. If the target floating-point instruction is a floating-point division instruction, the target instruction function corresponding to the inefficient instruction set can be a division-by-zero detection function. If the target floating-point instruction is a floating-point comparison instruction, the target instruction function corresponding to the inefficient instruction set can be an illegal operation detection function.
[0053] In this embodiment of the invention, by taking the instruction type of the target floating-point instruction as the starting point, and based on the correspondence between the instruction type of the floating-point instruction and the instruction function of the exception detection instruction, the target instruction function implemented by the inefficient instruction set can be quickly and accurately located.
[0054] Optionally, step 102 may include the following steps:
[0055] Step 401: If the target instruction function is a division-by-zero detection function or an illegal operation detection function, determine the instruction optimization rule that matches the target instruction function as the first optimization rule.
[0056] In this embodiment of the invention, when the target instruction function is a division-by-zero detection function or an illegal operation detection function, the instruction optimization rule matching the target instruction function is determined as the first optimization rule. The first optimization rule is used to instruct the generation of simple instructions such as bit selection operation instructions, first arithmetic instructions, first logical left shift instructions, and second arithmetic instructions.
[0057] Step 402: Based on the first optimization rule, generate a bit selection operation instruction, a first operation instruction, a first logical left shift instruction, and a second operation instruction, and determine the target optimization instruction;
[0058] Wherein, the bit selection operation instruction is used to obtain a specified exponent bit of the floating-point operation result corresponding to the target floating-point instruction; the first operation instruction is used to determine whether the value of the specified exponent bit is 1, and obtain a first operation result; the first logical left shift instruction is used to obtain the valid bit of the floating-point operation result when the first operation result indicates that the value of the specified exponent bit is 1; the second operation instruction is used to determine whether the valid bit is zero, and the second operation result of the second operation instruction is used to indicate whether the target floating-point instruction has experienced a division-by-zero exception or an illegal operation exception.
[0059] In this embodiment of the invention, since the division-by-zero detection function or the illegal operation detection function needs to judge the exponent bias and the significant digits, a bit selection operation instruction, a first operation instruction, a first logical left shift instruction, and a second operation instruction can be generated based on the first optimization rule to realize the instruction function of judging the exponent bias and the significant digits.
[0060] Based on the precision and result of the floating-point operation corresponding to the target floating-point instruction, a bit selection operation instruction is generated. This bit selection operation instruction can be a `bstrpick` instruction. The bit selection operation instruction is used to extract the exponent bits corresponding to the precision of the floating-point operation result. For example, for a single-precision floating-point number, the exponent bits are located at bits 30 to 23. The `bstrpick` instruction, according to specific bit selection rules, accurately extracts the exponent bits within this range from the binary representation of the floating-point operation result. For a double-precision floating-point number, the exponent bits are located at bits 62 to 52. Similarly, the `bstrpick` instruction is used to extract the exponent bits at the corresponding positions according to the bit layout rules of double-precision floating-point numbers.
[0061] Based on the specified exponent bits extracted by the bit selection operation instruction and the first mask, a first operation instruction is generated. The first mask can have the same number of bits as the specified exponent bits, with each bit being 1. The first operation instruction can be an AND operation instruction. The first operation instruction is used to perform a bitwise AND operation on the extracted exponent bit data and a mask of all 1s to obtain a first calculation result. If the first calculation result is equal to the first mask, it is determined that the extracted specified exponent bits are all 1s; if they are not equal, it is determined that the specified exponent bits are not all 1s, indicating that no floating-point exception has occurred.
[0062] If the first calculation result is equal to the first mask, a first logical left shift instruction is generated based on the floating-point operation result and the first shift bit. This first logical left shift instruction can be an sll instruction, used to shift the floating-point operation result to the left by the first shift bit to obtain the significant bits of the floating-point operation result. The shift bit must ensure that only significant bits remain in the floating-point operation result after the left shift. The specific first shift bit can be determined according to the normalized form of the floating-point number, so that in the left-shifted result, the significant bits of the original floating-point number are near the least significant bit, while other insignificant bits (such as redundant parts of the exponent and sign bits in certain cases) are shifted out.
[0063] Based on the significant bit and 0, a second arithmetic instruction is generated. This second arithmetic instruction can be an AND operation instruction. The second arithmetic instruction performs a bitwise AND operation between the significant bit and 0 to obtain the second operation result. The second operation result indicates whether the target floating-point instruction has encountered a division-by-zero exception or an illegal operation exception. If the second operation result is 0, then the result after left shifting is determined to be 0, meaning that if all exponent bits are 1, the significant bit is 0. If the second operation result is not 0, then the result after left shifting is determined to be not 0, and the significant bit is not 0. A significant bit of 0 indicates a division-by-zero exception has occurred; a significant bit of non-zero indicates an illegal operation exception has occurred.
[0064] Understandably, after generating the bit selection operation instruction, the first operation instruction, the first logical left shift instruction, and the second operation instruction, it is necessary to fully consider the overall program operation requirements and data interaction characteristics. Based on the instruction set architecture, auxiliary instructions such as storage instructions and load instructions should be generated in a targeted manner. The generated logical operation instructions, storage instructions, and load instructions should be reasonably combined and coordinated to construct a complete optimized instruction system that is functionally equivalent to the inefficient instruction set, thereby obtaining the target optimized instruction.
[0065] In this embodiment of the invention, by generating bit selection operation instructions, a first operation instruction, a first logical left shift instruction, and a second operation instruction, a division-by-zero detection function or an illegal operation function can be implemented through a set of simple instructions, thereby improving the instruction execution speed of the floating-point exception judgment process while ensuring code correctness.
[0066] Optionally, step 102 may include the following steps:
[0067] Step 501: If the target instruction function is an overflow detection function or an underflow detection function, determine the instruction optimization rule that matches the target instruction function as the second optimization rule.
[0068] In this embodiment of the invention, when the target instruction function is an overflow detection function or an underflow detection function, the instruction optimization rule matching the target instruction function is determined as the second optimization rule. The second optimization rule is used to instruct the generation of a second logical left shift instruction, an XOR operation instruction, a circular right shift instruction, and a conditional judgment instruction.
[0069] Step 502: Based on the second optimization rule, generate a second logical left shift instruction, an XOR operation instruction, a circular right shift instruction, and a condition judgment instruction, and determine the target optimization instruction;
[0070] The second logical left shift instruction is used to logically left shift the first immediate value to obtain the second mask; the XOR operation instruction is used to XOR the second mask and the floating-point operation result corresponding to the target floating-point instruction to obtain the third operation result; the circular right shift instruction is used to circularly right shift the third operation result to obtain the fourth operation result; the condition judgment instruction is used to compare the third operation result and the fourth operation result to obtain the exception judgment result; the exception judgment result is used to characterize whether the target floating-point instruction has an overflow exception or an underflow exception.
[0071] In this embodiment of the invention, since the overflow detection function or the underflow detection function needs to truncate the floating-point number to generate an integer for judgment, a second logical left shift instruction, an XOR operation instruction, a circular right shift instruction, and a condition judgment instruction can be generated based on the second optimization rule to realize the instruction function of truncating the floating-point number to generate an integer and making a judgment.
[0072] Based on the first immediate value and the second shift bit, a second logical left shift instruction is generated. This second logical left shift instruction logically shifts the first immediate value to the left to obtain a second mask. Specifically, the second logical left shift instruction can be a `slli.d` instruction, the first immediate value can be 1, and the second logical left shift instruction shifts the first immediate value to the left by the second shift bit. The specific second shift bit can be determined according to the normalized form of the floating-point number. For example, for a double-precision floating-point number, the first immediate value can be 1, and the second left shift bit can be 63 bits, resulting in a left-shifted value of 0x80000000000000000. This result serves as the second mask for subsequent XOR operations.
[0073] Based on the second mask and the floating-point operation result of the target floating-point instruction, an XOR operation instruction is generated. This XOR operation instruction performs a bitwise XOR operation between the generated second mask and the floating-point operation result to obtain the third operation result. The XOR operation instruction can be an XOR instruction. In floating-point arithmetic, for double-precision floating-point numbers, when an overflow occurs, the most significant bit and its related bits of the output floating-point operation result will exhibit a specific pattern. By performing a bitwise XOR operation with the second mask, for floating-point operation results that have overflowed, the XOR result will become all 0s or all 1s. This is because when overflow occurs, the most significant bit and its related bits of the floating-point operation result will exhibit a specific numerical pattern, and XORing with the second mask will result in all bits being the same (all 0s or all 1s).
[0074] Based on the result of the third operation and the third shift bit, a circular right shift instruction is generated. This instruction shifts the result of the third operation to the right by the third shift bit to obtain the fourth operation result. The circular right shift instruction can be a `rotri` instruction. The circular right shift operation cyclically shifts each bit of the result of the third operation obtained from the XOR operation. The third shift bit can be set according to actual needs, for example, shifting by 1 bit. If the result after the XOR operation is all 0s or all 1s, the result will remain unchanged during the circular right shift, regardless of the shift bit.
[0075] Based on the results of the third and fourth operations, a conditional judgment instruction is generated. This instruction can be a `bne` instruction. It compares the results before and after the circular right shift to determine the exception. The exception result indicates whether the target floating-point instruction has experienced an overflow or underflow exception. If the results before and after the circular right shift are equal (i.e., the third and fourth results are equal), then a floating-point overflow has occurred; if they are not equal (i.e., the third and fourth results are not equal), then no overflow has occurred.
[0076] Understandably, after generating the second logical left shift instruction, XOR operation instruction, circular right shift instruction, and condition judgment instruction, it is necessary to fully consider the overall program operation requirements and data interaction characteristics. Based on the instruction set architecture, auxiliary instructions such as store instructions and load instructions should be generated in a targeted manner. The generated logical operation instructions, condition judgment instructions, store instructions, and load instructions should be reasonably combined and coordinated to construct a complete optimized instruction system that is functionally equivalent to the inefficient instruction set, thereby obtaining the target optimized instruction.
[0077] In this embodiment of the invention, by generating a second logical left shift instruction, an XOR operation instruction, a circular right shift instruction, and a condition judgment instruction, an overflow detection function can be achieved through a set of simple instructions, thereby improving the instruction execution speed of the floating-point exception judgment process while ensuring code correctness.
[0078] Optionally, step 102 may include the following steps:
[0079] Step 601: If the target instruction function is a non-precise detection function and the target floating-point instruction is a floating-point conversion instruction, determine the instruction optimization rule that matches the target instruction function as the third optimization rule; the target floating-point instruction is used to perform a target conversion type conversion operation on the original data.
[0080] In this embodiment of the invention, since floating-point numbers are stored internally in binary scientific notation, their representation range and precision are limited by the mantissa width and exponent range. This results in most floating-point numbers being unable to accurately represent decimal fractions (e.g., 0.1 is an infinite repeating decimal in binary). Therefore, most inaccuracies (such as small rounding errors in floating-point operations) are inherent characteristics of floating-point numbers and usually do not require additional detection. However, in floating-point conversion operations (such as FCVT precision conversion, FFINT integer to floating-point, FTINT floating-point to integer, FRINT floating-point rounding), if the conversion process triggers explicit rounding due to the precision limitations or insufficient value range of the target format (e.g., losing significant bits when converting from double precision to single precision, or the floating-point number exceeding the integer representation range), detectable inaccuracies may occur. In this case, reverse conversion verification (e.g., converting the converted result back to the original format and comparing it with the original value) is needed to identify whether inaccurate rounding has occurred. Such detection is typically only applied in scenarios requiring strict numerical consistency (e.g., financial calculations or high-precision simulations). Therefore, when the target instruction function is a non-precise detection function and the target floating-point instruction is a floating-point conversion instruction, the instruction optimization rule matching the target instruction function is determined to be the third optimization rule. The third optimization rule is used to indicate the generation of data conversion instructions and comparison instructions. The target floating-point instruction is used to perform target conversion type conversion operations on the original data to obtain the floating-point operation result. The target conversion type can include double-precision to single-precision, single-precision to double-precision, floating-point to integer, integer to floating-point, and floating-point to integer form of floating-point, etc.
[0081] Step 502: If the target conversion type is the first conversion type, generate a data conversion instruction and a comparison instruction based on the third optimization rule, and determine the target optimization instruction;
[0082] The data conversion instruction is used to perform an inverse conversion operation on the floating-point operation result corresponding to the target floating-point instruction, which is the opposite of the target conversion type, to obtain a conversion result; the comparison instruction is used to compare the original data with the conversion result to obtain a comparison result, which is used to characterize whether the target floating-point instruction has experienced a precision anomaly.
[0083] In this embodiment of the invention, when the target conversion type is a first conversion type, a data conversion instruction is generated based on the original data corresponding to the target floating-point instruction and the floating-point operation result corresponding to the target floating-point instruction. A first comparison instruction is then generated based on the conversion result generated by the data conversion instruction and the original data. The first conversion type may include double-precision to single-precision, single-precision to double-precision, floating-point to integer, and integer to floating-point. The first comparison instruction may be an fcmp instruction. The data conversion instruction is used to perform a conversion operation of the inverse conversion type (opposite to the target conversion type) on the floating-point operation result corresponding to the target floating-point instruction to obtain the conversion result. The inverse conversion type (opposite to the target conversion type) refers to converting the conversion result of the target conversion type's conversion operation back to the original format or form. For example, when the target conversion type is double-precision to single-precision, the inverse conversion type is single-precision to double-precision; when the target conversion type is single-precision to double-precision, the inverse conversion type is double-precision to single-precision; when the target conversion type is floating-point to integer, the inverse conversion type is integer to floating-point.
[0084] The data conversion instruction performs the inverse conversion operation corresponding to the type of the floating-point operation on the result, obtaining the conversion result. The first comparison instruction compares the original data corresponding to the target floating-point instruction with the conversion result, obtaining a comparison result. The comparison result is used to characterize whether the target floating-point instruction has experienced a precision anomaly. If the comparison result shows that the conversion result is not equal to the original data, it is determined that the target floating-point instruction has experienced a precision anomaly. This is because in normal precise floating-point operations, if no inaccuracy anomaly occurs, the conversion result should remain equal to the original data while conforming to the operation rules; a difference indicates that inaccuracy, such as precision loss, has occurred during the operation.
[0085] Step 503: If the target conversion type is the second conversion type, generate a second comparison instruction based on the third optimization rule, and determine the target optimization instruction.
[0086] In this embodiment of the invention, when the target conversion type is the second conversion type, a second comparison instruction can be generated based on the third optimization rule. The second comparison instruction is generated based on the original data and the floating-point operation result. This second comparison instruction can be an `fcmp` instruction. The second comparison instruction is used to compare the original data corresponding to the target floating-point instruction and the floating-point operation result corresponding to the target floating-point instruction to obtain a comparison result. The comparison result is used to characterize whether the target floating-point instruction has experienced a precision anomaly. If the comparison result shows that the floating-point operation result is not equal to the original data, then it is determined that the target floating-point instruction has experienced a precision anomaly. This is because in normal precise floating-point operations, if no inaccuracy anomaly occurs, the operation result should remain equal to the original data while conforming to the operation rules; a discrepancy indicates that inaccuracy, such as precision loss, has occurred during the operation.
[0087] Understandably, after generating data conversion and comparison instructions, it is necessary to fully consider the overall program operation requirements and data interaction characteristics. Based on the instruction set architecture, auxiliary instructions such as storage and loading instructions should be generated in a targeted manner. The generated logical operation instructions, storage instructions, and loading instructions should be reasonably combined and coordinated to build a complete optimized instruction system that is functionally equivalent to the inefficient instruction set, thereby obtaining the target optimized instructions.
[0088] In this embodiment of the invention, by generating comparison instructions, non-precise detection functions can be achieved through simple instructions, thereby improving the instruction execution speed of the floating-point exception judgment process while ensuring code correctness.
[0089] For example, taking the LoongArch architecture as an example, this embodiment of the invention provides an instruction optimization method, assuming that the set of inefficient instructions in the assembly instruction sequence is as follows:
[0090] main:
[0091] la$al,input
[0092] fld.d$f1,$al,0 / / Load floating-point numbers into f1
[0093] movfcsr2gr$a2,$fcsr
[0094] movgr2fcsr$fcsr0,$zero / / Save the current floating-point control status register fcsr0 and then reset it to zero
[0095] ftintrz.wd$f2,$f1
[0096] movfr2gr.d$al,$f2 / / Performs floating-point truncation and stores the result in a1
[0097] movfcsr2gr$a3,$fcsr0
[0098] movgr2fcsr$fcsr0,$a2 / / Restore the contents of fcsr0
[0099] bstrpick.w$a3,$a3,26,25
[0100] beq$a3,$zero,.done / / Retrieve bits 26-25 of fcsr0 and check for overflow and underflow. If not, jump to done.
[0101] done:
[0102] The target instruction functions of this inefficient instruction set are overflow detection and underflow detection. The inefficient instruction set is used to determine whether the floating-point operation result of the target floating-point instruction has an overflow or underflow exception. Based on the second optimization rule, second logical left shift instructions, XOR operation instructions, circular right shift instructions, and conditional judgment instructions can be generated, and the target optimized instructions can be determined. The target optimized instructions can be as follows:
[0103] main:
[0104] la$al,input
[0105] fld.d$f1,$al,0 / / Load floating-point numbers into f1
[0106] ftintrz.ld$f2,$f1 / / This function truncates the double-precision floating-point number in $f1 into integers and stores the result in $f2 as a double-precision floating-point number.
[0107] movfr2gr.d$al,$f2 / / 2. Convert the double-precision floating-point number in the floating-point register $f2 to an integer and move the result to the general-purpose register $al.
[0108] li.d$a2,1 / / Load the immediate value 1 into the general-purpose register $a2
[0109] slli.d$a2,$a2,63 / / Second logical left shift instruction: Shifts the number in general-purpose register $a2 (i.e., the immediate value 1) left by 63 bits.
[0110] `xor $a2, $a2, $a1` / / XOR operation instruction: performs an XOR operation between the numbers in `$a2` and the result `$a1`.
[0111] rotri.d$al,$a2,1 / / Circular right shift instruction: Shifts the data in general-purpose register $a2 one bit to the right and stores the result in general-purpose register $al.
[0112] `movfr2gr.s$a3,$f2` / / Converts the single-precision floating-point number in the floating-point register $f2 to an integer and moves the result to the general-purpose register $a3, storing the lower 32 bits of the truncated floating-point result in a3.
[0113] bne$al,$a2,.done / / Conditional instruction: Checks if the values in general-purpose registers $al and $a2 are not equal; if not, jumps to completion.
[0114] done:
[0115] The optimized target instruction can also achieve overflow and underflow detection functions equivalent to the aforementioned inefficient instruction set. Furthermore, since it no longer involves the floating-point control status register, it avoids the use of inefficient instructions such as movgr2fcsr and movfcsr2gr. Through efficient and simple instruction combinations, it improves instruction execution efficiency and response speed, reduces performance loss caused by inefficient instructions, and provides better support for high-performance computing and embedded systems.
[0116] Figure 2 This is a schematic diagram of the structure of an instruction optimization device provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the device may specifically include:
[0117] The first identification module 701 is used to identify an inefficient instruction set in the assembly instruction sequence and obtain the target instruction function corresponding to the inefficient instruction set; the inefficient instruction set is used to perform a floating-point exception judgment operation that matches the target instruction function.
[0118] The first generation module 702 is used to generate a target optimized instruction that is functionally equivalent to the target instruction and is used for floating-point exception judgment, based on the instruction optimization rules that match the function of the target instruction.
[0119] The first substitution module 703 is used to replace the inefficient instruction set in the assembly instruction sequence based on the target optimized instruction.
[0120] Optionally, the first identification module 701 includes:
[0121] The first acquisition module is used to acquire the instruction execution time corresponding to each instruction in the assembly instruction sequence;
[0122] The first acquisition module is used to acquire the instruction type corresponding to the instruction whose execution time is greater than a preset threshold; the instruction type includes floating-point type and potential anomaly detection type.
[0123] The first determining module is used to determine the inefficient instruction set based on the relative positional relationship, data flow relationship, and control flow relationship between the instructions of each potential anomaly detection type and the instructions of each floating-point type in the assembly instruction sequence.
[0124] Optionally, the first identification module 701 includes:
[0125] The second determining module is used to determine the target instruction function corresponding to the inefficient instruction set based on the instruction type of the target floating-point instruction corresponding to the inefficient instruction set.
[0126] Optionally, the first generation module 702 includes:
[0127] The third determining module is used to determine the instruction optimization rule that matches the target instruction function as the first optimization rule when the target instruction function is a division-by-zero detection function or an illegal operation detection function.
[0128] The first generation submodule is used to generate a bit selection operation instruction, a first operation instruction, a first logical left shift instruction, and a second operation instruction based on the first optimization rule, and to determine the target optimization instruction;
[0129] Wherein, the bit selection operation instruction is used to obtain a specified exponent bit of the floating-point operation result corresponding to the target floating-point instruction; the first operation instruction is used to determine whether the value of the specified exponent bit is 1, and obtain a first operation result; the first logical left shift instruction is used to obtain the valid bit of the floating-point operation result when the first operation result indicates that the value of the specified exponent bit is 1; the second operation instruction is used to determine whether the valid bit is zero, and the second operation result of the second operation instruction is used to indicate whether the target floating-point instruction has experienced a division-by-zero exception or an illegal operation exception.
[0130] Optionally, the first generation module 702 includes:
[0131] The fourth determining module is used to determine, when the target instruction function is an overflow detection function or an underflow detection function, an instruction optimization rule that matches the target instruction function as the second optimization rule;
[0132] The second generation submodule is used to generate a second logical left shift instruction, an XOR operation instruction, a circular right shift instruction, and a condition judgment instruction based on the second optimization rule, and to determine the target optimization instruction;
[0133] The second logical left shift instruction is used to logically left shift the first immediate value to obtain the second mask; the XOR operation instruction is used to XOR the second mask and the floating-point operation result corresponding to the target floating-point instruction to obtain the third operation result; the circular right shift instruction is used to circularly right shift the third operation result to obtain the fourth operation result; the condition judgment instruction is used to compare the third operation result and the fourth operation result to obtain the exception judgment result; the exception judgment result is used to characterize whether the target floating-point instruction has an overflow exception or an underflow exception.
[0134] Optionally, the first generation module 702 includes:
[0135] The fifth determining module is used to determine, when the target instruction function is a non-precise detection function and the target floating-point instruction is a floating-point conversion instruction, the instruction optimization rule that matches the target instruction function as the third optimization rule; the target floating-point instruction is used to perform a target conversion type conversion operation on the original data;
[0136] The third generation submodule is used to generate a data conversion instruction and a first comparison instruction based on the third optimization rule when the target conversion type is the first conversion type, and to determine the target optimization instruction.
[0137] The fourth generation submodule is used to generate a second comparison instruction based on the third optimization rule and determine the target optimization instruction when the target conversion type is the second conversion type.
[0138] The data conversion instruction is used to perform an inverse conversion operation on the floating-point operation result corresponding to the target floating-point instruction, which is opposite to the target conversion type, to obtain a conversion result; the first comparison instruction is used to compare the original data with the conversion result to obtain a comparison result; the second comparison instruction is used to compare the original data with the floating-point operation result to obtain a comparison result; and the comparison result is used to characterize whether the target floating-point instruction has experienced a precision anomaly.
[0139] The present invention also provides an electronic device, see [link to relevant documentation]. Figure 3 It includes: a processor 801, a memory 802, and a computer program 8021 stored in the memory and executable on the processor, wherein the processor implements the instruction optimization method of the foregoing embodiments when executing the program.
[0140] The present invention also provides a readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to execute the instruction optimization method of the foregoing embodiments.
[0141] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0142] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this invention is not directed to any particular programming language. It should be understood that the contents of the invention described herein can be implemented using various programming languages, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the invention.
[0143] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0144] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.
[0145] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0146] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the sorting device according to the present invention. The present invention can also be implemented as a device or apparatus program for performing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0147] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0148] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0149] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0150] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0151] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An instruction optimization method, characterized in that, The method includes: Obtain the execution time of each instruction in the assembly instruction sequence; Obtain the instruction type corresponding to the instruction whose execution time is greater than a preset threshold; Based on the relative positional relationship, data flow relationship, and control flow relationship between instructions of each potential anomaly detection type and instructions of each floating-point type in the assembly instruction sequence, an inefficient instruction set is determined, and the target instruction function corresponding to the inefficient instruction set is obtained; the inefficient instruction set is used to perform floating-point anomaly judgment operations that match the target instruction function. Based on the instruction optimization rules that match the function of the target instruction, a target optimized instruction that is equivalent to the function of the target instruction and is used for floating-point exception detection is generated. Based on the target optimized instructions, the inefficient instruction set in the assembly instruction sequence is replaced.
2. The method according to claim 1, characterized in that, The function of obtaining the target instruction corresponding to the inefficient instruction set includes: Based on the instruction type of the target floating-point instruction corresponding to the inefficient instruction set, the target instruction function corresponding to the inefficient instruction set is determined.
3. The method according to claim 1, characterized in that, The step of generating a target optimized instruction that is functionally equivalent to the target instruction and used for floating-point exception detection, based on instruction optimization rules that match the function of the target instruction, includes: If the target instruction function is a division-by-zero detection function or an illegal operation detection function, the instruction optimization rule that matches the target instruction function is determined as the first optimization rule; Based on the first optimization rule, a bit selection operation instruction, a first operation instruction, a first logical left shift instruction, and a second operation instruction are generated, and the target optimization instruction is determined. Wherein, the bit selection operation instruction is used to obtain a specified exponent bit of the floating-point operation result corresponding to the target floating-point instruction; the first operation instruction is used to determine whether the value of the specified exponent bit is 1, and obtain a first operation result; the first logical left shift instruction is used to obtain the valid bit of the floating-point operation result when the first operation result indicates that the value of the specified exponent bit is 1; the second operation instruction is used to determine whether the valid bit is zero, and the second operation result of the second operation instruction is used to indicate whether the target floating-point instruction has experienced a division-by-zero exception or an illegal operation exception.
4. The method according to claim 1, characterized in that, The step of generating a target optimized instruction that is functionally equivalent to the target instruction and used for floating-point exception detection, based on instruction optimization rules that match the function of the target instruction, includes: If the target instruction function is an overflow detection function or an underflow detection function, the instruction optimization rule that matches the target instruction function is determined as the second optimization rule; Based on the second optimization rule, a second logical left shift instruction, an XOR operation instruction, a circular right shift instruction, and a condition judgment instruction are generated, and the target optimization instruction is determined. The second logical left shift instruction is used to logically left shift the first immediate value to obtain the second mask; the XOR operation instruction is used to XOR the second mask and the floating-point operation result corresponding to the target floating-point instruction to obtain the third operation result; the circular right shift instruction is used to circularly right shift the third operation result to obtain the fourth operation result; the condition judgment instruction is used to compare the third operation result and the fourth operation result to obtain the exception judgment result; the exception judgment result is used to characterize whether the target floating-point instruction has an overflow exception or an underflow exception.
5. The method according to claim 1, characterized in that, The step of generating a target optimized instruction that is functionally equivalent to the target instruction and used for floating-point exception detection, based on instruction optimization rules that match the function of the target instruction, includes: When the target instruction function is a non-precise detection function and the target floating-point instruction is a floating-point conversion instruction, the instruction optimization rule that matches the target instruction function is determined as the third optimization rule; the target floating-point instruction is used to perform a target conversion type conversion operation on the original data; If the target conversion type is the first conversion type, a data conversion instruction and a first comparison instruction are generated based on the third optimization rule, and the target optimization instruction is determined. If the target conversion type is the second conversion type, a second comparison instruction is generated based on the third optimization rule, and the target optimization instruction is determined. The data conversion instruction is used to perform an inverse conversion operation on the floating-point operation result corresponding to the target floating-point instruction, which is opposite to the target conversion type, to obtain a conversion result; the first comparison instruction is used to compare the original data with the conversion result to obtain a comparison result; the second comparison instruction is used to compare the original data with the floating-point operation result to obtain a comparison result; and the comparison result is used to characterize whether the target floating-point instruction has experienced a precision anomaly.
6. An instruction optimization apparatus, characterized in that, The device includes: The first identification module is used to obtain the instruction execution time corresponding to each instruction in the assembly instruction sequence; obtain the instruction type corresponding to the instruction whose execution time is greater than a preset threshold; determine the inefficient instruction set based on the relative positional relationship, data flow relationship, and control flow relationship between instructions of each potential anomaly detection type and instructions of each floating-point type in the assembly instruction sequence, and obtain the target instruction function corresponding to the inefficient instruction set; the inefficient instruction set is used to perform floating-point anomaly judgment operation matching the target instruction function; The first generation module is used to generate a target optimized instruction that is functionally equivalent to the target instruction and is used for floating-point exception detection, based on instruction optimization rules that match the function of the target instruction. The first substitution module is used to replace the inefficient instruction set in the assembly instruction sequence based on the target optimized instructions.
7. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the instruction optimization method as described in any one of claims 1-5.
8. A readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the instruction optimization method according to any one of claims 1-5.
Citation Information
Patent Citations
Fault detection in instruction translations
CN103902251A