A method for optimizing floating-point instruction conversion calculation accuracy based on extended control words

By adding virtual exception flags and establishing mapping tables in the ARM system, the precise conversion and exception handling of x86 floating-point number calculation instructions are realized, which solves the problems of floating-point number operation accuracy loss and improper exception handling in cross-instruction set execution, and improves the reliability of program cross-architecture execution.

CN119690516BActive Publication Date: 2025-05-13北京麟卓信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510208820.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

When executing across instruction sets, the differences in floating-point number operations between x86 and ARM instruction sets cannot be fully considered, resulting in poor accuracy loss or abnormal handling of floating-point number operations results.

Method used

By adding virtual exception flag bits to the control word calculated by ARM floating point number, establishing a rounding mode mapping table and an exception type mapping table, building a virtual exception instruction table, and executable file through dynamic instruction conversion loading in the ARM system, realizing the precise conversion and exception handling of x86 floating point calculation instructions.

Benefits of technology

It reduces the accuracy loss of program cross-architecture execution, improves the reliability of program cross-architecture execution, and ensures the accuracy of floating-point number calculation results and the integrity of exception handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690516B_ABST
    Figure CN119690516B_ABST
Patent Text Reader

Abstract

The invention discloses a method for optimizing calculation precision of floating-point instruction conversion based on extended control words. By extending an ARM floating-point calculation control word, pre-establishing a rounding mode mapping table, an exception type mapping table and a virtual exception instruction table, the conversion of an x86 floating-point control word loading instruction is realized based on the rounding mode mapping table and the exception type mapping table. For an x86 floating-point calculation instruction, the ARM instruction sequence to which the instruction should be converted is preliminarily determined according to whether the instruction is a high-precision calculation-related instruction. Then, according to whether the instruction exists in a virtual exception instruction table, it is determined whether to add an instruction group for simulating exception triggering to the instruction, so as to complete the conversion of the x86 floating-point calculation instruction, realize the conversion and execution of the x86 floating-point calculation-related instruction to the ARM system, reduce the precision loss of program cross-architecture execution, and improve the reliability of program cross-architecture execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer software development, and in particular relates to a floating-point instruction conversion calculation accuracy optimization method based on an extended control word. Background Art

[0002] In the field of instruction set conversion, the x86 instruction set and the ARM instruction set are widely used in different computer systems. However, they have many differences in floating-point operations, such as different rounding rules and exception handling mechanisms. These differences bring challenges to the cross-instruction set execution of programs, especially in ensuring the accuracy of the precision of floating-point operation results. Existing conversion methods often fail to fully consider the above differences, which may cause the converted instructions to suffer from precision loss or improper exception handling when performing floating-point operations. Summary of the invention

[0003] In view of this, the present invention provides a floating-point instruction conversion calculation accuracy optimization method based on extended control words, which realizes the conversion and execution of x86 floating-point calculation related instructions to ARM system.

[0004] The present invention provides a floating-point instruction conversion calculation accuracy optimization method based on an extended control word, which specifically comprises the following steps:

[0005] Step 1, adding a virtual exception flag bit to the control word of ARM floating point calculation to indicate the virtual exception related to x86 floating point calculation that does not exist in the ARM system; establishing a rounding mode mapping table and an exception type mapping table for the x86 control word and the ARM control word respectively, wherein the exception type mapping table contains the correspondence between the virtual exception flag bit and the x86 floating point calculation exception; adding the virtual exception and the virtual exception handler to the ARM exception vector table, constructing a virtual exception instruction table consisting of the correspondence between the virtual exception and the x86 floating point calculation instruction; loading and executing the executable file in the ARM system through dynamic instruction conversion;

[0006] Step 2, obtain the current instruction to be converted, if the current instruction to be converted is an x86 floating point control word load instruction, execute step 3, if the current instruction to be converted is an x86 floating point calculation instruction, execute step 4;

[0007] Step 3, determining, according to the rounding mode mapping table, that the flag bit and the exception flag bit of the x86 rounding mode correspond to the first destination flag bit and the second destination flag bit, and converting the current instruction to be converted into an ARM instruction sequence consisting of an instruction to read the ARM control word, an instruction to modify the control word to set the first destination flag bit and the second destination flag bit, and a write-back instruction according to the exception type mapping table;

[0008] Step 4: Convert the current instruction to be converted into an ARM instruction sequence with the same function;

[0009] Step 5. If the current instruction to be converted is in the virtual exception instruction table, construct an abnormal ARM instruction group that simulates the triggering of the virtual exception corresponding to the current instruction to be converted, and convert the current instruction to be converted into an ARM instruction sequence formed by the ARM instruction sequence obtained in step 4 and the abnormal ARM instruction group; otherwise, keep the ARM instruction sequence obtained in step 4 unchanged.

[0010] Furthermore, step 4 also includes: recording an ARM instruction having the same function as the current instruction to be converted as a first ARM instruction; if the current instruction to be converted is a high-precision basic calculation instruction, splitting the operand into a low-order part and a high-order part, and converting it into an ARM instruction sequence consisting of a first ARM instruction with the low-order part as an operand, a second ARM instruction that accumulates the operation result of the low-order part to the high-order part, and the first ARM instruction with the high-order part as an operand; if the current instruction to be converted is a high-precision complex calculation instruction, converting it into an ARM instruction sequence consisting of a first ARM instruction and a second ARM instruction group constructed based on Newton's iteration method.

[0011] Furthermore, the second ARM instruction group includes iterative calculations for improving calculation accuracy.

[0012] Furthermore, the processing process of the abnormal ARM instruction group in step 5 includes: obtaining the instruction or operand required to be detected according to the known abnormal triggering condition, and then judging whether it meets the condition for triggering the exception, if so, calling the virtual exception handler, otherwise executing the processing flow normally.

[0013] Furthermore, the virtual exception flag in step 1 is implemented by extending the control word FPSCR of ARM floating point calculation.

[0014] Furthermore, the processing process of the virtual exception handling program in step 1 includes: saving the current state of the ARM system, building the exception handling logic, and after the exception handling is completed, restoring the previously saved state of the ARM system and continuing to execute the program.

[0015] Furthermore, the exception handling logic is to record error information, generate prompt information or try to repair the error.

[0016] Furthermore, saving the current state of the ARM system includes: saving register values ​​and a program counter.

[0017] Furthermore, the x86 floating point control word loading instruction in step 2 is FLDCW or LDMXCSR.

[0018] Furthermore, it also includes: saving the correspondence between the ARM instruction sequence obtained in the conversion process and the current instruction to be converted in the instruction conversion cache, and querying the instruction conversion cache when performing the conversion again. If there is a corresponding conversion result, it is directly converted, otherwise execute step 2. Beneficial Effects

[0019] The present invention extends the control word of ARM floating-point number calculation, pre-establishes a rounding mode mapping table, an exception type mapping table and a virtual exception instruction table, realizes the conversion of the x86 floating-point number control word loading instruction based on the rounding mode mapping table and the exception type mapping table, preliminarily determines the ARM instruction sequence to be converted to the x86 floating-point number calculation instruction according to whether the x86 floating-point number calculation instruction is a high-precision calculation related instruction, and then determines whether to add an instruction group for simulating exception triggering according to whether the instruction exists in the virtual exception instruction table, so as to complete the conversion of the x86 floating-point number calculation instruction, realize the conversion and execution of the x86 floating-point number calculation related instruction to the ARM system, reduce the precision loss of the program cross-architecture execution, and improve the reliability of the program cross-architecture execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flow chart of a method for optimizing floating-point instruction conversion calculation accuracy based on an extended control word provided by the present invention. DETAILED DESCRIPTION

[0021] The present invention is described in detail below with reference to the accompanying drawings and with reference to the embodiments.

[0022] The invention provides a floating-point instruction conversion calculation precision optimization method based on extended control words. The core idea is: extending the control word of ARM floating-point calculation, pre-establishing a rounding mode mapping table, an exception type mapping table and a virtual exception instruction table, realizing the conversion of x86 floating-point control word loading instructions based on the rounding mode mapping table and the exception type mapping table, preliminarily determining the ARM instruction sequence to be converted into for the x86 floating-point calculation instruction according to whether it is a high-precision calculation related instruction, and then determining whether to add an instruction group for simulating exception triggering according to whether it exists in the virtual exception instruction table, so as to complete the conversion of the x86 floating-point calculation instruction.

[0023] The present invention provides a floating-point instruction conversion calculation accuracy optimization method based on extended control words, the specific process is as follows: Figure 1 As shown, the specific steps include:

[0024] Step 1, add a virtual exception flag bit in the control word of ARM floating-point calculation, the virtual exception flag bit is used to indicate the state of the virtual exception, and the virtual exception is the x86 floating-point calculation exception that does not exist in the ARM system; for the rounding mode flag bit and the exception flag bit in the x86 control word and the ARM control word related to the floating-point calculation, respectively establish a rounding mode mapping table and an exception type mapping table, wherein the exception type mapping table also includes the correspondence between the virtual exception flag bit and the x86 floating-point calculation exception; add the virtual exception and its corresponding virtual exception handler to the ARM exception vector table, and construct a virtual exception instruction table consisting of the correspondence between the virtual exception and the x86 floating-point calculation instruction.

[0025] Among them, the present invention is a virtual exception handler constructed for virtual exceptions, and the processing process may include: saving the current state of the ARM system, including register values, program counters, etc.; building exception handling logic, such as recording error information, generating prompt information, trying to repair errors, etc.; after the exception handling is completed, restoring the previously saved state of the ARM system and continuing to execute the program.

[0026] The virtual exception flag is implemented by extending the control word FPSCR of ARM floating-point calculation.

[0027] Step 2: Load and execute the executable file in the ARM system through dynamic instruction conversion.

[0028] Step 3, obtain the current instruction to be converted. If the current instruction to be converted exists in the instruction conversion cache, convert it into an ARM instruction sequence in the instruction conversion cache and then execute step 7; otherwise, when the current instruction to be converted is an x86 floating-point control word load instruction, execute step 4; when the current instruction to be converted is an x86 floating-point calculation instruction, execute step 5; when the current instruction to be converted is other instructions, convert the current instruction to be converted into an ARM instruction and then execute step 7.

[0029] Among them, floating point control word loading instructions include FLDCW, LDMXCSR, etc.

[0030] In the floating-point operation process of the x86 architecture, the x87FPU (floating-point unit) control word and the MXCSR (multimedia extension control and status register) control word are mainly involved. Among them, the x87FPU control word is a 16-bit register with a default value of 0x037F, which is used to control the floating-point operation behavior of the x87FPU, including precision, rounding mode, and exception masking. The processing instructions of the x87FPU control word include the floating-point control word load instruction FLDCW and the floating-point control word store instruction FSTCW. The MXCSR control word is a 32-bit register, which is mainly used to control the floating-point operation behavior of instruction sets such as SSE (streaming SIMD extensions) and AVX (advanced vector extensions), including rounding control, exception mask bits, and status flag bits. The processing instructions of the MXCSR control word include the floating-point control word load instruction LDMXCSR and the floating-point control word store instruction STMXCSR.

[0031] Specifically, in the process of floating-point number calculation, usually before executing the floating-point number calculation, the FLDCW instruction is first executed to load the 16-bit value in the memory into the control word register to complete the setting of the control word, so as to realize the control of the floating-point number calculation; when it is necessary to perform floating-point number calculations with different requirements, it may be necessary to temporarily change the setting of the floating-point number control word. Therefore, in order to restore to the original configuration after completing the specific calculation, it is necessary to store the current control word before changing the setting. It can be seen that in the present invention, only the floating-point number control word loading instruction needs to be specially converted to realize the conversion of x86 floating-point number calculation into ARM floating-point number calculation.

[0032] Step 4, obtain the operand of the x86 floating-point control word load instruction, parse the operand to extract the x86 rounding mode flag and exception flag, find the corresponding ARM rounding mode flag according to the rounding mode mapping table and record it as the first destination flag, find the corresponding ARM exception flag according to the exception type mapping table and record it as the second destination flag, then convert the current instruction to be converted into an instruction sequence consisting of an instruction to read the ARM control word, an instruction to modify the control word to set the first destination flag and the second destination flag, and a write-back instruction to write the modified ARM control word back to the control word register, save the conversion result in the instruction conversion cache, and execute step 7.

[0033] Step 5. If the current instruction to be converted does not belong to a high-precision calculation instruction, the current instruction to be converted is converted into an ARM instruction sequence with the same function, and the conversion result is saved in the instruction conversion cache before executing step 6; otherwise, if it is a high-precision basic calculation instruction, the operand is split into a low-order part and a high-order part, and the ARM instruction with the same function as the current instruction to be converted is recorded as the first ARM instruction, and the current instruction to be converted is converted into an ARM instruction sequence consisting of a first ARM instruction with the low-order part as an operand, a second ARM instruction that accumulates the operation result of the low-order part to the high-order part, and the first ARM instruction with the high-order part as an operand, and the conversion result is saved in the instruction conversion cache before executing step 6; if it is a high-precision complex calculation instruction, the current instruction to be converted is converted into an ARM instruction sequence formed by a first ARM instruction and a second ARM instruction group constructed based on the Newton iteration method, and the conversion result is saved in the instruction conversion cache before executing step 6.

[0034] Among them, the second ARM instruction group constructed based on Newton's iteration method includes iterative calculation for improving calculation accuracy.

[0035] Step 6. If the current instruction to be converted is not in the virtual exception instruction table, execute step 7; otherwise, record the ARM instruction sequence that can simulate the virtual exception corresponding to the current instruction to be converted as the third ARM instruction group, convert the current instruction to be converted into an ARM instruction sequence formed by the ARM instruction sequence obtained in step 5 and the third ARM instruction group, save the conversion result in the instruction conversion cache and execute step 7.

[0036] Among them, the third ARM instruction group realizes the simulation of exception triggering by constructing an instruction sequence. Specifically, the instructions or operands required to be detected are obtained according to the known exception triggering conditions, and then it is determined whether they meet the conditions for triggering the exception. If they meet the conditions, the virtual exception handling program constructed by the present invention is called, otherwise the processing flow is executed normally.

[0037] Taking the illegal opcode exception as an example, the first step is to perform an instruction legitimacy check: in the decoding stage, determine whether the currently parsed x86 instruction is a legal instruction, which can be determined by comparing it with the maintained legal instruction table; if the instruction is not in the legal instruction table, it is determined to be an illegal opcode and triggers the corresponding virtual exception handler.

[0038] Step 7: If the executable file is executed and the conversion is completed, then the process ends; otherwise, proceed to step 3. Example

[0039] In this embodiment, a floating-point instruction conversion calculation accuracy optimization method based on an extended control word provided by the present invention is adopted to realize efficient execution of an x86 architecture executable file on an ARM system, including the following steps:

[0040] S1. Instruction conversion engine instruction parsing and information extraction.

[0041] Receive and use the instruction decoder to parse x86 floating-point calculation instructions, obtain operands, opcodes and control information, store and read FPU control words through instructions such as FSTCW, and obtain rounding mode, precision control and exception masking information. The specific steps are as follows:

[0042] S1.1. The instruction conversion engine receives x86 floating-point calculation instructions, such as floating-point addition instruction FADD, floating-point multiplication instruction FMUL, floating-point division instruction FDIV, etc. The instruction decoder is used to parse the input x86 instructions to obtain instruction operands, operation codes, and control information that may affect rounding and exception handling.

[0043] For the FADDst(0),st(1) instruction, the instruction decoder is used to decompose it into the opcode FADD, operands st(0) and st(1), where st(0) and st(1) are the top and next top elements of the x86 floating-point register stack.

[0044] For the FMULst(0),st(1) instruction, the opcode FMUL and operands st(0) and st(1) are also parsed to make it clear that the operands are elements from the floating-point register stack.

[0045] For the FLD[mem_addr] instruction, the opcode FLD and the memory address [mem_addr] are parsed, indicating that a floating-point number is loaded from the memory address mem_addr to the top of the floating-point register stack.

[0046] S1.2, for control information in the instruction, such as the FPU control word, the FPU control word is stored in the memory by executing the FSTCW instruction, and then the storage location is read to obtain the current rounding mode, precision control and exception masking information.

[0047] Execute the FSTCW[control_word_addr] instruction to store the FPU control word to the memory address control_word_addr.

[0048] The stored control word is then loaded into the register through instructions such as MOVeax,[control_word_addr] and MOVZXecx,ax for subsequent analysis of rounding mode, precision control, and exception masking information.

[0049] S2, rounding mode mapping: Construct a rounding mode mapping table, covering the mapping of various x86 rounding modes to ARM rounding modes, including considerations for special cases and extended instruction sets; for each instruction to be converted, look up the mapping table based on the extracted rounding mode and use the corresponding ARM rounding mode. The specific steps are as follows:

[0050] Build a detailed rounding mode mapping table to map x86 rounding modes, such as round to nearest even, round toward zero, round up, round down, etc., to the corresponding ARM rounding modes. This mapping table is based on an in-depth analysis of the x86 and ARM instruction sets, and considers not only basic rounding modes but also rounding behaviors of special cases and extended instruction sets.

[0051] For each x86 floating point calculation instruction to be converted, the corresponding ARM rounding mode is found using the mapping table according to the rounding mode extracted from step 1.

[0052] For the FADDst(0),st(1) instruction, assume that the rounding mode extracted from step 1 is round towards zero. First, st(0) and st(1) are mapped to the ARM S0 and S1 registers, and then converted using the round towards zero logic in the above mapping table.

[0053] Executions of VLDRS0,[st(0)_mem] and VLDRS1,[st(1)_mem] assume that the values ​​of st(0) and st(1) are loaded from memory into S0 and S1, followed by execution of VADD.F32S2,S0,S1, followed by additional comparison and adjustment logic for rounding toward zero:

[0054] VADD.F32S2,S0,S1

[0055] VCMP.F32S2,#0.0

[0056] VMRSAPSR_nzcv,FPSCR

[0057] BGEround_zero_case

[0058] VSUBS.F32S2,S2,#0.1

[0059] round_zero_case:

[0060] If there is no direct corresponding x86 rounding mode in ARM, perform the following conversion steps.

[0061] Assume that the x86 rounding mode is: alternate rounding according to a specific flag bit, that is, different rounding modes are used under different conditions. This is implemented in ARM in the following way:

[0062] First, execute VADD.F32S2,S0,S1 to perform basic addition operations.

[0063] Then check the flag bit and use VMRSAPSR_nzcv, FPSCR to get the status flag bit.

[0064] Depending on the status of the flag bit, different branch instructions are used to perform different rounding operations, for example:

[0065] VADD.F32S2,S0,S1

[0066] VMRSAPSR_nzcv,FPSCR

[0067] TSTR0,#flag_mask / / Assume that flag_mask is the flag bit mask to be checked

[0068] BEQround_even_case

[0069] BNEround_up_case

[0070] round_even_case:

[0071] / / Round to the nearest even number logic, you can use the default VFP rounding mode

[0072] Bend_rounding

[0073] round_up_case:

[0074] VADDS.F32S2,S2,#0.1 / / Round up logic

[0075] end_rounding

[0076] By creating a detailed rounding mode mapping table, accurate conversion from x86 rounding mode to ARM rounding mode is achieved, which not only covers basic rounding modes, but also handles special cases and rounding behaviors of extended instruction sets, ensuring the accuracy of rounding results.

[0077] S3. Exception handling mechanism conversion: Analyze the x86 exception generation conditions and processing flow, determine the exception by checking the FPU control word; design exception handling logic for ARM, use VMRS to map the FPSCR status to the general register to check the exception flag, insert a custom exception handler for instructions that may be abnormal, and design auxiliary code simulation for x86 exception types that do not exist in ARM.

[0078] S3.1. Floating point exception handling mechanism for x86, including the conditions for exception generation, such as overflow, underflow, invalid operation, etc., and the corresponding processing flow. By checking the exception flag bit in the FPU control word, determine under what circumstances the exception will be triggered.

[0079] Execute FSTSWax to store the FPU control word into the ax register, and then check the bits in the ax register, such as the overflow bit, underflow bit, divide by zero bit, etc.

[0080] For the FDIVst(0),st(1) instruction, after executing this instruction, FSTSWax is executed, and then the division by zero bit in ax is checked:

[0081] FDIVst(0),st(1)

[0082] FSTSWax

[0083] TESTax,0x04 / / Check division by zero

[0084] JNZdivide_by_zero_handler

[0085] S3.2. Implement the corresponding exception handling logic in the generated ARM code. For ARM's VFP or Neon instruction set, use the VMRS instruction to map the state of the FPSCR to a general register to check the exception flag. According to the x86 exception conditions, add the corresponding exception checking and handling logic to the ARM instructions.

[0086] For the VDIV.F32S0,S0,S1 instruction, the converted ARM code is as follows:

[0087] VDIV.F32S2,S0,S1

[0088] VMRSR0,FPSCR

[0089] TSTR0,#divide_by_zero_mask / / Assume that divide_by_zero_mask is the divide by zero flag mask

[0090] BNEdivide_by_zero_arm_handler

[0091] Here divide_by_zero_arm_handler is a custom ARM exception handler.

[0092] S3.3. For each x86 floating-point operation that may generate an exception, insert a custom exception handler into the converted ARM instruction sequence.

[0093] The complete exception handling for the above VDIV.F32S2,S0,S1 instructions is as follows:

[0094] VDIV.F32S2,S0,S1

[0095] VMRSR0,FPSCR

[0096] TSTR0,#divide_by_zero_mask

[0097] BNEdivide_by_zero_arm_handler

[0098] / / Execute subsequent instructions normally

[0099] Bnormal_execution

[0100] divide_by_zero_arm_handler:

[0101] / / Save the current context, for example, save registers to the stack

[0102] PUSH{R0-R12,LR}

[0103] / / Set the error code or other exception information, assuming that R1 is used to store the error code

[0104] MOVR1,#error_code_divide_by_zero

[0105] / / Call the exception handling function, assuming it is handle_exception

[0106] BLhandle_exception

[0107] / / Restore the context

[0108] POP{R0-R12,LR}

[0109] normal_execution:

[0110] S3.4. For x86 exception types that do not exist in ARM, generate additional auxiliary code to simulate these exceptions.

[0111] For example, x86 has a special exception type that triggers an exception when some precision is lost. In ARM, for the VADD.F32S2,S0,S1 instruction, additional precision checking logic is required:

[0112] VADD.F32S2,S0,S1

[0113] VCMP.F32S2,S3 / / S3 is the expected result of using high-precision calculation

[0114] BGTprecision_loss_arm_handler / / Assuming the result is greater than the expected result, the precision may be lost

[0115] / / Execute subsequent instructions normally

[0116] Bnormal_execution

[0117] precision_loss_arm_handler:

[0118] / / Simulate x86 precision loss exception handling logic

[0119] MOVR1,#error_code_precision_loss

[0120] BLhandle_exception

[0121] normal_execution:

[0122] Comprehensive exception handling logic is designed for ARM, including additional auxiliary code for exception types that do not exist in ARM, which ensures the completeness and accuracy of exception handling and avoids missing or incorrect exception handling after conversion.

[0123] S4. Instruction conversion and precision adjustment: Convert x86 floating-point calculation instructions into ARM equivalent instructions based on rounding and exception handling results, and perform precision adjustment on complex operations, such as using multiple operations to simulate extended precision, and add precision verification and adjustment steps to instructions that store results.

[0124] Convert the x86 floating point calculation instructions to ARM equivalent instructions based on the results of steps 2 and 3. For simple conversions, use the existing ARM instruction set mapping directly.

[0125] For FADDst(0),st(1) directly convert to VADD.F32S2,S0,S1, assuming that st(0) and st(1) are mapped to S0 and S1 and the result is stored in S2.

[0126] For FMULst(0),st(1), it is converted to VMUL.F32S2,S0,S1.

[0127] During the conversion process, for some complex instructions or operations, the precision of ARM instructions needs to be adjusted.

[0128] For high-precision operations, such as x86's extended precision operations, multiple operations may be needed to simulate them in ARM. For example:

[0129] For extended precision operation of FADDst(0),st(1):

[0130] First, split the operand into high and low parts and store them in different ARM registers. Assume that st(0) is split into S0_high and S0_low, and st(1) is split into S1_high and S1_low.

[0131] Perform addition of the high bits: VADD.F32S2_high,S0_high,S1_high.

[0132] Then perform the addition of the low order parts, taking the carry into account:

[0133] VADD.F32S2_low,S0_low,S1_low

[0134] VADCS.F32S2_high,S2_high,#0.0 / / Process carry

[0135] For the FSQRT (square root operation) instruction, an iterative algorithm is needed to increase the precision:

[0136] First use VSQRT.F32S2,S0 to perform an initial square root calculation.

[0137] Then use Newton's method to adjust the accuracy:

[0138] VSQRT.F32S2,S0

[0139] VMUL.F32S3,S2,S2 / / S3=S2*S2

[0140] VSUBS.F32S3,S3,S0 / / S3=S3-S0

[0141] VDIV.F32S3,S3,S2 / / S3=S3 / S2

[0142] VADD.F32S2,S2,S3 / / S2=S2+S3 / 2 / / Newton iteration formula

[0143] For instructions involving floating-point result storage, such as FSTP, after being converted to instructions such as VSTR or VMOV in ARM, an accuracy verification step is added.

[0144] For the FSTP[result] instruction, convert it to VSTRS2,[result], and then add precision verification:

[0145] VSTRS2,[result]

[0146] VCMP.F32S2,S3 / / S3 is the expected result obtained by simulating x86 operations using high-precision software algorithms

[0147] BGTadjust_result

[0148] BLEnormal_result

[0149] adjust_result:

[0150] / / The precision deviation is positive, so make a subtraction adjustment

[0151] VSUB.F32S2,S2,S4 / / S4 is a smaller correction

[0152] normal_result:

[0153] In addition to basic instruction conversion, precision verification and adjustment steps are added, and the precision of the converted results is corrected by adding adjustment instructions to ensure the accuracy of floating-point operation results.

[0154] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A floating-point instruction conversion calculation accuracy optimization method based on extended control words, characterized in that: The specific steps include: Step 1: Add a virtual exception flag bit to the control word of ARM floating point calculation to indicate virtual exceptions related to x86 floating point calculation that do not exist in the ARM system; For the x86 control word and the ARM control word, a rounding mode mapping table and an exception type mapping table are established respectively, wherein the exception type mapping table contains the correspondence between the virtual exception flag and the x86 floating point calculation exception; the virtual exception and the virtual exception handler are added to the ARM exception vector table, and a virtual exception instruction table consisting of the correspondence between the virtual exception and the x86 floating point calculation instruction is constructed; and an executable file is loaded and executed in the ARM system through dynamic instruction conversion; Step 2, obtain the current instruction to be converted, if the current instruction to be converted is an x86 floating point control word load instruction, execute step 3, if the current instruction to be converted is an x86 floating point calculation instruction, execute step 4; Step 3, determining, according to the rounding mode mapping table, that the flag bit and the exception flag bit of the x86 rounding mode correspond to the first destination flag bit and the second destination flag bit, and converting the current instruction to be converted into an ARM instruction sequence consisting of an instruction to read the ARM control word, an instruction to modify the control word to set the first destination flag bit and the second destination flag bit, and a write-back instruction according to the exception type mapping table; Step 4: Convert the current instruction to be converted into an ARM instruction sequence with the same function; Step 5: If the current instruction to be converted is in the virtual exception instruction table, construct an abnormal ARM instruction group that simulates triggering the virtual exception corresponding to the current instruction to be converted, and convert the current instruction to be converted into an ARM instruction sequence formed by the ARM instruction sequence obtained in step 4 and the abnormal ARM instruction group; Otherwise, keep the ARM instruction sequence obtained in step 4 unchanged; The step 4 also includes: recording the ARM instruction with the same function as the current instruction to be converted as the first ARM instruction; if the current instruction to be converted is a high-precision basic calculation instruction, splitting the operand into a low-order part and a high-order part, and converting it into an ARM instruction sequence consisting of the first ARM instruction with the low-order part as the operand, the second ARM instruction that accumulates the operation result of the low-order part to the high-order part, and the first ARM instruction with the high-order part as the operand; if the current instruction to be converted is a high-precision complex calculation instruction, converting it into an ARM instruction sequence consisting of the first ARM instruction and the second ARM instruction group constructed based on the Newton iteration method.

2. The floating-point instruction conversion calculation accuracy optimization method according to claim 1, characterized in that: The second ARM instruction group includes iterative calculations for improving calculation accuracy.

3. The method for optimizing floating-point instruction conversion calculation accuracy according to claim 1, characterized in that: The processing process of the abnormal ARM instruction group in step 5 includes: obtaining the instruction or operand required to be detected according to the known abnormal triggering condition, and then judging whether it meets the condition for triggering the exception, if so, calling the virtual exception handler, otherwise executing the processing flow normally.

4. The method for optimizing floating-point instruction conversion calculation accuracy according to claim 1, characterized in that: The virtual exception flag in step 1 is implemented by extending the control word FPSCR of ARM floating point calculation.

5. The floating-point instruction conversion calculation accuracy optimization method according to claim 1, characterized in that: The processing process of the virtual exception handling program in step 1 includes: saving the current state of the ARM system, building the exception handling logic, and after the exception handling is completed, restoring the previously saved state of the ARM system and continuing to execute the program.

6. The floating-point instruction conversion calculation accuracy optimization method according to claim 5, characterized in that: The exception handling logic is to record error information, generate prompt information or try to repair the error.

7. The method for optimizing floating-point instruction conversion calculation accuracy according to claim 5, characterized in that: The saving of the current state of the ARM system includes: saving register values ​​and a program counter.

8. The method for optimizing floating-point instruction conversion calculation accuracy according to claim 1, characterized in that: The x86 floating point control word loading instruction in step 2 is FLDCW or LDMXCSR.

9. The method for optimizing floating-point instruction conversion calculation accuracy according to claim 1, characterized in that: Also includes: The correspondence between the ARM instruction sequence obtained during the conversion process and the current instruction to be converted is saved in the instruction conversion cache. When performing the conversion again, the instruction conversion cache is queried first. If there is a corresponding conversion result, it is converted directly, otherwise, step 2 is executed.

Citation Information

Patent Citations

  • Floating point rounding processors, methods, systems, and instructions

    CN104011647A

  • Prefetch instruction conversion optimization method based on memory access mode virtualization

    CN119440626A