Input Generation Method, System and Medium for Differences in Floating-Point Compilation Optimization Results

Through the input generation method for floating-point compilation optimization, the driver and Markov chain Monte Carlo sampling technology are used to generate inputs that can trigger the greatest difference in results, solving the problem of floating-point program results changes caused by the compilation optimization options, and achieving evaluation of precision constraints and floating-point calculation optimization.

CN116185414BActive Publication Date: 2025-06-24NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211096328.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-06-24
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

Compilation optimization options will cause changes in the results of floating-point programs. It is difficult for the existing technology to effectively evaluate whether floating-point program compilation optimization meets the accuracy constraints of the output, and further optimizes the floating-point computing behavior to avoid accuracy loss.

Method used

It provides an input generation method for the difference in floating-point compilation optimization results. By constructing a driver to call the executable version of the floating-point program under different compilation optimization options, and divide the input fields for random sampling and Markov chain Monte Carlo sampling, it generates an input that can trigger the greatest difference in results.

Benefits of technology

It can quickly generate inputs, so that the output results of floating-point programs under different compilation optimization options are as large as possible, evaluate whether the compilation optimization meets the accuracy constraints, and optimize floating-point calculations through the input execution information to avoid the loss of accuracy introduced by aggressive compilation optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116185414B_ABST
    Figure CN116185414B_ABST
Patent Text Reader

Abstract

The present invention discloses an input generation method, system and medium for the difference in floating-point compilation optimization results. The present invention includes constructing a driver program that can output the difference in execution results under different compilation optimization options; dividing the input domain of the floating-point program into subspaces and randomly sampling and screening valuable samples; using the difference in results generated by executing the valuable samples by the driver program as its fitness value, and further searching by Markov chain Monte Carlo sampling to obtain samples that trigger greater differences in results under different compilation optimization options. The present invention can quickly generate inputs to make the output results of the floating-point program as different as possible under different compilation optimization options. The generated inputs can not only evaluate whether the floating-point program meets the result accuracy constraints during compilation optimization, but also further optimize the floating-point calculation method in the program according to the execution information of the inputs, avoiding the loss of floating-point precision introduced when the program performs aggressive compilation optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the reliability assurance technology of computer high-performance computing, and particularly to an input generation method, system and medium for the difference of floating-point compilation optimization results. Based on the generated input, the present invention can trigger the difference in the running results of floating-point programs under different compilation optimization options, and can be used to evaluate whether the floating-point program compilation optimization meets the accuracy constraints of the output, and further optimize the floating-point calculation behavior in the program according to the execution information of the input, reducing the accuracy loss caused by compilation optimization. Background Art

[0002] Floating-point numbers are the de facto standard for representing real numbers in code, and floating-point operations play an important role in numerical calculations. Due to limited precision (finite binary bit encoding), rounding errors inevitably exist in floating-point operations, and the rounding errors will accumulate and propagate during program execution, resulting in abnormal running results of floating-point programs. In safety-critical software, the loss of floating-point precision may lead to catastrophic consequences, such as data chaos in the Vancouver Stock Exchange, Patriot missile launch failures, and brake problems in the Toyota Prius, etc.

[0003] According to the IEEE-754 standard, a floating-point number can be expressed as: (-1) S ×M×2 E , where S ∈ {0, 1} is the sign bit, where 0 represents a positive number and 1 represents a negative number. M = m0.m1m2…m n is the mantissa, where m0 is the hidden bit, m1m2…m n is an n-bit decimal. E = e - bias is a p-bit exponent, where e represents the biased exponent and the offset bias = 2 p-1 - 1.

[0004] Table 1 shows the format of IEEE floating-point numbers. For example, a single-precision floating-point number (32-bit float type) has 1 sign bit, 8 exponent bits (bias is 127), and 23 mantissa bits.

[0005] Table 1: IEEE 754 Floating-Point Number Format.

[0006] Sign bit Exponent Mantissa Half precision 1 5 10 Single precision 1 8 23 Double precision 1 11 52

[0007] Due to the limited number of binary encoding bits, floating-point numbers can only represent a subset of real numbers, and rounding errors are inevitable. There are various rounding modes in a computer to convert a real number into a floating-point number. For example, the round-to-nearest mode converts a real number into the floating-point number representation closest to it. We use R and F to represent the set of real numbers and the set of floating-point numbers respectively.

[0008] Assume that the floating - point representation of a real number \(r(x\in R)\) is \(r\) f (\(r\) f \(\in F\)), the rounding error can be expressed as \(r - r\) f .

[0009] Given a floating - point program \(FP\), we use \(FP(x)\) to represent the output produced by the floating - point program \(FP\) when executing the floating - point input \(x\), and use \(FP\) * (\(x\)) to represent the ideal mathematical calculation result of the floating - point program \(FP\) when executing the floating - point input \(x\). Obviously, the absolute error of the floating - point program \(FP\) at the input \(x\) is \(\vert FP\) * (\(x)-FP(x)\vert\).

[0010] The ULP error \(Err\) ulp is widely used to evaluate the loss of floating - point precision, and its calculation formula is as follows:

[0011]

[0012] In the above formula, \(Err\) ulp (\(FP(x),FP\) * (\(x\)) represents the ULP error \(Err\) * between the output \(FP(x)\) produced by the floating - point program \(FP\) when executing the floating - point input \(x\) and the ideal mathematical calculation result \(FP\) ulp , and its calculation method is the absolute error \(\vert FP\) * (\(x)-FP(x)\vert\) divided by the ULP value \(ULP(FP\) * (\(x\)) f ) of the ideal floating - point result \(FP\) * (\(x\) f ). \(ULP(x)\) represents the distance between the two closest floating - point numbers to \(x\), and existing tool interfaces can directly call to calculate the ULP value. Further, calculating the logarithm of 2 can obtain the number of error bits of the program floating - point calculation result relative to the ideal exact calculation result, that is:

[0013] \(Err\) bits =\(\log_2(Err\) ulp )), (2)

[0014] Monte Carlo refers to a class of methods for solving problems through random sampling and statistics. For example, randomly scatter points into a square with a side length of 1, and solve for the value of \(\pi\) by statistically calculating the probability that the points fall into the inscribed circle of the square. Since the probability that a point falls into the inscribed circle is the ratio of the area of the inscribed circle to the area of the square, \(\pi / 4\), after statistically obtaining the probability through a large number of samples and then multiplying by 4, we can get \(\pi\). A Markov Chain describes a stochastic process \(\{X_0,X_1,X_2,\cdots\}\), where the state \(X\)t+1 Only depends on state X t , and is independent of other historical states. Markov Chain Monte Carlo (MCMC) describes a process of sampling in the input space C according to its probability distribution p(C). The Metropolis-Hastings algorithm is a widely used MCMC sampling algorithm. The probability of sampling sample X * based on sample state X is p(X → X * ), and its calculation formula is as follows, where q(X * |X) represents the proposal distribution, that is, the probability of migrating from sample X to sample X * . In most application backgrounds, the proposal distribution is symmetric, that is, q(X * |X) = q(X|X * ). Therefore, the probability of sampling sample X * based on sample state X is simplified to the Metroplis ratio min(1, p(X * ) / p(X)). Obviously, MCMC sampling will sample more in the regions with dense probability distributions, and there is a certain probability of sampling from high-density regions to low-density regions. The probability of sampling sample X * based on sample state X is p(X → X * ), and its functional expression is:

[0015]

[0016] Considering that the probability density distribution p(C) of the samples is often unknown, while the importance of the samples can be directly evaluated. Related research gives the fitness function fit(x) based on the samples to convert it into the sample probability density distribution function p(x):

[0017]

[0018] where x is the sample, z is the normalization partition function, and σ is a constant. Obviously, the formula can ensure that more samples are taken in the regions with high fitness values during sampling. The Monte Carlo jumps in MCMC sampling effectively avoid the local optimum problem of the search.

[0019] In science and high-performance computing, the compiler is one of the core tools that developers rely on to seek program optimization. Compilation optimization improves performance by converting a piece of code into another piece of code with equivalent functionality. However, there are a large number of optimization options in the compiler, and developers usually do not understand the true meaning of the compiler optimization options. In fact, a large number of compilation optimization options can change the floating-point calculation behavior, such as -freciprocal-math and -fassociative-math. In addition, a compilation optimization option may have different optimization meanings in different compilers. For example, the GNU compiler GCC and the Intel compiler ICC perform different code optimizations under the O3 option.

[0020] The difference in the compilation optimization results of floating-point programs refers to the inconsistency in the execution results of floating-point programs after being compiled with different compilation optimization options. For example, for the double-precision floating-point variables a = 1.1e300 and b = 1.0e-6, the expression a + b - a will return the result 0 under the compilation optimization option "-O0", while it will return the exact result 1e-6 under the compilation optimization option "-O3 -ffast-math". This is because in the -O0 mode, the addition and subtraction operations are performed sequentially, and for 1.1e300 + 1.0e-6, since the double-precision mantissa has only 52 bits, 1.0e-6 will be discarded after the exponent alignment. While -O3 -ffast-math will optimize a + b - a into the expression b. When developing the hydrodynamics application Laghos at the Lawrence Livermore National Laboratory in the United States, it was found that when using the IBM compiler xlc to compile and run under the optimization options "-O2" and "-O3", the energy results calculated by the program for specific inputs were 129664.923 and 144174.934 respectively. Through debugging, it was found that this was caused by the "-O3" option optimizing a floating-point division operation in the program into a combination operation of floating-point reciprocal and multiplication. Rapidly generating floating-point inputs to maximize the output differences of programs under different compilation optimization options is of great significance for evaluating whether the compilation optimization of floating-point programs meets the accuracy constraints of the output and further optimizing the floating-point calculation behavior in the program based on the execution information of the input to avoid accuracy loss.

[0021] Currently, there is relatively little research on the result differences caused by the compilation optimization of floating-point programs. The tool Flit developed by the University of Utah and the Lawrence Livermore National Laboratory in the United States evaluates whether there are result differences between different compilation optimization versions of floating-point programs based on random sampling. The tool Plinear developed by the University of California, Davis can, when an input triggers a compilation optimization result difference problem, locate the statement position that causes the inconsistency in the compilation optimization results of floating-point programs by searching for different compilation optimization options of the function and improving the accuracy of floating-point calculation statements within the function. Summary of the Invention

[0022] Technical problems to be solved by the present invention: Aiming at the problem that compilation optimization options may cause changes in the results of floating-point programs, the present invention provides an input generation method, system and medium for the differences in floating-point compilation optimization results. For a given floating-point program FP, different compilation optimization options flag1 and flag2 as a comparison, the present invention can quickly generate inputs to make the compilation execution results of the floating-point program FP under flag1 and flag2 as different as possible. The inputs generated based on the present invention can not only evaluate whether the floating-point program meets the accuracy constraints of the program output results during compilation optimization, but also further optimize the floating-point calculation method according to the execution information of the input to avoid the loss of floating-point precision caused by aggressive compilation optimization.

[0023] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0024] An input generation method for the differences in floating-point compilation optimization results, including:

[0025] S101, According to the given floating-point program FP, the compilation optimization options flag1 and flag2 as a comparison, construct a driver program FP0 through program compilation and linking. The driver program FP0 is used to call the executable versions FP1 and FP2 generated by compiling the floating-point program FP under the compilation optimization options flag1 and flag2 respectively, and take the result difference of the executable versions FP1 and FP2 for the same input as the output;

[0026] S102, Divide the input domain C of the floating-point program FP into a specified number of subspaces, randomly sample a specified number k of samples in each subspace, and save those samples in each subinterval that make the output of the driver program FP0 not zero or the floating-point instruction sequences executed when calling and executing the executable versions FP1 and FP2 are different as the set of inputs to be refined;

[0027] S103, Based on the set of inputs to be refined, take the result difference of an input under different compilation optimization options flag1 and flag2 as its fitness value, and use Markov chain Monte Carlo sampling to further search for samples that trigger greater result differences and output them as results.

[0028] Optionally, the result difference in step S101 refers to the number of error bits between the floating-point results obtained by the executable versions FP1 and FP2 for the same input, and the number of error bits refers to the number of different bits in the binary codes of the two floating-point results.

[0029] Optionally, step S102 includes:

[0030] S201, Initialize the refined input set tmp_list as empty; divide the program input domain C into a set of subspaces confs;

[0031] S202, Determine whether the set of subspaces confs is not empty. If it holds, jump to S203; otherwise, use the samples in the refined input set tmp_list as the obtained refined input set, and jump to S103;

[0032] S203, Traverse and extract a subspace c from the set of subspaces confs i , from the subspace c i randomly sample k samples I1, I2, …, I k and save them as a list list = [I1, I2, …, I k ;

[0033] S204, If the list list is not empty, jump to S205; otherwise, jump to S202;

[0034] S205, Traverse and extract a sample I from the list list j , use the sample I j as input to execute the driver program FP0, and record the floating-point instruction sequence tf1 when the executable version FP1 executes the sample I j , the floating-point instruction sequence tf2 when the executable version FP2 executes the sample I j , and the output output of the driver program FP0. The output output corresponds to the result difference when the executable versions FP1 and FP2 execute the sample I j ;

[0035] S206, Determine whether the output output is not 0, or whether there are differences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. If it holds, save the sample I j in the refined input set tmp_list; jump to S204.

[0036] Optionally, when dividing the program input domain C into the set of subspaces confs in step S201, for the input domain C with a value range of (-∞, +∞), the set of subspaces confs obtained by evenly dividing according to the floating-point number distribution includes 2000 subintervals.

[0037] Optionally, the determination of differences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 in step S206 includes: calculating the difference diff(tf1, tf2) between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. If the difference diff(tf1, tf2) is not equal to 0, it is determined that there are differences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2.

[0038] Optionally, the difference diff(tf1, tf2) between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 includes: first, calculating the set M of matching sequences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2, then using the number of instructions in the set M of matching sequences multiplied by 2 and divided by the sum of the number of instructions in the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 as the similarity ratio ratio(tf1, tf2), and then using the result of 1.0 - ratio(tf1, tf2) as the calculated difference diff(tf1, tf2).

[0039] Optionally, the set M of matching sequences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 includes:

[0040] S301, adding the sequence pair (tf1, tf2) composed of the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 to the set QS to be calculated, and initializing the set M of matching sequences to be empty;

[0041] S302, if the set QS to be calculated is not empty, taking out a sequence pair (t1, t2) from it, then calculating the longest matching sequence ms between the floating-point instruction sequence t1 and the floating-point instruction sequence t2, adding the longest matching sequence ms to the set M of matching sequences, and jumping to execute S303; otherwise, determining that the calculation of the set M of matching sequences is completed;

[0042] S303, deducting the longest matching sequence ms from the floating-point instruction sequence t1, using to represent the left part after deducting the longest matching sequence ms from the floating-point instruction sequence t1, to represent the right part after deducting the longest matching sequence ms from the floating-point instruction sequence t1; removing the longest matching sequence ms from the floating-point instruction sequence t2, using to represent the left part after deducting the longest matching sequence ms from the floating-point instruction sequence t2, to represent the right part after deducting the longest matching sequence ms from the floating-point instruction sequence t2; if the left part after deducting the longest matching sequence ms from the floating-point instruction sequence t1 and the left part after deducting the longest matching sequence ms from the floating-point instruction sequence t2 are both not empty, adding the sequence pair to the set QS to be calculated; if the right part after deducting the longest matching sequence ms from the floating-point instruction sequence t1 and the right part

[0043] after deducting the longest matching sequence ms from the floating-point instruction sequence t2 are both not empty, adding the sequence pair to the set QS to be calculated; jumping to S302.

[0043] Optionally, step S103 includes:

[0044] S401, initialize the list variation_list to be empty;

[0045] S402, determine whether the input set tmp_list to be refined is not empty. If not, directly use the input I corresponding to the largest execution result res in the list variation_list as the input that can trigger the maximum difference in the compilation optimization result, and end and exit; otherwise, jump to S403;

[0046] S403, traverse and take out a sample I from the input set tmp_list to be refined i , use the calculation function FP0 of the compilation optimization result difference as the fitness function fit, and start Markov chain Monte Carlo sampling from the sample I i , find a sample I with a larger fitness value through Markov chain Monte Carlo sampling t , record the sample I t The difference in the calculation result obtained by executing the driver program FP0 is res t , and the input I t and its corresponding execution result res t to form a result pair (I t , res t ) and append it to the list variation_list, and return to S402.

[0047] In addition, the present invention also provides an input generation system for the difference in floating-point compilation optimization results, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the input generation method for the difference in floating-point compilation optimization results.

[0048] In addition, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is used to be programmed or configured by a microprocessor to execute the input generation method for the difference in floating-point compilation optimization results.

[0049] Compared with the prior art, the present invention mainly has the following advantages: For a given floating-point program FP, different compilation optimization options flag1 and flag2 as a comparison, the present invention can quickly generate an input so that the output results of the floating-point program FP compiled and executed under the options flag1 and flag2 are as different as possible. The input generated based on the present invention can not only evaluate whether the floating-point program meets the result accuracy constraint during compilation optimization, but also further optimize the floating-point calculation method according to the execution information of the input, and avoid the loss of floating-point precision caused by aggressive compilation optimization. Description of the Drawings

[0050] Figure 1 This is a schematic diagram of the basic process of the method according to the embodiments of the present invention.

[0051] Figure 2 This is a schematic diagram of the specific implementation process of the method according to the embodiments of the present invention. Detailed implementation manners

[0052] As Figure 1 and Figure 2 shown, the input generation method for the difference in floating-point compilation optimization results in this embodiment includes:

[0053] S101. According to the given floating-point program FP, construct a driver program FP0 through program compilation and linking with the compilation optimization options flag1 and flag2 for comparison. The driver program FP0 is used to call the executable versions FP1 and FP2 generated by compiling the floating-point program FP under the compilation optimization options flag1 and flag2 respectively, and take the difference in the calculation results of the executable versions FP1 and FP2 for the same input as the output (return value);

[0054] See Figure 2 , where constructing the driver program FP0 is expressed as:

[0055] FP0 ← contruct(FP, flag1, flag2)

[0056] where contruct represents the process of compiling and linking the program to construct the driver program FP0. This process includes the executable versions FP1 and FP2 of the floating-point program FP obtained by compiling through the compilation optimization options flag1 and flag2 respectively, and calculates the difference in the calculation results between the results r1 (obtained by the floating-point program FP1) and r2 (obtained by the floating-point program FP2) by calling the executable versions FP1 and FP2 to calculate the same input x (the method supports programs with multiple inputs) respectively, as well as the floating-point instruction sequences tf1 (obtained by executing the floating-point program FP1) and tf2 (obtained by executing the floating-point program FP2), and takes the difference in the calculation results between the results r1 (obtained by the floating-point program FP1) and r2 (obtained by the floating-point program FP2) as the output (return value).

[0057] S102. Divide the input domain C of the floating-point program FP into a specified number of subspaces, randomly sample a specified number k of samples in each subspace, and save the samples in each subinterval that make the output of the driver program FP0 non-zero or the floating-point instruction sequences executed when calling and executing the executable versions FP1 and FP2 different as the set of inputs to be refined;

[0058] S103. Based on the set of inputs to be refined, the difference in results of an input under different compilation optimization options flag1 and flag2 is used as its fitness value. Markov Chain Monte Carlo (MCMC) sampling is employed to further search for samples that trigger a greater difference in results and output them as the result.

[0059] In this embodiment, the difference in calculation results in step S101 refers to the number of error bits between the floating-point results obtained by the computable executable versions FP1 and FP2 for the same input. The number of error bits refers to the number of different bits in the binary encodings of the two floating-point results. The calculation method can be found in equations (1) and (2) mentioned in the background art.

[0060] See Figure 2 , step S102 in this embodiment includes:

[0061] S201. Initialize the set of inputs to be refined tmp_list as empty; divide the program input domain C into a set of subspaces confs;

[0062] See Figure 2 , in this embodiment, initializing the set of inputs to be refined tmp_list as empty is expressed as:

[0063] tmp_list ← []

[0064] That is, assign the empty data [] to the set of inputs to be refined tmp_list.

[0065] It should be noted that the step of initializing the set of inputs to be refined tmp_list as empty can be implemented in step S102, or see Figure 2 , the step of initializing the set of inputs to be refined tmp_list as empty can be advanced to the very beginning of the entire program, that is, after inputting the floating-point program FP, after the compilation optimization options flag1 and flag2 for comparison, and before constructing the driver program FP0 for unified initialization. In addition, it can also be at other positions in step S101, or after "divide the program input domain C into a set of subspaces confs" in step S201. In principle, as long as the set of inputs to be refined tmp_list is initialized as empty before executing S202.

[0066] See Figure 2 , in this embodiment, dividing the program input domain C into a set of subspaces confs is expressed as:

[0067] confs ← partition(C)

[0068] Among them, partition is the operation of dividing the program input domain C into subspaces, and the subspace division is evenly divided according to the floating-point number distribution.

[0069] S202. Determine whether the subspace set confs is not empty. If it is, jump to S203; otherwise, use the samples in the input set tmp_list to be refined as the obtained input set to be refined, and jump to S103.

[0070] S203. Traverse and retrieve a subspace c from the subspace set confs i , and randomly sample k samples I1, I2, …, I i from the subspace c k and save them as a list list = [I1, I2, …, I k ;

[0071] See Figure 2 , in this embodiment, traversing and retrieving a subspace c from the subspace set confs i is expressed as:

[0072] c i ← confs.pop()

[0073] where pop() represents the pop (traverse and retrieve) operation on the subspace set confs. After popping a subspace, the subspace will be removed from the subspace set confs. In this embodiment, randomly sampling k samples I1, I2, …, I i from the subspace c k and saving them as a list list = [I1, I2, …, I k is expressed as:

[0074] list = [I1, I2, …, I k ← random(c i , k)

[0075] where random(c i , k) represents the operation of randomly sampling k samples I1, I2, …, I i from the subspace c k . In the method implementation, we set the value of k to 81.

[0076] S204. If the list list is not empty, jump to S205; otherwise, jump to S202.

[0077] S205. Traverse and retrieve a sample I from the list list j , use the sample I j as the input to execute the driver program FP0, and record the floating-point instruction sequence tf1 when the executable version FP1 executes the sample I j , and the floating-point instruction sequence when the executable version FP2 executes the sample I jThe floating-point instruction sequence tf2 at a certain time, and the output output of the driver FP0. The output output corresponds to the execution samples I of the executable versions FP1 and FP2 j The result difference;

[0078] See Figure 2 In this embodiment, a sample I will be traversed and taken out from the list list j It is expressed as:

[0079] I j ← list.pop()

[0080] where pop() represents the pop (traverse and take out) operation on the list list, which has the same meaning as the pop (traverse and take out) operation of the subspace set confs. The sample I j As the call to the driver FP0 is expressed as:

[0081] output, tf1, tf2 ← run(FP0, I j )

[0082] where run represents the call and execution.

[0083] S206. Determine whether the output output is not 0 (output ≠ 0), or whether there is a difference between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. If it holds, then the sample I j is saved in the input set tmp_list to be refined (tmp_list.append(I j )); Jump to S204. Among them, append represents the append / new operation of the input set tmp_list to be refined.

[0084] In this embodiment, when the program input domain C is divided into the subspace set confs in step S201, for the input domain C with a value range of (-∞, +∞), the subspace set confs obtained by evenly dividing according to the floating-point number distribution includes 2000 sub-intervals.

[0085] In this embodiment, the determination of the difference between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 in step S206 includes: calculating the difference diff(tf1, tf2) between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. If the difference diff(tf1, tf2) is not equal to 0 (expressed as diff(tf1, tf2) ≠ 0), then it is determined that there is a difference between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. Therefore, see Figure 2 to determine whether the output output is not 0, or whether there is a difference between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 is expressed as:

[0086] (output ≠ 0 ∨ diff(tf1, tf2) ≠ 0)

[0087] Among them, "∨" represents the logical OR.

[0088] In this embodiment, calculating the difference diff(tf1, tf2) between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 includes: First, calculate the set M of matching sequences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. Then, multiply the number of instructions in the set M of matching sequences by 2 and divide by the sum of the number of instructions in the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 to obtain the similarity ratio ratio(tf1, tf2). Then, use the result of 1.0 - ratio(tf1, tf2) as the calculated difference diff(tf1, tf2).

[0089] In order to obtain the set M of matching sequences, in this embodiment, the idea of recursion is adopted. Calculating the set M of matching sequences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 includes:

[0090] S301, Add the sequence pair (tf1, tf2) composed of the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 to the set QS to be calculated, and initialize the set M of matching sequences to be empty;

[0091] S302, If the set QS to be calculated is not empty, take out a sequence pair (t1, t2) from it, and then calculate the longest matching sequence ms between the floating-point instruction sequence t1 and the floating-point instruction sequence t2, and add the longest matching sequence ms to the set M of matching sequences; If the set QS to be calculated is empty, it is determined that the calculation of the set M of matching sequences is completed; Otherwise, jump to S303;

[0092] S303, Deduct the longest matching sequence ms from the floating-point instruction sequence t1, where represents the left part of the floating-point instruction sequence t1 after deducting the longest matching sequence ms, represents the right part of the floating-point instruction sequence t1 after deducting the longest matching sequence ms; Remove the longest matching sequence ms from the floating-point instruction sequence t2, where represents the left part of the floating-point instruction sequence t2 after deducting the longest matching sequence ms, represents the right part of the floating-point instruction sequence t2 after deducting the longest matching sequence ms; If the left part of the floating-point instruction sequence t1 after deducting the longest matching sequence ms and the left part of the floating-point instruction sequence t2 after deducting the longest matching sequence ms are both not empty, add the sequence pair to the set QS to be calculated; If the right part of the floating-point instruction sequence t1 after deducting the longest matching sequence ms and the right part of the floating-point instruction sequence t2 after deducting the longest matching sequence ms are not empty, add the sequence pair to the set QS to be calculated; jump to S302. For example, we use the sequenceMatcher class in the python difflib module to implement such a calculation function. For example, assume that two instruction sequences tf1 and tf2 are <fadd,fsub,fmul,fdiv,fmov> and <fsub,fmul,fmov> respectively. Initially, QS will contain (tf1,tf2) and M is empty. Take (t1,t2) from QS, then calculate the longest matching sequence ms = <fsub,fmul> and add it to the set M. At this time and are respectively <fadd>and <>. And and are <fdiv,fmov> and <fmov>, all are non-empty. Add to QS. Continue to take out the newly added sequence pair to be calculated from QS (<fdiv, fmov>, <fmov>), calculate the longest matching sequence ms = <fmov>and add it to M. At this time and are both empty, so QS becomes an empty set. Since M = {<fsub, fmul>, <fmov>}, so the similarity ratio is ratio(tf1,tf2)=(2*3) / (5+3)=0.75, and the difference value is 1-0.75=0.25.

[0093] In this embodiment, step S103 includes:

[0094] S401, initialization list variation_list is empty;

[0095] S402, determine whether the set of inputs to be refined tmp_list is not empty. If not, directly use the input I corresponding to the largest execution result res in the list variation_list as the input that can trigger the maximum difference in the compilation optimization results, and end and exit; otherwise, jump to S403;

[0096] S403, traverse and take out a sample I from the input set to be refined (i.e., the input set to be refined tmp_list mentioned above) i (I i ←tmp_list.pop()), use the result difference calculation function FP0 as the fitness function fit, and then i Start Markov chain Monte Carlo sampling and find samples with larger fitness value I through Markov chain Monte Carlo sampling t , and record I t The difference in the calculation results triggered res t , Figure 2 It is expressed as: (I t ,res t )←MCMC(I i ,FP0,fit), input I t And its corresponding difference result res t Append to the list variation_list (variation_list.append((I t ,res t ))), return to S402.

[0097] It should be noted that Markov chain Monte Carlo sampling is a known method. The key is how to determine its fitness function fit. In this embodiment, the calculation function FP0 of the result difference is used as the fitness function fit. i Sample more in regions with higher fitness values and use Monte Carlo jumps to avoid local optima in sampling searches. Specifically, we use the Basinhopping algorithm implemented in the Scipy library (version 1.5.4) as our Markov chain Monte Carlo sampling tool. To effectively guide the Markov chain Monte Carlo sampling to obtain inputs that can trigger differences in compilation optimization results, we use FP0(x) as the fitness value of input x because the output of the driver FP0 corresponds to the number of error bits Err between the calculation results of different compilation optimization versions. bits , and the larger its value indicates that the difference in compilation optimization results triggered by the input x is greater.

[0098] To verify the method of this embodiment, in this embodiment, it is specifically implemented based on the LLVM compiler, the sequenceMatcher class in the pythondifflib module, and the Scipy library (version 1.5.4). First, we use the clang compiler in LLVM to compile the input floating-point program FP under different compilation optimization options flag1 and flag2 to generate corresponding shared libraries FP1.so and FP2.so (corresponding to different compilation optimization versions of the program), and then use the CDLL module in Python ctypes to load the shared libraries to generate a test driver FP0; then, we implemented an LLVM optimization pass (PASS) in LLVM to collect the floating-point instruction sequences generated during program execution; we use the ratio interface in the sequenceMatcher class in the python difflib module to calculate the similarity and difference of the instruction sequences; we use the Basinhopping method in the Scipy library as the MCMC sampling engine. We applied our tool to 158 functional programs with input and output in the open-source GNU Scientific Library (GSL). The experiments show that compared with the existing random search and binary search, the method of this embodiment can generate inputs for 100% of the programs to make the difference in compilation optimization results greater under a given time overhead, and on the premise of generating inputs of the same quality (that is, the difference degree of the results triggered by the final inputs generated by the random search / binary search and the method of this embodiment is the same), the method of this embodiment can obtain at least an 11-fold time speedup. Therefore, the experimental results show the effectiveness and efficiency of the method of this embodiment.

[0099] It should be noted that the "compilation optimization option" involved in this embodiment refers to a compilation optimization parameter or a combination of multiple compilation optimization parameter options, which can be specifically selected only according to the needs of compilation optimization. In addition, this embodiment involves a typical situation of comparing compilation optimization options. For the comparison of multiple compilation optimization options, they are split into pairs of compilation optimization options, and then the method of this embodiment is adopted for each pair of obtained compilation optimization options one by one.

[0100] In addition, this embodiment also provides an input generation system for the difference in floating-point compilation optimization results, including a microprocessor and a memory connected to each other. The aforementioned microprocessor is programmed or configured to execute the aforementioned input generation method for the difference in floating-point compilation optimization results.

[0101] In addition, this embodiment also provides a computer-readable storage medium. A computer program is stored in the aforementioned computer-readable storage medium, and the aforementioned computer program is used to be programmed or configured by a microprocessor to execute the aforementioned input generation method for the difference in floating-point compilation optimization results.

[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes

[0103] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention< / fmov> < / fmov> < / fmov> < / fmov> < / fadd>

Claims

1. An input generation method for the difference in floating-point compilation optimization results, characterized in that, Including: S101. According to a given floating-point program FP, compilation optimization options flag1 and flag2 for comparison are used to construct a driver program FP0 through program compilation and linking. The driver program FP0 is used to call the executable versions FP1 and FP2 generated by compiling the floating-point program FP under the compilation optimization options flag1 and flag2 respectively, and the difference in the calculation results of the executable versions FP1 and FP2 for the same input is used as the output; S102. Divide the input domain C of the floating-point program FP into a specified number of subspaces. Randomly sample a specified number k of samples in each subspace, and save those samples in each subinterval that make the output of the driver program FP0 non-zero or for which the floating-point instruction sequences executed when calling and executing the executable versions FP1 and FP2 are different as the set of inputs to be refined; S103. Based on the set of inputs to be refined, use the difference in the results of an input under different compilation optimization options flag1 and flag2 as its fitness value, and adopt Markov chain Monte Carlo sampling to further search for samples that trigger a greater difference in results and output them as the result.

2. The input generation method for the difference in floating-point compilation optimization results according to claim 1, wherein The difference in the calculation results in step S101 refers to the number of error bits between the floating-point results obtained by the executable versions FP1 and FP2 for the same input. The number of error bits refers to the number of bits that are different in the binary encodings of the two floating-point results.

3. The input generation method for the difference in floating-point compilation optimization results according to claim 1, wherein Step S102 includes: S201. Initialize the set of inputs to be refined tmp_list to be empty; divide the program input domain C into a set of subspaces confs; S202. Determine whether the set of subspaces confs is not empty. If it is, jump to S203; otherwise, use the samples in the set of inputs to be refined tmp_list as the obtained set of inputs to be refined, and jump to S103; S203, traverse and retrieve a subspace c from the subspace set confs i , from the subspace c i randomly sample k samples I1, I2, …, I k and save them as a list list = [I1, I2, …, I k ; S204. If the list list is not empty, jump to S205; otherwise, jump to S202; S205, traverse and retrieve a sample I from the list list j , and use the sample I j as the input to execute the driver program FP0, and record the floating-point instruction sequence tf1 when the executable version FP1 executes the sample I j , the floating-point instruction sequence tf2 when the executable version FP2 executes the sample I j , and the output output of the driver program FP0. The output output corresponds to the result difference when the executable versions FP1 and FP2 execute the sample I j ; S206, determine whether it holds that the output output is not 0, or there are differences between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. If it holds, save the sample I j in the refinement input set tmp_list; jump to S204.

4. The input generation method for the difference in floating-point compilation optimization results according to claim 3, characterized in that When dividing the program input domain C into the set of subspaces confs in step S201, for the input domain C with a value range of (-∞, +∞), the set of subspaces confs obtained by evenly dividing according to the floating-point number distribution includes 2000 subintervals.

5. The input generation method for the difference in floating-point compilation optimization results according to claim 3, wherein The determination that the floating-point instruction sequences tf1 and tf2 are different in step S206 includes: calculating the difference diff(tf1, tf2) between the floating-point instruction sequences tf1 and tf2. If the difference diff(tf1, tf2) is not equal to 0, it is determined that the floating-point instruction sequences tf1 and tf2 are different.

6. The input generation method for the difference in floating-point compilation optimization results according to claim 5, characterized in that The difference diff(tf1, tf2) between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 includes: First, calculate the set of matching sequences M between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2. Then, multiply the number of instructions in the set of matching sequences M by 2 and divide it by the sum of the number of instructions in the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 to obtain the similarity ratio ratio(tf1, tf2). Then, take the result of 1.0 - ratio(tf1, tf2) as the calculated difference diff(tf1, tf2).

7. The input generation method for the difference in floating-point compilation optimization results according to claim 6, characterized in that The set of matching sequences M between the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 includes: S301, Add the sequence pair (tf1, tf2) composed of the floating-point instruction sequence tf1 and the floating-point instruction sequence tf2 to the set QS to be calculated, and initialize the set of matching sequences M to be empty; S302, If the set QS to be calculated is not empty, take out a sequence pair (t1, t2) from it, then calculate the longest matching sequence ms between the floating-point instruction sequence t1 and the floating-point instruction sequence t2 in it, and add the longest matching sequence ms to the set of matching sequences M, then jump to execute S303; Otherwise, determine that the calculation of the set of matching sequences M is completed; S303, deduct the longest matching sequence ms from the floating-point instruction sequence t1, It represents the floating-point instruction sequence t1 minus the left part of the longest matching sequence ms. represents the right part of the floating-point instruction sequence t1 after deducting the longest matching sequence ms; remove the longest matching sequence ms from the floating-point instruction sequence t2, It represents the floating-point instruction sequence t2 minus the left part of the longest matching sequence ms. It means the right part after deducting the longest matching sequence ms from the floating-point instruction sequence t2; if the left part after deducting the longest matching sequence ms from the floating-point instruction sequence t1 The floating-point instruction sequence t2 minus the left part of the longest matching sequence ms Are not empty, the sequence Add to the set to be calculated QS; if the floating-point instruction sequence t1 deducts the right part of the longest matching sequence ms The right part of the floating-point instruction sequence t2 minus the longest matching sequence ms Are not empty, the sequence Add to the to-be-computed set QS; jump to S302.

8. The input generation method for the difference in floating-point compilation optimization results according to claim 1, characterized in that, Step S103 includes: S401, Initialize the list variation_list to be empty; S402, Determine whether it is true that the set tmp_list to be refined is not empty. If it is not true, directly take the input I corresponding to the largest execution result res in the list variation_list as the input that can trigger the largest difference in the compilation optimization result obtained by searching, and end and exit; Otherwise, jump to S403; S403, retrieve a sample I from the input set tmp_list to be refined i , take the calculation function FP0 of the compilation optimization result difference as the fitness function fit, and start Markov chain Monte Carlo sampling from the sample I i , find a sample I with a larger fitness value through Markov chain Monte Carlo sampling t , record the calculation result difference res obtained by executing the driver program FP0 for the sample I t t t , append the result pair (I t , res t t ) composed of the input I and its corresponding execution result res to the list variation_list, and return to S402​​​ 9. An input generation system for the difference in floating-point compilation optimization results, comprising a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the input generation method for the difference in floating-point compilation optimization results according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program therein, characterized in that, The computer program is used to be programmed or configured by the microprocessor to execute the input generation method for the difference in floating-point compilation optimization results according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Compilation method supporting multi-format semi-precision floating points

    CN114217804A

  • Method and device for selecting compiler way in operating time

    CN1255674A