Automatic program optimization method and system for synthesizing and eliminating intermediate data structure by using induction program

By summarizing the program synthesis method, the intermediate data structure is automatically eliminated, which solves the problem of insufficient expression ability of the existing technology when eliminating intermediate data structures and achieves more efficient program optimization effects.

CN120832167APending Publication Date: 2025-10-24PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410499005.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies lack expressiveness when eliminating intermediate data structures and are unable to handle most optimization tasks, especially those involving complex recursion and large-scale computation.

Method used

The inductive program synthesis method is adopted to obtain the intermediate data structure of user input, rewrite the relevant program fragments using efficient program fragments with constant time complexity, and automatically optimize the program to eliminate the intermediate data structure by combining program analysis, iterative calculation and quantifier elimination method.

Benefits of technology

It significantly improves the expressive ability of eliminating intermediate data structures, can solve more optimization tasks than existing methods, and improves the efficiency and effectiveness of program optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832167A_ABST
    Figure CN120832167A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic program optimization method and system for synthesizing and eliminating an intermediate data structure by utilizing an induction program. The method comprises the following steps: acquiring a low-efficiency program which is input by a user and contains an intermediate data structure, and appointing the intermediate data structure which needs to be eliminated; a specified intermediate data structure is replaced by a scalar type, and related program fragments in the original program are rewritten using efficient program fragments of constant time complexity. The invention provides a novel induction program synthesis method aiming at the problem of eliminating an intermediate data structure, the induction program synthesis problem can be efficiently decomposed into simpler subtasks, so that the confronted efficiency problem is overcome, and the method has stronger expression ability than an existing fusion method, and can be widely applied to the field of data fusion. And the time for a programmer to carry out efficient system development can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer technology and program optimization, and specifically relates to an automatic program optimization method and system for eliminating intermediate data structures by using inductive program synthesis. BACKGROUND

[0002] Intermediate data structures are a common factor leading to low efficiency of functional programs. As an example, consider a classic algorithmic problem, maximum suffix sum (abbreviated as mts). Given a list of integers as input, the goal of this problem is to select an optimal suffix of the input list to maximize the sum of all elements in this suffix. For example, when the input list is [1, -2, 3, -1, 2], the expected output of maximum suffix sum is 4, which corresponds to the sum of all elements in the suffix [3, -1, 2] of length 3.

[0003] If a functional language is used to write a program solving the maximum suffix sum problem, a standard way of writing this program is to combine a series of common list manipulation functions, such as: mts xs = maximum (map sum (tails xs)). Here, the function tails returns a nested list containing all suffixes; the function map applies the function sum to each suffix, thus returning a list containing all suffix sums; and finally the function maximum returns the maximum value among all suffix sums, which is the expected output of the maximum suffix sum problem.

[0004] The above program has the advantages of simplicity and ease of implementation: all four functions used in this program can be found in a common list function library (such as the Data.List library in the Haskell language). However, the disadvantage of this program is its low efficiency. Given a list of length n, the time complexity of this program is O(n 2 ), which is unacceptable in performance-sensitive scenarios. The reason for the low efficiency of this program is the intermediate data structure built by the function tails, i.e., the nested list containing all suffixes. The size of this data structure is already O(n 2 ), so the above program needs to spend O(n 2 ) time to calculate the maximum suffix sum from this data structure.

[0005] In order to eliminate the damage of intermediate data structure to program efficiency, how to automatically eliminate the intermediate data structure becomes an important problem in the field of program optimization. The existing methods (named fusion methods) in this problem are all based on deductive reasoning. Specifically, the fusion methods rely on domain experts to pre-design a series of syntax-based rewriting rules. In each optimization process, they will constantly try to apply their rewriting rules to the original program, thereby gradually rewriting the original program into an optimized program that uses less intermediate data structure. However, these methods have a significant defect in expressiveness: they can only handle optimization tasks that can be described by rewriting rules. On the contrary, many optimization tasks that need to eliminate intermediate data structures need to take advantage of the specific properties of specific tasks. These specific properties are difficult to predict, so even domain experts are difficult to define these specific properties into rewriting rules in advance. It is found in the evaluation process that the best fusion method can only handle less than half of the optimization tasks in the data set. SUMMARY

[0006] In view of the above problems, the present application provides an automatic method and system for eliminating intermediate data structure with significantly stronger expressiveness and high efficiency.

[0007] The technical solutions adopted by the present application are as follows:

[0008] An automatic program optimization method for eliminating intermediate data structure by using inductive program synthesis, comprising the following steps:

[0009] Obtaining an inefficient program containing intermediate data structure input by a user, and specifying the intermediate data structure to be eliminated;

[0010] Replacing the specified intermediate data structure with a scalar type, and rewriting the related program fragments in the original program using an efficient program fragment with constant time complexity.

[0011] Further, the specified intermediate data structure to be eliminated is specified using a keyword, and the keyword is Packed.

[0012] Further, the rewriting of the related program fragments in the original program using the efficient program fragment with constant time complexity comprises:

[0013] Obtaining the related program fragments in the original program by using program analysis technology;

[0014] Synthesizing the optimized program by using program synthesis technology;

[0015] The program synthesis technique comprises: extracting local specifications about all rewritten program fragments by tracking the running process of the input program; decomposing the local specifications into specifications involving only single program fragments by iterative computation; further decomposing the decomposed specifications to obtain simplified specifications by quantifier elimination; and synthesizing the target program from the simplified specifications.

[0016] Further, the local specifications are extracted by the following steps:

[0017] Random inputs are generated, and all input-output behaviors of the relevant program fragments in the original program on the random inputs are collected;

[0018] The collected input-output behaviors are converted into local specifications involving only the program to be synthesized by using auxiliary programs.

[0019] Further, the iterative computation comprises:

[0020] Starting from the auxiliary program with empty output, each program fragment is iteratively verified against the current auxiliary program, and more outputs are incrementally synthesized for the auxiliary program when the verification fails until a legal auxiliary program is obtained.

[0021] Further, the quantifier elimination method comprises:

[0022] The existential quantifier of the program fragment with constant time complexity in the specification is eliminated by using the following two properties: (1) the program fragment with constant time complexity can only access constant-size input; (2) almost all scalar operations can be completed in constant time.

[0023] Further, the synthesis of the target program from the simplified specifications comprises:

[0024] The intermediate data structure in the local specification is replaced by the corresponding scalar value by using the auxiliary program, so that the local specification is simplified to a "example programming" problem, and an existing solver is used for solving.

[0025] An automatic program optimization system for eliminating intermediate data structures by inductive program synthesis comprises an input end and an output end; at the input end, an inefficient program containing intermediate data structures input by a user is obtained, and the intermediate data structures to be eliminated are specified; at the output end, the specified intermediate data structures are replaced by scalar types, and the relevant program fragments in the original program are rewritten by using program fragments with constant time complexity.

[0026] The beneficial effects of the present application are as follows:

[0027] In the design stage, in order to overcome the expression ability defect of the fusion method, another technical route different from the deductive reasoning is adopted, that is, the inductive program synthesis. Specifically, the technical feature of the inductive program synthesis is to search for a legal program in an infinite program space. In theory, it can solve any optimization task as long as the optimized program exists in the program space. Therefore, the inductive program synthesis has a significantly stronger expression ability than the deductive reasoning method.

[0028] However, due to efficiency problems, the existing inductive program synthesis technology cannot be directly used to eliminate the intermediate data structure. Specifically, the existing inductive program synthesis method can only generate small-scale programs at the expression level, while in the problem of eliminating the intermediate data structure, complex recursion and large-scale calculation are often involved, which cannot be handled by the existing inductive program synthesis method. In order to overcome the efficiency problem, the present application proposes a novel inductive program synthesis method for the problem of eliminating the intermediate data structure. The present application can efficiently decompose the inductive program synthesis problem into simpler subtasks, thereby overcoming the efficiency problem it faces.

[0029] In the problem of eliminating the intermediate data structure, the method provided by the present application has a significantly stronger expression ability than the existing fusion method. Evaluation shows that the method provided by the present application can solve at least 76% more optimization tasks than the existing fusion method. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 . The step flow chart of the method of the present application.

[0031] Figure 2 . The input program.

[0032] Figure 3 . The output program.

[0033] Figure 4 . The relevant program segment in the input program.

[0034] Figure 5 . The input and output behavior of the code segment t2 when the input is list [2].

[0035] Figure 6 . The reduction of?compress1.

[0036] Figure 7 . The reduction of?compress2.

[0037] Figure 8 . The split reduction.

[0038] Figure 9 . The reduction after eliminating?comb. DETAILED DESCRIPTION

[0039] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention will be further described in detail below by taking the maximum suffix and this problem as an example.

[0040] An automatic program optimization system takes a user-provided, less efficient program as input and automatically generates a more efficient program with the same behavior as the input and output. The system can also require the user to provide additional information (e.g., the const keyword in C++) for optimization purposes.

[0041] The present invention provides an automatic program optimization method using inductive program synthesis to eliminate intermediate data structures, such as Figure 1 As shown, the following steps are included:

[0042] Get user input of inefficient programs that contain intermediate data structures and specify the intermediate data structures that need to be eliminated;

[0043] Replace the specified intermediate data structure with a scalar type and rewrite the relevant program fragments in the original program using efficient program fragments with constant time complexity.

[0044] Corresponding to the above method, the present invention also provides a program optimization system for automatically eliminating intermediate data structures. The input and output of the system are described as follows:

[0045] On the input side, the user provides an inefficient program that contains intermediate data structures and uses the keyword Packed to specify the intermediate data structures that need to be eliminated. Figure 2 Shows a possible input to the maximum suffix sum problem.

[0046] Figure 2 The upper half shows a part of a list function library, which includes the definition of the data structure list (List) and nested list (NList), as well as the implementation of the function tails. Specifically, the function tails accepts a list as input and outputs a nested list containing all the suffixes of the input list. The implementation of this function is divided into two cases: when the input list is empty (ie Nil), the result is a nested list containing only empty lists; otherwise, tails will recursively calculate all the suffixes of the tail list of the input list (ie tails t), and then insert the input list itself into the beginning of the recursive result. For convenience, it is assumed that the function library also contains other common list functions, such as maximum, map, sum, etc., but the specific implementation of these functions is omitted.

[0047] Figure 2The lower half of the figure shows the inefficient program for computing the maximum suffix sum provided by the user. In this program, the output of the function tails is marked as Packed, indicating that the output of tails (i.e., the nested list containing all suffixes) is an intermediate data structure that needs to be eliminated.

[0048] At the output, the system replaces all the specified intermediate data structures with scalar types (e.g., integer, Boolean, and tuple types) and rewrites the original program fragments with efficient program fragments of constant time complexity. Figure 3 The output program for the maximum suffix sum example is shown. Compared with the input program, Figure 2 In the output program, the output of the function tails' is rewritten as a scalar tuple consisting of the maximum suffix sum and the list elements. Meanwhile, the system also rewrites three operations related to the intermediate data structure (the red part in Figure 3 from top to bottom, corresponding to (1) the initialization of the function tails when it takes an empty list, (2) the computation of the complete result from the recursive results of tails when it takes a non-empty list, and (3) the computation of the maximum suffix sum from the results of tails by the input function mts.

[0049] To automatically replace the user-specified intermediate data structures and rewrite the related program fragments, the invention first applies existing program analysis techniques to obtain all the related computations in the source program, i.e., the related program fragments. Figure 4 The analysis result on the maximum suffix sum example is shown, where three related computations are marked as?t1,?t2, and?t3. The "?" indicates that the system will generate (or rewrite) the program fragments. Then, the invention uses an innovative program synthesis technique to synthesize the optimized program. The program synthesis technique can be divided into four stages.

[0050] A. Extract local specifications about all the rewritten program fragments by tracing the execution of the input program.

[0051] B. Decompose the local specifications into specifications involving only a single program fragment by iterative computation.

[0052] C. Further decompose the decomposed specifications to obtain simplified specifications by quantifier elimination.

[0053] D. Synthesize the target program from the simplified specifications using existing program synthesis techniques.

[0054] The innovative part of the invention, i.e., the program synthesis part, will be introduced in detail below.

[0055] A. Extract local specifications

[0056] After the relevant program fragment (i.e. Figure 4 ) is analyzed, the system randomly generates a series of lists, runs the original program (i.e. input program) with these lists as input, and records the input-output behavior of the relevant program fragment during the running process. For example, when the input list is [2], the input-output behavior of the relevant program fragment?t2 can be collected as shown in Figure 5 , where the blue color represents the intermediate data structure that needs to be eliminated.

[0057] In order to obtain the correct optimized program (i.e. semantically equivalent to the original program), the code fragment for rewriting should have consistent input-output behavior with the corresponding relevant program fragment. For example, according to the input-output behavior collected in Figure 5 , the rewriting result of?t2 should be [[2], []] when the input list is [2] and the replaced [[]].

[0058] In order to formally describe this restriction, the system introduces an auxiliary unknown program?compress. By definition,?compress receives the intermediate data structure in the original program as input and needs to return the corresponding replacement result. In the example of maximum suffix sum, the expected?compress program should take the list containing all suffixes (i.e. the intermediate data structure in Figure 2 ) as input and produce a tuple consisting of the corresponding maximum suffix sum and the list elements (i.e. the replaced data structure in Figure 3 ). An implementation of this program is shown below:

[0059] ?compress ts = (maximum (map sum ts), sum (head ts))

[0060] After introducing the auxiliary program?compress, Figure 5 , the input-output behavior can be converted into local reductions about?compress and?t2, as shown below:

[0061] ?t2 ([2],?compress [[[]]) =?compress [[2], []]

[0062] In actual running process, the system will generate a large number of random inputs, collect all input-output behaviors of the relevant program fragment in the original program on these inputs, and convert all these input-output behaviors into local reductions. In this way, the system simplifies the original correctness reduction about the entire optimized program (i.e. semantically equivalent to the original program) to a local reduction involving only the program to be synthesized (i.e. the code fragment for replacement and the?compress program).

[0063] B. Iterative synthesis

[0064] After obtaining the local specifications, the system will first synthesize the helper program?compress iteratively. Specifically, the system will maintain a candidate helper program (referred to as the candidate program) and iteratively update it according to the local specifications. Initially, the candidate program is the trivial program?compress ts = () that does not preserve any values. Then, in each round, the system will check whether the current candidate program provides enough information for the replacement code snippet to complete the computation. If not, the system will extend the output of?compress with those insufficient local specifications. Below, we will still use the maximum suffix sum as an example to show the process of this iterative synthesis in detail.

[0065] First, in the first round of iteration, by substituting the candidate program?compress ts = () into the local specifications, the system can find that the information provided by the candidate program cannot make?t3 produce the expected output. Specifically, suppose that the input list [2] and [2, -1] are considered when extracting the local specifications. Then, we can obtain the following local specifications for?t3.

[0066] ?t3(?compress[[2], []]) = 2?t3(?compress[[2, -1], [2], []]) = 1

[0067] Substituting the current candidate program, the above left-hand side specification requires?t3 to output 2 when the input is (), while the right-hand side specification requires?t3 to produce the same result under the same input. Here, a contradiction is derived because an arbitrary program (without considering randomness and nondeterminism) can only produce the same output for the same input.

[0068] After discovering the contradiction, the system will try to provide more inputs for?t3 based on the candidate program to eliminate the contradiction in the local specifications of?t3. Specifically, the system will synthesize a supplementary program?compress1 (where the ellipsis corresponds to other collected local specifications for?t3) from the specification shown in Figure 6 which represents the additional inputs that the local specifications of?t3 require on top of the candidate?compress.

[0069] We will introduce the specific process of solving this specification in subsection C. Here, we only show the expected synthesis result of the system, which is shown below. This program takes the original intermediate data structure (containing the list of all suffixes) as input and returns the corresponding maximum suffix sum.

[0070] ?t3(ts) = maximum(mapsum ts)

[0071] By incorporating this program into the original candidate program, a new candidate program is obtained that produces the maximum suffix sums. The system then proceeds to the second round of iteration and examines this new candidate program.

[0072] In the second round of iteration, the system can find that the information provided by the current candidate program does not allow?t2 to produce the expected output. Specifically, if the input lists [2] and [2, -1] are considered when extracting the local specification, the following local specification for?t2 can be obtained.

[0073] ?t2([2],?compress[[2], []]) =?compress[[2], [], 2]

[0074] ?t2([2, -1],?compress[[-1], []]) =?compress[[2, -1], [-1], [], 1]

[0075] Substituting the current candidate program, the above specification requires?t2 to return the maximum suffix sum of the list [2] (i.e., 2) when the input is ([2], 0) and the maximum suffix sum of the list [2, -1] (i.e., 1) when the input is ([2, -1], 0). At this point, because the second component of the input is the same but the expected output is different,?t2 must compute the maximum suffix sum of the entire array from the first input (i.e., the entire input list) to the second input. It can be shown that there does not exist any constant-time expression that can accomplish this computation, thus creating a contradiction with the constant-time complexity requirement of?t2.

[0076] Similar to the first round of iteration, after discovering the contradiction, the system attempts to add more inputs to?t2 to eliminate the contradiction in its local specification. Specifically, the system synthesizes the supplementary program?compress2 from the specification shown in Figure 7

[0077] Here, the system expects the synthesized result to be?compress2 ts = sum(head ts), which takes the original intermediate data structure as input and returns the corresponding list sum. By incorporating this result into the candidate program obtained from the first round, the expected candidate program can be obtained, as shown below, where the left and right components come from the first and second rounds of iteration, respectively.

[0078] ?compress ts = (maximum(map sum ts), sum(head ts))

[0079] Finally, the system still performs a new round of iteration with the above?compress as the candidate program. At this point, there is no longer any contradiction, meaning that the current candidate program is a legal candidate program.​

[0080] C. Quantifier Elimination

[0081] Next, this article will introduce how this system Figure 6 and Figure 7 The main challenge of this synthesis task is to efficiently synthesize ? compress1 and ? compress2 in the specification shown. and Existing program synthesis techniques are unable to efficiently handle this type of program synthesis problem. To address this challenge, our system leverages the efficiency requirements of ?t2 and ?t3 (i.e., the time complexity must be constant) to eliminate second-order existential quantifiers, thereby simplifying these originally difficult synthesis tasks into simpler ones that can be solved by existing program synthesis techniques.

[0082] Here, this article will take compress2 as an example (the specification is as follows Figure 7 ) shows the technical details of quantifier elimination.

[0083] Note that t2 also has an efficiency requirement here: to ensure that an efficient program is ultimately obtained, this system restricts all program fragments used for rewriting to have a constant time complexity.

[0084] To eliminate the second-order existential quantifiers on ?t2, this system first decomposes ?t2 according to its efficiency requirements. Specifically, any expression with constant time complexity can only access constant-sized inputs and can therefore be equivalently written as follows.

[0085] ? tip:=? comb(?extract inp)

[0086] Extract accepts the entire input and returns a tuple of scalars in constant time, representing all inputs accessible to the expression t. Combination, on the other hand, accepts only the output of extract and completes the final computation in constant time. For example, in the above t2 reduction, the expected result is compress2 ts = sum(headts). The corresponding t2 and split results are shown below.

[0087] ? t2(xs, (mts, sum)):=(maxmts(headxs+sum), headxs+sum)

[0088] ? extract(xs, (mts, sum)): = (head xs, mts, sum)

[0089] comb(head, mts, sum) := (max mts (head + sum), head + sum)

[0090] Applying this splitting, the expression for?t2 can be expanded into the form shown in Figure 8 , i.e., the reduced form after splitting.

[0091] Each row in this reduced form presents the input-output behavior of an unknown procedure?comb, e.g., in the first and second rows,?comb is required to produce outputs 2 and 1 respectively under the same input. Note that no matter what the final value of?comb is,?comb cannot produce different results under the same input. Therefore, a necessary condition for?comb in the above reduced form is that for any two rows, if their outputs of?comb are different, their inputs of?comb must also be different. Using this property,?comb in the above reduced form can be eliminated to the form shown in Figure 9 , i.e., the reduced form after eliminating?comb.

[0092] From the analysis in the last paragraph, it is known that Figure 9 forms a necessary condition for the reduced form in Figure 8 . In fact, the effect of this reduced form in practice is close to being a sufficient condition. Specifically, the only constraint on the value range of?comb is constant time. However, because the input-output types of?comb are both scalars, and most scalar functions can be implemented in constant time, almost all possible input-output behaviors can be produced within the value range of?comb. The exception here is some complex scalar functions, such as the prime number test, prime factorization, and other number theory functions. However, these functions are almost impossible to appear in the context of program optimization.

[0093] Based on the above sufficiency analysis, the system approximately uses the reduced form in Figure 9 as a sufficient and necessary condition for the reduced form in Figure 8 , and then uses the reduced form in Figure 9 to synthesize?compress2. This reduced form involves two unknown procedures,?extract and?compress2. One common feature of these two procedures is that their target procedure sizes are often small.

[0094] First,?compress2 is only introduced by the system to simplify the program synthesis problem. It will not appear in the result program, so its running efficiency is not important, and it can be implemented very compactly by combining library functions, like the input procedure in the maximum suffix sum problem.

[0095] Second, the role of?extract is to extract the scalar value from the input, which does not need to do any complex calculation itself. Therefore, its size is usually small.

[0096] Because Figure 9 Because the unknown programs involved in the invariants are small, the system can use an enumeration-based program synthesis method to synthesize these unknown programs efficiently. Specifically, the system enumerates all possible (?compress2,?extract) in the order of the size of the sum from small to large, checks whether they satisfy the invariant in turn, and takes?compress2 in the first group of programs that satisfy the invariant as the synthesis result. Figure 9

[0097] In summary, the steps of the quantifier elimination of the system are as follows: (1) use the property that the output type of the program fragment for rewriting (such as?t2) is scalar to split it into?extract and?comb, (2) use the feature that almost all scalar calculations can be completed in constant time to eliminate the part of?comb from the invariant, and (3) use the feature that?extract and the auxiliary program have small sizes to synthesize them using an enumeration-based program synthesis.

[0098] D. Application of existing program synthesis techniques

[0099] After successfully synthesizing the auxiliary function, the system will use the synthesis result to simplify the original local invariant into the input-output examples of the program fragment for rewriting, as follows.

[0100] ?t2([2],?compress[[[]]) =?compress[[2], []] ->?t2([2], (o, o)) = (2, 2)

[0101] ?compress ts = (maximum (mapsum ts), sum (head ts))

[0102] Synthesizing a program from input-output examples is a classic problem called "example programming". Here, because?t2 is a recursive-independent expression that mainly involves scalar calculations, there already exists an existing example programming solver that can synthesize it efficiently. In implementation, the system uses the example programming solver PolyGen to synthesize each program fragment for rewriting from the final input-output examples.

[0103] Evaluation of the invention:

[0104] Finally, in order to evaluate the effectiveness of the invention, the invention creates a data set with 290 optimization problems about the elimination of intermediate data structures. The sources of these data sets are as follows:​

[0105] Source 1: There are 16 optimization problems from existing research on fusion problems. The inventors considered five classic papers on fusion problems and collected all the sample problems in them as part of the dataset.

[0106] Source 2: There are 178 optimization problems from existing research on structural recursive synthesis. Specifically, the input of a structural recursive synthesis problem is two different data structure types A and B, a recursive program rec on type A, and a transformation program trans from type B to type A; and the goal of the problem is to synthesize a recursive program on type B whose input-output behavior is exactly the same as the function composition of rec and trans. This problem can be reduced to an intermediate data structure elimination problem, where the input program is the composition of rec and trans, and the intermediate data structure is the translation result of trans. According to this reduction, the inventors used the Synduce dataset on structural recursive synthesis problems as part of the evaluation.

[0107] Source 3: There are 97 optimization problems from existing research on algorithm synthesis. Specifically, the input of an algorithm synthesis problem is an algorithm template and a reference program; and the goal of the problem is to rewrite the algorithm template and obtain a program whose input-output behavior is exactly the same as the reference program. This problem can be regarded as an intermediate data structure elimination problem in the algorithm template. According to this connection, the inventors used the AutoLifter dataset on algorithm synthesis problems as part of the evaluation.

[0108] In the dataset of 290 optimization problems created by the invention, the invention solved 264 of them (91.0%) at an average speed of 38.08 seconds, far exceeding the prior art: the existing fusion system can at most solve only 150 problems (51.7%), and the existing general inductive program synthesis technology can only solve 85 problems (29.3%).

[0109] The key points of the invention include:

[0110] 1) Extract local reductions by introducing auxiliary functions?compress.

[0111] 2) Separate sub-problems about different program fragments by iterative synthesis.

[0112] 3) Efficiently solve the sub-problems decomposed by variable elimination and enumeration-based synthesis.

[0113] 4) Efficiently synthesize program fragments for rewriting by example programming techniques.

[0114] Practical application scenarios of the invention:

[0115] The main application scenario of the present application is compilation optimization. In the actual development process of a software system, program optimization is often an important source of overhead and risk. It not only requires programmers to have higher algorithm knowledge and to refactor most of the code, but also often reduces the readability of the program and destroys the modularity of the program, thereby greatly increasing the risk of defects. Through the present application, when optimization is needed, the programmer only needs to annotate the intermediate data structure at the bottleneck, and the present application can quickly eliminate these data structures to achieve automatic optimization, thereby greatly reducing the time of the programmer to develop an efficient system.

[0116] Another embodiment of the present application provides a computer device (computer, server, smart phone, etc.), which comprises a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing each step in the method of the present application.

[0117] Another embodiment of the present application provides a computer readable storage medium (such as ROM / RAM, magnetic disk, optical disk), which stores a computer program, and the computer program is executed by a computer to realize each step of the method of the present application.

[0118] The specific embodiments of the present application disclosed above are intended to help understand the content of the present application and to implement the same, and those skilled in the art can understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present application. The present application should not be limited to the content disclosed by the embodiments of the present application, and the protection scope of the present application is defined by the scope of the claims.

Claims

1. An automatic program optimization method for eliminating intermediate data structures using inductive program synthesis, characterized by, The method comprises the following steps: acquiring an inefficient program containing intermediate data structures input by a user and specifying intermediate data structures to be eliminated; replacing the specified intermediate data structures with scalar types and rewriting relevant program fragments in the original program using high-efficiency program fragments with constant time complexity.

2. The method of claim 1, wherein, The specified intermediate data structures to be eliminated are specified using a keyword, and the keyword is Packed.

3. The method of claim 1, wherein, The rewriting of the relevant program fragments in the original program using high-efficiency program fragments with constant time complexity comprises: obtaining relevant program fragments in the original program using program analysis techniques; synthesizing an optimized program using program synthesis techniques. The program synthesis techniques comprise extracting local specifications about all rewritten program fragments by tracing the running process of the input program, decomposing the local specifications into specifications involving only a single program fragment by iterative calculation, further decomposing the decomposed specifications into simplified specifications using a quantifier elimination method, and synthesizing a target program from the simplified specifications.

4. The method of claim 3, wherein, The local specifications are extracted using the following steps: generating random inputs and collecting all input-output behaviors of the relevant program fragments in the original program under the random inputs; converting the collected input-output behaviors into local specifications involving only the program to be synthesized using auxiliary programs.

5. The method of claim 3, wherein, The iterative calculation comprises: starting from an auxiliary program with no output, iteratively checking the current auxiliary program for each program fragment, and incrementally synthesizing more outputs for the auxiliary program when the verification fails until a legal auxiliary program is obtained.

6. The method of claim 3, wherein, The quantifier elimination method comprises: eliminating existential quantifiers about program fragments with constant time complexity in the specifications using the following two properties: (1) a program fragment with constant time complexity can only access constant-size inputs; (2) almost all scalar operations can be completed within constant time.

7. The method of claim 3, wherein, The synthesis of the target program from the simplified specifications comprises: replacing intermediate data structures in the local specifications with corresponding scalar values using auxiliary programs, thereby simplifying the local specifications into "example programming" problems, and solving the problems using existing solvers.

8. An automatic program optimization system that eliminates intermediate data structures using inductive program synthesis, characterized by, The device comprises an input end and an output end; at the input end, an inefficient program containing intermediate data structures input by a user is acquired, and intermediate data structures to be eliminated are specified; at the output end, the specified intermediate data structures are replaced with scalar types, and relevant program fragments in the original program are rewritten using high-efficiency program fragments with constant time complexity.

9. A computer device, comprising: The device comprises a memory and a processor, the memory stores a computer program configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by a computer to implement the method according to any one of claims 1-7.