Code compiling optimization method and device based on algorithm representation, equipment and medium
Through the algorithm characterization method, code compilation optimization items are automatically selected, which solves the low efficiency problem of traditional methods and achieves efficient and accurate code compilation optimization.
Patent Information
- Application Number
- CN202510672525.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional code compilation methods are inefficient in selecting the best combination of compilation optimization items and require a lot of computing costs and manual intervention, resulting in low compilation optimization efficiency.
Through an algorithm-based characterization method, the target source code file is obtained, problem instantiation and random sampling are performed, and candidate compilation optimization models are screened and fine-tuned using similarity and true solution quality scores. The target compilation optimization model is generated, and the optimal compilation optimization items are automatically selected.
It improves the efficiency and accuracy of code compilation and optimization, reduces computing costs, and realizes an automated compilation and optimization process without human intervention.
Smart Images

Figure CN120687097A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of code compilation and testing technology, and in particular to a code compilation optimization method and apparatus, device, and medium based on algorithm characterization. Background Art
[0002] Currently, traditional code compilation methods typically use a compiler to first generate binary code from the source code, and then perform code optimization processing to obtain computer-executable binary code. For example, the GCC (GNU Compiler Collection) compiler scans a program written in C language, first compiling the source code into initial binary code. This initial binary code is then optimized based on GCC's internal compilation optimization options to reduce the code size, ultimately producing executable binary code that can run on Linux systems. However, since GCC provides over 200 compilation optimization options (such as conditional statement optimization, loop structure optimization, and expression optimization), each compilation optimization option may affect the final compiled binary code. Different source codes require different compilation optimization options. To find the optimal compilation optimization option combination, the compiler needs to evaluate each possible compilation optimization option combination. How to select the best compilation optimization option combination among these numerous compilation optimization option combinations has become a challenging research problem.
[0003] To address this issue, the search for the optimal combination of compilation optimization options is typically treated as a black-box optimization problem. By combining functions from evolutionary algorithms with evaluation criteria based on expert experience, each combination of compilation optimization options is evaluated to determine if it is the optimal solution. However, this compilation optimization process requires compilation performance testing for each different combination of compilation optimization options. Each performance test relies on a function for evaluation, resulting in high computational costs. Furthermore, manual intervention in the compilation optimization evaluation is required, resulting in low code compilation optimization efficiency. Therefore, improving the efficiency of code compilation optimization has become a pressing issue. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a code compilation optimization method and device based on algorithm representation, an electronic device and a storage medium, aiming to improve the efficiency of code compilation optimization.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application proposes a code compilation optimization method based on algorithm representation, the method comprising:
[0006] Obtaining a target source code file, and compiling the target source code file to obtain an initial binary file;
[0007] Instantiate the problem according to at least two code logic blocks in the initial binary file to obtain an optimization problem instance;
[0008] Performing a first random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions;
[0009] Calculating similarities between the plurality of candidate compilation optimization models and the optimization problem instance based on the at least two first reference problem solutions, and screening the plurality of candidate compilation optimization models based on the similarities to obtain an initial compilation optimization model; wherein the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance;
[0010] Obtaining true solution quality scores of the at least two first reference problem solutions, and fine-tuning the initial compilation optimization model based on the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model; wherein the target compilation optimization model is used to represent the fine-tuned intelligent algorithm;
[0011] Performing a second random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two second reference problem solutions;
[0012] Reconstructing each of the at least two second reference problem solutions using the target compilation optimization model to obtain a target reconstructed solution; the target reconstructed solution is used to indicate whether to enable or disable a candidate compilation optimization item;
[0013] A target compilation optimization item is determined from the plurality of candidate compilation optimization items according to the target reconstruction solution, and compilation optimization is performed on the initial binary file according to the target compilation optimization item to generate a target binary file.
[0014] In some embodiments, the target compilation optimization model includes a target encoder, a target reconstructor, and a target scorer;
[0015] The reconstructing each of the at least two second reference problem solutions using the target compiled optimization model to obtain a target reconstructed solution includes:
[0016] encoding each of the at least two second reference problem solutions by the target encoder to obtain a second problem solution vector;
[0017] Performing a quality score on the second problem solution vector using the target scorer to obtain a second problem solution quality score;
[0018] Screening the at least two second reference problem solutions according to the quality score of the second problem solution to obtain a retained problem solution;
[0019] The target reconstructor is used to reconstruct the retained problem solution to obtain the target reconstructed solution.
[0020] In some embodiments, screening multiple candidate compilation optimization models based on the similarity to obtain an initial compilation optimization model includes:
[0021] sorting the plurality of candidate compilation optimization models according to the similarities to obtain sorted candidate compilation optimization models;
[0022] The candidate compilation optimization model with the highest similarity is selected from the ranked candidate compilation optimization models and determined as the initial compilation optimization model.
[0023] In some embodiments, before screening multiple candidate compilation optimization models based on the similarity to obtain an initial compilation optimization model, the method further includes:
[0024] Obtaining the initial candidate compilation optimization model; the initial candidate compilation optimization model includes a candidate encoder, a candidate reconstructor, and a candidate scorer;
[0025] Encoding each of the at least two first reference problem solutions using the candidate encoder to obtain a candidate first problem solution vector;
[0026] Reconstructing the candidate first problem solution vector using the candidate reconstructor to obtain a candidate reconstructed solution;
[0027] Using the candidate scorer to perform a quality score on the candidate first problem solution vector to obtain a candidate problem solution quality score;
[0028] Obtaining a true reconstruction solution corresponding to each of the at least two first reference problem solutions, and calculating target loss values of the true reconstruction solution, the candidate reconstruction solutions, the quality scores of the candidate problem solutions, and the quality score of the true solution according to a preset target loss function;
[0029] The candidate model parameters of the candidate compilation optimization model are updated based on the target loss value until the target loss value meets a preset model training condition, thereby obtaining the candidate compilation optimization model.
[0030] In some embodiments, the initial compilation optimization model includes an initial encoder and an initial reconstructor;
[0031] The fine-tuning of the initial compilation optimization model according to the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model includes:
[0032] Encoding each of the at least two first reference problem solutions by the initial encoder to obtain an initial first problem solution vector;
[0033] Reconstructing the initial first problem solution vector by the initial reconstructor to obtain an initial first reconstructed solution;
[0034] constructing a problem solution data pair for the initial first reconstructed solution corresponding to each of the at least two first reference problem solutions at the same quality level according to the true solution quality score;
[0035] The mean square error value of the problem solution data pair is calculated according to a preset fine-tuning objective function, and the reconstructor parameters of the initial reconstructor are updated based on the mean square error value until the mean square error value meets the preset fine-tuning training conditions, thereby obtaining the target compiled optimization model.
[0036] In some embodiments, the initial compilation optimization model further includes an initial scorer;
[0037] The constructing of a problem solution data pair for the initial first reconstructed solution corresponding to each of the at least two first reference problem solutions at the same quality level according to the true solution quality score includes:
[0038] Performing a quality score on the initial first problem solution vector by the initial scorer to obtain an initial solution quality score;
[0039] Performing a quality grade classification on the initial first reconstructed solution according to the initial solution quality score to obtain a reconstructed solution quality grade;
[0040] Performing a quality grade classification on each of the at least two first reference problem solutions according to the true solution quality score to obtain a first problem solution quality grade;
[0041] The problem solution data pair of the initial first reconstructed solution corresponding to each of the at least two first reference problem solutions at the same quality level is constructed based on the reconstructed solution quality level and the first problem solution quality level.
[0042] In some embodiments, the calculating the similarity between the plurality of candidate compilation optimization models and the optimization problem instance based on the at least two first reference problem solutions includes:
[0043] Inputting each of the at least two first reference problem solutions into the candidate compilation optimization model for quality scoring to obtain a predicted solution quality score;
[0044] Obtaining a true solution quality score corresponding to each of the at least two first reference problem solutions, and calculating a Pearson correlation coefficient and a Spearman rank correlation coefficient between the true solution quality score and the predicted solution quality score;
[0045] The similarity between the candidate compilation optimization model and the optimization problem instance is determined according to the Pearson correlation coefficient and the Spearman rank correlation coefficient.
[0046] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application proposes a code compilation optimization device based on algorithm representation, the device comprising:
[0047] A code compilation module is used to obtain a target source code file, compile the target source code file, and obtain an initial binary file;
[0048] A problem instantiation module, configured to instantiate a problem according to at least two code logic blocks in an initial binary file to obtain an optimization problem instance;
[0049] A first problem solution sampling module is used to perform a first random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions;
[0050] a model screening module, configured to calculate similarities between the plurality of candidate compilation optimization models and the optimization problem instance based on the at least two first reference problem solutions, and screen the plurality of candidate compilation optimization models based on the similarities to obtain an initial compilation optimization model; wherein the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance;
[0051] a model fine-tuning module, configured to obtain true solution quality scores of the at least two first reference problem solutions, and fine-tune the initial compilation optimization model based on the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model; wherein the target compilation optimization model is used to represent the fine-tuned intelligent algorithm;
[0052] A second problem solution sampling module is used to perform a second random sampling on multiple candidate problem solutions of the optimization problem instance to obtain at least two second reference problem solutions;
[0053] a problem solution reconstruction module, configured to reconstruct each of the at least two second reference problem solutions using a target compilation optimization model to obtain a target reconstructed solution; the target reconstructed solution is used to indicate whether to enable or disable a candidate compilation optimization item;
[0054] The compilation optimization module is used to determine a target compilation optimization item from multiple candidate compilation optimization items according to the target reconstruction solution, and to compile and optimize the initial binary file according to the target compilation optimization item to generate a target binary file.
[0055] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0056] To achieve the above-mentioned purpose, a fourth aspect of an embodiment of the present application proposes a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method of the above-mentioned first aspect is implemented.
[0057] The code compilation optimization method and device, electronic device and storage medium based on algorithm representation proposed in the present application first compile the acquired target source code file to obtain an initial binary file, and instantiate the problem based on at least two code logic blocks in the initial binary file, which can clarify the optimization requirements and locate the code parts that need to be optimized; secondly, a first random sampling is performed on multiple candidate problem solutions of the optimization problem instance. Since each candidate compilation optimization model represents an intelligent algorithm that can solve the optimization problem instance, the candidate compilation optimization model is screened and fine-tuned by using the similarity and true solution quality score between the candidate compilation optimization model determined by the first reference problem solution and the optimization problem instance, thereby screening out the intelligent algorithm most relevant to the optimization problem instance, avoiding the use of irrelevant intelligent algorithms for invalid function evaluation, reducing the computational cost, and realizing the optimization of relevant models. Fine-tuning eliminates the need to adjust all intelligent algorithm parameters, helping to improve the efficiency of subsequent code compilation optimization. Furthermore, a second random sampling of multiple candidate problem solutions for the optimization problem instance can further explore the problem solution space, and reconstruct the second reference problem solution through the target compilation optimization model to obtain a target reconstructed solution. This can generate a higher-quality problem solution through the fine-tuned intelligent algorithm, improving the accuracy of finding a better combination of compilation optimization items. Finally, the target compilation optimization items are screened out through the target reconstructed solution, and the initial binary file is compiled and optimized according to the target compilation optimization items to generate a target binary file with optimal performance and size. This eliminates the need to perform compilation performance testing on each different combination of compilation optimization items, and also realizes the automation of the code compilation optimization process, eliminating the need for manual intervention in the code compilation process, significantly improving the efficiency and accuracy of code compilation. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flowchart of a code compilation optimization method based on algorithm representation provided by an embodiment of the present application;
[0059] Figure 2 is another flow chart of the code compilation optimization method based on algorithm representation provided in an embodiment of the present application;
[0060] Figure 3 yes Figure 1 Flowchart of step S104 in FIG.
[0061] Figure 4 yes Figure 3 Flowchart of step S301 in FIG.
[0062] Figure 5 yes Figure 1 Flowchart of step S105 in FIG.
[0063] Figure 6 yes Figure 5 Flowchart of step S503 in FIG.
[0064] Figure 7 yes Figure 1 Flowchart of step S107 in FIG.
[0065] Figure 8 is a schematic diagram of the deep learning model training process provided in an embodiment of the present application;
[0066] Figure 9 This is a schematic diagram of the structure of a code compilation optimization device based on algorithm representation provided in an embodiment of the present application;
[0067] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0069] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0071] The embodiments of the present application provide a code compilation optimization method and device based on algorithm representation, an electronic device and a storage medium, aiming to improve the efficiency of code compilation optimization.
[0072] The code compilation optimization method and device based on algorithm representation, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the code compilation optimization method based on algorithm representation in the embodiments of the present application is described.
[0073] The code compilation optimization method based on algorithmic representation provided in the embodiment of the present application relates to the field of code compilation test technology. The code compilation optimization method based on algorithmic representation provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the code compilation optimization method based on algorithmic representation, etc., but is not limited to the above forms.
[0074] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0075] Figure 1 This is an optional flowchart of the code compilation optimization method based on algorithm representation provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S108.
[0076] Step S101: Obtain a target source code file, compile the target source code file, and obtain an initial binary file.
[0077] Step S102 : instantiate the problem according to at least two code logic blocks in the initial binary file to obtain an optimization problem instance.
[0078] Step S103 : performing a first random sampling on multiple candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions.
[0079] Step S104 : Calculate similarities between multiple candidate compilation optimization models and the optimization problem instance based on at least two first reference problem solutions, and screen the multiple candidate compilation optimization models based on the similarities to obtain an initial compilation optimization model; wherein the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance.
[0080] Step S105 , obtaining the true solution quality scores of at least two first reference problem solutions, and fine-tuning the initial compilation optimization model based on the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model; wherein the target compilation optimization model is used to characterize the fine-tuned intelligent algorithm.
[0081] Step S106 : performing a second random sampling on multiple candidate problem solutions of the optimization problem instance to obtain at least two second reference problem solutions.
[0082] Step S107 , reconstructing each of the at least two second reference problem solutions using the target compilation optimization model to obtain a target reconstructed solution; the target reconstructed solution is used to indicate whether to enable or disable the candidate compilation optimization item.
[0083] Step S108 , determining a target compilation optimization item from a plurality of candidate compilation optimization items according to the target reconstruction solution, and compiling and optimizing the initial binary file according to the target compilation optimization item to generate a target binary file.
[0084] In the steps S101 to S108 shown in the embodiment of the present application, the target source code file is first compiled to obtain an initial binary file, and the problem is instantiated according to at least two code logic blocks in the initial binary file, so as to clarify the optimization requirements and locate the code parts that need to be optimized; secondly, a first random sampling is performed on multiple candidate problem solutions of the optimization problem instance. Since each candidate compilation optimization model represents an intelligent algorithm that can solve the optimization problem instance, the candidate compilation optimization model is screened and fine-tuned by using the similarity and true solution quality score between the candidate compilation optimization model determined by the first reference problem solution and the optimization problem instance, so that the intelligent algorithm most relevant to the optimization problem instance can be screened out, thereby avoiding the use of irrelevant intelligent algorithms for invalid function evaluation, reducing the computing cost, and realizing the fine-tuning of relevant intelligent algorithms. All intelligent algorithm parameters need to be adjusted, which helps to improve the efficiency of subsequent code compilation optimization; then, a second random sampling of multiple candidate problem solutions of the optimization problem instance can further explore the problem solution space, and reconstruct the second reference problem solution through the target compilation optimization model to obtain the target reconstructed solution. The fine-tuned intelligent algorithm can generate higher quality problem solutions, thereby improving the accuracy of finding better combinations of compilation optimization items; finally, the target compilation optimization items are screened out through the target reconstructed solution, and the initial binary file is compiled and optimized according to the target compilation optimization items to generate a target binary file with optimal performance and size. There is no need to perform compilation performance testing on each different combination of compilation optimization items, and the code compilation optimization process is automated, eliminating the need for manual intervention in the code compilation process, significantly improving the efficiency and accuracy of code compilation.
[0085] In step S101 of some embodiments, specifically, the target source code file refers to a code text written based on a high-level programming language (such as C language, C++ language, etc.), and the target source code file contains program logic and logic implementation functions.
[0086] For example, the target source code file may include but is not limited to linear algebra operation logic (such as matrix multiplication, eigenvalue solution, etc.) and differential equation solution logic, etc.
[0087] Specifically, the initial binary file refers to a binary code file executable by a computer, which also includes linear algebra operation logic and differential equation solving logic.
[0088] Specifically, the target source code file may be compiled using a GCC (GNU Compiler Collection) compiler to obtain an initial binary code file.
[0089] In step S102 of some embodiments, specifically, a code logic block refers to an algorithm code fragment in a numerical calculation in a program, and is usually composed of a group of code statements.
[0090] For example, the code logic block for the linear algebra operation code may be a matrix multiplication function and a Gaussian elimination function.
[0091] Specifically, an optimization problem instance is an optimization problem consisting of a specific source code, multiple candidate compilation optimization items of the compiler, and an optimization target.
[0092] For example, an optimization problem instance may consist of a code for computing matrix multiplication, 186 candidate compilation optimization items of the compiler GCC, and the optimization goal of minimizing the size of the compiled binary file.
[0093] Specifically, since different GCC compilation optimization items have different sizes of binary files generated by compiling source code, an optimization problem instance can be generated by combining different GCC candidate compilation optimization items with at least two code logic blocks.
[0094] In step S103 of some embodiments, specifically, the candidate problem solution refers to a decision variable corresponding to the optimization problem instance, which is used to indicate activation or deactivation of a candidate compilation optimization item corresponding to the candidate problem solution.
[0095] Specifically, in the compilation optimization problem of binary files, the decision variable is usually a 0 / 1 vector, that is, each decision variable is represented by a vector consisting of 0 or 1. For each candidate compilation optimization item of the compiler, when the value is assigned to 0, it means that the candidate compilation optimization item is not used and is turned off. When the value is assigned to 1, it means that the candidate compilation optimization item is used and turned on.
[0096] Specifically, since there are a large number of candidate compilation optimization items, by performing first random sampling on multiple candidate problem solutions, at least two candidate compilation items can be randomly sampled from the multiple candidate compilation optimization items as at least two first reference problem solutions.
[0097] Specifically, the number of samples of the solution to the first reference problem may be 64, which is not limited here.
[0098] For example, the first random sampling is to select the four candidate compilation optimization items of conditional statement optimization, loop structure optimization, expression optimization and function inline optimization. Each candidate compilation optimization item can be represented as a binary bit, such as [1,0,1,0], which means that the conditional statement optimization and expression optimization are enabled, but the loop structure optimization and function inline optimization are disabled.
[0099] In this embodiment, by performing a first random sampling on multiple candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions, the number of candidate problem solutions can be reduced, which helps to improve the efficiency of the subsequent code compilation optimization process.
[0100] See also Figure 2 In some embodiments, before step S104, the code compilation optimization method based on algorithm characterization further includes but is not limited to steps S201 to S206:
[0101] Step S201 : obtaining an initial candidate compilation optimization model; the initial candidate compilation optimization model includes a candidate encoder, a candidate reconstructor, and a candidate scorer.
[0102] Step S202 : Encode each of the at least two first reference problem solutions using a candidate encoder to obtain a candidate first problem solution vector.
[0103] Step S203: Reconstruct the candidate first problem solution vector using a candidate reconstructor to obtain a candidate reconstructed solution.
[0104] Step S204: Use a candidate scorer to perform quality scoring on the candidate first problem solution vector to obtain a quality score of the candidate problem solution.
[0105] Step S205, obtain the true reconstruction solution corresponding to each of the at least two first reference problem solutions, and calculate the target loss values of the true reconstruction solution, candidate reconstruction solution, candidate problem solution quality score and true solution quality score according to a preset target loss function.
[0106] Step S206 , updating the candidate model parameters of the candidate compilation optimization model based on the target loss value until the target loss value satisfies the preset model training condition, thereby obtaining the candidate compilation optimization model.
[0107] In step S201 of some embodiments, specifically, the initial candidate compilation optimization model refers to an untrained deep learning model, which includes a candidate encoder, a candidate reconstructor, and a candidate scorer; wherein the candidate encoder and the candidate reconstructor are VAE (Variational Autoencoder) networks; and the candidate scorer is a multi-layer perceptron structure.
[0108] Specifically, the initial candidate compilation optimization model is used to represent the candidate intelligent algorithms applicable to different optimization problem instances in the form of weight parameters in a neural network model, and solve the optimization problem instance based on the neural network model, that is, each initial candidate compilation optimization model corresponds to each candidate intelligent algorithm, and the candidate intelligent algorithm solves the optimization problem instance through the candidate compilation optimization model.
[0109] Specifically, the candidate encoder is used to implement encoding for the first reference problem solution to convert the first reference problem solution into a vector of fixed length; the candidate reconstructor is used to reconstruct the candidate first problem solution vector so that the model can learn the structural information of the first reference problem solution; the candidate scorer is used to perform quality scoring on the candidate first problem solution vector to evaluate the effectiveness of the first reference problem solution in solving the optimization problem instance.
[0110] In step S202 of some embodiments, specifically, the candidate first problem solution vector refers to information of the candidate first reference problem solution represented in vector form, including a mean vector and a standard deviation vector.
[0111] In step S203 of some embodiments, specifically, the candidate reconstructed solution refers to structural information of the problem solution learned by the model from the first reference problem solution.
[0112] In step S204 of some embodiments, specifically, the quality score of the candidate problem solution refers to the candidate score of the binary file size generated by compiling the first reference problem solution, that is, the quality score of the candidate problem solution can be the candidate fitness value of the candidate problem solution, and the candidate fitness value is used to characterize the quality of the candidate problem compilation. The larger the candidate fitness value, the better the quality of the candidate problem solution, and the smaller the binary file size compiled by the candidate problem solution.
[0113] For example, if the solutions to the first reference problem are [1,0,1,0] and [0,1,0,1], [1,0,1,0] means enabling conditional selection statement optimization and expression optimization, and disabling loop structure optimization and function inlining optimization; while [0,1,0,1] means enabling loop structure optimization and function inlining optimization, and disabling conditional selection statement optimization and expression optimization. Since the optimization goal is to generate the smallest binary file, the scoring criterion can be the size of the binary file generated after compilation. The smaller the file size, the higher the score. For the vector [1,0,1,0], the size of the binary file generated after compilation is 10KB, and the corresponding candidate problem solution quality score is 0.85. For the vector [0,1,0,1], the size of the binary file generated after compilation is 12KB, and the corresponding candidate problem solution quality score is 0.78.
[0114] In step S205 of some embodiments, specifically, the true reconstruction solution refers to the optimal reconstruction solution corresponding to the first reference problem solution learned by the known model.
[0115] Specifically, the true solution quality score refers to the actual score of the size of the binary file generated by compiling the solution to the first reference problem. The true solution quality score can be the actual true fitness value of the solution to the first reference problem, which is used to characterize the quality of the actual compilation of the solution to the first reference problem.
[0116] Specifically, the objective loss function can be expressed by the following formula:
[0117] min∑MSE(x,x ′ )+λ1MSE(y,y ′ )+λ2D KL (N(μ,σ 2 ),N(0,1))
[0118] Among them, MSE represents the mean square error, D KL represents the KL (Kullback-Leibler) divergence, x represents the solution to the first reference problem, x ′ represents the candidate reconstruction solution, λ1 represents the first weight parameter, which is used to balance the influence of the reconstruction error in the target loss function, λ2 represents the second weight parameter, which is used to control the influence of the KL divergence in the target loss function, y represents the quality score of the true solution, y ′ Represents the quality score of the candidate problem solution, μ and σ represent the mean vector and standard deviation vector of the multivariate Gaussian distribution obeyed by the candidate first problem solution vector output by the encoder in the VAE, and N(0,1) represents a standard normal distribution.
[0119] In this embodiment, the target loss values of the true reconstruction solution, candidate reconstruction solution, candidate problem solution quality score and true solution quality score are calculated according to the preset target loss function, which can evaluate the performance of the initial candidate compilation optimization model and help to optimize the model performance later.
[0120] In step S206 of some embodiments, specifically, the preset model training condition may be that the target loss value is lower than the loss threshold or the model reaches a maximum number of iterations.
[0121] For example, the target loss value falls below the loss threshold of 0.2, or the model reaches the maximum number of iterations, 100.
[0122] In this embodiment, the candidate model parameters of the candidate compilation optimization model are updated based on the target loss value until the target loss value meets the preset model training conditions, and the candidate compilation optimization model is obtained. It is possible to train a candidate intelligent algorithm that better understands and generates high-quality solutions, which facilitates the subsequent improvement of the accuracy of the code compilation optimization items.
[0123] See also Figure 3 In some embodiments, step S104 includes but is not limited to steps S301 to S302:
[0124] Step S301 : Input each of the at least two first reference problem solutions into a candidate compilation optimization model for quality scoring to obtain a predicted solution quality score.
[0125] Step S302 : obtaining a true solution quality score corresponding to each of the at least two first reference problem solutions, and calculating a Pearson correlation coefficient and a Spearman rank correlation coefficient between the true solution quality score and the predicted solution quality score.
[0126] Step S303 : determining the similarity between the candidate compilation optimization model and the optimization problem instance according to the Pearson correlation coefficient and the Spearman rank correlation coefficient.
[0127] In step S301 of some embodiments, specifically, the predicted solution quality score refers to a score of the size of a binary file generated by compiling and solving each first reference problem.
[0128] Specifically, the method of inputting each of the at least two first reference problem solutions into the candidate compilation optimization model for quality scoring is consistent with the process of using the initial candidate compilation optimization model to perform quality scoring on the first reference problem solutions, and will not be repeated here.
[0129] In step S302 of some embodiments, specifically, the Pearson correlation coefficient is used to measure the strength of the linear relationship between two variables.
[0130] Specifically, the Spearman rank correlation coefficient is used to measure the monotonic relationship between two variables, which is not required to be linear.
[0131] Specifically, the values of the Pearson correlation coefficient and the Spearman rank correlation coefficient range from -1 to 1, where 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no correlation.
[0132] For example, if there is a strong positive correlation between the true solution quality score and the predicted solution quality score, both the Pearson and Spearman correlation coefficients will be close to 1.
[0133] In step S303 of some embodiments, specifically, if the Pearson correlation coefficient and the Spearman rank correlation coefficient of the candidate compilation optimization model are both greater than zero, and the Pearson correlation coefficient and the Spearman rank correlation coefficient are both the largest with respect to other candidate compilation optimization models, it indicates that the candidate compilation optimization model has the highest similarity to the optimization problem instance, that is, the candidate intelligent algorithm has the highest similarity to the optimization problem instance, and the candidate intelligent algorithm is capable of solving new optimization problem instances corresponding to similar optimization problem instances.
[0134] In this embodiment, the similarity between the candidate compilation optimization model and the optimization problem instance is determined based on the Pearson correlation coefficient and the Spearman rank correlation coefficient, which provides effective data support for subsequent screening of the intelligent algorithm with the highest similarity to the optimization problem instance.
[0135] See also Figure 4 In some embodiments, step S104 further includes but is not limited to steps S401 to S402:
[0136] Step S401 : sorting multiple candidate compilation optimization models according to similarity to obtain sorted candidate compilation optimization models.
[0137] Step S402 : Filter out the candidate compilation optimization model with the highest similarity from the sorted candidate compilation optimization models and determine it as the initial compilation optimization model.
[0138] In step S401 of some embodiments, specifically, the sorted candidate compilation optimization models refer to models sorted in descending order based on similarity.
[0139] Specifically, multiple candidate compilation optimization models are sorted from largest to smallest based on the similarity between each candidate compilation optimization model and the optimization problem instance to obtain sorted candidate compilation optimization models.
[0140] In step S402 of some embodiments, specifically, a candidate compilation optimization model with the largest Pearson correlation coefficient and Spearman rank correlation coefficient is screened out from multiple candidate compilation optimization models and determined as the initial compilation optimization model. This can screen out the intelligent algorithm that is most similar to the optimization problem instance, avoid using irrelevant intelligent algorithms for invalid function evaluation, reduce computing costs, and the screened intelligent algorithm can also generate corresponding high-quality solutions for similar optimization instance problems. There is no need to perform compilation performance testing for each different combination of compilation optimization items, which helps to improve the efficiency of subsequent code compilation optimization.
[0141] See also Figure 5 In some embodiments, step S105 includes but is not limited to steps S501 to S504:
[0142] Step S501 : Encode each of at least two first reference problem solutions using an initial encoder to obtain an initial first problem solution vector.
[0143] Step S502 : reconstructing the initial first problem solution vector by an initial reconstructor to obtain an initial first reconstructed solution.
[0144] Step S503 : constructing a problem solution data pair for each of the initial first reconstructed solutions corresponding to at least two first reference problem solutions of the same quality level according to the true solution quality score.
[0145] Step S504 , calculating the mean square error (MSE) of the problem solution data pair according to the preset fine-tuning objective function, and updating the reconstructor parameters of the initial reconstructor based on the MSE until the MSE meets the preset fine-tuning training conditions, thereby obtaining the target compiled optimization model.
[0146] In step S501 of some embodiments, specifically, the initial compilation optimization model includes an initial encoder and an initial reconstructor, and the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance, and the intelligent algorithm can also generate corresponding high-quality solutions for similar optimization instance problems.
[0147] Specifically, the initial encoder and the initial reconstructor are also VAE networks, wherein the initial encoder is used to encode the solution to the first reference problem to convert the solution to the first reference problem into a vector of fixed length; the initial reconstructor is used to reconstruct the candidate first problem solution vector so that the model can learn the structural information of the first reference problem solution.
[0148] Specifically, the initial first problem solution vector refers to information of the candidate first reference problem solution represented in vector form, and also includes a mean vector and a standard deviation vector.
[0149] In step S502 of some embodiments, specifically, the initial first reconstructed solution refers to structural information of the problem solution learned by the initial compilation optimization model from the first reference problem solution.
[0150] See also Figure 6 In some embodiments, step S503 includes but is not limited to steps S601 to S604:
[0151] Step S601: An initial scorer is used to perform a quality score on the initial first problem solution vector to obtain an initial solution quality score.
[0152] Step S602 : classifying the quality of the initial first reconstructed solution according to the initial solution quality score to obtain a reconstructed solution quality grade.
[0153] Step S603 , classifying the quality of each of the at least two first reference problem solutions according to the true solution quality score to obtain a first problem solution quality level.
[0154] Step S604 : constructing a problem solution data pair of an initial first reconstructed solution corresponding to each of at least two first reference problem solutions of the same quality level based on the reconstructed solution quality level and the first problem solution quality level.
[0155] In step S601 of some embodiments, specifically, the initial compiled optimization model further includes an initial scorer, which is used to perform quality scoring on the initial first reconstructed solution to evaluate the effectiveness of the first reference problem solution in solving the optimization problem instance.
[0156] Specifically, the initial solution quality score is used to evaluate the actual score of the size of the binary file generated by compiling the first reference problem solution after the initial compilation optimization model learns the structural information of the first reference problem solution.
[0157] In step S602 of some embodiments, specifically, the reconstructed solution quality levels are divided from high to low according to the initial solution quality scores, and initial first reconstructed solutions with the same initial solution quality scores are at the same reconstructed solution quality level.
[0158] In this embodiment, the reconstructed solution quality level may be S1 specific levels, which is not limited here.
[0159] In step S603 of some embodiments, specifically, the first problem solution quality level is divided from high to low according to the true solution quality score, and first reference problem solutions with the same true solution quality score are at the same first problem solution quality level.
[0160] In this embodiment, the first problem solution quality level may also be S1 specific levels, which is not limited here.
[0161] In this embodiment, there is a one-to-one mapping relationship between the quality level of the first problem solution and the quality level of the reconstructed solution.
[0162] In this embodiment, since one first problem solution quality level may correspond to at least an initial first reconstructed solution, the number of initial first reconstructed solutions is greater than the number of first reference problem solutions.
[0163] In step S604 of some embodiments, specifically, the problem solution data pair refers to a mapping data pair between the initial first reconstructed solution and the first reference problem solution at each same quality level.
[0164] Specifically, a set of problem solution data pairs from the first reference problem solution to the initial first reconstructed solution can be obtained by constructing a Cartesian product between each initial first reconstructed solution and the first reference problem solution of the same quality level.
[0165] In this embodiment, a problem solution data pair of the initial first reconstructed solution corresponding to each first reference problem solution in at least two first reference problem solutions of the same quality level is constructed based on the quality level of the reconstructed solution and the quality level of the first problem solution. This can construct a mapping between different quality levels, and use the high-quality first reference problem solution as the input data of the initial compilation optimization model. The reconstructed data obtained after mapping can also have a high-quality solution for the new optimization problem of the same category as the first reference problem solution, thereby realizing the model's automatic learning of problem solutions of different categories, and generating high-quality solutions for new optimization problems of the same category. This solves the problem of needing to perform compilation performance testing on each different combination of compilation optimization items, and is helpful to subsequently improve the efficiency of code compilation optimization.
[0166] In step S504 of some embodiments, specifically, the fine-tuning objective function can be expressed by the following formula:
[0167]
[0168] Where R represents the problem solution data pair from the first reference problem solution to the initial first reconstruction solution x represents the solution to the first reference problem, represents the initial first reconstruction solution, and MSE represents the mean square error of the problem data pair.
[0169] Specifically, the preset fine-tuning training condition may be that the mean square error value of the problem data pair is less than a preset error threshold. The smaller the mean square error value is, the better the performance of the initial compiled optimization model is.
[0170] For example, the error threshold may be 0.15.
[0171] In this embodiment, the mean square error value of the problem solution data pair is calculated according to the preset fine-tuning objective function, and the reconstructor parameters of the initial reconstructor are updated based on the mean square error value until the mean square error value meets the preset fine-tuning training conditions, thereby obtaining a target compilation optimization model. The model update can be achieved through mapping between levels, which has better generalization performance than mapping between data in scenarios where fine-tuning data is insufficient. In the fine-tuning process, only the weight parameters of the decoder part are modified, while the encoder and scorer parts remain unchanged. There is no need to adjust all model parameters, which reduces the model training cost and helps to improve the efficiency of subsequent code compilation optimization.
[0172] In step S106 of some embodiments, specifically, due to the large number of candidate compilation optimization items, by performing second random sampling on multiple candidate problem solutions, at least two candidate compilation items can be randomly sampled from the multiple candidate compilation optimization items as at least two second reference problem solutions.
[0173] Specifically, since the second reference problem solution can cover a wider range of problem solution combinations and can more comprehensively evaluate the performance of the target compilation optimization model, the number of samples of the second reference problem solution is greater than the number of samples of the first reference problem solution.
[0174] See also Figure 7 In some embodiments, step S107 includes but is not limited to steps S701 to S704:
[0175] Step S701 : Encode each of at least two second reference problem solutions using a target encoder to obtain a second problem solution vector.
[0176] Step S702: Perform a quality score on the second problem solution vector using a target scorer to obtain a quality score of the second problem solution.
[0177] Step S703 : Screen at least two second reference problem solutions according to the second problem solution quality scores to obtain a retained problem solution.
[0178] Step S704: reconstruct the retained problem solution through the target reconstructor to obtain a target reconstructed solution.
[0179] In step S701 of some embodiments, specifically, the target compilation optimization model includes a target encoder, a target reconstructor, and a target scorer, wherein the target compilation optimization model is a fine-tuned compilation optimization model.
[0180] Specifically, the target compilation optimization model refers to a neural network model generated after the initial compilation optimization model is quickly fine-tuned. The target compilation optimization model is used to represent the fine-tuned intelligent algorithm and can adapt to new optimization problem instances. That is, the target compilation optimization model is actually a representation of an intelligent algorithm that is more suitable for new optimization problem instances. The target compilation optimization model is used to perform high-quality solutions to the characteristics of new optimization problem instances.
[0181] Specifically, the target compiler is used to encode the second reference problem solution and has converted the second reference problem solution into a fixed-length vector representation.
[0182] Specifically, the second problem solution vector refers to information of the candidate second reference problem solution represented in a vector form.
[0183] In step S702 of some embodiments, specifically, the target scorer is used to perform a quality score on the second problem solution vector to evaluate the effectiveness of the second reference problem solution in solving the optimization problem instance.
[0184] Specifically, the second problem solution quality score is used to evaluate the quality score of the size of the binary file generated by the target compilation optimization model for compiling the second reference problem solution.
[0185] In step S703 of some embodiments, specifically, the retained problem solution refers to a portion of the second reference problem solution retained from at least two second reference problem solutions.
[0186] Specifically, a predetermined number of solutions are screened from at least two second reference problem solutions according to the second problem solution quality scores from high to low to obtain retained problem solutions, and the size of the binary file corresponding to the retained problem solution is relatively small.
[0187] For example, for a code snippet for solving a numerical differential equation, candidate compilation optimization items with high computational accuracy and short running time can be retained based on the quality score, so that the size of the compiled and optimized binary file is smaller.
[0188] In this embodiment, at least two second reference problem solutions are screened according to the quality scores of the second problem solutions to obtain retained problem solutions, which can reduce the amount of data for subsequent compilation and optimization while retaining better optimization solutions, thereby helping to improve the efficiency of subsequent code compilation and optimization.
[0189] In step S704 of some embodiments, specifically, the target reconstructor is used to reconstruct the solution to the retained problem so that the model learns the structural information of the solution to the retained problem.
[0190] Specifically, since the retained problem solution is a high-quality solution during compilation optimization, the target reconstructed solution after model reconstruction is also a high-quality solution.
[0191] Specifically, the retained problem solution is reconstructed through the target reconstructor to obtain the target reconstructed solution, which can generate a higher quality problem solution and improve the accuracy of finding a better combination of compilation optimization items.
[0192] See Figure 8 , provides a Mixture-of-Experts as General-purpose Optimizers (MEGO) for generating high-quality target reconstruction solutions. Figure (a) shows the training of expert models M1, M2, and M3 (i.e., candidate compilation optimization models) corresponding to each type of optimization problem based on a training set of optimization problem instances. E1, E2, and E3 represent training sets of optimization problem instances of different categories, respectively. xi represents the i candidate problem solutions corresponding to the optimization problem training set, and yi represents the candidate fitness value of the i candidate problem solutions, i.e., the quality score of the candidate problem solution.
[0193] Figure (b) shows the neural network model structure and training objectives (i.e., target loss function) of each expert model. The candidate problem solution is input as input data x into the encoder to generate a problem solution vector Z containing a mean vector μ and a standard deviation vector σ. The problem solution vector Z is decoded by the decoder, and the reconstructed solution data is output. The problem solution vector Z is scored by the scorer, and the predicted fitness value y', i.e., the quality score of the candidate problem, is output.
[0194] Figure (c) uses a routing strategy to find an expert model that is relevant to the optimization problem instance. First, through a random first sampling solution, s first sampling reference problem solutions can be obtained, and the quality of the new optimization problem instance Inew corresponding to the optimization problem instance of the same category is scored to obtain the objective function score of Inew (that is, the true fitness value corresponding to the first sampling reference problem solution). According to the true fitness value corresponding to the first sampling reference problem solution, the s first sampling reference problem solutions are sorted and divided into quality levels to obtain the true quality levels corresponding to the s first sampling reference problem solutions. And calculate the predicted fitness value generated by each expert model for each of the s first sample reference problem solutions, sort and divide the s first sample reference problem solutions into quality levels according to the predicted fitness value, and obtain the predicted quality level corresponding to the s first sample reference problem solutions The Pearson correlation coefficient ρ1 and the Spearman rank correlation coefficient ρ2 between the predicted quality level and the true quality level are calculated. If ρ1>0 and ρ2>0 are satisfied, it means that the expert model is correlated with the new optimization problem instance Inew. If ρ1>0 and ρ2>0 are not satisfied, it means that the expert model is not correlated with the new optimization problem instance Inew.
[0195] Figure (d) is used to represent the fine-tuning of each relevant expert model, and the training goal is the fine-tuning objective function. Specifically, by inputting multiple first sampling reference problem solutions xi into the relevant expert model, the reference problem reconstruction solution and the predicted reference solution quality score corresponding to each first sampling reference problem solution xi are obtained, the true reference solution quality score corresponding to each first sampling reference problem solution xi is obtained, and the reference problem reconstruction solutions are sorted based on the predicted reference solution quality score, and the first sampling reference problem solutions are sorted based on the true reference quality score. Based on the sorted, the Cartesian product of the first sampling reference problem solutions and the reference problem reconstruction solutions at the same quality level is constructed to obtain the problem solution data pair, and the relevant expert model is fine-tuned according to the problem solution data pair and the fine-tuning objective function to obtain the fine-tuned expert model.
[0196] Figure (e) shows that a high-quality reconstruction solution is generated based on the fine-tuned expert model. Specifically, the second reference problem solution is randomly sampled for a second time and input into the fine-tuned expert model, and the structural information of the second reference problem solution is learned through the decoder to generate a target reconstruction solution corresponding to the second reference problem solution. After learning the structural information corresponding to the high-quality second reference problem solution based on the fine-tuned expert model, the final second problem solution quality score corresponding to the second reference problem solution is output.
[0197] In step S108 of some embodiments, specifically, the target compilation optimization item refers to a compilation optimization item enabled during compilation optimization.
[0198] Specifically, the target reconstruction solution is used to indicate whether to enable or disable a candidate compilation optimization item, and a compilation optimization item that best meets the current optimization goal can be determined from among numerous candidate compilation optimization items.
[0199] Furthermore, the target binary file is a machine code file that can be executed by a computer after code compilation and optimization.
[0200] Specifically, the initial binary file is compiled and optimized using a GCC encoder and target compilation optimization items to generate a target binary file.
[0201] In this embodiment, a target compilation optimization item is determined from multiple candidate compilation optimization items according to a target reconstruction solution, and the initial binary file is compiled and optimized according to the target compilation optimization item to generate a target binary file. By generating a target binary file with optimal performance and volume, there is no need to perform compilation performance testing on each different combination of compilation optimization items. The automation of the code compilation optimization process is also achieved, and there is no need for manual intervention in the code compilation process, which significantly improves the efficiency and accuracy of code compilation.
[0202] The embodiment of the present application first compiles the acquired target source code file to obtain an initial binary file, and instantiates the problem based on at least two code logic blocks in the initial binary file, so as to clarify the optimization requirements and locate the code parts that need to be optimized; secondly, a first random sampling is performed on multiple candidate problem solutions of the optimization problem instance. Since each candidate compilation optimization model represents an intelligent algorithm that can solve the optimization problem instance, the candidate compilation optimization model is screened and fine-tuned by using the similarity and true solution quality score between the candidate compilation optimization model determined by the first reference problem solution and the optimization problem instance, so that the intelligent algorithm most relevant to the optimization problem instance can be screened out, thereby avoiding the use of irrelevant intelligent algorithms for invalid function evaluation, reducing computing costs, and achieving fine-tuning of relevant models without the need for all intelligent algorithm parameters. All of them are adjusted, which helps to improve the efficiency of subsequent code compilation optimization; then, a second random sampling of multiple candidate problem solutions of the optimization problem instance can further explore the problem solution space, and reconstruct the second reference problem solution through the target compilation optimization model to obtain the target reconstructed solution, which can generate higher quality problem solutions through the fine-tuned intelligent algorithm, and improve the accuracy of finding a better combination of compilation optimization items; finally, the target compilation optimization items are screened out through the target reconstructed solution, and the initial binary file is compiled and optimized according to the target compilation optimization items to generate a target binary file with optimal performance and size. There is no need to perform compilation performance testing on each different combination of compilation optimization items, and the code compilation optimization process is automated, without the need for manual intervention in the code compilation process, which significantly improves the efficiency and accuracy of code compilation.
[0203] See also Figure 9 The embodiment of the present application further provides a code compilation optimization device based on algorithm representation, which can implement the above-mentioned code compilation optimization method based on algorithm representation. The device includes:
[0204] A code compilation module is used to obtain a target source code file, compile the target source code file, and obtain an initial binary file;
[0205] A problem instantiation module, configured to instantiate a problem according to at least two code logic blocks in an initial binary file to obtain an optimization problem instance;
[0206] A first problem solution sampling module is used to perform a first random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions;
[0207] a model screening module, configured to calculate similarities between a plurality of candidate compilation optimization models and the optimization problem instance based on at least two first reference problem solutions, and to screen the plurality of candidate compilation optimization models based on the similarities to obtain an initial compilation optimization model; wherein the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance;
[0208] a model fine-tuning module, configured to obtain true solution quality scores of at least two first reference problem solutions and fine-tune the initial compilation optimization model based on the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model; wherein the target compilation optimization model is used to represent the fine-tuned intelligent algorithm;
[0209] A second problem solution sampling module is used to perform a second random sampling on multiple candidate problem solutions of the optimization problem instance to obtain at least two second reference problem solutions;
[0210] a problem solution reconstruction module, configured to reconstruct each of the at least two second reference problem solutions using a target compilation optimization model to obtain a target reconstructed solution; the target reconstructed solution is used to indicate whether to enable or disable a candidate compilation optimization item;
[0211] The compilation optimization module is used to determine a target compilation optimization item from multiple candidate compilation optimization items according to the target reconstruction solution, and to compile and optimize the initial binary file according to the target compilation optimization item to generate a target binary file.
[0212] The specific implementation of the code compilation optimization device based on algorithm representation is basically the same as the specific embodiment of the code compilation optimization method based on algorithm representation mentioned above, and will not be repeated here.
[0213] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor implements the aforementioned algorithm-representation-based code compilation optimization method when executing the computer program. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0214] See also Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0215] The processor 1001 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0216] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store the processing system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002, and the processor 1001 calls and executes the code compilation optimization method based on algorithm representation in the embodiments of this application;
[0217] Input / output interface 1003, used to implement information input and output;
[0218] Communication interface 1004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0219] Bus 1005 , which transmits information between various components of the device (e.g., processor 1001 , memory 1002 , input / output interface 1003 , and communication interface 1004 );
[0220] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .
[0221] An embodiment of the present application also provides a storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned code compilation optimization method based on algorithm representation.
[0222] The memory, as a non-transient storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0223] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0224] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0225] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0226] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0227] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0228] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0229] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0230] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0231] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0232] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0233] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A code compilation optimization method based on algorithm characterization, characterized in that: The method comprises: Obtaining a target source code file, and compiling the target source code file to obtain an initial binary file; Instantiate the problem according to at least two code logic blocks in the initial binary file to obtain an optimization problem instance; Performing a first random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions; Calculating similarities between the plurality of candidate compilation optimization models and the optimization problem instance based on the at least two first reference problem solutions, and screening the plurality of candidate compilation optimization models based on the similarities to obtain an initial compilation optimization model; wherein the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance; Obtaining true solution quality scores of the at least two first reference problem solutions, and fine-tuning the initial compilation optimization model based on the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model; wherein the target compilation optimization model is used to represent the fine-tuned intelligent algorithm; Performing a second random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two second reference problem solutions; Reconstructing each of the at least two second reference problem solutions using the target compilation optimization model to obtain a target reconstructed solution; the target reconstructed solution is used to indicate whether to enable or disable a candidate compilation optimization item; A target compilation optimization item is determined from the plurality of candidate compilation optimization items according to the target reconstruction solution, and compilation optimization is performed on the initial binary file according to the target compilation optimization item to generate a target binary file.
2. The method according to claim 1, characterized in that The target compilation optimization model includes a target encoder, a target reconstructor and a target scorer; The reconstructing each of the at least two second reference problem solutions using the target compiled optimization model to obtain a target reconstructed solution includes: encoding each of the at least two second reference problem solutions by the target encoder to obtain a second problem solution vector; Performing a quality score on the second problem solution vector using the target scorer to obtain a second problem solution quality score; Screening the at least two second reference problem solutions according to the quality score of the second problem solution to obtain a retained problem solution; The target reconstructor is used to reconstruct the retained problem solution to obtain the target reconstructed solution.
3. The method according to claim 1, characterized in that The screening of multiple candidate compilation optimization models based on the similarity to obtain an initial compilation optimization model includes: sorting the plurality of candidate compilation optimization models according to the similarities to obtain sorted candidate compilation optimization models; The candidate compilation optimization model with the highest similarity is selected from the ranked candidate compilation optimization models and determined as the initial compilation optimization model.
4. The method according to claim 3, characterized in that Before screening a plurality of candidate compilation optimization models based on the similarity to obtain an initial compilation optimization model, the method further includes: Obtaining the initial candidate compilation optimization model; the initial candidate compilation optimization model includes a candidate encoder, a candidate reconstructor, and a candidate scorer; Encoding each of the at least two first reference problem solutions using the candidate encoder to obtain a candidate first problem solution vector; Reconstructing the candidate first problem solution vector using the candidate reconstructor to obtain a candidate reconstructed solution; Using the candidate scorer to perform a quality score on the candidate first problem solution vector to obtain a candidate problem solution quality score; Obtaining a true reconstruction solution corresponding to each of the at least two first reference problem solutions, and calculating target loss values of the true reconstruction solution, the candidate reconstruction solutions, the quality scores of the candidate problem solutions, and the quality score of the true solution according to a preset target loss function; The candidate model parameters of the candidate compilation optimization model are updated based on the target loss value until the target loss value meets a preset model training condition, thereby obtaining the candidate compilation optimization model.
5. The method according to claim 3, characterized in that The initial compilation optimization model includes an initial encoder and an initial reconstructor; The fine-tuning of the initial compilation optimization model according to the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model includes: Encoding each of the at least two first reference problem solutions by the initial encoder to obtain an initial first problem solution vector; Reconstructing the initial first problem solution vector by the initial reconstructor to obtain an initial first reconstructed solution; constructing a problem solution data pair for the initial first reconstructed solution corresponding to each of the at least two first reference problem solutions at the same quality level according to the true solution quality score; The mean square error value of the problem solution data pair is calculated according to a preset fine-tuning objective function, and the reconstructor parameters of the initial reconstructor are updated based on the mean square error value until the mean square error value meets the preset fine-tuning training conditions, thereby obtaining the target compiled optimization model.
6. The method according to claim 5, characterized in that The initial compilation optimization model also includes an initial scorer; The constructing of a problem solution data pair for the initial first reconstructed solution corresponding to each of the at least two first reference problem solutions at the same quality level according to the true solution quality score includes: Performing a quality score on the initial first problem solution vector by the initial scorer to obtain an initial solution quality score; Performing a quality grade classification on the initial first reconstructed solution according to the initial solution quality score to obtain a reconstructed solution quality grade; Performing a quality grade classification on each of the at least two first reference problem solutions according to the true solution quality score to obtain a first problem solution quality grade; The problem solution data pair of the initial first reconstructed solution corresponding to each of the at least two first reference problem solutions at the same quality level is constructed based on the reconstructed solution quality level and the first problem solution quality level.
7. The method according to any one of claims 1 to 6, characterized in that The calculating, based on the at least two first reference problem solutions, similarities between the plurality of candidate compilation optimization models and the optimization problem instance comprises: Inputting each of the at least two first reference problem solutions into the candidate compilation optimization model for quality scoring to obtain a predicted solution quality score; Obtaining a true solution quality score corresponding to each of the at least two first reference problem solutions, and calculating a Pearson correlation coefficient and a Spearman rank correlation coefficient between the true solution quality score and the predicted solution quality score; The similarity between the candidate compilation optimization model and the optimization problem instance is determined according to the Pearson correlation coefficient and the Spearman rank correlation coefficient.
8. A code compilation and optimization device based on deep learning, characterized in that: The device comprises: A code compilation module is used to obtain a target source code file and compile the target source code file to obtain an initial binary file; A problem instantiation module, configured to instantiate a problem according to at least two code logic blocks in the initial binary file to obtain an optimization problem instance; A first problem solution sampling module is used to perform a first random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two first reference problem solutions; a model screening module, configured to calculate similarities between the plurality of candidate compilation optimization models and the optimization problem instance based on the at least two first reference problem solutions, and screen the plurality of candidate compilation optimization models based on the similarities to obtain an initial compilation optimization model; wherein the initial compilation optimization model is used to represent an intelligent algorithm for solving the optimization problem instance; a model fine-tuning module, configured to obtain true solution quality scores of the at least two first reference problem solutions, and fine-tune the initial compilation optimization model based on the at least two first reference problem solutions and the true solution quality scores to obtain a target compilation optimization model; wherein the target compilation optimization model is used to represent the fine-tuned intelligent algorithm; A second problem solution sampling module is used to perform a second random sampling on a plurality of candidate problem solutions of the optimization problem instance to obtain at least two second reference problem solutions; a problem solution reconstruction module, configured to reconstruct each of the at least two second reference problem solutions using the target compilation optimization model to obtain a target reconstructed solution; the target reconstructed solution is used to indicate whether to enable or disable a candidate compilation optimization item; The compilation optimization module is configured to determine a target compilation optimization item from the plurality of candidate compilation optimization items according to the target reconstruction solution, and to perform compilation optimization on the initial binary file according to the target compilation optimization item to generate a target binary file.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the code compilation optimization method based on algorithm representation according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the code compilation optimization method based on algorithm representation according to any one of claims 1 to 7 is implemented.