Code optimization method and device, electronic equipment, storage medium and computer product

By combining the target model of time series network and reinforcement learning network, the problem of too long iteration time in the compilation automatic tuning framework is solved, and the efficiency and accuracy of code optimization are achieved.

CN120335819AInactive Publication Date: 2025-07-18INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510798625.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, there is a problem that the iteration time is too long when performing code optimization based on the compilation automatic tuning framework.

Method used

The target model combined with the time series network and the reinforcement learning network is adopted to determine the probability distribution of the optimal solution through feature extraction and policy parameter update, thereby selecting the best optimization task for code optimization.

Benefits of technology

It effectively reduces the time-consuming of code optimization and improves the efficiency and accuracy of code optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335819A_ABST
    Figure CN120335819A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a code optimization method and device, electronic equipment, a storage medium and a computer product, and the method comprises the steps: carrying out feature extraction based on a to-be-optimized code, and obtaining a target feature; determining probability distribution of an optimal solution of at least one preset optimization task in a current state based on the target feature and the target strategy parameter; wherein the target strategy parameter is obtained by performing parameter updating on an initial strategy parameter based on a sample feature in combination with a target model; the target model comprises a time sequence network and a reinforcement learning network; the initial strategy parameters define actions which should be taken by the intelligent agent in a given state and guide the decision-making process of the intelligent agent in the environment; determining a target optimization task from the preset optimization tasks based on the probability distributions; and performing code optimization based on the target optimization task. According to the method and the device, code optimization time consumption can be reduced by reducing iteration time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a code optimization method, apparatus, electronic device, storage medium, and computer product. Background Art

[0002] With the continuous increase in the complexity of software development and the increasing requirements for program performance, how to efficiently generate machine code has become an important topic in compiler design. Compilation auto-tuning technology has received extensive attention in recent years by automatically selecting appropriate compiler options to improve the execution efficiency of the code corresponding to the program. Such technology aims to reduce the time and effort required for manual configuration of compilation parameters, while maximizing the running speed of the target program or minimizing resource consumption.

[0003] The method based on the heuristic search algorithm is a current compilation auto-tuning framework. Such algorithms adopt a space search strategy to explore possible combinations of compiler settings and evaluate the performance of the program under each configuration. Although the above algorithms can find relatively optimal solutions to a certain extent, when the search space is large, due to the randomness of selection, the randomness of the results is high and the required iteration time is very long. Furthermore, it leads to a long time-consuming for current code optimization. Summary of the Invention

[0004] The present application aims to at least solve one of the technical problems existing in the related art. For this purpose, the present application provides a code optimization method, apparatus, electronic device, storage medium, and computer product, which are used to solve the problem that the required iteration time is very long when performing code optimization based on the compilation auto-tuning framework currently, and to reduce the time-consuming of code optimization.

[0005] The code optimization method according to the first aspect embodiment of the present application includes: Performing feature extraction based on the code to be optimized to obtain target features; Based on the target features and target policy parameters, determining the probability distribution of the optimal solution of at least one preset optimization task in the current state; wherein, the target policy parameters are obtained by updating the initial policy parameters based on sample features in combination with a target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that the agent should take in a given state and guide the decision-making process of the agent in the environment; Determining a target optimization task from each preset optimization task based on each probability distribution; Performing code optimization based on the target optimization task.

[0006] According to an embodiment of the present application, the target policy parameters are obtained by the following method: Obtaining sample code and initial policy parameters; Extract features from the sample code to obtain sample features; Train the target model based on the sample features and the initial policy parameters; Determine the updated policy parameters of the target model after training as the target policy parameters.

[0007] According to an embodiment of the present application, the training of the target model based on the sample features and the initial policy parameters includes: Perform state transformations for a preset number of time steps based on the initial policy parameters and determine the calculated values of the advantage function obtained in the state transformations; Train the target model based on the initial policy parameters, each state in the state transformations, and each calculated value of the advantage function.

[0008] According to an embodiment of the present application, the determination of the target optimization task from each preset optimization task based on each probability distribution includes: Determine the preset optimization task with the largest value of the probability distribution among each preset optimization task as the target optimization task.

[0009] According to an embodiment of the present application, the extraction of features from the code to be optimized to obtain target features includes: Convert the code to be optimized into an intermediate representation form to obtain intermediate data; Extract features from the intermediate data to obtain target features.

[0010] According to an embodiment of the present application, the code optimization based on the target optimization task includes: Optimize the intermediate data obtained based on the code to be optimized according to the target optimization task to obtain optimized data; Convert the optimized data into code form to obtain the target code.

[0011] According to the code optimization device of the second aspect embodiment of the present application, it includes: An extraction module for extracting features from the code to be optimized to obtain target features; A first determination module for determining the probability distribution of the optimal solution of at least one preset optimization task in the current state based on the target features and the target policy parameters; wherein, the target policy parameters are obtained by updating the initial policy parameters based on the sample features in combination with the target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that the agent should take in a given state and guide the decision-making process of the agent in the environment; A second determination module, configured to determine a target optimization task from each preset optimization task based on each probability distribution; An optimization module, configured to perform code optimization based on the target optimization task.

[0012] An electronic device according to an embodiment of the third aspect of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the code optimization method as described in any one of the above is implemented.

[0013] A storage medium according to an embodiment of the fourth aspect of the present application is a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the code optimization method as described in any one of the above is implemented.

[0014] A computer program product according to an embodiment of the fifth aspect of the present application includes a computer program. When the computer program is executed by a processor, the code optimization method as described in any one of the above is implemented.

[0015] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects: By combining sample features with a target model to update the initial policy parameters to obtain target policy parameters, since the target model includes a time series network and a reinforcement learning network, and the time series network can capture the evolution of the state over time, it is possible to effectively utilize past information and experience to guide the reinforcement learning network to make current decisions, complete the update of the initial policy parameters to obtain target policy parameters, reduce the number and time of ineffective explorations, effectively reduce the iteration time, and then after feature extraction is performed on the code to be optimized to obtain target features, based on the target features and the target policy parameters, accurately determine the probability distribution of the optimal solution of at least one preset optimization task in the current state, so that the target optimization task can be determined from each preset optimization task based on each probability distribution, and code optimization can be performed based on the target optimization task, so the time-consuming of code optimization can be reduced.

[0016] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of the code optimization method provided by the embodiments of the present application.

[0019] Figure 2 It is a schematic structural diagram of the electronic device provided by the present application. Specific Embodiments

[0020] The following further describes the embodiments of the present application in detail with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0021] In the description of the embodiments of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the embodiments of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0022] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.

[0023] In the embodiments of the present application, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature can be that the first feature is directly above or obliquely above the second feature, or simply indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or simply indicates that the first feature has a lower horizontal height than the second feature.

[0024] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0025] This application proposes a code optimization method, device, electronic device, storage medium, and computer product.

[0026] Figure 1 It is a schematic flowchart of the code optimization method provided by the embodiments of this application. As Figure 1 shown, the code optimization method includes: Step 110, perform feature extraction based on the code to be optimized to obtain target features.

[0027] Step 120, determine the probability distribution of the optimal solution of at least one preset optimization task in the current state based on the target features and target policy parameters; wherein, the target policy parameters are obtained by updating the initial policy parameters based on the sample features in combination with the target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that the agent should take in a given state and guide the decision-making process of the agent in the environment.

[0028] Step 130, determine the target optimization task from each preset optimization task based on each probability distribution.

[0029] Step 140, perform code optimization based on the target optimization task.

[0030] It should be noted that the execution subject of the code optimization method provided by the embodiments of this application can be a computer device, etc. The computer device can be, for example, a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc. It should be noted that all the data that needs to be obtained in this application is obtained through regular channels after being authorized by relevant users.

[0031] A code optimization device can be set or connected in the computer device of the present application, whereby the code optimization device can be controlled to execute the code optimization method of the present application.

[0032] Specifically, the present application can pre-obtain a certain number of codes as sample codes according to requirements, and obtain preset policy parameters as initial policy parameters. Among them, a policy is a rule for the model to select actions according to the current state. The initialization of policy parameters is the starting point of the training process. Usually, these parameters are randomly initialized and then gradually adjusted through training to optimize the objective function. Therefore, the initial policy parameters define the actions that the agent should take in a given state and guide the decision-making process of the agent in the environment.

[0033] Furthermore, feature extraction can be performed based on the sample codes to obtain sample features.

[0034] Since the conditions of many PASS applications are related to the current program state, and the current program state is related to the previously applied PASS, the application of the previous PASS also has a certain impact on the selection of the current PASS. To remember the impact of the previously applied PASS on the current PASS, the present application uses a time series network to remember the relationship between successive PASSes.

[0035] Therefore, the present application can fuse a time series network in a reinforcement learning network to construct a target model. Among them, reinforcement learning includes an action space, an observation space, and a reward.

[0036] Among them, an action is a decision or action required in the current environmental state. In the present application, the action space is part of the optimization PASS of LLVM (Low-Level Virtual Machine). In LLVM, PASS refers to a modular component that analyzes and transforms the intermediate representation code and is used to perform various optimization tasks. LLVM provides a rich set of PASSes.

[0037] LLVM is an open-source compiler and toolchain technology collection.

[0038] An observation space is a view of the current state. In the present application, the observation space is selected as the number of each instruction after the program is converted into an intermediate representation.

[0039] A reward is a judgment index for the previously selected one. Generally speaking, the higher the reward value of a choice, the higher the possibility of being selected in the current state. In the present application, the reward is the degree of optimization of the corresponding target after adopting the selected optimization PASS. For example, in code volume optimization, the reward value can adopt the reduction ratio of the code volume after adopting the optimization PASS relative to not adopting this PASS.

[0040] In this application, the reinforcement learning network can adopt the Proximal Policy Optimization (PPO) algorithm in the policy gradient algorithm. The time series network can adopt the Long Short-Term Memory (LSTM) network and the Attention mechanism. The Attention mechanism is an architecture widely used in deep learning to help the model better focus on the most important parts of the input data.

[0041] The PPO algorithm replaces the KL divergence in the policy gradient with a truncated surrogate loss function, and the surrogate loss function is as follows: ; where ρ t is the importance sampling coefficient, θ is the policy parameter, is a custom hyperparameter, t represents time, t is the advantage function, and its calculation formula is as follows: ; ; where, represents the expected value, s t represents the state of the agent at time step t, R(s t ) is the reward when taking action a t , γ is the discount factor, Q π , V π are the action value function and the state value function when adopting policy π respectively, and T represents time.

[0042] Among them, the KL divergence is an asymmetric measure to measure the difference between two probability distributions, and is also called relative entropy.

[0043] It should be noted that clip is a key operation in the PPO algorithm to limit the amplitude of policy update. Its purpose is to prevent the policy from being updated too much, thus ensuring the stability of training.

[0044] is the average of all sampled data, ensuring that the loss function can reflect the optimization goal of the entire data set.

[0045] It should be noted that in the actual production process, due to different production environments having different requirements for code optimization goals, in addition to the basic code size and code running time, power consumption and the memory occupancy rate during operation are also limiting conditions for embedded devices. At the same time, in some cases, the combined requirements of different optimization goals also need to be considered. To handle this situation, in the reward setting of this application, a hyperparameter is set for each target reward to adjust the proportion of different rewards in optimization. This parameter is defined by the user himself / herself to target a specific complex usage environment.

[0046] According to the above situation, the reward set for learning in this application is R t , and the code size, code running time, power consumption, and memory occupancy are respectively R c 、R t 、R p 、R m . At the same time, a hyperparameter is set for each parameter (for example, hyperparameters α, β, γ, ζ are respectively set for parameters R c 、R t 、R p 、R m ), and thus the reward for learning can be obtained as follows: R t =αR c +βR t +γR p +ζR m .

[0047] When there is only one of these hyperparameters, the result of single-target optimization can be obtained. When multiple hyperparameters are not zero, the multiple hyperparameters are normalized to train a multi-target optimization model defined by the user.

[0048] To cope with a more complex actual production environment, a single metric may seem relatively ineffective in some cases. Therefore, some parameters are added during learning to enable the model to handle more complex situations. Since in reinforcement learning, the model parameters are mainly adjusted according to the rewards for the current actions, it is a more appropriate choice to select the reward module for adjustment.

[0049] In the selection of PASS, when different PASS are used crosswise, there may be a weakening of the effect caused by the destruction of conditions, showing a certain marginal diminishing effect. However, the resources consumed may not necessarily decrease. Adding an importance coefficient to the reward module would be a better choice. During the production process, the trade-off between running time and code size often needs to be considered. After adding the importance coefficient, the proportion of rewards for the two can be weighed, thus avoiding choices that consume a large amount of running time but result in relatively little reduction in code size. Moreover, this importance coefficient can be defined by the user himself / herself and trained independently to meet his / her own usage requirements.

[0050] In the actual programming process, this application uses the -Oz compilation optimization option of LLVM as the benchmark for code size optimization, the -Os option as the benchmark for code running speed, and the power consumption and memory occupancy rate are respectively set as the benchmarks of the -Oz and -Os options. Therefore, the setting of the reward R is as follows: R=(c t -c0) / c0; where c t represents the optimized target quantity, and c0 represents the benchmark target quantity.

[0051] This application can use a time series network to memorize the influence of the current PASS on the future selected PASS. Among them, the time series network needs to be initialized before training, and at the same time, the existing policy parameters are used as the input of the time series network, and the corresponding output is the policy parameter for the next round of training.

[0052] This application adds a time series network to the reinforcement learning network, so that the policy parameters during training can consider the influence of previous selections and take the influence into account during training. In addition, in order to cope with various situations in actual production, a combined reward function for different goals is proposed so that the framework can face different complex situations.

[0053] Furthermore, based on the sample features and the initial policy parameters, the target model can be trained so that the target model can update the initial policy parameters and at the same time update the hidden layer data of the time series network.

[0054] After the training of the target model is completed, the updated policy parameters of the target model can be determined as the target policy parameters. Among them, during the training process, the sample data can be divided into a training set and a test set. The training can be carried out through the training set, and the target model obtained by training is evaluated through the test set. When the target model passes the evaluation, the training is determined to be completed.

[0055] Furthermore, the target policy parameters can be saved on the local computer device or server according to the usage situation.

[0056] Furthermore, when code optimization is required, the code to be optimized can be obtained and determined as the code to be optimized.

[0057] Furthermore, the code to be optimized can be transformed and then feature extraction can be carried out, and the extracted features are defined as the target features. Among them, the target features are multi-dimensional integer feature vectors.

[0058] Furthermore, the extracted target features can be input into the time series network of the target model to obtain the hidden state or feature representation.

[0059] Further, the output of the time series network is used as the input of the reinforcement learning network, and the action probability distribution of each preset optimization task in the current state is calculated through the policy network in the reinforcement learning network.

[0060] Furthermore, the action probability distribution of each preset optimization task can be output through the policy network, representing the optimal probability of selecting each preset optimization task in the current state.

[0061] Further, the target optimization task can be determined from each preset optimization task by comparing the probability distributions.

[0062] Further, the intermediate data obtained based on the code to be optimized can be optimized through the target optimization task, and then the optimized data is converted into code form to obtain the target code. Thus, the code optimization is completed.

[0063] According to the code optimization method of the embodiments of the present application, the initial policy parameters are updated by combining the sample features with the target model to obtain the target policy parameters. Since the target model includes a time series network and a reinforcement learning network, and the time series network can capture the evolution of the state over time, it can effectively utilize past information and experience to guide the reinforcement learning network to make current decisions, complete the update of the initial policy parameters to obtain the target policy parameters, reduce the number and time of ineffective explorations, effectively reduce the iteration time. Furthermore, after feature extraction is performed on the code to be optimized to obtain the target features, based on the target features and the target policy parameters, the probability distribution of the optimal solutions of at least one preset optimization task in the current state can be accurately determined. Thus, the target optimization task can be determined from each preset optimization task based on the probability distributions, and code optimization can be performed based on the target optimization task. Therefore, the time consumed for code optimization can be reduced.

[0064] Based on the above embodiments, the target policy parameters are obtained in the following manner: Obtain the sample code and the initial policy parameters; Perform feature extraction on the sample code to obtain sample features; Train the target model based on the sample features and the initial policy parameters; Determine the updated policy parameters of the target model after training as the target policy parameters.

[0065] Further, training the target model based on the sample features and the initial policy parameters includes: Perform state transformation for a preset number of time steps based on the initial policy parameters and determine the calculated values of the advantage function obtained in the state transformation; Train the target model based on the initial policy parameters, each state in the state transformation, and each calculated value of the advantage function.

[0066] Specifically, the present application can obtain pre-set policy parameters that define the actions that an agent should take in a given state and guide the decision-making process of the agent in the environment as the initial policy parameters.

[0067] In addition, it can obtain a random program generated by csmith or other data sets that allow execution and record the corresponding states as sample codes. Among them, csmith is a tool for generating random C programs, mainly used to test the correctness and stability of compilers.

[0068] Furthermore, the present application can convert the sample code into the form of Intermediate Representation (IR) through LLVM. The main role of IR is to convert the abstract syntax tree of a high-level language into an intermediate form that is closer to the hardware execution model while retaining the semantic information of the program. This intermediate form is neither as abstract as a high-level language nor as dependent on specific hardware as machine code.

[0069] Furthermore, it can count the quantities of various types of data in the converted data, such as the number of basic blocks, phi nodes, arithmetic instructions, etc. Thus, feature extraction is achieved, and a multi-dimensional integer feature vector of each individual program code is obtained as the sample feature.

[0070] Furthermore, state transformations can be performed for a preset number K of time steps under the initial policy parameters, and the calculated values of the advantage function obtained in the state transformations are determined. Among them, K can be set according to actual needs.

[0071] Furthermore, the initial policy parameters, each state in the state transformation, and each calculated value of the advantage function can be obtained.

[0072] Furthermore, the various data obtained above can be put into a rollout buffer. The target model is trained with the data in the rollout buffer.

[0073] Specifically, each data is input into the first fully connected layer of the reinforcement learning network in the target model, and this layer is responsible for preliminary feature extraction and transformation.

[0074] Furthermore, the data output by the first fully connected layer passes through a time series network, and this layer is responsible for capturing the time dependence of the state and providing time context information for the prediction of the policy and value.

[0075] Furthermore, the data output by the time series network passes through the second fully connected layer of the reinforcement learning network in the target model, and this layer is responsible for generating the final policy parameters and value estimation according to the output of the time series network.

[0076] During the training process, the policy loss and value loss are calculated based on the output of the target model and the actual reward value. The policy loss is usually calculated by comparing the action probability distribution output by the model with the actually taken action, while considering the estimated value of the advantage function. The value loss is calculated by comparing the difference between the value estimate output by the model and the actual return (or target value).

[0077] Furthermore, the calculated loss function can be used for backpropagation of the model parameters to calculate the gradients.

[0078] Update the parameters of the model according to the gradients, including the parameters of the first fully connected layer, the time series network, and the second fully connected layer.

[0079] In the PPO algorithm, a clipped objective function is used to limit the magnitude of the policy update to maintain the stability of the policy.

[0080] Optimize the policy network and value network simultaneously, so that the policy network can learn better action selection strategies, and the value network can more accurately estimate the value of the state.

[0081] Repeat the above steps until the target model converges or reaches the preset number of training epochs.

[0082] In this application, the initial policy parameters are updated to obtain the target policy parameters by combining the sample features with the target model. Since the target model includes a time series network and a reinforcement learning network, and the time series network can capture the evolution of the state over time, it can effectively utilize past information and experience to guide the reinforcement learning network to make current decisions, complete the update of the initial policy parameters to obtain the target policy parameters, reduce the number and time of ineffective explorations, effectively reduce the iteration time, so that after feature extraction is performed on the code to be optimized based on the target features and the target policy parameters, the probability distribution of the optimal solution of at least one preset optimization task in the current state can be accurately determined, and thus the target optimization task can be determined from each preset optimization task based on each probability distribution, and code optimization can be performed based on the target optimization task, so the time-consuming of code optimization can be reduced.

[0083] Based on the above embodiments, feature extraction is performed on the code to be optimized to obtain target features, including: Convert the code to be optimized into an intermediate representation form to obtain intermediate data; Perform feature extraction on the intermediate data to obtain target features.

[0084] Specifically, this application can use the compilation toolchain (or compiler) LLVM to convert the code to be optimized as source code into the form of IR, and determine the converted data as intermediate data.

[0085] Furthermore, the quantities of various types of data in the intermediate data can be counted, such as the number of basic blocks, the number of phi nodes, the number of arithmetic instructions, etc.

[0086] Furthermore, the quantities of various types of data in the intermediate data can be integrated into a multi-dimensional integer feature vector, thereby completing feature extraction and determining the multi-dimensional integer feature vector as the target feature.

[0087] Since the structure of the IR in this application is relatively simple and regular, first converting the code to be optimized into the form of intermediate representation can improve the efficiency of feature extraction and facilitate various analysis and optimization operations of the compiler, such as constant propagation, dead code elimination, loop optimization, etc. Therefore, it helps to speed up the code optimization speed and can reduce the time consumed for code optimization.

[0088] Based on the above embodiments, determining the target optimization task from each preset optimization task based on each probability distribution includes: Determining the preset optimization task with the largest value of the probability distribution among each preset optimization task as the target optimization task.

[0089] Specifically, this application can compare the probability distributions corresponding to each preset optimization task to determine the magnitude relationship of each probability distribution.

[0090] Furthermore, after the comparison is completed, the preset optimization task with the largest value of the probability distribution among each preset optimization task can be determined as the target optimization task.

[0091] By determining the preset optimization task with the largest value of the probability distribution among each preset optimization task as the target optimization task, this application can accurately determine the best optimization task, and then perform code optimization through the best optimization task, which can improve the accuracy of code optimization.

[0092] Based on the above embodiments, performing code optimization based on the target optimization task includes: According to the target optimization task, optimizing the intermediate data obtained based on the code to be optimized to obtain optimized data; Converting the optimized data into code form to obtain the target code.

[0093] Specifically, after obtaining the target optimization task, this application can apply the target optimization task to the intermediate data obtained based on the code to be optimized through the compiler LLVM, thereby optimizing the intermediate data through the target optimization task to obtain optimized data.

[0094] Furthermore, converting the optimized data from the intermediate representation form back to the code form can obtain the optimized code and use it as the target code.

[0095] The present application optimizes the code through the best optimization tasks, which can improve the accuracy of code optimization. Moreover, by optimizing the data in the intermediate representation through the optimal optimization tasks and then converting it into the code form, the code optimization speed can be increased.

[0096] The code optimization device provided by the present application will be described below. The code optimization device described below can be correspondingly referred to the code optimization method described above.

[0097] Furthermore, the present application also provides a code optimization device.

[0098] The code optimization device includes: An extraction module, configured to perform feature extraction based on the code to be optimized to obtain target features; A first determination module, configured to determine the probability distribution of the optimal solutions of at least one preset optimization task in the current state based on the target features and target policy parameters; wherein, the target policy parameters are obtained by updating the initial policy parameters based on the sample features in combination with the target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that the agent should take in a given state and guide the decision-making process of the agent in the environment; A second determination module, configured to determine the target optimization task from each preset optimization task based on each probability distribution; An optimization module, configured to perform code optimization based on the target optimization task.

[0099] In the code optimization device of the present application, the initial policy parameters are updated to obtain the target policy parameters by combining the sample features with the target model. Since the target model includes a time series network and a reinforcement learning network, and the time series network can capture the evolution of the state over time, it can effectively utilize past information and experience to guide the reinforcement learning network to make current decisions, complete the update of the initial policy parameters to obtain the target policy parameters, reduce the number and time of ineffective explorations, effectively reduce the iteration time, and then, after performing feature extraction on the code to be optimized to obtain the target features, determine the probability distribution of the optimal solutions of at least one preset optimization task in the current state accurately based on the target features and target policy parameters, so that the target optimization task can be determined from each preset optimization task based on each probability distribution, and code optimization can be performed based on the target optimization task, thus reducing the time-consuming of code optimization.

[0100] In one embodiment, the extraction module is specifically configured to: Convert the code to be optimized into an intermediate representation form to obtain intermediate data; Perform feature extraction on the intermediate data to obtain target features.

[0101] In one embodiment, the second determination module is specifically configured to: Determine the preset optimization task with the largest value of the probability distribution among the preset optimization tasks as the target optimization task.

[0102] In one embodiment, the optimization module is specifically configured to: Optimize the intermediate data obtained based on the code to be optimized according to the target optimization task to obtain optimized data; Convert the optimized data into code form to obtain the target code.

[0103] Figure 2 An entity structure diagram of an electronic device is illustrated. As Figure 2 shown, the electronic device may include: a processor 210, a communications interface 220, a memory 230, and a communication bus 240. Among them, the processor 210, the communications interface 220, and the memory 230 complete mutual communication through the communication bus 240. The processor 210 may call the logical instructions in the memory 230 to execute the following method: perform feature extraction based on the code to be optimized to obtain target features; Based on the target features and the target policy parameters, determine the probability distribution of the optimal solution of at least one preset optimization task in the current state; wherein, the target policy parameters are obtained by updating the initial policy parameters based on the sample features in combination with the target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that the agent should take in a given state and guide the decision-making process of the agent in the environment; Determine the target optimization task from the preset optimization tasks based on each probability distribution; Perform code optimization based on the target optimization task.

[0104] In addition, when the logical instructions in the above-mentioned memory 230 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0105] In another aspect, an embodiment of this application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments. For example, it includes: performing feature extraction based on the code to be optimized to obtain target features; Based on the target features and target policy parameters, determining the probability distribution of the optimal solution of at least one preset optimization task in the current state; wherein, the target policy parameters are obtained by parameter-updating the initial policy parameters based on sample features in combination with a target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that an agent should take in a given state and guide the decision-making process of the agent in the environment; Determining a target optimization task from each preset optimization task based on each probability distribution; Performing code optimization based on the target optimization task.

[0106] In another aspect, an embodiment of this application also provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments. For example, it includes: performing feature extraction based on the code to be optimized to obtain target features; Based on the target features and target policy parameters, determining the probability distribution of the optimal solution of at least one preset optimization task in the current state; wherein, the target policy parameters are obtained by parameter-updating the initial policy parameters based on sample features in combination with a target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that an agent should take in a given state and guide the decision-making process of the agent in the environment; Determining a target optimization task from each preset optimization task based on each probability distribution; Optimize the code based on the target optimization task.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that makes contributions to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application.

Claims

1. A code optimization method, characterized in that, Including: Performing feature extraction on the code to be optimized to obtain target features; Based on the target features and target policy parameters, determining the probability distribution of the optimal solutions of at least one preset optimization task in the current state; wherein, the target policy parameters are obtained by updating the initial policy parameters based on sample features in combination with a target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that an agent should take in a given state and guide the decision-making process of the agent in the environment; Determining a target optimization task from each preset optimization task based on each probability distribution; Performing code optimization based on the target optimization task.

2. The code optimization method according to claim 1, wherein The target policy parameters are obtained in the following manner: Obtaining sample code and initial policy parameters; Performing feature extraction on the sample code to obtain sample features; Training the target model based on the sample features and the initial policy parameters; Determining the policy parameters updated by the target model after training as the target policy parameters.

3. The code optimization method according to claim 2, wherein The training of the target model based on the sample features and the initial policy parameters includes: Performing state transformation for a preset number of time steps based on the initial policy parameters and determining the calculated values of the advantage function obtained in the state transformation; Training the target model based on the initial policy parameters, each state in the state transformation, and each calculated value of the advantage function.

4. The code optimization method according to claim 1, wherein, The determining of the target optimization task from each preset optimization task based on each probability distribution includes: Determining the preset optimization task with the largest value of the probability distribution among each preset optimization task as the target optimization task.

5. The code optimization method according to claim 1, wherein The performing of feature extraction on the code to be optimized to obtain target features includes: Converting the code to be optimized into an intermediate representation form to obtain intermediate data; Performing feature extraction on the intermediate data to obtain target features.

6. The code optimization method according to claim 1, wherein The performing of code optimization based on the target optimization task includes: Optimizing the intermediate data obtained based on the code to be optimized according to the target optimization task to obtain optimized data; Converting the optimized data into code form to obtain target code.

7. A code optimization device, characterized in that, Including: An extraction module for performing feature extraction on the code to be optimized to obtain target features; A first determination module for determining the probability distribution of the optimal solutions of at least one preset optimization task in the current state based on the target features and target policy parameters; wherein, the target policy parameters are obtained by updating the initial policy parameters based on sample features in combination with a target model; the target model includes a time series network and a reinforcement learning network; the initial policy parameters define the actions that an agent should take in a given state and guide the decision-making process of the agent in the environment; A second determination module for determining a target optimization task from each preset optimization task based on each probability distribution; An optimization module for performing code optimization based on the target optimization task.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the code optimization method according to any one of claims 1-6.

9. A storage medium, the storage medium being a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the code optimization method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the code optimization method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Code optimization item acquisition method and device, storage medium and electronic equipment

    CN110727437A

  • Compilation optimization model training method and compilation optimization method

    CN118626093A

  • Code optimization method for heterogeneous HPC platform and related device

    CN119917107A

  • Task-oriented machine learning and a configurable tool thereof on a computing environment

    US20210397941A1