A method for automatic configuration of optimization parameters from Simulink model to C language

Through the reinforcement learning-based method, the compilation parameters of Simulink model to C language are automatically configured, which solves the problem of low efficiency in the existing technology of compilation parameter optimization and achieves more efficient code generation and execution performance.

CN114995818BActive Publication Date: 2025-05-23DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210395425.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-05-23
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively optimize the compilation parameters of the Simulink model to the C language, resulting in the generation of C code not reaching optimal performance in terms of time complexity.

Method used

Using reinforcement learning-based methods, optimized parameters are automatically configured by constructing reinforcement learning agents and deep reinforcement learning algorithms, combining the graph structure of the Simulink model and the timing relationship during the compilation process.

Benefits of technology

Effectively capture information from Simulink model, optimize compilation parameters, improve the quality and compilation effect of generated code, and the execution time of generated C code is shorter than that recommended by Matlab.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114995818B_ABST
    Figure CN114995818B_ABST
Patent Text Reader

Abstract

The invention discloses an automatic configuration method for optimizing parameters from a Simulink model to a C language, comprising: using an existing random generation tool to generate a Simulink model, constructing a reinforcement learning agent, inputting a graph structure into the reinforcement learning agent, wherein the reinforcement learning agent selects an action to be executed by the Simulink model in the next step according to the input information, and transmits the action to be executed to the Simlink model; using the selected parameter sequence in the process of compiling the current Simulink model into the C language, updating the reinforcement learning agent, updating the reinforcement learning agent according to a time acceleration ratio, inputting a new Simulink model into the reinforcement learning agent that has been updated to recommend optimization parameters, and the execution time of the C language compiled by the parameters obtained by the method will be shorter than the execution time of the C language compiled by using the parameters recommended by Matlab.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software testing, and in particular to a method for automatically configuring optimization parameters from a Simulink model to a C language based on reinforcement learning. Background Art

[0002] Cyber-physical system (CPS) is a controllable, reliable and scalable networked physical device system that deeply integrates computing, communication and control capabilities based on environmental perception. It is mainly used in various intelligent systems. CPS development tool chain has been widely used in the design, simulation and verification of CPS data flow models. Typical CPS development tools, such as Simulink from MathWorks, include model-based design tools, simulators, compilers and automatic code generators. Developers can quickly use this tool to design complex systems based on graphical models, and combine it with automatic code generation tools to automatically generate target code from data flow models, saving a lot of time and labor costs for software development and improving development efficiency.

[0003] The role of a compiler is to convert a programming language into an executable program. Current work mainly uses machine learning-based methods to solve many different compiler parameter optimization problems. These technologies improve the quality of the executable program obtained after compilation and can effectively improve the performance of the compiler. This provides ideas for parameter optimization in the process of converting Simulink models to C code. Due to the new target architecture, the compiler optimization space continues to grow. There are still many difficulties in optimizing the parameters in the process of converting Simulink models to C code using existing technologies. First of all, other compilers target programming languages ​​such as C or C++ when compiling code, while the Simulink model itself is a graph that contains programming logic and timing relationships. Previous research on compiler optimization cannot capture this information in the Simulink model itself. In addition, when using the Embedded Coder tool that comes with Matlab to generate C language code for the Simulink model, the tool will automatically recommend parameters that can be optimized during compilation for the Simulink model, but the code generated using these recommended parameters cannot achieve optimal performance in terms of time complexity. Summary of the invention

[0004] In view of the problems existing in the prior art, the present invention discloses a method for automatically configuring optimization parameters from a Simulink model to a C language based on reinforcement learning, which specifically includes the following steps:

[0005] Use the existing random generation tool to generate Simulink models and build a model seed library. During the seed library generation process, record the weight of each component in the Simulink model and the weight of the edges between components, and finally generate the node and edge weight matrix, and use the weight matrix to build the graph structure;

[0006] Build a reinforcement learning agent, combine the deep reinforcement learning algorithm with the automatic configuration of optimization parameters during the Simulink model to C language compilation process, optimize the parameters during the conversion of the Simulink model to C code, and model the temporal relationship of the optimization process as a Markov decision process;

[0007] Open a model in the seed library in sequence, input the graph structure to the reinforcement learning agent, the reinforcement learning agent selects the next action to be performed by the Simulink model according to the input information, and transmits the action to be performed to the Simulink model, the Simulink model executes the action, and repeats this step to optimize the model parameters;

[0008] The selected parameter sequence is used in the process of compiling the current Simulink model into C language. After the compilation is completed, the execution time of the C language is calculated, and the execution time is compared with the optimization parameters recommended by Matlab to obtain the time acceleration ratio, which is used as a reference for rewards. The larger the time acceleration ratio, the shorter the execution time of the C language generated by the obtained parameters, and the better the model effect.

[0009] The reinforcement learning agent is updated according to the time acceleration ratio, and the new Simulink model is input into the updated reinforcement learning agent to recommend optimization parameters.

[0010] When using deep reinforcement learning algorithms to optimize parameters in the process of converting Simulink models to C code: the Markov decision is recorded as a four-tuple, including the state set, action set, state transition strategy, and reward function.

[0011] The state represents the state of the graph structure composed of the current Simulink model. The reinforcement learning agent performs the next action selection according to the state of the current model. Each action is a set of parameter optimization sequences. The mapping learning relationship from the environment state to the action is defined as the state transition strategy.

[0012] Furthermore, the action selection method is as follows: when the reinforcement learning agent selects an action according to the input of the state feature, an ε-greedy strategy is adopted. Specifically, a threshold ε is set. In the action selection process, there is a probability of ε to randomly select an action, and a probability of 1-ε is to execute the action predicted by the neural network. The initial ε is set to a larger value. In the continuous iteration process, the ε value gradually decays, and the action predicted by the neural network is executed with a larger probability.

[0013] Furthermore, when the reinforcement learning agent is updated: the state of each Simulink model is read from the seed library into the current reinforcement learning model in turn. The model uses an agent to guide the selection of the optimization parameter sequence and feedback is given for the action with the greatest behavior value. The reinforcement learning agent obtains a reward after executing the action and obtains the next state after selecting the action. The current state, action, reward and next state are re-entered into the reinforcement learning model to calculate the loss function of the reinforcement learning model to update the parameters of the current value network. Every fixed number of steps, the parameters of the current value network are copied to the target value network to update it.

[0014] Due to the adoption of the above technical scheme, the present invention provides a method for automatically configuring optimization parameters from a Simulink model to a C language based on reinforcement learning. In the process of compiling a Simulink model diagram into C language, reinforcement learning is used to automatically configure the optimal parameters for this process, which can effectively capture the information of the CPS model itself and achieve a balance between the development and exploration of the optimization space, thereby improving the quality of the generated code, improving the compilation effect, and solving the parameter optimization problem in the process of compiling a Simulink model diagram into C language. In addition, the execution time of the C language compiled by the parameters obtained by this method will be shorter than the execution time of the C language compiled by using the parameters recommended by Matlab. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0017] In order to make the technical solutions and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention:

[0018] like Figure 1 The method for automatically configuring optimization parameters from a Simulink model to C language based on reinforcement learning specifically includes the following steps:

[0019] S1: Build the initial model seed library components:

[0020] Before using reinforcement learning to generate code for Simulink models to automatically recommend optimization parameters, a model seed library is constructed through existing model generation methods. During the generation of each Simulink model, information such as node weights and edge weights is recorded to form a graph structure. During the process of automatically recommending optimization parameters using reinforcement learning, the seed model library is traversed and the above information is used to construct a graph. The constructed graph is input into reinforcement learning to provide an initial model for the agent.

[0021] S2: Building a reinforcement learning agent

[0022] The process of using reinforcement learning to optimize the parameters in the process of converting Simulink models to C code can be modeled as a Markov decision process. We usually record Markov decisions as a four-tuple, which includes a state set, an action set, a state transition strategy, and a reward function. We use the reinforcement learning algorithm for the optimization of converting Simulink models to C code. The elements in the algorithm are defined as follows:

[0023] a) State: The state of the model is the graph structure of the current Simulink model.

[0024] b) Action: The reinforcement learning agent guides the model to perform the next action based on the current state of the model, and each action is a set of parameter optimization sequences.

[0025] c) Strategy: Reinforcement learning is the mapping learning from environment state to action, and this mapping relationship is called strategy. In layman's terms, the thinking process of how an intelligent agent chooses an action is called strategy.

[0026] d) Reward: After the reinforcement learning agent selects a set of optimization parameters, it applies this set of parameters to generate C code for the current Simulink model and calculates the execution time of the code. Then, it uses Matlab to convert the parameters recommended for the Simulink model into C language and calculate the execution time and the time acceleration ratio. It then gives appropriate rewards based on the acceleration ratio. If the acceleration ratio is greater than 1, the reward is positive; if the acceleration ratio is less than 1, the reward is negative.

[0027] S3: Training a reinforcement learning agent:

[0028] Open a model from the seed library in sequence. The initial state of the agent is the Simulink model graph structure read in for the first time. Then select the action to be performed, that is, select a set of parameters from all parameters for recommendation. After the agent performs the action, it switches to the next state according to the state transition strategy, that is, the next Simulink model, and then performs the next action, that is, the next set of recommended parameters. The state is transferred to the next graph model, and so on. It is worth noting that there is no logical relationship between models, so this is a degenerate reinforcement learning model, that is, our model is a mode-free model, and the state transition strategy is random transition.

[0029] Our state and action space is high-dimensional, so as the dimension increases, the amount of computation increases exponentially. Deep neural networks are very effective in extracting complex features, so we use an algorithm that combines deep learning with reinforcement learning, namely the Deep Q-network (DQN) algorithm. The agent in the DQN algorithm selects the best action for the current state based on the state-action value function (the function represents the total utility that can be obtained by being in a certain state and taking relevant actions immediately, and then following the optimal strategy operation). Then enter the next state according to the selected action. Since there is no connection between Simulink models, the state transition strategy is random transition.

[0030] The state-action value function is updated based on the reward obtained by selecting an action in the current state. The update is done gradually, similar to gradient descent, which can reduce the impact of estimation errors and finally converge to the optimal function. The transfer rules are as follows:

[0031] Q(S t ,A t )←Q(S t ,A t )+α[R t+1 +γmax a Q(S t+1 ,a)-Q(S t ,A t )]

[0032] Where St represents the state at time t, At represents the action performed at time t, St+1 represents the state after performing action At at time t, Rt+1 represents the immediate reward obtained by transferring from state S to St1, and Q(St, At) represents the total reward that can be obtained by performing action A in state St. γ is the decay value, which is a constant that satisfies 0≤γ<1. The closer γ is to 1, the more far-sighted it is and the more it will focus on the rewards of subsequent states. When γ is close to 0, it will become short-sighted and only consider the impact of the current reward. α is the learning rate, which represents the impact of the new value on the updated value.

[0033] S4: Computing Reinforcement Learning Rewards

[0034] During the execution of the action, the recommended parameters are input into Simulink, compiled into C language and executed, and the execution time of C language is calculated. The execution time of C language converted using the parameters recommended by Matlab is divided by this time to calculate the time acceleration ratio, and appropriate rewards are given according to the acceleration ratio. If the acceleration ratio is greater than 1, the reward is positive; if the acceleration ratio is less than 1, the reward is negative.

[0035] S5: Update the reinforcement learning agent

[0036] We used two DQN networks with the same structure, Q_net and TQ_net. The only difference between the two networks is the parameters. Q_net is the network we want to train, and the parameters in TQ_net are the parameters of the old Q_net. After each fixed number of training steps, the parameters of Q_net are assigned to TQ_net, which improves the stability of training. The loss function of Q_net is as follows

[0037]

[0038]

[0039] DQN is an offline learning method. It can learn from the current experience as well as the past experience. Therefore, randomly adding previous experience during the learning process will make the neural network more efficient. The experience pool solves the problems of correlation and non-static distribution. At time t, after the reinforcement learning agent takes action, the generated sample is recorded as (s t ,a t ,r t ,s t+1 ), we will t ,a t ,r t ,s t+1 ) are stored in the experience pool. At fixed time steps, random sampling is performed from the memory bank to disrupt the correlation, and then training is performed. The memory bank + random sampling destroys the continuity of the samples, making the training more effective.

[0040] The specific implementation methods are as follows:

[0041] Step 1: Set a threshold ε, which is a value between 0 and 1. The optimization parameter sequence used in the training process is not all calculated based on the neural network. There is a probability of ε to randomly select actions, and a probability of 1-ε to execute the actions predicted by the neural network. The ε value decays with the number of training times.

[0042] Step 2: We use an agent to guide the selection of optimization parameter sequences. In order to capture the logical relationship and timing information of the current Simulink model, we construct the Simulink model into a graph structure. The agent takes the graph structure of the Simulink model as input, selects the next action of the model in the action library and outputs it to the model. The state is the specific information of the model, that is, the graph structure, and the action refers to the selection of a set of optimization parameter sequences to be applied to the current model.

[0043] Step 3: Calculate the execution time of the generated C language and calculate the speedup ratio

[0044] By compiling a large number of Simulink models, we count the parameters recommended by Matlab when converting these models into C language and take the intersection. We use these parameters as our action space. In the previous step, the reinforcement learning agent will recommend a set of recommended optimization parameter sequences based on the current Simulink model. We input this set of parameters into Matlab for SIL simulation, compile it into C language, and calculate the execution time of the code. We compare it with the optimization parameters recommended by Matlab itself and calculate the time acceleration ratio. This acceleration ratio is used as a reference for our reward. The larger the acceleration ratio, the shorter the execution time of the C language generated by the parameters calculated using our method, and the better the model effect.

[0045] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for automatic configuration of optimization parameters from Simulink model to C language based on reinforcement learning, Features include: Use the existing random generation tool to generate Simulink models and build a model seed library. During the seed library generation process, record the weight of each component in the Simulink model and the weight of the edges between components, and finally generate the node and edge weight matrix, and use the weight matrix to build the graph structure; Build a reinforcement learning agent, combine the deep reinforcement learning algorithm with the automatic configuration of optimization parameters during the Simulink model to C language compilation process, optimize the parameters during the conversion of the Simulink model to C code, and model the temporal relationship of the optimization process as a Markov decision process; Open a model in the seed library in sequence, input the graph structure to the reinforcement learning agent, the reinforcement learning agent selects the next action to be performed by the Simulink model according to the input information, and transmits the action to be performed to the Simulink model, the Simulink model executes the action, and repeats this step to optimize the model parameters; The selected parameter sequence is used in the process of compiling the current Simulink model into C language. After the compilation is completed, the execution time of the C language is calculated, and the execution time is compared with the optimization parameters recommended by Matlab to obtain the time acceleration ratio, which is used as a reference for the reward; The reinforcement learning agent is updated according to the time acceleration ratio, and the new Simulink model is input into the updated reinforcement learning agent to recommend optimization parameters.

2. The method according to claim 1, Features: When using deep reinforcement learning algorithms to optimize parameters in the process of converting Simulink models to C code: the Markov decision is recorded as a four-tuple, including the state set, action set, state transition strategy, and reward function.

3. The method according to claim 2, Features: The state represents the state of the graph structure composed of the current Simulink model. The reinforcement learning agent performs the next action selection according to the state of the current model. Each action is a set of parameter optimization sequences. The mapping learning relationship from the environment state to the action is defined as the state transition strategy.

4. The method according to claim 3, Features: The action selection method is as follows: when the reinforcement learning agent selects an action based on the input of the state feature, it adopts the ε-greedy strategy. Specifically, a threshold ε is set. During the action selection process, there is a probability of ε to randomly select an action, and a probability of 1-ε to execute the action predicted by the neural network.

5. The method according to claim 3, Features: When updating the reinforcement learning agent: the state of each Simulink model is read from the seed library into the current reinforcement learning model in turn. The model uses an agent to guide the selection of the optimization parameter sequence and feedback the action with the greatest behavior value. The reinforcement learning agent obtains a reward after executing the action and obtains the next state after selecting the action. The current state, action, reward and next state are re-entered into the reinforcement learning model to calculate the loss function of the reinforcement learning model to update the parameters of the current value network. Every fixed number of steps, the parameters of the current value network are copied to the target value network to update it.