Multi-dimensional Communication Performance Topology Mapping Optimization Guidance Method for Parallel Applications Based on Q-Learning
By dynamically adjusting the weight of multi-dimensional optimization indicators based on Q-Learning, the problem of difficulty in optimizing multi-dimensional communication performance in large-scale parallel applications in the existing technology is solved, and a better topological mapping solution and the effect of reducing communication overhead is achieved.
Patent Information
- Application Number
- CN202111294103.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing topology mapping optimization methods are difficult to effectively optimize multi-dimensional communication performance in large-scale parallel applications, and the weight setting of multi-dimensional indicator optimization functions depends on manual experience and are not universal.
The optimization guidance method for multi-dimensional communication performance topology mapping is adopted based on Q-Learning, and the optimization direction is guided through the Q-table update process, and the weight of multi-dimensional optimization indicators is dynamically adjusted.
It realizes the rapid and accurate solution of the weight adjustment problem of multi-dimensional optimization function, obtains better topological mapping solutions, and reduces the communication overhead of parallel applications.
Smart Images

Figure CN114048044B_ABST
Abstract
Description
Technical Field:
[0001] The present invention discloses a multi-dimensional communication performance topology mapping optimization guidance method based on Q-Learning, which relates to the communication performance optimization of large-scale parallel applications in high-performance computing and belongs to the field of computer technology. Background Art:
[0002] With the continuous increase in the scale of high-performance computing parallel applications, the communication overhead between parallel application processes has become a performance bottleneck. Since there is no need to change communication hardware, communication protocols, etc., constructing an optimization mapping (referred to as topology mapping for short) method from the parallel application communication topology to the network physical topology has high cost performance and good universality, and has become a research hotspot in the field of reducing the communication overhead of parallel applications.
[0003] The existing topology mapping methods mainly include topology mapping methods based on heuristic search and topology mapping methods based on communication locality optimization.
[0004] 1) Topology mapping methods based on heuristic search
[0005] The topology mapping method based on heuristic search refers to using a heuristic algorithm or an improved algorithm of a heuristic algorithm to search for a better topology mapping algorithm in the solution space. Although the algorithm overhead of this kind of algorithm is large, it has strong universality. Currently, the heuristic algorithms applied to topology mapping optimization mainly include simulated annealing algorithm, Pair-Exchange, neighborhood exchange algorithm, etc. However, due to the single dimension of the objective function or the fixed multi-index weights, the communication performance optimization effect of these algorithms is limited.
[0006] 2) Topology mapping methods based on communication locality optimization
[0007] The topology mapping method based on communication locality optimization is to distribute processes with strong communication requirements to adjacent nodes, thereby enhancing communication locality and improving the performance of parallel applications. Graph partitioning is a typical representative based on communication locality optimization. Graph partitioning is to convert the parallel application topology and the network physical topology into graphs, and then use mathematical knowledge such as graph theory to solve the optimal topology mapping. This type of optimization method depends on a specific network structure and cannot adapt to diverse network topologies.
[0008] In summary, most of the existing topology mapping optimization methods adopt a single - dimension index optimization function. For the few topology mapping methods that utilize multi - dimension index optimization functions, the weights of each index in the optimization function are directly set to fixed values. However, on the one hand, the single - dimension index optimization function can no longer meet the communication performance optimization requirements of large - scale complex parallel applications; on the other hand, the multi - dimension index optimization functions in the existing methods face problems in setting the optimization direction and tuning. Although setting fixed weights can obtain a relatively optimal topology mapping scheme to a certain extent, the weight setting depends on manual experience and has poor universality. Summary of the Invention:
[0009] The main objective of the present invention is to accurately guide the optimization direction (adjust the weights of each index in the multi - dimension optimization index) of the communication topology mapping method for parallel applications based on multi - dimension optimization indexes, and then obtain a more optimal multi - dimension communication performance topology mapping scheme for parallel applications. Aiming at the problem of how to construct the quantitative relationship among the candidate multi - dimension communication performance topology mapping scheme for parallel applications, the simulation running time of parallel applications, and the optimization direction, the present invention proposes an optimization guidance method for the multi - dimension communication performance topology mapping of parallel applications based on Q - Learning, abbreviated as DAQ (Optimized Direction Adjustment Method based Q - Learning). This method guides the optimization direction of the multi - dimension communication performance topology mapping of parallel applications by utilizing the Q - table update process in the Q - Learning algorithm. It includes the following steps:
[0010] (1) Determine the Q - table structure, state variables, reward mechanism, update function, actions in the DAQ method, and the cut - off condition of the DAQ method.
[0011] (2) According to the historical information recorded in the Q - table, use the reward mechanism to calculate the reward value obtained from the current state transition, execute the action with the maximum reward value, and iteratively update the Q - table until the cut - off condition of the DAQ method is reached, and output the state variable value in the Q - table at this time.
[0012] (3) Generate the corresponding optimization direction for the multi - dimension communication performance topology mapping of parallel applications based on the state variables in (2).
[0013] Among them, step 1) includes the following steps:
[0014] Step (1.1) Determine the Q-table structure. The numbers 0, 1, 2, 3... in the first column of the Q-table represent the state variables in the Q-table. As the Q-table is updated, the number of its states increases dynamically. The second, third, and fourth columns in the Q-table are the execution overheads of the parallel application of the multi-dimensional communication performance topology mapping scheme under this state weight, which is called the Q-value.
[0015] Step (1.2) Determine the state variables. The state State in the Q-table is the weight vector <w1, w2, w3> of the multi-index optimization function, where w1, w2, and w3 respectively correspond to the weights of the message delay overhead function f1, the link load balancing degree function f2, and the network congestion degree function f3 in the multi-index optimization.
[0016] Step (1.3) Determine the reward mechanism. The reward mechanism of the DAQ method is the opposite of the communication overhead of the parallel application, that is, R = -T i . To ensure that the state selected each time approaches the minimum parallel application communication overhead.
[0017] Step (1.4) Determine the update function. The update function of the DAQ method is:
[0018] Q(s,a)←Q(s,a)+α[R+γmax a′ Q(s′,a′)-Q(s,a)]
[0019] where s is the state, a is the selected action, Q(s,a) is the Q-value in the Q-table when the state is s and the action is a, α is the learning rate, γ is the decay factor, and R is the reward value.
[0020] Step (1.5) Determine the actions in the DAQ method. Each state State has three actions to choose from. The three actions are to add the action step size δ to w1, w2, or w3, and normalize the vector after adding the action step size.
[0021] Step (1.6) Determine the termination condition of the DAQ method. ① If the parallel application simulation execution overhead T obtained in consecutive n iterations is equal, then the DAQ stops; ② If the DAQ still does not meet condition ① after N iterations, then the DAQ stops.
[0022] Among them, step 2) includes the following steps:
[0023] Step (2.1) Initialize the Q-table so that the Q-value in the current Q-table is 0.
[0024] Step (2.2) randomly selects an initial state State0, and for the values of the state <w1, w2, w3>, it satisfies the conditions w1 + w2 + w3 = 1, w1 ≥ 0, w2 ≥ 0, w3 ≥ 0.
[0025] Step (2.3) If the cut-off condition in the DAQ algorithm is not satisfied, execute steps (2.3.1) to (2.3.5).
[0026] Step (2.3.1) Select the action with the maximum reward value in the current state, add the step size δ to the w in the State value, and perform vector normalization to obtain <w′1, w′2, w′3>. The process of transforming the state vector is the execution of the action. i Add the step size δ to the w in the State value, and perform vector normalization to obtain <w′1, w′2, w′3>. The process of transforming the state vector is the execution of the action.
[0027] Step (2.3.2) Use the selected behavior to obtain the next state. If the action is as shown in step (2.3.1), the value of the next state is <w′1, w′2, w′3>.
[0028] Step (2.3.3) Generate a corresponding parallel application multi-dimensional communication performance topology mapping scheme according to the state vector.
[0029] Step (2.3.4) Calculate the parallel application multi-dimensional communication overhead of the scheme in (2.3.3) and update the Q-table.
[0030] Step (2.3.5) Update State to the next state.
[0031] The advantages of the present invention include:
[0032] A method for optimizing and guiding the topology mapping of parallel application multi-dimensional communication performance based on Q-Learning proposed by the present invention has the following advantages compared with the prior art:
[0033] Existing topology mapping methods either do not consider the problem of index weight tuning or directly set the weights of each index in the optimization function to fixed values, which cannot meet the requirements of multi-dimensional communication performance optimization for large-scale parallel applications. In view of the above problems, this patent proposes a method that can accurately guide the optimization direction of the topology mapping method for multi-dimensional optimization indicators. This method is based on the Q-table structure, state variables, reward mechanism, update function, actions in the DAQ method, and cut-off conditions in the DAQ method of the Q-learning algorithm. By continuously checking the stop condition, performing the action with the maximum reward value on the existing state, and iteratively updating the Q-table, the weights of the optimization indicators of the best multi-dimensional optimization function can be obtained. The present invention can quickly and accurately solve the problem of weight adjustment of the multi-dimensional optimization function, thereby obtaining a better topology mapping scheme and reducing the communication overhead of parallel applications. Brief Description of the Drawings:
[0034] Figure 1 It is the structure diagram of the Q-table in the DAQ method.
[0035] Figure 2 It is the flowchart of the DAQ method. Specific implementation manner:
[0036] The present invention will be further described in detail below with reference to the accompanying drawings.
[0037] As Figure 1 shown, it is the structure diagram of the Q-table in the DAQ method proposed in this patent. The numbers 0, 1, 2, 3... in the first column of the Q-table represent the State numbers in the Q-table. Each State represents a three-dimensional vector <w0, w1, w2>. w1, w2, and w3 are the weights of the message delay overhead function f1, the link load balancing degree function f2, and the network congestion degree function f3 respectively. In addition, as the Q-table is updated, the number of states in the Q-table increases dynamically. The 0s in the second, third, and fourth columns in the figure are all obtained by taking the weights in the state State i in the parallel application multi-dimensional communication performance topology mapping scheme M i under which the simulated execution overhead T i of the parallel application, which is called Q-value and is initialized to 0 here. Each state State has three actions to choose from. The three actions are to add an action step δ to w1, w2, or w3 respectively, and normalize the vector after adding the action step, then the Q-table state after executing the action can be obtained.
[0038] As Figure 2 shown, it is the flowchart of the DAQ method proposed in this patent.
[0039] First, the structure of the Q-table needs to be designed, as Figure 1 shown. In addition, design state variables, reward mechanisms, update functions, actions in the DAQ method, and the cut-off condition of the DAQ method. The state variable of the Q-table is the weight vector <w1, w2, w3> of the multi-index optimization function, where w1, w2, and w3 respectively correspond to the weights of the message delay overhead function f1, the link load balancing degree function f2, and the network congestion degree function f3 in the multi-index optimization. The reward mechanism of the Q-table is the opposite of the communication overhead of the parallel application, that is, R = -T i . To ensure that the state selected each time approaches the minimum parallel application communication overhead. The update function of the Q-table is:
[0040] Q(s, a) ← Q(s,a) + α[R + γmax a′Q(s′, a′) - Q(s, a)
[0041] where s is the state, a is the selected action, Q(s, a) is the Q - value in the Q - table when the state is s and the action is a, α is the learning rate, γ is the decay factor, and R is the reward value. The actions in the DAQ method are to add the action step size δ to w1, w2, or w3 of the state variable respectively, and normalize the vector after adding the action step size. The stopping conditions of the DAQ method are: ① If the parallel application simulation execution overhead T obtained in n consecutive iterations is equal, then DAQ stops; ② If DAQ has not reached condition ① after N iterations, then DAQ stops.
[0042] Then initialize the Q - table, set the Q - value of the first row in the Q - table to 0, and then select an initial state <w1, w2, w3>. Next, determine whether the stopping condition of the algorithm is reached. If the condition is met, the algorithm stops, outputs the optimal State value, and generates a parallel application multi - dimensional communication performance topology mapping scheme accordingly. If the algorithm does not reach the stopping condition, loop through the following operations ① - ⑤. ① Select the action with the largest reward value in the current state, add the step size δ to w in the State value, and normalize the vector. ② Use the selected action to obtain the next state. ③ Generate a corresponding parallel application multi - dimensional communication performance topology mapping scheme according to the state vector. ④ Calculate the communication overhead of the parallel application under this scheme and update the value of the Q - table. ⑤ Update the current state to the next state. Finally, generate the corresponding parallel application multi - dimensional communication performance topology mapping optimization direction according to the state variables in the Q - table. i Finally, it should be noted that: The present invention may have many other application scenarios. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the present invention.
[0043] Finally, it should be noted that: The present invention may have many other application scenarios. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the present invention.
Claims
1. An optimized guidance method for multi-dimensional communication performance topology mapping of parallel applications based on Q-Learning, abbreviated as DAQ (Optimized Direction Adjustment Method based Q-Learning), guides the optimization direction of multi-dimensional communication performance topology mapping of parallel applications by using the Q-table update process in the Q-Learning algorithm. It is characterized in that, It includes the following steps: (1) Determine the Q-table structure, state variables, reward mechanism, update function, actions in the DAQ method, and the cut-off condition of the DAQ method; (2) According to the historical information recorded in the Q-table, use the reward mechanism to calculate the reward value obtained from the current state transition, execute the action with the maximum reward value, and iteratively update the Q-table until the cut-off condition of the DAQ method is reached, and output the state variable value in the Q-table at this time as the weight for multi-index optimization; (3) Generate the corresponding parallel application multi-dimensional communication performance topology mapping optimization direction based on the state variables in step (2); The specific process of step (1) includes: Step (1.1) Determine the Q-table structure. The first column of numbers 0, 1, 2, 3... in the Q-table represents the state variables in the Q-table. As the Q-table is updated, the number of its states increases dynamically. The remaining columns in the Q-table are the parallel application execution overheads of various dimensional communication performance topology mapping schemes under this state variable, which is called Q-value; Step (1.2) Determine the state variables. The state State in the Q-table is the weight vector <w1, w2, w3> of the multi-index optimization function, where w1, w2, and w3 respectively correspond to the weights of the message delay overhead function f1, the link load balancing degree function f2, and the network congestion degree function f3 in multi-index optimization; Step (1.3) determines the reward mechanism. The reward mechanism of the DAQ method is the opposite of the parallel application communication overhead, i.e., R = -T i , to ensure that the state selected each time approaches the minimum parallel application communication overhead; Step (1.4) Determine the update function. The update function of the DAQ method is: Q(s,a) ← Q(s,a) + α[R + γ max a′ Q(s′,a′) - Q(s,a)] where s is the state, a is the selected action, Q(s, a) is the Q-value in the Q-table when the state is s and the action is a, α is the learning rate, γ is the decay factor, and R is the reward value; Step (1.5) Determine the actions in the DAQ method. Each state State has three actions to choose from. The three actions are to add the action step size δ to w1, w2, or w3 respectively, and normalize the vector after adding the action step size; Step (1.6) Determine the cut-off condition of the DAQ method. ① If the parallel application simulation execution overhead T obtained from consecutive n iterations is equal, the DAQ stops; ② If the DAQ has not reached condition ① after N iterations, the DAQ stops.
2. The method according to claim 1, wherein The specific process of step (2) includes: Step (2.1) Initialize the Q-table so that the Q-value in the current Q-table is 0; Step (2.2) Randomly select an initial state State0, and for the values <w1, w2, w3> of the state, satisfy the conditions w1 + w2 + w3 = 1, w1 ≥ 0, w2 ≥ 0, w3 ≥ 0; Step (2.3) If the cut-off condition in the DAQ algorithm is not met, execute steps (2.3.1) to (2.3.5); Step (2.3.1) selects the action with the maximum reward value in the current state, adds the action step size δ to the w in the State value, and performs vector normalization to obtain <w′1, w′2, w′3>. The process of transforming the state vector is the execution of the action; i Add the action step size δ to the w in the State value, and perform vector normalization to obtain <w′1, w′2, w′3>. The process of transforming the state vector is the execution of the action; Step (2.3.2) Use the selected behavior to obtain the next state. If the action is as shown in step (2.3.1), the value of the next state is <w′1, w′2, w′3>; Step (2.3.3) generates a corresponding parallel application multi-dimensional communication performance topology mapping scheme according to the state vector; Step (2.3.4) calculates the parallel application multi-dimensional communication overhead of the scheme in step (2.3.3) and updates the Q-table; Step (2.3.5) updates the State to the next state.
Citation Information
Patent Citations
Dynamic optimization method of LTE and WiFi coexistence competition window value based on Q-learning algorithm
CN108924944A
Power line communication network topology control method, device, equipment and medium
CN112714064A