Grid optimization method based on double-delay depth deterministic strategy gradient algorithm
The TD3 algorithm is used to autonomously optimize the mesh, solving the problems of excessive manual intervention and low efficiency in traditional methods. It achieves efficient and automated mesh optimization, which is suitable for scenarios such as finite element analysis, semiconductor device simulation, and fluid mechanics simulation.
Patent Information
- Application Number
- CN202510745340.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional grid optimization methods rely heavily on manual intervention, have low optimization efficiency, and are prone to falling into local optimality, making it difficult to meet the needs of high-precision, large-scale simulation in modern industry.
A grid optimization method based on the TD3 deep reinforcement learning algorithm is adopted. By constructing an Actor-Dual Critic network architecture and a delayed policy update mechanism, the state space, continuous action space and compound reward function are defined to achieve autonomous optimization of the grid.
It improves the computational efficiency and stability of mesh optimization, realizes efficient and automated mesh optimization, supports seamless integration of industrial standard format files, generates quality assessment reports, and enhances the mesh optimization effect of simulation software.
Smart Images

Figure CN120633411A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of engineering simulation grid optimization, and in particular relates to a grid optimization method based on a Twin Delayed Deep Deterministic policy gradient (TD3) algorithm. Background Art
[0002] In the field of engineering simulation, meshing quality is one of the core factors that determine the accuracy and efficiency of numerical simulation. High-quality meshes have uniform aspect ratios, moderate internal angles, and good orthogonality. These properties ensure that the interpolation function can accurately approximate the true solution and avoid integral errors caused by the ill-conditioned Jacobian matrix. In nonlinear problems, high-quality meshes can reduce numerical oscillations and help the solver converge faster. Commonly used mesh optimization methods are divided into heuristic rule-based algorithms and optimization-based algorithms. Algorithms based on heuristic rules, such as Laplacian smoothing and angle-based methods, do not explicitly construct the objective function and perform mesh optimization based on empirical rules or heuristic rules. This type of method is simple to calculate but has problems such as unsatisfactory optimization results and the possibility of generating invalid elements. Optimization-based algorithms, such as the centroidal Voronoi tessellation (CVT) and optimal Delaunay tessellation (ODT), explicitly construct objective functions (such as maximizing mesh quality and minimizing element distortion) and use optimization methods to optimize the mesh by maximizing or minimizing the objective function. However, these methods have high computational complexity and are prone to local optimality or slow convergence when processing large-scale meshes.
[0003] In recent years, deep reinforcement learning technology has provided new ideas for grid optimization due to its autonomous decision-making and continuous space exploration capabilities. TD3 is an advanced algorithm in deep reinforcement learning. By improving the Deep Deterministic Policy Gradient (DDPG) algorithm and introducing technologies such as dual critic networks, delayed policy updates, and target policy smoothing, it effectively solves the problem of action over-estimation and significantly improves training stability. This paper proposes a grid optimization method based on the TD3 deep reinforcement learning algorithm. By constructing a reinforcement learning environment containing a normalized node coordinate state space, a continuous action space, and a compound reward function, combined with the Actor-Dual Critic network architecture and the delayed policy update mechanism, autonomous optimization of the grid is achieved. This method can effectively improve the computational efficiency and automation level of grid optimization in engineering simulation, and meet the urgent needs of modern industry for high-precision, large-scale simulation. Summary of the Invention
[0004] This paper aims to overcome the problems of traditional mesh optimization methods, such as high manual intervention, low optimization efficiency, and susceptibility to local optimality. It proposes a mesh optimization method based on a double-delayed deep deterministic policy gradient algorithm. This method achieves autonomous mesh optimization, significantly improving optimization efficiency and stability compared to traditional mesh optimization methods. It provides a more efficient and automated mesh optimization solution for scenarios such as finite element analysis in engineering simulation, semiconductor device simulation, and fluid dynamics simulation.
[0005] To achieve the above object, the technical solution adopted by the present invention is: S1. Initialize the mesh data to be optimized: read the mesh file to be optimized in the industry standard format, extract the initial mesh data to be optimized, the initial mesh data to be optimized includes node identifiers, node coordinates and connection relationships between nodes; S2. Model the grid optimization reinforcement learning environment: Define a state space containing the normalized coordinates of the node to be optimized and the normalized coordinates of the nodes directly connected to it; define a continuous action space containing the movement direction and distance of the node to be optimized; and define a reward function containing rewards for reaching a grid quality threshold, rewards for improving grid quality, and penalties for nodes exceeding boundaries. S3. Build a grid optimization network model based on the TD3 algorithm: Build a TD3 deep reinforcement learning network architecture, consisting of an actor network (policy network), a dual critic network (Q-value evaluation network), and a target network. The actor network outputs a continuous action vector based on the current state, and the dual critic network uses independent parameters to calculate the Q-value of each state-action pair. S4. Iteratively train the grid optimization network model in step S3. The specific steps of iterative training are as follows: S41. Initialization: Set the learning rate of the Actor and Critic networks, the discount factor for calculating future rewards, the maximum number of iterations, etc., and initialize the network; Set the exploration noise parameter to explore the action space during training and avoid premature convergence to a suboptimal solution; Set the strategy noise parameter to add noise to the target strategy's actions to prevent overfitting; Set the policy network delay update frequency. When the number of Critic network updates reaches the set number, the Actor network will be updated once. Initialize the experience replay buffer to store the experience data generated during the learning process, set the number of samples in the experience replay buffer, and randomly select a specified number of samples from the experience replay buffer for calculation each time the strategy is updated; S42. Generating Actions via Actor Networks: In state Under this circumstance, the Actor network generates actions according to the current strategy , the agent performs an action Transition to a new state And get rewards according to the reward function , store experience data to the experience replay buffer; S43. Calculate the target action and target Q value: Randomly select a specified number of samples from the experience replay buffer for calculation, use the target actor network to calculate the target action, and use the target critic network to calculate the target Q value; S44. Update Network: Update the Critic network parameters. When the number of updates reaches the set number, update the Actor network parameters. After each update of the Critic network and Actor network parameters, update the corresponding target network parameters. When updating the target network parameters, you can use the soft update method. The target network parameters are gradually synchronized through the adjustable soft update coefficient to improve training stability. The specific soft update method can be: in, is the adjustable soft update coefficient, are the parameters of the target Actor network, It is Parameters of the target critic network; S45. Repeat steps S42 to S44. After several rounds of training, the network can automatically learn the grid optimization strategy through the reward value provided by the reward function. , the policy network can directly calculate the optimal action ; S5. Obtain optimized mesh data: Use the network trained in step S4 to optimize the mesh. Output the optimized mesh file in the same format as the mesh file to be optimized in step S1. Generate a mesh quality assessment report and visualization results.
[0006] Furthermore, in steps S1 and S5, the reading and output of industrial standard format mesh files are supported, including Gmsh's msh file, ABAQUS's inp file, and ANSYS's cdb file, and topology verification is performed on the mesh data during reading and output.
[0007] Furthermore, in step S2, the reward function is: in, is the reward item for grid quality reaching the threshold, is the node out-of-bounds penalty term, It is a reward item for improving grid quality. is the quality of the optimized mesh points, is the quality of the mesh point before optimization. The quality of the mesh point is the minimum value among all adjacent element qualities. The calculation formula of the element quality should be specifically defined according to the different element types. Taking the triangular element as an example, the calculation formula of the triangular element quality can be: in, is the area of the triangle unit, It is a triangle unit The length of the edge, in the reward function 、 、 Are all positive numbers, the node threshold is a value range of The specific values of the four can be adjusted according to the network training situation in step S4.
[0008] Furthermore, in step S5, the mesh quality assessment report includes mesh optimization time, average unit quality, minimum unit quality, unit information, and unit quality compliance percentage.
[0009] The present invention has at least the following beneficial technical effects: 1. Achieve efficient and automated grid optimization through the TD3 algorithm, breaking through the bottleneck of traditional methods.
[0010] This method uses the TD3 algorithm to achieve autonomous mesh optimization, achieving improved results, efficiency, and stability compared to traditional heuristic and optimization-based algorithms. By using a trained neural network, the mesh file to be optimized is input and the optimized mesh is output without manual intervention.
[0011] 2. Compatible with industrial standards and seamlessly integrated with simulation processes.
[0012] Mesh file input and output supports industry standard formats such as msh, inp, and cdb. Optimization results can be directly used in simulation software such as ANSYS and COMSOL. Upon completion of optimization, a quality assessment report is generated, providing key indicators such as mesh optimization time, average element quality, minimum element quality, element information, and the percentage of element quality meeting standards, helping engineers quickly verify mesh reliability.
[0013] 3. The TD3 algorithm introduces three major improvement technologies to increase the training convergence speed.
[0014] The TD3 algorithm builds on the DDPG algorithm by introducing technologies such as a dual critic network, delayed policy updates, and target policy smoothing. This effectively addresses the problem of action overestimation and significantly improves training stability. The TD3 algorithm supports a continuous action space, significantly reducing suboptimal solutions caused by discrete actions. It also uses an experience replay mechanism to enhance data utilization and avoid local optimality traps. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of a grid optimization method based on a double-delayed deep deterministic policy gradient algorithm provided by an embodiment of the present invention; Figure 2 Schematic diagram of the structure of a double-delayed deep deterministic policy gradient algorithm network provided by an embodiment of the present invention; Figure 3 3 is a comparison diagram before and after grid optimization provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below with reference to the accompanying drawings.
[0017] The grid optimization method of the present invention is implemented in the following steps: S1. Initialize the mesh data to be optimized: read the mesh file to be optimized in the industry standard format, extract the initial mesh data to be optimized, the initial mesh data to be optimized includes node identifiers, node coordinates and connection relationships between nodes; S2. Model the grid optimization reinforcement learning environment: Define a state space containing the normalized coordinates of the node to be optimized and the normalized coordinates of the nodes directly connected to it; define a continuous action space containing the movement direction and distance of the node to be optimized; and define a reward function containing rewards for reaching a grid quality threshold, rewards for improving grid quality, and penalties for nodes exceeding boundaries. S3. Build a grid optimization network model based on the TD3 algorithm: Build a TD3 deep reinforcement learning network architecture, consisting of an actor network (policy network), a dual critic network (Q-value evaluation network), and a target network. The actor network outputs a continuous action vector based on the current state, and the dual critic network uses independent parameters to calculate the Q-value of each state-action pair. S4. Iteratively train the grid optimization network model in step S3. The specific steps of iterative training are as follows: S41. Initialization: Set the learning rate of the Actor and Critic networks, the discount factor for calculating future rewards, the maximum number of iterations, etc., and initialize the network; Set the exploration noise parameter to explore the action space during training and avoid premature convergence to a suboptimal solution; Set the strategy noise parameter to add noise to the target strategy's actions to prevent overfitting; Set the policy network delay update frequency. When the number of Critic network updates reaches the set number, the Actor network will be updated once. Initialize the experience replay buffer to store the experience data generated during the learning process, set the number of samples in the experience replay buffer, and randomly select a specified number of samples from the experience replay buffer for calculation each time the strategy is updated; S42. Generating Actions via Actor Networks: In state Under this circumstance, the Actor network generates actions according to the current strategy , the agent performs an action Transition to a new state And get rewards according to the reward function , store experience data to the experience replay buffer; S43. Calculate the target action and target Q value: Randomly select a specified number of samples from the experience replay buffer for calculation, use the target actor network to calculate the target action, and use the target critic network to calculate the target Q value; S44. Update Network: Update the Critic network parameters. When the number of updates reaches the set number, update the Actor network parameters. After each update of the Critic network and Actor network parameters, update the corresponding target network parameters. When updating the target network parameters, you can use the soft update method. The target network parameters are gradually synchronized through the adjustable soft update coefficient to improve training stability. The specific soft update method can be: in, is the adjustable soft update coefficient, are the parameters of the target Actor network, It is Parameters of the target critic network; S45. Repeat steps S42 to S44. After several rounds of training, the network can automatically learn the grid optimization strategy through the reward value provided by the reward function. , the policy network can directly calculate the optimal action ; S5. Obtain optimized mesh data: Use the network trained in step S4 to optimize the mesh. Output the optimized mesh file in the same format as the mesh file to be optimized in step S1. Generate a mesh quality assessment report and visualization results.
[0018] In steps S1 and S5, the grid files in industrial standard formats are supported for reading and output, including Gmsh's msh files, ABAQUS's inp files, and ANSYS's cdb files, and topology verification is performed on the grid data during reading and output.
[0019] In step S2, the reward function is: in, is the reward item for grid quality reaching the threshold, is the node out-of-bounds penalty term, It is a reward item for improving grid quality. is the quality of the optimized mesh points, is the quality of the mesh point before optimization. The quality of the mesh point is the minimum value among all adjacent element qualities. The calculation formula of the element quality should be specifically defined according to the different element types. Taking the triangular element as an example, the calculation formula of the triangular element quality can be: in, is the area of the triangle unit, It is a triangle unit The length of the edge, in the reward function 、 、 Are all positive numbers, the node threshold is a value range of The specific values of the four can be adjusted according to the network training situation in step S4.
[0020] In step S5, the mesh quality assessment report includes mesh optimization time, average unit quality, minimum unit quality, unit information, and unit quality compliance percentage.
[0021] It should be noted that the specific implementation methods listed in this embodiment are only used to illustrate the core design concept and implementation of the present invention, and are not intended to limit the scope of patent protection. Those skilled in the relevant fields should recognize that, under the premise of complying with the principles of the present invention, the technical means described in the embodiment can be adaptively adjusted or replaced with equivalent technologies. Such modifications and equivalent replacements of technical solutions should be deemed to fall within the scope of protection defined by the claims of the present invention.
Claims
1. A grid optimization method based on the Twin Delayed Deep Deterministic policy gradient (TD3) algorithm, characterized in that: Follow these steps to achieve this: S1. Initialize the grid data to be optimized: Reading a grid file to be optimized in an industrial standard format and extracting initial grid data to be optimized, wherein the initial grid data to be optimized includes node identifiers, node coordinates, and connection relationships between nodes; S2. Modeling a Grid-Optimized Reinforcement Learning Environment: Define the state space, which includes the normalized coordinates of the current node to be optimized and the normalized coordinates of the nodes directly connected to it; define the continuous action space, which includes the movement direction and movement distance of the current node to be optimized; define the reward function, which includes the reward for grid quality reaching the threshold, the reward for grid quality improvement, and the penalty for nodes exceeding the boundary; S3. Constructing a grid optimization network model based on the TD3 algorithm: Build a TD3 deep reinforcement learning network architecture, consisting of an actor network (policy network), a dual critic network (Q-value evaluation network), and a target network. The actor network outputs a continuous action vector based on the current state, and the dual critic network uses independent parameters to calculate the Q value of each state-action pair. S4. Iteratively train the grid optimization network model in step S3; S5. Get optimized mesh data: Using the network trained in step S4, optimize the grid; Output the optimized mesh file in the same format as the mesh file to be optimized in step S1, and generate a mesh quality assessment report and visualization results.
2. The grid optimization method based on the double-delayed deep deterministic policy gradient algorithm according to claim 1, characterized in that: In steps S1 and S5, the grid files in industrial standard formats are supported for reading and output, including Gmsh's msh files, ABAQUS's inp files, and ANSYS's cdb files, and topology verification is performed on the grid data during reading and output.
3. The grid optimization method based on the double-delayed deep deterministic policy gradient algorithm according to claim 1, characterized in that: In step S2, the reward function is: in, is the reward item for grid quality reaching the threshold, is the node out-of-bounds penalty term, It is a reward item for improving grid quality. is the quality of the optimized grid points, It is the quality of the grid point before optimization. The quality of the grid point is the minimum value among all adjacent unit masses. The calculation formula of unit quality should be specifically defined according to different unit categories.
4. The grid optimization method based on the double-delayed deep deterministic policy gradient algorithm according to claim 1, characterized in that: In step S4, the specific steps of the iterative training are: S41. Initialization: Set the learning rate of the Actor and Critic networks, the discount factor for calculating future rewards, the maximum number of iterations, etc., and initialize the network; Set the exploration noise parameter to explore the action space during training and avoid premature convergence to a suboptimal solution; Set the strategy noise parameter to add noise to the target strategy's actions to prevent overfitting; Set the policy network delay update frequency. When the number of Critic network updates reaches the set number, the Actor network will be updated once. Initialize the experience replay buffer to store the experience data generated during the learning process, set the number of samples in the experience replay buffer, and randomly select a specified number of samples from the experience replay buffer for calculation each time the strategy is updated; S42. Generating Actions via Actor Networks: In state Under this circumstance, the Actor network generates actions according to the current strategy , the agent performs an action Transition to a new state And get rewards according to the reward function , store experience data to the experience replay buffer; S43. Calculate the target action and target Q value: Randomly select a specified number of samples from the experience replay buffer for calculation, use the target actor network to calculate the target action, and use the target critic network to calculate the target Q value; S44. Update Network: Update the Critic network parameters. When the number of updates reaches the set number, update the Actor network parameters. After each update of the Critic network and Actor network parameters, update the corresponding target network parameters. S45. Repeat steps S42 to S44. After several rounds of training, the network can automatically learn the grid optimization strategy through the reward value provided by the reward function. , the policy network can directly calculate the optimal action .
5. The grid optimization method based on the double-delayed deep deterministic policy gradient algorithm according to claim 1, characterized in that: In step S5, the mesh quality assessment report includes mesh optimization time, average unit quality, minimum unit quality, unit information, and unit quality compliance percentage.