Laser power parameter optimization method based on deep reinforcement learning
Through a method based on deep reinforcement learning, a DQN network model is established to correlate the melt pool state and laser power to generate a laser power adjustment strategy, which solves the problem of inaccurate process parameter adjustment in laser powder bed melt additive manufacturing, realizes efficient optimization of process parameters, and improves the quality stability of workpieces.
Patent Information
- Application Number
- CN202510250395.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
AI Technical Summary
In the process of melt additive manufacturing of laser powder beds, it is difficult for the prior art to accurately adjust process parameters, resulting in workpieces being prone to defects such as pores, deformation and unfusion. The process parameter adjustment method is highly limited and is only applicable to specific workpieces.
Using a method based on deep reinforcement learning, the DQN network model is designed, the melt pool state is correlated with the laser power, and the laser power adjustment strategy is generated using the ε-greedy algorithm, and the network model parameters are optimized through feedback to realize real-time adjustment of the laser power to optimize the process parameters.
It effectively reduces the overheating phenomenon caused by changes in complex geometric structures and laser scanning trajectory, reduces the fluctuation of the melting pool depth, and significantly improves the quality stability of the laser powder bed melting process.
Smart Images

Figure CN120145855A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of simulation and artificial intelligence. Specifically, it relates to a method for optimizing laser power parameters based on deep reinforcement learning. Background Art
[0002] Laser powder bed fusion is a common additive manufacturing technology for manufacturing metal components, featuring flexible design and high resource utilization efficiency. However, the formation principle of LPBF is different from that of traditional processing technologies, and various types of pore defects are likely to form. Therefore, corresponding process parameters need to be formulated to improve the mechanical properties of workpieces and the processing quality.
[0003] During the process of laser powder bed fusion additive manufacturing of complex geometric parts, in order to reduce workpiece defects and improve workpiece quality, operators set process parameters based on experience. The process parameters include: scanning speed, laser power, powder spreading speed, etc. Currently, the adjustment of process parameters mainly relies on manual experience, which is laborious and cumbersome, and the resulting defects are not well controlled. To simplify the adjustment of process parameters, the laser power parameter with a relatively large correlation with the melt pool depth is selected for adjustment, while other process parameters remain unchanged.
[0004] In recent years, with the development of computer vision technology, using machine vision to directly observe the melt pool, obtain geometric information of the melt pool characteristics, and perform closed-loop control on process parameters has become an important research direction in additive manufacturing technology. In actual production, the method of processing melt pool images based on image processing algorithms to obtain melt pool morphology and size parameters has problems of poor generalization performance and low accuracy. In actual production, the melt pool depth needs to be measured by damaging the workpiece. A large number of experiments are required to obtain large sample data, which is laborious and has low accuracy, resulting in inaccurate adjustment of process parameters. In addition, the process parameter adjustment method has limitations and can only be applied to specific workpieces. Summary of the Invention
[0005] Aiming at the defects and deficiencies existing in the prior art in the dynamic adjustment strategy of laser power, the present invention proposes a method for optimizing laser power parameters based on deep reinforcement learning, aiming to solve the problems of defects such as pores, deformation, and lack of fusion in the additive manufacturing process. By designing and training a DQN network model, the melt pool state is associated with the laser power, and the ε-greedy algorithm is adopted to generate a laser power adjustment strategy. After executing the adjustment strategy, the generated results are fed back to optimize the parameters of the DQN network model. Therefore, the network model can be updated in real time to improve the ability to handle the laser power adjustment problem under complex working conditions.
[0006] The technical solution for implementing the present invention is as follows: A method for optimizing laser power parameters based on deep reinforcement learning, which adjusts the laser power in real time during the path scanning process through a deep reinforcement learning algorithm to reduce the molten pool depth fluctuation and heat accumulation, includes the following steps:
[0007] S1. Construct an LPBF simulation model, establish a distribution model for the three-dimensional temperature field of the molten pool using a moving heat source, simulate the temperature distribution on the laser scanning path, and realize the simulation calculation of the molten pool temperature field;
[0008] S2. Import the STL file of the part, perform layer slicing on the STL file, and obtain the laser scanning path of each layer after setting the scanning parameters;
[0009] S3. In the LPBF simulation model, simulate and calculate the part forming process according to the laser scanning path, extract the molten pool state and state characteristics in real time, construct a state space, an action space, and construct a reward function composed of two types of parameters, obtain the reward value through environmental interaction, and guide the action strategy to adjust in the direction of maximizing the reward value;
[0010] S4. Construct a deep reinforcement learning DQN network model, including a current Q network and a target Q network;
[0011] S5. Estimate the value function Q based on the DQN network model, obtain a molten pool state probability π through the ε-greedy algorithm, and then generate the optimal laser adjustment action for the current molten pool state characteristics;
[0012] S6. Feedback on the executed optimal laser adjustment action, integrate the data generated during the agent interaction process, store it as an experience sample in the experience pool, and use it to train the DQN network model and optimize the network model parameters;
[0013] S7. The agent learns and executes the strategy of adjusting the laser power action, making the molten pool state characteristics close to the target state characteristics.
[0014] In the above S1, the process of building the LPBF simulation environment is as follows:
[0015] S1.1. To improve the simulation calculation efficiency and provide support for reinforcement learning, the complex multi-scale effect is simplified to the continuous temperature distribution of the material: only the conduction mode of heat transfer is considered, the thermal properties are independent of temperature, and the powder bed is modeled as a solid continuum; the above process is modeled as two-dimensional conduction related to the moving heat source on the powder, and the temperature field equation is given by the following partial differential equation:
[0016]
[0017] where T(x,t) is a function of position x and time t, representing the temperature at a certain point; is the rate of change of temperature T(x,t) with respect to time t; is the thermal diffusivity, defined by the formula where k is the thermal conductivity of the powder, ρ is the mass density of the powder, and c p is the heat capacity coefficient of the powder; is the second spatial derivative of T(x,t), i.e., the degree of temperature diffusion in the x direction; Θ(x,t) represents the normalized heat source, a parameter normalized according to density and specific heat capacity where Q(x,t) represents the rate of heat generation per unit volume;
[0018] After solving with the Green's function, we get:
[0019] T(x,t) = T(x) = T Q (x) + T D (x)
[0020] The first term coefficient T Q (x) represents the influence of the heat source on the temperature, and the second term coefficient T Q (x) represents the thermal diffusion effect;
[0021] Define simulation parameters according to the workpiece material properties. The simulation parameters include the absorption rate and the thermal conductivity;
[0022] S1.2. The influence of the heat source is solved through the Eagar-Tsai model. When solving the temperature field of the molten pool with the Eagar-Tsai model, the heat source is parameterized as a circular area moving on the workpiece surface, and the heat is distributed in a Gaussian manner. The specific calculation model is as follows:
[0023]
[0024] where θ(x,y,z,t) represents the heat source intensity per unit volume of powder at spatial position (x,y,z) and time t, η is the absorption rate of the material, P is the laser power, σ is the diameter of the laser, v is the laser scanning speed, δ(z) represents the spatial coordinate for z = 0, δ(z) ≠ 0, and the heat source only acts on the workpiece surface;
[0025] The calculation model for the temperature distribution at coordinate (x,y,z) is T l as follows:
[0026]
[0027] In the formula, T l represents the temperature at a certain moment, P represents the laser power, τ is the normalized time, related to the time scale of the heat conduction process, i.e., the duration of the heat source action; ξ is the integration variable, which is a parameter related to the normalized time. Δt is the action time of the laser at the position (x, y, z). exp represents the exponential function, which describes the distribution characteristics of the heat source in space;
[0028] S1.3. The heat diffusion effect occurs near the domain boundary. The distance equal to 4σ from the center of the heat source G is defined as the domain boundary, and the diffusion length scale For the region outside the boundary, a virtual heat source is simulated at the same distance on the other side of the boundary. Adjust the following equation to implement the virtual heat source T l :
[0029]
[0030] represents the thermal diffusion effect in the x direction, represents the thermal diffusion effect in the y direction, represents the thermal diffusion effect in the z direction.
[0031] The process of slicing the STL file of the part in S2 to generate the scanning path planning is as follows:
[0032] S2.1. The part STL model is composed of triangular patches. Before performing pre-slicing, first extract the Z-axis coordinate range (Zmin, Zmax) of the triangular patches, and sort all the triangular patches from Zmin to Zmax; set the single-layer thickness, generate multiple slicing planes perpendicular to the Z-axis, with the starting coordinate of Zmin and the spacing of the single-layer thickness, that is, multiple sub-layers. Respectively find the intersection lines of each slicing plane and the triangular patches, and connect them in sequence to obtain the contour curve of the single-layer sliced graph, and finally complete the slicing of the model by multiple slicing planes;
[0033] S2.2. After the slicing is executed, according to the actual set scanning parameters, the scanning parameters include the island width, the shadow line angle, and the filling spacing, and define the scanning vector and the side length; fill the sliced contour based on the Z-shaped scanning algorithm to generate the laser scanning path planning, and then obtain the scanning path planning file for each layer.
[0034] The process of constructing the state and action space and the reward function in S3 is as follows:
[0035] S3.1. The state space is defined as the record of the temperature field state in a specific view and direction. Set an ω×ω X-Y central rectangular area centered at the current scanning position. The X-Z and Y-Z cross-sections extend ω downward along the horizontal plane, record the temperature field distribution on the X-Y, Y-Z, and X-Z cross-sections, and obtain the real-time molten pool temperature field M i , Respectively represent the temperature field distributions of the molten pool in the x-y, x-z, and y-z cross-sections;
[0036]
[0037] The element in the x-th row and y-th column of the matrix represents the temperature of the molten pool at the (x, y) coordinates in the x-y cross-section. The representation methods for the molten pool states in the x-z and y-z cross-sections are the same; the molten pool state S i =[M i-2 , M i-1 , M i , where M i-2 and M i-1 represent the molten pool states at the first two time steps;
[0038] S3.2. Molten pool depth D i As a key feature of the molten pool state, the specific extraction method is as follows:
[0039] Interpolate the temperature field along the axis perpendicular to the scanning direction, find the position with the highest temperature on the horizontal plane, and interpolate downward along the Z-axis at this point to find the critical point where the temperature is closest to the melting point of the material. The molten pool depth is the vertical distance from this point to the horizontal plane;
[0040] S3.3. Construct a continuous action space A = {a i |i = 1, 2, … n} for adjusting the laser power by deep reinforcement learning. a i represents the action at the i-th step. The selection range of the action parameter is [λ 1 , λ 2 , where λ 1 and λ 2 are real numbers, and the step size is c. From this, the formula for adjusting the laser power with the action parameter is deduced:
[0041] P i+1 = a i+1 ×(P max - P min ) + P min
[0042] where P i+1 is the laser power at the (i + 1)-th step, a i+1 represents the action parameter at the (i + 1)-th step, P max represents the maximum laser power in the previous i steps, and P min represents the minimum laser power in the previous i steps;
[0043] The reward function is composed of the absolute error between the target molten pool depth and the current molten pool depth, and a regularization term is combined to avoid abnormal behaviors; the regularization term prevents the molten pool depth from fluctuating violently by punishing the gap between the maximum and minimum molten pool depths observed during the period. The expression of the reward function R is:
[0044]
[0045] Wherein, D target is the set target molten pool depth, D melt is the molten pool depth obtained by simulation calculation, max i D melt is the maximum molten pool depth, min i D melt is the minimum molten pool depth, ∑ i represents the summation of the numerical values of the first i iterative time steps.
[0046] The construction processes of the current Q network and the target Q network in the S4 are specifically as follows:
[0047] The current Q network is used to select the action a of the agent according to the current state s, and the target network is used to calculate the maximum value function Q of the next state s';
[0048] The current network and the target network have the same structure, including the following parts: an input layer for receiving the current state characteristics of the system; a plurality of hidden layers including at least one convolutional layer and a fully connected layer for extracting state characteristics and estimating the value function; and an output layer for providing the optimal value function of the action. The iterative formula of the value function Q is expressed as:
[0049] Q(S i+1 ,A i+1 )←Q(Si,A i )+α[R i +γmaxQ(S i+1 ,A i+1 )-Q(S i ,A i )]
[0050] Where α is the learning rate, γ is the reward discount factor, and its value range is [0,1]. When it is 0, it means only considering the influence of the current action on the current situation and not considering the influence on subsequent steps. When it is 1, it means that the current action has an equal influence on each subsequent step; taking the state S i as the input, the Q values under different actions A i are output, that is, a set of Q values Q(S i ,A i ) containing all actions is output:
[0051] Q(S i ,A i )={Q{s 1 ,a 1},Q{s 1 ,a 2},…,Q{s1 , a m}, Q{s 2 , a 1}, … Q{s n , a m}}
[0052] Among them, s n is the nth state in the state space, and a m is the mth state in the action space.
[0053] The process of generating the S5 action selection strategy is as follows:
[0054] The ε-greedy strategy defines an exploration rate ε, which takes values between 0 and 1; when selecting an action each time, a random number between 0 and 1 is generated. If the random number is less than ε, a random action is selected; if it is greater than or equal to ε, the optimal action is selected based on the learned information;
[0055] Specifically, the ε-greedy algorithm calculates the probability π of the agent selecting each action in a given state through the following formula:
[0056]
[0057] π(a i | s) represents the probability of selecting action a in state s i , ε is the exploration parameter, indicating that a non-optimal action is selected with a certain probability, N represents the total number of executable actions, represents the probability of selecting action a from N actions with probability ε, and Q(s i , a i ) represents the expected reward of selecting action a in state s i , and a i is the optimal action with the maximum Q value in the current state s: * i
[0058] The process of training the DQN model in S6 is as follows:
[0059] S6.1. Select and execute action a through the ε-greedy strategy, and integrate the current moment molten pool state s i , action a i , reward function value r i , and the next moment molten pool state s i+1 into an experience sample (s i , a i , r i , s i+1 ), and store it in the experience pool Ω = {(S i , Ai , R i , S i+1 ) | i = 1, 2, …, ξ};
[0060] S6.2. Randomly select a batch of samples from the experience pool for learning. For each sample (S i , A i , R i , S i+1 ) calculate the target Q value:
[0061] y i = R i + γ maxQ(S i+1 , A'; μ - )
[0062] y t is the target Q value, γ is the discount factor, A' is the possible action at the next moment of S i+1 , maxQ(S i+1 , A'; μ - ) is the maximum target value generated by executing action A' at the next moment of state S i+1 ; μ - is the target Q network parameter;
[0063] S6.3. Based on the loss function, perform gradient descent to train the current Q network to minimize the mean square error between the predicted Q value of the network and the target, and update the current network model parameters μ and the exploration rate ε. The expression of the loss function in the iteration process:
[0064]
[0065] μ i represents the current network parameter in the i-th iteration step, represents the target network parameter in the i-th iteration step. After each iteration, it is μ i that is updated, not μ i - , and the target network parameter is set to the current network parameter μ of the (i - 1)-th iteration i-1 ; Q(s i , a i , μ i ) represents the neural network approximate value function, (s, a, r, s i+1 ) ~ U(Ω) represents the experience data of the agent, s represents the current bath state in the i-th iteration step, a represents the action selected in the current i-th step, s i+1 represents the next bath state, and a represents the action to be executed in the next step;
[0066] S6.4. Regularly update the weight parameters of the target Q value network Update the exploration rate ε;
[0067] S6.5. Repeat the above S6.1 - S6.4 until the DQN network model converges or reaches the preset number of training rounds.
[0068] The process of the S7 agent executing the laser power optimization strategy is as follows:
[0069] S7.1. The learning process of the power control agent is divided into four parts in each iteration: input feature generation, action decision-making, reliability evaluation and reward calculation, and training of the deep Q network:
[0070] First, extract the molten pool state feature s i and the state feature Ω i , and initialize the state space set S = {s i |i = 0, 1, 2, …, ξ}; then generate a policy through the action prediction module and select an appropriate action; subsequently, evaluate the generated policy and optimize the network parameters μ i through batch training; when entering the next iteration, update the initial state s i+1 , and regenerate a policy based on the current network with the latest parameters μ i ; in the first iteration, the decision-making process uses the current network structure initialized randomly with parameters μ 0 ;
[0071] S7.2. After the training is completed, input the molten pool state set S new of other printing layers and the state feature Ω new for model testing. The agent generates and executes laser power adjustment actions according to the input; for each time step, the agent optimized by DQN will use the input feature Ω i to predict the prior information, and select an action with a higher return based on the ε - greedy strategy to dynamically adjust the laser power; finally, by executing a series of actions with high returns, the agent generates a complete laser power control strategy.
[0072] The present invention also discloses a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer-executable instruction is executed, the method is implemented.
[0073] Compared with the prior art, the present invention has the following remarkable advantages:
[0074] (1) The present invention introduces the DQN deep reinforcement learning method, and maximizes the target reward by iteratively optimizing the policy network to generate an accurate laser power control strategy. This strategy can effectively reduce the overheating phenomenon caused by complex geometric structures and changes in laser scanning trajectories, and at the same time reduce the volatility of the molten pool depth, thereby significantly improving the quality stability of the laser powder bed fusion process.
[0075] (2) The present invention has the potential to be widely extended to other closed-loop control fields, especially in scenarios where complex physical effects need to be simulated. By introducing a deep network to model the physical effects not covered by traditional models, not only can the accuracy of the environment's description of the actual working conditions be improved, but also a computationally efficient and accurate reinforcement learning environment can be constructed, thus providing a solid foundation for optimizing control strategies. This method can dynamically adapt to complex working conditions, making the control system more robust and adaptable in diverse application scenarios, while providing a new solution for improving the efficiency and stability of the overall manufacturing process. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a system flow chart of the laser power parameter optimization method based on DQN.
[0077] Figure 2 It is a schematic diagram of the laser scanning path planning based on the cuboid model.
[0078] Figure 3 It is a schematic diagram of the real-time molten pool state extraction based on single-layer thermal simulation.
[0079] Figure 4 It is a schematic diagram of the DQN network structure.
[0080] Figure 5 It is a schematic diagram of the laser power optimization result based on the DQN algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0082] The technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions conflicts with each other or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0083] Next, the specific implementation manners, as well as the technical difficulties and inventive points of the present invention, will be further introduced in conjunction with the design examples of the present invention.
[0084] As Figure 1 shown, the laser power parameter optimization method based on DQN deep reinforcement learning is as follows:
[0085] Step 1: Build an LPBF (Laser Powder Bed Fusion) simulation model. Use a moving heat source to establish a distribution model for the three-dimensional temperature field of the molten pool, simulate the temperature distribution along the laser scanning path, and achieve the simulation calculation of the temperature field of the molten pool. To improve the simulation calculation efficiency and provide support for reinforcement learning, complex multi-scale effects are simplified to the continuous temperature distribution of the material: only the conduction mode of heat transfer is considered, the thermal properties are independent of temperature, and the powder bed is modeled as a solid continuum; the above process is modeled as two-dimensional conduction related to the moving heat source on the powder, and the temperature field equation is given by the following partial differential equation:
[0086]
[0087] where T(x, t) is a function of position x and time t, representing the temperature at a certain point; is the rate of change of temperature T(x, t) with respect to time t; is the thermal diffusivity, defined by the formula where k is the powder thermal conductivity, ρ is the powder mass density, and c p is the powder heat capacity coefficient; is the second spatial derivative of T(x, t), that is, the degree of temperature diffusion along the x direction; Θ((x), t) represents the normalized heat source, the parameter normalized according to density and specific heat capacity where Q(x, t) represents the heat generation rate per unit volume.
[0088] After solving by the Green's function, we get:
[0089] T(x, t) = T(x) = T Q (x) + T D (x)
[0090] The first term T Q (x) represents the influence of the heat source on the temperature, and the second term T D (x) represents the heat diffusion effect.
[0091] Define simulation parameters according to the workpiece material properties. The simulation parameters include absorptivity and thermal conductivity.
[0092] The influence of the heat source T Q (x) can be solved by the Eagar-Tsai model. When the Eagar-Tsai model solves the temperature field of the molten pool, the heat source is parameterized as a circular area moving on the workpiece surface, and the heat is distributed in a Gaussian manner. The specific calculation model is as follows:
[0093]
[0094] Among them, θ(x, y, z, t) represents the heat source intensity per unit volume of powder at the spatial position (x, y, z) and time t, η is the absorption rate of the material, P is the laser power, σ is the diameter of the laser, v is the scanning speed of the laser, δ(z) represents the spatial coordinate for z = 0, δ(z) ≠ 0, and the heat source only acts on the workpiece surface.
[0095] Calculation model T of the temperature distribution at the coordinate (x, y, z) l is as follows:
[0096]
[0097] In the formula, T l represents the temperature at a certain moment, P represents the laser power, τ is the normalized time, which is related to the time scale of the heat conduction process, that is, the duration of the heat source action; is the integration variable, which is a parameter related to the normalized time, Δt is the action time of the laser at the position (x, y, z), exp represents the exponential function, which describes the distribution characteristics of the heat source in space, and the specific parameters of the AlSi10Mg powder material are shown in Table 1:
[0098] Table 1 Thermal parameters of the powder material
[0099]
[0100] The heat diffusion effect T D (x) occurs near the domain boundary. The distance equal to 4σ from the heat source center G is defined as the domain boundary, and the diffusion length scale For the region outside the boundary, a virtual heat source is simulated at the same distance on the other side of the boundary, and the following equation is adjusted to implement the virtual heat source T l :
[0101]
[0102] represents the heat diffusion influence in the x direction, represents the heat diffusion influence in the y direction, represents the heat diffusion influence in the z direction.
[0103] Step 2: Select a cuboid specimen as the simulation model, and the specific dimensions are: L x = 10 mm, L x = 10 mm, L z= 3 mm, import the STL model of the specimen. The STL model of the part is composed of triangular facets. Extract the Z-axis coordinate range (0, 3) of the triangular facets, and sort all the triangular facets from 0 to 3; set the single-layer thickness to 30 μm, generate multiple slicing planes perpendicular to the Z-axis, with the starting coordinate (0, 0, 0) and the spacing equal to the single-layer thickness, that is, multiple slicing levels. Respectively find the intersection lines of each slicing plane and the triangular facets, and connect them in sequence. As Figure 2 shown, obtain the laser path planning contour curve of the single-layer slicing pattern, and finally complete the slicing of the model by multiple slicing levels.
[0104] After slicing is executed, set the scanning parameters. The scanning parameters include an island width of 2 mm, a single-layer scanning angle of 0°, an inter-layer rotation angle of 90°, and a filling spacing of 0.1 mm. The scanning vector gives priority to the X direction; fill the slicing contour based on the Z-shaped scanning algorithm to generate a laser scanning path plan, and then obtain the scanning path plan files for each layer.
[0105] Step 3: Extract the molten pool state and state features in real time, construct the state space and action space, and construct a reward function composed of two types of parameters. Obtain the reward value through environmental interaction, and guide the action strategy to adjust in the direction of maximizing the reward value.
[0106] The state space is defined as the record of the temperature field state in a specific view and direction: set a 160 μm × 160 μm X-Y central rectangular area centered at the current scanning position, and the X-Z and Y-Z cross-sections extend 160 μm downward along the horizontal plane. Record the temperature field distribution on the X-Y, Y-Z, and X-Z cross-sections to obtain the real-time molten pool temperature field M i , respectively represent the molten pool temperature field distributions on the x-y, x-z, and y-z cross-sections; set the initial laser power, scanning speed, and the simulation results at the initial moment are as Figure 3 shown.
[0107]
[0108] The element in the x-th row and y-th column of the matrix represents the molten pool temperature at the (x, y) coordinate on the x-y cross-section. The representation methods of the molten pool states on the x-z and y-z cross-sections are the same; the molten pool state is S i =[M i-2 ,M i-1 ,M i , M i-2 、M i-1 represent the molten pool states at the previous two time steps.
[0109] In order to extract the key feature, the molten pool depth D i, interpolate the temperature field along the axis perpendicular to the scanning direction to find the position with the highest temperature on the horizontal plane, and interpolate downward along the Z-axis at this point to find the critical point where the temperature is closest to the melting point of the material. The depth of the molten pool is the vertical distance from this point to the horizontal plane.
[0110] Construct a continuous action space A = {a i |i = 1, 2, …, ξ} for depth reinforcement learning to adjust the laser power. a i represents the action at the i-th step. The selection range of the action parameter is [-1, 1], and the step size is 0.1. From this, the formula for adjusting the laser power with the action parameter is deduced:
[0111] P i+1 = a i+1 ×(P max - P min ) + P min
[0112] where P i+1 is the laser power at the (i + 1)-th step, a i+1 represents the action parameter at the (i + 1)-th step, P max represents the maximum laser power in the previous i steps, and P min represents the minimum laser power in the previous i steps;
[0113] The reward function is composed of the absolute error between the target molten pool depth and the current molten pool depth, and a regularization term is combined to avoid the occurrence of abnormal behaviors; the regularization term prevents the molten pool depth from fluctuating violently by punishing the gap between the maximum and minimum molten pool depths observed within a period. The expression of the reward function R is:
[0114]
[0115] In the formula, D target is the set target molten pool depth, D melt is the molten pool depth calculated by simulation, max i D melt is the maximum molten pool depth, min i D melt is the minimum molten pool depth, and ∑ i represents the summation of the numerical values of the previous i iteration time steps.
[0116] Step 4: Construct a deep reinforcement learning DQN network model, including a current Q-network and a target Q-network; the current network is used to select the agent's action a according to the current state s, and the target network is used to calculate the maximum Q value of the next state s'; the structures of the current network and the target network are the same, mainly including the following parts: Input layer: 576 neurons, corresponding to including the molten pool state S at the current time step and the previous two time steps i. Fully connected layer: 2 layers, with 64 neurons in each layer, using the Tanh hyperbolic tangent function as the activation function; Output layer: 1 neuron, outputting the optimal value estimate Q of the action, and the network is constructed as Figure 4 shown. The iterative formula of the value function Q is expressed as:
[0117] Q(S i+1 ,A i+1 )←Q(S i ,A i )+α[R i +γmaxQ(S i+1 ,A i+1 )-Q(S i ,A i )]
[0118] where α is the learning rate, γ is the reward discount factor, and its value range is [0,1]. When it is 0, it means only considering the influence of the current action on the current situation and not considering the influence on subsequent steps. When it is 1, it means that the current action has an equal influence on each subsequent step. Using the state S i as the input, output the Q values under different actions A i , that is, output a set of Q values containing all actions:
[0119] Q(S i ,A i )={Q{s 1 ,a 1},Q{s 1 ,a 2},…,Q{s 1 ,a m},Q{s 2 ,a 1},…Q{s n ,a m}}
[0120] where, s n is the nth state in the state space set, and a m is the mth state in the action space set.
[0121] Step 5. Estimate the value function Q based on the DQN network model, and obtain a molten pool state probability π through the ε-greedy algorithm, and then generate the optimal laser adjustment action a * for the current molten pool state characteristics; The ε-greedy strategy defines an exploration rate ε, with a value between 0 and 1; When selecting an action each time, generate a random number between 0 and 1. If the random number is less than ε, select a random action; If it is greater than or equal to ε, select the optimal action according to the learned information.
[0122] Specifically, the ε-greedy algorithm calculates the probability π of the agent selecting each action in a given state through the following formula:
[0123]
[0124] π(a i |s) represents the probability of selecting action a in state s i , where ε is the exploration parameter, indicating that a non-optimal action is selected with a certain probability, N represents the total number of executable actions, represents the probability of selecting action a from N actions with probability ε, and Q(s i ,a i ) represents the expected reward of selecting action a i in state s i , and a * is the optimal action with the maximum Q value in the current state s:
[0125] Step 6: Provide feedback on the executed actions, collect s i , a i , r i , s i+1 during the agent's interaction process, integrate them into a data set, store them in the experience pool as experience samples, and use them to train the DQN network model to optimize the network model parameters. The specific parameter settings are listed in Table 2:
[0126] Table 2 Reinforcement Learning Training Parameter Table
[0127]
[0128] For each training episode (episode from 1 to 200), execute the following training process:
[0129] 1. Obtain the initial state, set the current step number = 1, and initialize the total reward for the current episode to 0;
[0130] 2. Within each episode, when the time step is less than the number of single-layer nodes N and the current state is not a terminal state, execute the following steps:
[0131] 1) Obtain the initial state s 0 or the current state s i , generate a random number, select and execute action a through the ε-greedy policy, and integrate the current molten pool state s i , action a i , the reward function value r i , and the next molten pool state s i+1 into an experience sample (s i , a i , r i,s i+1 ), store it in the experience pool Ω.
[0132] 2) When the number of samples in the experience pool D is greater than 128, randomly select 128 samples from the experience pool for learning. For each sample (S i ,A i ,R i ,S i+1 ), calculate the target Q value:
[0133] y i = r i
[0134] When the number of samples is less than 128:
[0135] y i = r i + γ maxQ(S i+1 , A'; μ - )
[0136] y t is the target Q value, γ is the discount factor, A' is the possible action at the next moment of S i+1 , maxQ(S i+1 , A'; μ - ) is the maximum target value generated by executing action A' at the next moment state S i+1 ;
[0137] 3) Based on the loss function, perform gradient descent to train the current Q network to minimize the mean square error between the predicted Q value of the network and the target pool, and update the current network model parameters μ and exploration rate ε. The expression of the loss function in the iteration process:
[0138]
[0139] μ i represents the current network parameters in the i-th iteration process, represents the target network parameters in the i-th iteration time step. After each iteration, it is μ i that is updated and not the target network parameters which is set to the current network parameters μ of the (i - 1)-th iteration i-1 ; Q(s i , a i , μ i ) represents the neural network approximate value function, (s, a, r, s i+1 ) ~ U(Ω) represents the agent's experience data, s represents the current pool state in the i-th iteration process, a represents the action selected in the current i-th step, s i+1 represents the next pool state, and a represents the action to be executed in the next step.
[0140] 4) Regularly update the target Q-value network weight parameters Update the exploration rate ε = max(ε - 2×10 -4 , 0.01), and update the state and time step: s i = s i+1 , i = i + 1.
[0141] 5) Repeat the above training processes 1) - 4) until the DQN network model converges or reaches the preset number of training rounds.
[0142] Step 7: The agent learns and executes the strategy of adjusting the laser power action, so that the molten pool state features are close to the target state features. The learning process of the agent is divided into four parts in each iteration: input feature generation, action decision-making, reliability evaluation and reward calculation, and training of the deep Q network.
[0143] First, extract the molten pool state feature s i and the state feature Ω i , and initialize the state space set S = {si|i = 0, 1, 2, …, ξ}; then generate a strategy through the action prediction module and select an appropriate action; subsequently, evaluate the generated strategy and optimize the network parameter μ i through batch training; when entering the next iteration, update the initial state s i+1 , and regenerate the strategy based on the current network with the latest parameter μ i ; in the first iteration, the decision-making process uses the current network structure initialized randomly with the parameter μ
[0144] ; after the training is completed, input the molten pool state set S 0 of other printing layers and the state feature Ω new for model testing. The agent generates and executes the laser power adjustment action according to the input; for each time step, the agent optimized by DQN will use the input feature Ω new to predict the prior information and select an action with a higher return based on the ε-greedy strategy to dynamically adjust the laser power; finally, by executing a series of actions with high returns, the agent generates a complete laser power control strategy, and the laser power optimization result is as i shown. Figure 5 shown.
Claims
1. A laser power parameter optimization method based on deep reinforcement learning, characterized in that: The deep reinforcement learning algorithm is used to adjust the laser power in real time during the path scanning process to reduce the fluctuation of the melt pool depth and heat accumulation, including the following steps: S1. Construct an LPBF simulation model, use a mobile heat source to establish a distribution model for the three-dimensional temperature field of the molten pool, simulate the temperature distribution on the laser scanning path, and realize simulation calculation of the molten pool temperature field; S2. Import the part STL file, perform layer-slicing on the STL file, and obtain the laser scanning path of each layer after setting the scanning parameters; S3. In the LPBF simulation model, the part forming process is simulated and calculated according to the laser scanning path, the molten pool state and state characteristics are extracted in real time, the state space and action space are constructed, and a reward function composed of two types of parameters is constructed. The reward value is obtained through environmental interaction, and the action strategy is guided to adjust in the direction of maximizing the reward value; S4. Build a reinforcement learning DQN network model, including the current Q network and the target Q network; S5. Estimate the value function Q based on the DQN network model, obtain a molten pool state probability π through the ε-greedy algorithm, and then generate the optimal laser adjustment action of the current molten pool state characteristics; S6. Provide feedback on the optimal laser adjustment action performed, integrate the data generated during the agent interaction process, and store them in the experience pool as experience samples for training the DQN network model and optimizing the network model parameters; S7. The intelligent agent learns and executes the strategy of adjusting the laser power action so that the state characteristics of the molten pool are close to the target state characteristics.
2. The laser power parameter optimization method based on deep reinforcement learning according to claim 1, characterized in that: In S1, an LPBF simulation model is constructed, and a distribution model of the three-dimensional temperature field of the molten pool is established by using a mobile heat source to simulate the temperature distribution on the laser scanning path and realize the simulation calculation of the molten pool temperature field. The specific steps are as follows: S1.
1. In order to improve the efficiency of simulation calculation and provide support for reinforcement learning, the complex multi-scale effects are simplified to the continuous temperature distribution of the material: only the conduction mode of heat transfer is considered, the thermal properties are independent of temperature, and the powder bed is modeled as a solid continuum; the above process is modeled as a two-dimensional conduction associated with a moving heat source on the powder, and the temperature field equation is given by the following partial differential equation: Where T(x,t) is a function of position x and time t, representing the temperature at a certain point; is the rate at which temperature T(x,t changes with time t; φ is the thermal diffusion coefficient, which is given by the formula Definition: k is the thermal conductivity of powder, ρ is the mass density of powder, c p is the powder heat capacity coefficient; is the spatial second-order derivative of T(x,t), that is, the degree of temperature diffusion along the x direction; Θ((x),t) represents the normalized heat source, which is the parameter normalized by density and specific heat capacity. Where Q(x,t) represents the heat generation rate per unit volume; After solving the Green function, we get: T(x,t)=T(x)=T Q (x)+T D (x) The first coefficient T Q (x) represents the effect of heat source on temperature, and the second coefficient T Q (x) represents the thermal diffusion effect; Define simulation parameters according to the material properties of the workpiece, including absorptivity and thermal conductivity; S1.
2. The influence of heat source is solved by Eagar-Tsai model. When the Eagar-Tsai model solves the temperature field of the molten pool, the heat source is parameterized as a circular area moving on the surface of the workpiece, and the heat is Gaussian distributed. The specific calculation model is as follows: Among them, θ(x,y,z,t) represents the heat source intensity per unit volume of powder at spatial position (x,y,z) and time t, η is the absorptivity of the material, P is the laser power, σ is the diameter of the laser, v is the scanning speed of the laser, δ(z) represents the spatial coordinate for z=0, δ(z)≠0, and the heat source only acts on the surface of the workpiece; Temperature distribution calculation model T at coordinates (x, y, z) l as follows: Where T l represents the temperature at a certain moment, P represents the laser power, and τ is the normalized time, which is related to the time scale of the heat conduction process, that is, the time duration of the heat source; is the integral variable, which is a parameter related to the normalized time, Δt is the action time of the laser at the position (x, y, z), and exp is an exponential function that describes the distribution characteristics of the heat source in space; S1.3, the heat diffusion effect occurs near the domain boundary, and the distance from the heat source center is equal to 4σ G Defined as domain boundaries, the diffusion length scale For the area outside the boundary, a virtual heat source is simulated at the same distance on the other side of the boundary. The following equation is adjusted to realize the virtual heat source T l : represents the thermal diffusion effect in the x direction, It represents the heat diffusion effect in the y direction. It represents the heat diffusion effect in z direction.
3. The laser power parameter optimization method based on deep reinforcement learning according to claim 2, characterized in that: In S2, the part STL file is imported, the STL file is sliced, and the scanning path planning files of each layer are obtained after setting the scanning parameters, as follows: S2.
1. The STL model of the part is composed of triangular facets. Before executing the front slicing, first extract the Z-axis coordinate range (Zmin, Zmax) of the triangular facets, and sort all the triangular facets from Zmin to Zmax; set the single layer thickness, generate multiple slicing planes perpendicular to the Z axis, with the starting coordinate of Zmin and the spacing of the single layer thickness, that is, multiple slicing planes, respectively, find the intersection line of each slicing plane and the triangular facet, connect them in sequence, obtain the contour curve of the single-layer slicing figure, and finally complete the slicing of the model by multiple slicing planes; S2.
2. After slicing is performed, the scanning parameters are set according to the actual situation. The scanning parameters include island width, shadow line angle and filling spacing, and the scanning vector and side length are defined; the slice contour is filled based on the Z-scanning algorithm, and the laser scanning path planning is generated, and then the scanning path planning files of each layer are obtained.
4. The laser power parameter optimization method based on deep reinforcement learning according to claim 3, characterized in that: In S3, the melt pool state and state features are extracted in real time, the state space and action space are constructed, and a reward function composed of two types of parameters is constructed. The reward value is obtained through environmental interaction, and the action strategy is guided to adjust in the direction of maximizing the reward value, as follows: S3.
1. The state space is defined as the record of the temperature field state in a specific view and direction. A ω×ω XY central rectangular area is set at the current scanning position as the center. The XZ and YZ sections extend downward along the horizontal plane ω. The temperature field distribution on the XY, YZ and XZ cross sections is recorded to obtain the real-time molten pool temperature field. i , They represent the temperature field distribution of the molten pool in the xy, xz, and yz cross sections respectively; The element in the matrix row x and column y represents the temperature of the molten pool at the coordinate (x, y) on the xy section. The molten pool state of the xz and yz sections is represented in the same way. The molten pool state S i =[M i-2 ,M i-1 ,M i ],M i-2 、M i-1 represents the molten pool state in the first two time steps; S3.2, Molten pool depth D i As the key feature of the molten pool state, the specific extraction method is as follows: Interpolate the temperature field along the vertical axis of the scanning direction to find the position with the highest temperature on the horizontal plane, and interpolate downward along the Z axis of this point to find the critical point with the temperature closest to the melting point of the material. The depth of the molten pool is the vertical distance from this point to the horizontal plane. S3.
3. Constructing a continuous action space A = {a i |i=1,2,…n},a i represents the action of step i, the action parameter selection range is [λ1, λ2], where λ1 and λ2 are real numbers, and the step length is c. The action parameter adjustment formula for laser power is derived as follows: P i+1 =a i+1 ×(P max -P min )+P min Where P i+1 is the laser power at step i+1, a i+1 represents the action parameter of step i+1, P max represents the maximum laser power in the previous i steps, P min represents the minimum value of laser power in the previous i steps; S3.4, the reward function is composed of the absolute error between the target melt pool depth and the current melt pool depth, and is combined with a regularization term to avoid abnormal behavior; the regularization term prevents drastic fluctuations in the melt pool depth by punishing the difference between the maximum and minimum melt pool depths observed during the cycle. The reward function R is expressed as: Where D target is the target molten pool depth, D melt To calculate the molten pool depth for simulation, max i D melt is the maximum molten pool depth, min i D melt is the minimum molten pool depth, ∑ i It means to sum the values of the previous i steps.
5. The laser power parameter optimization method based on deep reinforcement learning according to claim 4, characterized in that: In S4, a reinforcement learning DQN network model is constructed, including the current Q network and the target Q network, as follows: The current Q network is used to select the agent's action a according to the current state s, and the target network is used to calculate the maximum value function Q of the next state s'; The current network and the target network have the same structure, including the following parts: an input layer for receiving the current state features of the system; multiple hidden layers, including at least one convolutional layer and a fully connected layer, for extracting state features and estimating the value function; and an output layer for providing the optimal value function Q of the action.
6. The laser power parameter optimization method based on deep reinforcement learning according to claim 5, characterized in that: The iterative formula of the value function Q is expressed as: Q(S i+1 ,A i+1 )←Q(S i ,A i )+α[R i +γmaxQ(S i+1 ,A i+1 )-Q(S i ,A i )] Where α is the learning rate, γ is the reward discount factor, and the value range is [0,1]. When it is 0, it means that only the impact of the current action on the current step is considered, and the impact on the subsequent steps is not considered. When it is 1, it means that the current action has an equal impact on each subsequent step. i As input, output different actions A i The Q value under this condition is to output a Q value set Q(S i ,A i ): Q(S i ,A i )={Q{s1,a1},Q{s1,a2},…,Q{s1,a m },Q{s2,a1},…Q{s n ,a m }} Among them, s n is the nth state in the state space, a m is the mth state in the action space.
7. The laser power parameter optimization method based on deep reinforcement learning according to claim 6, characterized in that: In S5, the value function Q is estimated based on the DQN network model, and a molten pool state probability π is obtained through the ε-greedy algorithm, and then the optimal laser adjustment action a of the current molten pool state characteristics is generated. * , as follows: The ε-greedy strategy defines an exploration rate ε, which is between 0 and 1. Each time an action is selected, a random number between 0 and 1 is generated. If the random number is less than ε, a random action is selected. If it is greater than or equal to ε, the optimal action is selected based on the learned information. Specifically, the ε-greedy algorithm calculates the probability π of the agent choosing each action in a given state by the following formula: π(a i |s) means selecting action a in state s i The probability of, ε is the exploration parameter, indicating that a non-optimal action is selected under a certain probability, N represents the total number of executable actions, represents the probability of selecting action a from N actions with probability ε, Q(s i ,a i ) indicates state s i Next select action a i The expected reward, a * is the optimal action with the maximum Q value in the current state s:
8. The laser power parameter optimization method based on deep reinforcement learning according to claim 7, characterized in that: In S6, feedback is provided on the executed actions, and s is collected during the agent interaction process. i 、a i 、r i 、s i+1 , integrated into a data set, stored in the experience pool as experience samples, used to train the DQN network model and optimize the network model parameters, as follows: S6.1, select and execute action a through the ε-greedy strategy, and change the current molten pool state s i 、Action i , reward function value r i , the next moment molten pool state s i+1 Integrated into experience samples (s i , a i , r i ,s i+1 ), stored in the experience pool Ω = {(S i ,A i ,R i ,S i+1 )|i=1,2,…,ξ}; S6.2, randomly select a batch of samples from the experience pool for learning, and for each sample (S i ,A i ,R i ,S i+1 ) Calculate the target Q value: and i =R i +γmaxQ(S i+1 ,A;μ - ) y t is the target Q value, γ is the discount factor, and A' is the next moment S i+1 Possible actions, maxQ(S i+1 ,A';μ - ) is the next moment state S i+1 The maximum target value generated by executing action A' under - is the target Q network parameter; S6.
3. Perform gradient descent training on the current Q network based on the loss function to minimize the mean square error between the Q value predicted by the network and the target melting pool, update the current network model parameters μ and the exploration rate ε, and the loss function expression of the iterative process is: μ i represents the current network parameters in the i-th iteration, Represents the target network parameter of the iteration step of step i. After each iteration, μ is updated i Without updating Target network parameter μ i - Set to the current network parameter μ for the i-1th step iteration i-1 ; Q(s i ,a i ,μ i ) represents the neural network approximate value function, )s,a,r,s i+1 )~U(Ω) represents the agent’s experience data, s represents the melting pool state of the current i-step iteration process, a represents the action selected in the current i-step, and s i+1 represents the next step of the melt pool state, and a represents the action to be performed in the next step; S6.
4. Regularly update the target Q value network weight parameter μ i - =μ i , update the exploration rate ε; S6.
5. Repeat the above S6.1 to S6.4 until the DQN network model converges or reaches the preset training rounds.
9. The laser power parameter optimization method based on deep reinforcement learning according to claim 8, characterized in that: In S7, the agent learns and executes the strategy of adjusting the laser power action so that the state characteristics of the molten pool are close to the target state characteristics, as follows: S7.
1. The learning process of the power control agent is divided into four parts in each iteration: input feature generation, action decision, reliability evaluation and reward calculation, and deep Q network training: First, extract the melt pool state feature s i and state characteristics Ω i , and initialize the state space set S = {s i |i=0,1,2,…,ξ}; then generate strategies through the action prediction module and select appropriate actions; then evaluate the generated strategies and optimize the network parameters μ through batch training i ; When entering the next iteration, update the initial state s i+1 , and based on the latest parameter μ i The current network regeneration strategy of In the first iteration, the decision process uses the randomly initialized current network structure with parameters μ0; S7.
2. After training is completed, input the melt pool state set S of other printing layers new and state characteristics Ω new To test the model, the agent generates and executes the laser power adjustment action based on the input; for each time step, the DQN-optimized agent uses the input feature Ω i Predict prior information and select actions with higher rewards based on the ε-greedy strategy to dynamically adjust the laser power; ultimately, by executing a series of high-reward actions, the agent generates a complete laser power control strategy.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer executable instructions are executed, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Method and device for optimizing preparation process of ceramic circuit board
CN120471242A
Power module control method and device based on artificial intelligence
CN121763781A
Artificial Intelligence-Based Power Module Control Method and Device
CN121763781B
Directional energy deposition laser power off-line planning method based on depth Q network
CN122113539A
Deep q-network based directed energy deposition laser power off-line planning method
CN122113539B