Deep reinforcement learning-based directed energy deposition process parameter control method and device, medium and terminal
By constructing a simulated energy deposition system and optimizing the printing strategy using a deep reinforcement learning model, the control challenges caused by the complexity of the environment in the laser directional energy deposition process were solved, achieving intelligent and efficient control of process parameters and improving printing quality and material utilization.
Patent Information
- Application Number
- CN202311561561.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-11-21
AI Technical Summary
In existing technologies, the environmental complexity of laser directional energy deposition processes makes it difficult to determine control strategies, results in high labor costs, and lacks matching simulators, leading to high equipment and material losses.
By constructing a simulated energy deposition system, a deep reinforcement learning model is used to optimize the printing strategy, generate the optimal printing strategy and control parameters, and combine the Rosenthal equation and the three-dimensional heat transfer equation for simulation to optimize the laser directional energy deposition process parameters.
It enables intelligent and efficient control of laser directional energy deposition process parameters, reduces manual trial and error time and sample preparation costs, improves printing uniformity and component strength, and reduces material waste and equipment wear.
Smart Images

Figure CN117324637B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of deep reinforcement learning, and in particular to a method and device for controlling process parameters of a directed energy deposition process based on deep reinforcement learning, a medium and a terminal. BACKGROUND
[0002] Additive Manufacturing (AM) plays a crucial role in the current industrial revolution, providing a unique ability to produce high-quality parts of complex shapes. Among the many metal additive manufacturing technologies, Laser Directed Energy Deposition (LDED) stands out due to its ability to deposit at specific locations and its unlimited three-dimensional printing capability. Laser Directed Energy Deposition process is an important technology in the field of high-end manufacturing, driving progress in the fourth industrial revolution. Its unique ability to manufacture complex metal parts has led to its wider application in various industries with strict performance requirements. It is particularly noteworthy that Laser Directed Energy Deposition process is widely used in industrial repair and prototyping in the fields of medical, automotive and aerospace.
[0003] The latest developments in artificial intelligence technology are triggering a revolution in the field of additive manufacturing. Laser Directed Energy Deposition is gaining market share in the field of metal additive manufacturing, highlighting its great potential. Therefore, research has shown that optimizing process parameters is crucial to determining part quality. Moreover, systematically optimizing process parameters is not only necessary but also has great advantages for industrial production. However, due to the complexity of the real production environment, it is difficult for humans to control, resulting in suboptimal process parameters.
[0004] In the prior art, artificial intelligence technology has been used to optimize process parameters in additive manufacturing, for example, using a Long Short-Term Memory (LSTM) model to enhance the tensile strength performance of printed parts in the Material Extrusion (MEX) process. Or use a combination of recurrent neural network and deep neural network model to establish the internal relationship between laser scanning strategy and heat history distribution in the Laser Directed Energy Deposition process, and optimize the process parameters in additive manufacturing by relying on data-driven methods. However, there are certain limitations, first of all, the amount of data required is large, secondly, the existing technology has not been able to build an efficient simulator for deep reinforcement learning and training, and there is no matching simulation environment to control the loss of equipment and materials, so that the process parameters cannot be dynamically adjusted based on the simulator, and the best process parameters and control strategy cannot be obtained. SUMMARY
[0005] In view of the above-mentioned disadvantages of the prior art, the purpose of the present application is to provide a deep reinforcement learning-based directed energy deposition process parameter control method, device, medium and terminal, which is used to solve the problems of difficult determination of control strategy due to complex environment in directed energy deposition process, high labor cost, and high equipment and material loss caused by the lack of a matching simulator in the existing reinforcement learning-based control strategy generation scheme, and finally generate a process control strategy suitable for process parameters in directed energy deposition technology.
[0006] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides a deep reinforcement learning-based directed energy deposition process parameter control method, comprising: obtaining specification parameters of a to-be-printed model and environmental variable parameters of a real energy deposition system, and constructing a simulated energy deposition system based on the specification parameters; obtaining state information of the simulated energy deposition system and inputting it into an initialized deep reinforcement learning model; based on the specification parameters of the to-be-printed model, performing optimization operation on the printing strategy in the deep reinforcement learning model through an optimization algorithm to generate an optimal printing strategy; and based on the optimal printing strategy, generating control parameters of the to-be-printed model and deploying the real energy deposition system according to the control parameters.
[0007] In some embodiments of the first aspect of the present application, the process of generating an optimal printing strategy based on the specification parameters of the to-be-printed model by performing optimization operation on the printing strategy in the deep reinforcement learning model through an optimization algorithm includes: obtaining a reward value and the state information set from the simulated energy deposition system, and performing feature extraction on the state information set of the simulated energy deposition system to generate a state feature set; inputting the reward value and the state feature set into an experience replay pool to store the current reward value and the state feature set; performing value evaluation operation on the state feature set in the experience replay pool, and extracting action parameters containing a value higher than a threshold from the state feature set; updating an action network based on the action parameters, and performing corresponding operations in the simulated energy deposition system according to the action network to generate an updated reward value and state feature set.
[0008] In some embodiments of the first aspect of the present application, the process of generating an optimal printing strategy based on the specification parameters of the to-be-printed model by performing optimization operation on the printing strategy in the deep reinforcement learning model through an optimization algorithm also includes: calculating a loss function through the control performance target of the to-be-printed model and the updated state feature set; based on the loss function, adjusting the hyperparameters in the deep reinforcement learning model to update the deep reinforcement learning model.
[0009] In some embodiments of the first aspect of the present application, the process of initializing the deep reinforcement learning model comprises initializing an action network, a critic network, and a reward function in the deep reinforcement learning model.
[0010] In some embodiments of the first aspect of the present application, the process of obtaining the specification parameters of the to-be-printed model and the environmental variable parameter set of the real energy deposition system comprises: selecting a partial differential equation set for optimizing the deep reinforcement learning model and the environmental variable parameter set according to the physical characteristics of the real energy deposition system; and setting a third type of boundary condition for solving the partial differential equation set based on the printing model parameters and the real energy deposition system.
[0011] In some embodiments of the first aspect of the present application, the partial differential equation set for optimizing the deep reinforcement learning model is selected according to the physical characteristics of the real energy deposition system, wherein the partial differential equation set comprises a Rosenzweig equation and a three-dimensional heat transfer equation.
[0012] In some embodiments of the first aspect of the present application, in the process of optimizing the printing strategy in the deep reinforcement learning model based on the specification parameters of the to-be-printed model by using an optimization algorithm, the optimization algorithm comprises a proximal optimization algorithm.
[0013] To achieve the above object and other related objects, the second aspect of the present application provides a deep reinforcement learning-based directional energy deposition process parameter control device, comprising: a simulation system construction module for obtaining the specification parameters of a to-be-printed model and the environmental variable parameter set of a real energy deposition system, and constructing a simulation energy deposition system therefrom; an initialization module for initializing a deep reinforcement learning model, obtaining the state information set of the simulation energy deposition system and inputting the same into the deep reinforcement learning model; obtaining the state information of the simulation energy deposition system and inputting the same into the initialized deep reinforcement learning model; a strategy optimization module for optimizing the printing strategy in the deep reinforcement learning model based on the specification parameters of the to-be-printed model by using an optimization algorithm, to generate an optimal printing strategy; and a strategy deployment module for generating the control parameters of the to-be-printed model based on the optimal printing strategy and deploying the real energy deposition system according to the control parameters.
[0014] To achieve the above object and other related objects, the third aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the deep reinforcement learning-based directional energy deposition process parameter control method.
[0015] To achieve the above object and other related objects, the fourth aspect of the present application provides an electronic terminal, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes the deep reinforcement learning-based directed energy deposition process parameter control method.
[0016] As described above, the deep reinforcement learning-based directed energy deposition process parameter control method, device, medium and terminal of the present application have the following beneficial effects: the present application constructs a simulation system based on laser directed energy deposition, and uses a deep reinforcement learning proximal optimization algorithm to explore an optimal strategy, thereby generating a control strategy containing laser power, speed and motion trajectory parameters, so that the intelligentization and efficiency of the optimal process parameter strategy selection in the energy deposition process are improved. Compared with the printing strategy obtained by traditional artificial experience control, the present application can shorten the artificial trial and error time and sample manufacturing cost, and the deep reinforcement learning strategy reduces the variability of the hardness of the obtained sample. Moreover, the present application can be applied to various real production environments, reduces the difficulty of artificial control, and thus efficiently obtains optimized process parameters and strategies. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 A flowchart of an embodiment of the deep reinforcement learning-based directed energy deposition process parameter control method of the present application is shown.
[0018] Figure 2 A schematic diagram of a laser directed energy deposition simulator in an embodiment of the deep reinforcement learning-based directed energy deposition process parameter control method of the present application is shown.
[0019] Figure 3 A schematic diagram of a single-layer heat distribution state in an embodiment of the deep reinforcement learning-based directed energy deposition process parameter control method of the present application is shown.
[0020] Figure 4 A flowchart of a deep reinforcement learning algorithm in an embodiment of the deep reinforcement learning-based directed energy deposition process parameter control method of the present application is shown.
[0021] Figure 5 A comparison diagram of the hardness distribution of a printed model sample in an embodiment of the deep reinforcement learning-based directed energy deposition process parameter control method of the present application is shown.
[0022] Figure 6 A structural schematic diagram of an embodiment of the deep reinforcement learning-based directed energy deposition process parameter control device of the present application is shown.
[0023] Figure 7A structure diagram of a directional energy deposition process parameter control electronic terminal based on deep reinforcement learning of the present application is shown. DETAILED DESCRIPTION
[0024] The advantages and effects of the present application can be easily understood by those skilled in the art from the description of the embodiments of the present application. The present application can also be implemented or applied by other different embodiments, and the details in the description can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0025] It should be noted that in the following description, reference is made to the accompanying drawings, which illustrate several embodiments of the present application. It is understood that other embodiments can be used and that mechanical, structural, electrical, and operational changes can be made without departing from the spirit and scope of the present application. The following detailed description is not to be interpreted as limiting, and the scope of the embodiments of the present application is defined solely by the appended claims. The terminology used here is only for the purpose of describing specific embodiments and is not intended to limit the present application. Spatially relative terms such as "upper", "lower", "left", "right", "below", "beneath", "bottom", "top", and the like can be used herein for ease of description to describe one element or feature's relationship to another element(s) or feature(s) as illustrated in the figures.
[0026] In the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting", "fixing", "holding" and the like should be interpreted broadly, for example, it can be fixedly connected, or detachably connected, or integrally connected; it can be mechanically connected, or electrically connected; it can be directly connected, or indirectly connected through an intermediate medium; it can be the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0027] Moreover, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including" when used herein, specify the presence of stated features, operations, elements, components, items, and / or groups but do not preclude the presence or addition of one or more other features, operations, elements, components, items, and / or groups thereof. As used herein, the terms "or" and "and / or" are to be interpreted as inclusive, i.e., as meaning one or any combination of the items listed. Thus, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C". Only when the combination of elements, functions, or operations are inherently mutually exclusive is an exception to this definition presented.
[0028] To solve the problems in the background art, the present application provides a deep reinforcement learning-based directed energy deposition process parameter control method, device, medium and terminal, aiming to solve the problems of difficult determination of control strategy, high labor cost due to complex environment in directed energy deposition process, and high equipment and material loss due to lack of matching simulator in the existing reinforcement learning-based control strategy generation scheme. At the same time, in order to make the purpose, technical scheme and advantages of the present application more clear and explicit, the technical scheme in the embodiments of the present application is further described in detail in the following embodiments and in conjunction with the drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0029] Before the present application is further described, the terms and phrases involved in the embodiments of the present application are explained, which are applicable to the following explanations:
[0030] <1> Additive manufacturing: Additive manufacturing is a manufacturing process, also known as 3D printing. It builds objects by layering materials, unlike traditional subtractive manufacturing, which adds materials layer by layer according to the design requirements, so it can manufacture more complex structures.
[0031] <2> Directed energy deposition process: Directed energy deposition (DED) is an additive manufacturing (3D printing) process that uses high-energy beams or electron beams to layer materials, thus manufacturing complex metal parts or components.
[0032] <3> Simulator: A simulator is a software or hardware system that simulates the behavior of a particular system, environment or process. In the field of computer science and engineering, simulators are often used to simulate the behavior of various systems in order to conduct experiments, tests and predictions.
[0033] <4> Additive manufacturing print parameters: Additive manufacturing print parameters are the parameters that control the quality and speed of printing during the 3D printing process, including layer height, printing speed, temperature, etc. Choosing the right printing parameters is very important to obtain high-quality 3D printed products.
[0034] <5> Rosenthal equation: The Rosenthal equation is one of the equations describing incompressible fluid dynamics in fluid mechanics. It is used to describe the motion of fluid under incompressible conditions and is one of the basic equations in fluid mechanics.
[0035] <6> Three-dimensional heat transfer equation: The three-dimensional heat transfer equation describes the conduction and convection processes of heat in three-dimensional space. It is a mathematical description of heat conduction and convection phenomena, used to solve engineering problems related to heat conduction and convection.
[0036] <7> Third type of boundary condition: In mathematics and physics, the third type of boundary condition specifies specific constraints on the boundaries of a system. These conditions are often used to describe the interaction of the system with the external environment, such as boundary temperature in heat conduction, boundary velocity in fluid dynamics, etc.
[0037] <8> Proximal optimization algorithm: Proximal optimization algorithm is a class of algorithms used for optimization problems, commonly applied in deep learning and reinforcement learning. These algorithms update model parameters through continuous iteration to minimize the loss function or maximize the reward function, so as to achieve the purpose of optimizing the model.
[0038] <9> Discount factor: In reinforcement learning, the discount factor is used to measure the value of future rewards. It can affect the agent's emphasis on long-term rewards, helping the agent make long-term planning.
[0039] <10> Shear function factor: The shear function factor describes the shear behavior in fluid dynamics. In fluid mechanics, the shear function factor is used to describe the behavior of fluid in the shear process, which is very important for understanding the motion and deformation of fluid.
[0040] <11> Learning rate: In machine learning, the learning rate is used to control the speed of updating model parameters during training. A suitable learning rate can affect the convergence speed and stability of the model, and is one of the important hyperparameters in the optimization process of model training.
[0041] The embodiment of the present application provides a deep reinforcement learning-based directed energy deposition process parameter control method, a deep reinforcement learning-based directed energy deposition process parameter control method system, and a storage medium storing an executable program for implementing the deep reinforcement learning-based directed energy deposition process parameter control method. In terms of implementation of the deep reinforcement learning-based directed energy deposition process parameter control method, the embodiment of the present application will illustrate an exemplary implementation scenario of the deep reinforcement learning-based directed energy deposition process parameter control.
[0042] As shown in Figure 1 A flowchart of a deep reinforcement learning-based directed energy deposition process parameter control method in the embodiment of the present application is shown. The deep reinforcement learning-based directed energy deposition process parameter control method in the embodiment mainly includes the following steps:
[0043] Step S11: Obtain the specification parameters of a to-be-printed model and the environmental variable parameter group of a real energy deposition system, so as to construct a simulation energy deposition system.
[0044] In an embodiment of the present application, the process of obtaining the specification parameters of a to-be-printed model and the environmental variable parameter group of a real energy deposition system includes: selecting a partial differential equation group and the environmental variable parameter group for optimizing the deep reinforcement learning model according to the physical characteristics of the real energy deposition system; and setting a third type of boundary condition for solving the partial differential equation group based on the printing model parameters and the real energy deposition system.
[0045] Further, the partial differential equation group for optimizing the deep reinforcement learning model is selected according to the physical characteristics of the real energy deposition system, wherein the partial differential equation group includes a Rosenthal equation and a three-dimensional heat transfer equation.
[0046] In an embodiment of the present application, the process of selecting a partial differential equation group and the environmental variable parameter group for optimizing the deep reinforcement learning model according to the physical characteristics of the real energy deposition system includes: in the case of ignoring the secondary effect of powder on the size of a molten pool in a laser-assisted direct energy deposition (LDED) process, an energy input process is described as three-dimensional deposition using a moving point heat source in a cubic domain. The Rosenthal equation is used in the present application to optimize the deep learning model.
[0047]
[0048] Formula 1 shows the Rosenthal equation, wherein T0 is the initial temperature, P is the laser power, λ is the absorption rate of the laser, k is the thermal conductivity, α is the thermal diffusivity, r is the distance from the point heat source, and is defined as V is the scanning speed.
[0049] Meanwhile, considering the intermittent nature of the deposition path, i.e., the change of the deposition direction will cause the laser beam to temporarily stop moving. Assuming heat transfer in an isotropic and homogeneous medium in the present application, the three-dimensional heat transfer equation is as follows:
[0050]
[0051] where, and are the second-order spatial derivatives of temperature in the x, y, and z directions, respectively, p is the mass density, c p is the specific heat capacity, is the rate of change of temperature at a point over time.
[0052] Further, according to the actual environment and the printed part, parameters such as the third boundary condition for solving the partial differential equation system, the convective heat transfer coefficient, and the absorption rate are set. The third boundary condition refers to the boundary condition given at the boundary point, which includes not only the numerical value of the temperature, velocity, or displacement of the boundary point, but also the derivative of these physical quantities, in order to describe heat conduction, fluid flow, or solid mechanics problems.
[0053] Exemplarily, since the convective heat transfer between the surrounding air and the sample in the energy deposition process in the present application mainly affects the thermal state of the model, the third boundary condition shown below is adopted.
[0054]
[0055] where, is the first-order spatial derivative of temperature in the n direction, h represents the convective heat transfer coefficient, T w is the boundary temperature, and T f is the fluid temperature.
[0056] It is worth noting that the advantage of the simulation environment constructed by the above partial differential equation and the third boundary condition is that the technical problem to be solved by the present application is the problem of large amount of data required in existing additive manufacturing. Through the simulation environment provided by the present application, the behavior of the actual system can be simulated using computer simulation, thereby reducing the need for a large amount of actual data. Through the simulation environment, the behavior of the system can be simulated through mathematical models and physical laws without the need for actual collection of a large amount of data. By adjusting the model parameters and initial conditions, the behavior of the system under various conditions can be simulated, thereby allowing extensive testing and analysis on a small-scale data set. Therefore, by constructing the simulation environment, the need for a large amount of actual data can be reduced, and research and analysis of the directed energy deposition process based on deep reinforcement learning can be performed more quickly.
[0057] Further, the present application simplifies the complex multi-factor effects into continuous temperature distribution of the material by constructing a simulation energy deposition system. Wherein the multi-factor effects affecting the printing process of the simulation energy deposition system are simplified into continuous temperature distribution of the material. Wherein, the parameters input into the simulation energy deposition system include material properties, number of paths and number of layers in the part, and initial temperature of the environment and air convection heat transfer coefficient. Wherein, the above parameters can be independently adjusted as environmental variables. At the same time, the present application also proposes a deep learning framework for generating a complex strategy for simultaneously controlling laser power, scanning speed and deposition path under the condition of maximizing the use of state information in the simulation environment.
[0058] The construction process of the simulation energy deposition system is described in detail above, and the following will be combined Figures 2 to 4 The deep learning framework is described above.
[0059] Step S12: Obtain the state information of the simulation energy deposition system and input it into the initialized deep reinforcement learning model.
[0060] In an embodiment of the present application, the process of initializing the deep reinforcement learning model includes initializing the action network, evaluation network and reward function in the deep reinforcement learning model.
[0061] Further, the process of initializing the reward function includes designing the reward function according to the target requirements, and adjusting the hyperparameters to control the controllability and shape of the simulation energy deposition system. Since the temperature directly and significantly affects the performance of the sample in the actual energy deposition system, the average temperature of the current layer is set as follows:
[0062]
[0063] Wherein, T m is the average temperature of the current layer, T target represents the target temperature, and δ∈[0,1] is the running tolerance coefficient.
[0064] During the deposition process, the thermal distribution is related to the inherent strain sample, thereby indirectly affecting the shape of the sample. Therefore, in order to reduce the influence of strain, the temperature centroid needs to be approximately equal to the center of the sample shape. Therefore, the reward function is defined as follows:
[0065] R s =2δ(1-d)T target (Formula 5)
[0066] Wherein, d∈[0,1] is the distance between the temperature centroid and the physical center.
[0067] Further, the balance Rp and R s The total reward function is defined as follows:
[0068] R(s t ,a t ,s t+1 )=βR p +(1-β)R s (Formula 6)
[0069] Where β is a hyperparameter of the fine-tuning strategy.
[0070] Step S13: Based on the specifications of the model to be printed, the printing strategy is optimized in the deep reinforcement learning model using an optimization algorithm to generate the optimal printing strategy.
[0071] In one embodiment of the present invention, the process of optimizing the printing strategy in a deep reinforcement learning model based on the specifications of the model to be printed, to generate the optimal printing strategy, includes: obtaining a reward value and the state information set from the simulated energy deposition system, and extracting features from the state information set of the simulated energy deposition system to generate a state feature set; inputting the reward value and the state feature set into an experience replay pool to store the current reward value and the state feature set; performing a value evaluation operation on the state feature set in the experience replay pool, and extracting action parameters containing values higher than a threshold from the state feature set; updating the action network based on the action parameters, and performing corresponding operations in the simulated energy deposition system according to the action network to generate the updated reward value and the state feature set.
[0072] Furthermore, the process of optimizing the printing strategy in the deep reinforcement learning model based on the specifications of the model to be printed, in order to generate the optimal printing strategy, also includes: calculating a loss function using the performance target of the model to be printed and the updated state feature set; and adjusting the hyperparameters in the deep reinforcement learning model based on the loss function to update the deep reinforcement learning model.
[0073] like Figure 2 The diagram shows a schematic of a laser-directed energy deposition simulator in one embodiment of the present invention. In the laser-directed energy deposition simulator, by controlling the movement of the laser beam and the input of energy, energy is focused on a specific area on the surface of the workpiece in the sample on the substrate to melt the material. Then, additional material is deposited layer by layer on the surface of the molten material through energy output. A suitable printing strategy is used to precisely control the heating source and the deposition process to ensure that the deposited material matches the shape and size required by the design.
[0074] Furthermore, such asFigure 3 The single-layer heat distribution state diagram shows that the printing of the to-be-printed model is realized by material surface layer-by-layer deposition on each single layer, wherein the calculation point x refers to a discrete point on the deposition layer, that is, the construction surface is discretized into a plurality of points by a mathematical method during the construction of the workpiece, and each point corresponds to a calculation point x. The deposition direction y refers to the direction of energy input in the energy deposition process. In the additive manufacturing process such as laser melting and electron beam melting, energy is usually irradiated from an energy source (such as a laser or an electron beam) to the calculation point x along a specific direction y, so that the material is melted or solidified. The selection of the deposition direction will affect the formation of the melting pool, the thermal history of the material and the performance of the final construction piece.
[0075] As shown in Figure 4 , a flowchart of a deep reinforcement learning algorithm in an embodiment of the present application is shown. Feature extraction is performed on the state information of the simulated energy deposition system, the feature information of the current state is combined with the corresponding reward value to form a data tuple, and the data tuple is stored in the experience replay pool. The experience replay pool is used to store the experience before the deep reinforcement learning model, and then these experiences are randomly sampled for learning during the training process. Subsequently, a data is randomly extracted from the experience replay pool and input into the network, and the corresponding action and evaluation are output respectively, so as to the deep reinforcement learning network and the simulated energy deposition system.
[0076] In an embodiment of the present application, based on the specification parameters of the to-be-printed model, the printing strategy is optimized in the deep reinforcement learning model by an optimization algorithm, wherein the optimization algorithm includes a proximal optimization algorithm.
[0077] Specifically, the process of using the proximal optimization algorithm to optimize the printing strategy in the deep reinforcement learning model includes: the proximal optimization algorithm improves the sampling efficiency by introducing importance sampling, thereby realizing the reuse of data. It realizes the best balance between convergence speed and practicability by using a clipping function. The loss function is as follows:
[0078]
[0079] Wherein, π θ represents a policy with parameters θ, represents the old policy k steps ago, t represents the current time step, A t corresponds to the advantage function under the current policy, and ∈ represents the hyperparameter of the clipping function.
[0080] In an embodiment of the present application, the hyperparameters in the deep reinforcement learning model are adjusted to update the deep reinforcement learning model. Exemplarily, the hyperparameters include but are not limited to: discount factor, clipping function ∈ factor, network layer number of the model, width, learning rate, etc. According to the simulation environment constructed by the partial differential equation set and the third boundary condition, and by the model training through the proximal optimization algorithm, the part size and the metal material attribute are also synchronized to the simulator.
[0081] Step S14: Based on the optimal printing strategy, the control parameters of the to-be-printed model are generated, and the real energy deposition system is deployed according to the control parameters.
[0082] In an embodiment of the present application, by adjusting the environmental parameters to be consistent with the industrial environment, the best training model is obtained by tuning, and the optimal printing strategy is generated. The generation of the optimal printing strategy includes but is not limited to one or more of the corresponding control of laser power, scanning speed and deposition path.
[0083] Further, the strategy output by the corresponding best deep reinforcement learning model is recorded, and the laser power, scanning speed and printing trajectory parameters in the strategy are migrated to the real device for production work.
[0084] As shown in Figure 5 , a sample hardness distribution comparison chart before and after using the deep reinforcement learning-based directed energy deposition process parameter control method provided by the present application is shown. The method provided by the present application can effectively improve the accuracy and uniformity of the model. It is worth noting that the deep reinforcement learning-based directed energy deposition process parameter control method provided by the present application can effectively improve the overall strength and durability of the component while improving the printing uniformity; it helps to reduce the risk of cracks and deformation, improve the stability of the component; it can also reduce the flaws and traces on the surface of the component, improve the quality and precision of the surface; finally, uniform printing can maximize the use of materials, thereby reducing waste and reducing costs.
[0085] As shown in Figure 6 , a structure schematic diagram of a deep reinforcement learning-based directed energy deposition process parameter control device in an embodiment of the present application is shown. In this embodiment, the deep reinforcement learning-based directed energy deposition process parameter control device 600 includes:
[0086] The simulation system construction module 601 is used to obtain the specification parameters of the to-be-printed model and the environmental variable parameter set of the real energy deposition system, and to construct a simulation energy deposition system.
[0087] The initialization module 602 is configured to initialize the deep reinforcement learning model, and obtain state information of the simulation energy deposition system and input the state information into the deep reinforcement learning model. The state information of the simulation energy deposition system is obtained and input into the initialized deep reinforcement learning model.
[0088] The strategy optimization module 603 is configured to optimize the printing strategy in the deep reinforcement learning model based on the specification parameters of the model to be printed, to generate an optimal printing strategy.
[0089] The strategy deployment module 604 is configured to generate control parameters of the model to be printed based on the optimal printing strategy, and deploy the real energy deposition system according to the control parameters.
[0090] It should be noted that the deep reinforcement learning-based directed energy deposition process parameter control device provided in the above embodiments is only used as an example for the division of the above program modules when performing deep reinforcement learning-based directed energy deposition process parameter control. In actual applications, the above processing can be completed by different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the deep reinforcement learning-based directed energy deposition process parameter control device and the deep reinforcement learning-based directed energy deposition process parameter control method provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.
[0091] The deep reinforcement learning-based directed energy deposition process parameter method provided in the embodiments of the present application can be implemented on the terminal side or the server side. As for the hardware structure of the deep reinforcement learning-based directed energy deposition process parameter terminal, please refer to Figure 7 , which is an optional hardware structure diagram of the deep reinforcement learning-based directed energy deposition process parameter terminal 700 provided in the embodiments of the present application. The terminal 700 can be a mobile phone, a computer device, a tablet device, a personal digital processing device, a factory background processing device, etc. The deep reinforcement learning-based directed energy deposition process parameter terminal 700 includes at least one processor 701, a memory 702, at least one network interface 704, and a user interface 706. Each component in the device is coupled together through a bus system 705. It can be understood that the bus system 705 is used to realize the connection and communication between the components. The bus system 705 includes a data bus, a power bus, a control bus, and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system in Figure 7 .
[0092] The user interface 706 can include a display, a keyboard, a mouse, a trackball, a pointing gun, a key, a button, a touchpad, a touch screen, or the like.
[0093] It can be understood that the memory 702 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present application is intended to include but not limited to these and any other suitable categories of memory.
[0094] The memory 702 in the embodiments of the present application is used to store various categories of data to support the operation of the deep reinforcement learning-based directional energy deposition process parameter terminal 700. Examples of these data include: any executable programs for operating on the deep reinforcement learning-based directional energy deposition process parameter terminal 700, such as an operating system 7021 and an application program 7022; the operating system 7021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 7022 can contain various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The deep reinforcement learning-based directional energy deposition process parameter method provided by the embodiments of the present application can be included in the application program 7022.
[0095] The method disclosed by the embodiments of the present application can be applied to the processor 701 or implemented by the processor 701. The processor 701 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 701. The processor 701 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor 701 can implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application. The general-purpose processor 701 can be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiments of the present application, the steps can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium in the memory, and the processor reads the information in the memory to complete the steps of the foregoing method in combination with the hardware.
[0096] In the example embodiments, the deep reinforcement learning-based directed energy deposition process parameter terminal 700 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), or the like for executing the foregoing deep reinforcement learning-based directed energy deposition process parameter method.
[0097] Those of ordinary skill in the art can understand that all or part of the steps of the foregoing method embodiments can be completed by computer program-related hardware. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the foregoing method embodiments are executed; and the foregoing storage medium includes ROM, RAM, magnetic or optical disc, and various media that can store program codes.
[0098] In the embodiments provided in the present application, the computer readable and writable storage medium can include a read-only memory, a random access memory, an EEPROM, a CD-ROM or other optical disk storage device, a magnetic disk storage device or other magnetic storage device, a flash memory, a U disk, a mobile hard disk, or any other medium capable of storing desired program code in the form of instructions or data structures and capable of being accessed by a computer. In addition, any connection can be appropriately referred to as a computer readable medium. For example, if instructions are sent from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology such as infrared, radio and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave is included in the definition of the medium. However, it should be understood that the computer readable and writable storage medium and the data storage medium do not include connections, carriers, signals or other transitory media, but are intended for non-transitory, tangible storage media. As used in the application, magnetic disks and optical disks include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, wherein magnetic disks typically magnetically copy data, and optical disks optically copy data with a laser.
[0099] In summary, the present application provides a method, device, terminal and medium for improving the efficiency of directed energy deposition process parameters based on deep reinforcement learning. The present application provides a method for improving the efficiency of directed energy deposition process parameters based on deep reinforcement learning. A simulation system based on laser directed energy deposition is constructed, and a deep reinforcement learning proximal optimization algorithm is used to explore the optimal strategy, thereby generating a control strategy containing laser power, speed and motion trajectory parameters, to improve the intelligent and efficient selection of optimal process parameter strategies in the directed energy deposition process. Compared with the printing strategy obtained by traditional manual experience control, the present application can shorten the manual trial and error time and sample cost, and the deep reinforcement learning strategy reduces the variability of the hardness of the obtained sample. It can be applied to various real production environments, reducing the difficulty of manual control, thereby efficiently obtaining optimized process parameters and strategies. Therefore, the present application effectively overcomes the various shortcomings in the prior art and has high industrial utilization value.
[0100] The above embodiments only exemplarily illustrate the principles and effects of the present application, and are not intended to limit the present application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical idea disclosed in the present application should be covered by the claims of the present application.
Claims
1. A method for controlling the parameters of a directional energy deposition process based on deep reinforcement learning, characterized in that, include: The process involves obtaining the specifications of the model to be printed and the environmental variable parameter set of the real energy deposition system, and then constructing a simulated energy deposition system. The process of obtaining the specifications of the model to be printed and the environmental variable parameter set of the real energy deposition system includes: selecting a set of partial differential equations and the environmental variable parameter set for optimizing the deep reinforcement learning model based on the physical characteristics of the real energy deposition system; and setting third-type boundary conditions for solving the set of partial differential equations based on the printed model parameters and the real energy deposition system. The set of partial differential equations includes the Rosenthal equations and the three-dimensional heat transfer equations. The state information of the simulated energy deposition system is obtained and input into the initialized deep reinforcement learning model; Based on the specifications of the model to be printed, an optimization algorithm is used to optimize the printing strategy in a deep reinforcement learning model to generate the optimal printing strategy. The specific process includes: obtaining a reward value and a state information set from the simulated energy deposition system, and extracting features from the state information set to generate a state feature set; inputting the reward value and the state feature set into an experience replay pool to store the current reward value and state feature set; performing a value evaluation operation on the state feature set in the experience replay pool, and extracting action parameters containing values above a threshold from the state feature set; updating the action network based on the action parameters, and performing corresponding operations in the simulated energy deposition system according to the action network to generate the updated reward value and state feature set; the optimization algorithm includes a proximal optimization algorithm. Based on the optimal printing strategy, control parameters for the model to be printed are generated, and the real energy deposition system is deployed according to the control parameters.
2. The method for controlling directional energy deposition process parameters based on deep reinforcement learning according to claim 1, characterized in that, Based on the specifications of the model to be printed, the process of optimizing the printing strategy in a deep reinforcement learning model using an optimization algorithm to generate the optimal printing strategy also includes: The loss function is calculated using the performance target of the model to be printed and the updated state feature set; Based on the loss function, the hyperparameters in the deep reinforcement learning model are adjusted to update the deep reinforcement learning model.
3. The method for controlling directional energy deposition process parameters based on deep reinforcement learning according to claim 1, characterized in that, The process of initializing a deep reinforcement learning model includes initializing the action network, evaluation network, and reward function in the deep reinforcement learning model.
4. A device for controlling the parameters of a directional energy deposition process based on deep reinforcement learning, characterized in that, include: The simulation system construction module is used to obtain the specifications of the model to be printed and the environmental variable parameter set of the real energy deposition system, and to construct the simulated energy deposition system accordingly. The process of obtaining the specifications of the model to be printed and the environmental variable parameter set of the real energy deposition system includes: selecting a set of partial differential equations and the environmental variable parameter set for optimizing the deep reinforcement learning model based on the physical characteristics of the real energy deposition system; and setting third-type boundary conditions for solving the set of partial differential equations based on the printed model parameters and the real energy deposition system. The set of partial differential equations includes the Rosenthal equations and the three-dimensional heat transfer equations. Initialization module: Used to initialize the deep reinforcement learning model, obtain the state information of the simulated energy deposition system and input it into the initialized deep reinforcement learning model; The strategy optimization module is used to optimize the printing strategy in a deep reinforcement learning model based on the specifications of the model to be printed, using an optimization algorithm to generate the optimal printing strategy. The specific process includes: obtaining a reward value and a state information set from the simulated energy deposition system, and extracting features from the state information set to generate a state feature set; inputting the reward value and the state feature set into an experience replay pool to store the current reward value and state feature set; performing a value evaluation operation on the state feature set in the experience replay pool, and extracting action parameters containing values above a threshold from the state feature set; updating the action network based on the action parameters, and performing corresponding operations in the simulated energy deposition system according to the action network to generate the updated reward value and state feature set; the optimization algorithm includes a proximal optimization algorithm. Strategy deployment module: used to generate control parameters for the model to be printed based on the optimal printing strategy and deploy the real energy deposition system according to the control parameters.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for controlling the process parameters of directional energy deposition based on deep reinforcement learning as described in any one of claims 1 to 3.
6. An electronic terminal, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory to cause the terminal to perform the deep reinforcement learning-based directional energy deposition process parameter control method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Intelligent coating track planning method based on deep reinforcement learning
CN115408813A