Process parameter optimization method and system for functionally graded material and medium
By constructing a quadratic term polynomial regression model and a multidimensional continuous reinforcement model based on a dual-delay deep deterministic strategy gradient algorithm, the problem of optimizing process parameters for functionally graded materials in arc additive manufacturing was solved, thereby improving the material forming quality and optimizing the parameter combination.
Patent Information
- Application Number
- CN202511796785.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies struggle to globally optimize the process parameters for arc additive manufacturing of functionally graded materials, especially during deposition processes with multi-component gradients. A single process parameter is insufficient to adapt to the continuous changes in material composition, resulting in poor forming quality.
Orthogonal experiments were used to obtain process parameters and forming response data. A quadratic polynomial regression model was constructed as a virtual simulation environment. A multidimensional continuous reinforcement model was constructed based on the dual-delay deep deterministic strategy gradient algorithm. The optimal combination of process parameters was searched through dynamic interaction between the intelligent agent and the virtual simulation environment.
The process parameters were optimized across compositional gradients, improving the forming quality of functionally graded materials, ensuring melt width consistency and deposition layer stability, and verifying the model's predictive accuracy.
Smart Images

Figure CN121706548A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of process parameter optimization technology, and in particular to a method, system and medium for optimizing process parameters of functionally graded materials. Background Technology
[0002] The fabrication of functionally graded materials (FJTs) using additive manufacturing is essentially a complex nonlinear physical process. In actual deposition, the forming quality of FJTs is influenced by the synergistic effects of multiple parameters. These can be mainly categorized into two types: (1) material property parameters, encompassing the physicochemical properties of the matrix and the deposited material; and (2) process control parameters, including hot welding current, welding speed, interlayer cooling time, and the number of deposition layers. To address the challenge of multi-parameter synergistic optimization, existing technologies have introduced optimization methods such as response surface methodology and genetic algorithms, combined with machine learning techniques to construct process parameter prediction models, providing a new approach for optimizing the additive manufacturing process window for FJTs.
[0003] However, current research on optimizing process parameters in the preparation of functionally graded materials (FJTs) mainly focuses on powder additive manufacturing, with limited research on optimizing process parameters in arc additive manufacturing and even less on optimizing parameters in plasma welding. Furthermore, due to the inherent multi-component gradient characteristics of FJTs, single process parameters are difficult to adapt to the continuously changing deposition process of material components, making it challenging to globally optimize the process parameters and composition distribution of FJTs. Summary of the Invention
[0004] The main objective of this invention is to provide a method, system, and medium for optimizing process parameters of functionally graded materials, which can achieve process parameter optimization across compositional gradients.
[0005] To achieve the above objectives, the present invention provides a method for optimizing process parameters of functionally graded materials, the method comprising the following steps:
[0006] Orthogonal experiments were conducted to obtain the process parameters and forming response data corresponding to different composition gradients.
[0007] Based on the process parameters and the forming response data, a quadratic polynomial regression model is constructed. The quadratic polynomial regression model is used to characterize the mapping relationship between process parameters and forming response data within different composition gradients.
[0008] The quadratic polynomial regression model is used as the virtual simulation environment;
[0009] A multidimensional continuous reinforcement model is constructed based on the dual-delay deep deterministic policy gradient algorithm;
[0010] Based on the multidimensional continuous reinforcement model, the intelligent agent interacts dynamically with the virtual simulation environment to search for the optimal combination of process parameters within different component gradients.
[0011] The optimal combination of process parameters is verified through a pre-set experiment to obtain the target combination of process parameters.
[0012] In some embodiments, the different component gradients include five component regions: 316L, 75wt%316L / 25wt%IN625, 50wt%316L / 50wt%IN625, 25wt%316L / 75wt%IN625, and IN625.
[0013] In some embodiments, the process parameters include wire feed speed, welding speed, and welding current.
[0014] In some embodiments, the forming response data includes melt width, melt height, dilution rate, and wetting angle.
[0015] In some embodiments, the quadratic polynomial regression model is as follows:
[0016] ;
[0017] In the formula, Indicates the output parameters. Indicates the intercept. Denotes the coefficient of the linear term. Denotes the coefficient of the quadratic term. Represents the coefficient of the interaction term. This indicates the error term.
[0018] In some embodiments, the multidimensional continuous reinforcement model includes a state space, an action space, a multi-objective reward function, an environment model mapping module, and a dual-delay deep deterministic policy gradient model; the state space, action space, multi-objective reward function, and environment model mapping module are all set in the virtual simulation environment; the dual-delay deep deterministic policy gradient model interacts dynamically with the virtual simulation environment through an agent.
[0019] In some embodiments, the multi-objective reward function is as follows:
[0020] ;
[0021] In the formula, , and These represent the penalty term weighting coefficients corresponding to melting height, dilution rate, and wetting angle, respectively. , and Let represent the penalty terms corresponding to melt height, dilution rate, and wetting angle, respectively; This represents the melt width error term, used to reflect the main optimization objective.
[0022] In some embodiments, the dual-delay deep deterministic policy gradient model includes a first network branch and a second network branch;
[0023] The first network branch includes an online network for action policy generation, an online network for first action policy evaluation, and an online network for second action policy evaluation; the online network for action policy generation is used to generate online action policies based on input data; the online network for first action policy evaluation and the online network for second action policy evaluation evaluate the online action policies using action value functions, respectively.
[0024] The second network branch includes an action policy generation target network, a first action policy evaluation target network, and a second action policy evaluation target network. The action policy generation target network is used to generate a target action policy based on input data. The first action policy evaluation target network and the second action policy evaluation target network evaluate the target action policy after adding Gaussian noise by using an action value function and combining the evaluation results of the first action policy evaluation online network and the second action policy evaluation online network.
[0025] The policy update frequency of the action policy generation online network is lower than the value update frequency of the first action policy evaluation target network and the second action policy evaluation target network.
[0026] The policy update frequency of the target network generated by the action policy is lower than the value update frequency of the target network evaluated by the first action policy and the target network evaluated by the second action policy.
[0027] To achieve the above objectives, another aspect of the present invention provides a process parameter optimization system for functionally graded materials, the system comprising:
[0028] The acquisition module is used to obtain the process parameters and forming response data corresponding to different composition gradients through orthogonal experiments;
[0029] The first construction module is used to construct a quadratic polynomial regression model based on the process parameters and the forming response data. The quadratic polynomial regression model is used to characterize the mapping relationship between process parameters and forming response data within different composition gradients.
[0030] The second construction module is used to use the quadratic polynomial regression model as a virtual simulation environment.
[0031] The third building module is used to construct a multidimensional continuous reinforcement model based on the dual-delay deep deterministic policy gradient algorithm.
[0032] The virtual simulation module is used to dynamically interact with the virtual simulation environment through an intelligent agent based on the multidimensional continuous reinforcement model, so as to search for the optimal combination of process parameters within different component gradients.
[0033] The verification module is used to verify the optimal combination of process parameters through preset experiments, thereby verifying the accuracy of the prediction model.
[0034] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0035] The present invention provides the following beneficial effects: The method in this application obtains process parameters and forming response data corresponding to different composition gradients through orthogonal experiments. Based on these data, a quadratic polynomial regression model is constructed to characterize the mapping relationship between process parameters and forming response data within different composition gradients. Using this quadratic polynomial regression model as a virtual simulation environment, and simultaneously constructing a multidimensional continuous reinforcement model based on a dual-delay deep deterministic strategy gradient algorithm, an intelligent agent dynamically interacts with the virtual simulation environment to search for the optimal combination of process parameters within different composition gradients, thereby achieving cross-composition gradient process parameter optimization. Finally, physical experiments are used to verify the optimal combination of process parameters and determine the accuracy of the model's predictions. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0037] Figure 1 This is a flowchart of a method for optimizing process parameters of functionally graded materials provided in an embodiment of this application;
[0038] Figure 2 This is a schematic diagram of the sampling location and melt pool outline of a single-layer, single-pass deposition layer provided in the embodiments of this application;
[0039] Figure 3 This is a schematic diagram showing the comparison between the predicted melting height and the actual melting height provided in the embodiments of this application;
[0040] Figure 4 This is a schematic diagram showing the comparison between the predicted and actual weld width values provided in the embodiments of this application;
[0041] Figure 5 This is a schematic diagram showing the comparison between the predicted dilution rate and the actual dilution rate provided in the embodiments of this application;
[0042] Figure 6 This is a schematic diagram showing the comparison between the predicted and actual wetting angle values provided in the embodiments of this application.
[0043] Figure 7 This is a schematic diagram of the multidimensional continuous reinforcement model provided in the embodiments of this application;
[0044] Figure 8 This is a schematic diagram of the dual-delay deep deterministic policy gradient model provided in the embodiments of this application.
[0045] Figure 9 This is a schematic diagram illustrating the relationship between the number of training iterations and the model convergence speed and stability within the prior pool size, as provided in the embodiments of this application.
[0046] Figure 10 This is a schematic diagram illustrating the relationship between the number of training iterations and the model convergence speed and stability in the target network update rate provided in the embodiments of this application;
[0047] Figure 11 This is a schematic diagram illustrating the relationship between the number of training iterations and the model convergence speed and stability in the discount factor provided in the embodiments of this application;
[0048] Figure 12 This is a schematic diagram illustrating the relationship between the number of training iterations and the model's convergence speed and stability in the learning rate, provided in an embodiment of this application.
[0049] Figure 13 This is a schematic diagram of the value distribution under a gradient of 316L provided in an embodiment of this application;
[0050] Figure 14 This is a schematic diagram of the value distribution under a gradient of 75wt%316L / 25wt%IN625 provided in the embodiments of this application;
[0051] Figure 15 This is a schematic diagram of the value distribution under a gradient of 50wt%316L / 50wt%IN625 provided in the embodiments of this application;
[0052] Figure 16 This is a schematic diagram of the value distribution under a gradient of 25wt%316L / 75wt%IN625 provided in the embodiments of this application;
[0053] Figure 17 This is a schematic diagram of the value distribution under gradient IN625 provided in the embodiments of this application;
[0054] Figure 18This is a schematic diagram of a process parameter optimization system for functionally graded materials provided in an embodiment of this application. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0056] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly, for example, as a fixed connection, a detachable connection, or an integral connection; a mechanical connection or an electrical connection; a direct connection or an indirect connection through an intermediate medium; or a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0057] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.
[0058] Before providing a detailed description of the embodiments of this application, the terms used in the embodiments of this application are explained as follows:
[0059] The "316L / IN625" in 316L / IN625 functionally graded materials refers to a composite gradient layer material composed of 316L stainless steel and IN625 nickel-based superalloy. This functionally graded material consists of two different matrix materials: one end is 316L stainless steel, and the other end is IN625 nickel-based superalloy. The phase composition of the gradient region of the 316L / IN625 functionally graded material is the same as that of the IN625 matrix, both being γ-Ni phase.
[0060] Arc additive manufacturing is an additive manufacturing technology based on the melting of wire by an electric arc, primarily used to fabricate functionally graded metal structures. Arc additive manufacturing uses gas metal arc welding (GMAW), gas tungsten arc welding (GTAW), or plasma welding techniques to melt and deposit metal wire along a programmed path, layer by layer, to form metal parts.
[0061] Plasma arc welding uses a compressed plasma arc as a heat source. Its working principle involves applying a compression ionization effect to the arc between the tungsten electrode and the base material through a plasma nozzle. After high-frequency arc ignition, the ion gas is compressed to form a high-energy-density heat source. This technology boasts advantages such as high material utilization, strong weld pool stability, concentrated energy density, and precise arc control. Its deposition rate (2–6 kg / h) falls between that of gas tungsten inert gas (GTW) and gas metal arc welding (GMAW), and it employs a dual-mode feeding system using both metal powder (coaxial multi-path powder feeding) and wire (independent filler wire mechanism). Compared to other processes, plasma arc welding exhibits significant advantages in the preparation of functionally graded materials due to its high deposition rate, stable deposition process, multiple material feed types, low equipment cost, and precise controllability of alloy composition.
[0062] The following is a detailed description of the embodiments of this application:
[0063] like Figure 1 As shown, this application provides a method for optimizing process parameters of functionally graded materials. The method of this embodiment includes, but is not limited to, the following steps:
[0064] Step S110: Obtain the corresponding process parameters and forming response data under different composition gradients through orthogonal experiments;
[0065] Step S120: Based on the process parameters and forming response data, construct a quadratic polynomial regression model, wherein the quadratic polynomial regression model is used to characterize the mapping relationship between process parameters and forming response data within different composition gradients.
[0066] Step S130: Use a quadratic polynomial regression model as the virtual simulation environment;
[0067] Step S140: Construct a multidimensional continuous reinforcement model based on the dual-delay deep deterministic policy gradient algorithm;
[0068] Step S150: Based on the multidimensional continuous reinforcement model, the intelligent agent interacts dynamically with the virtual simulation environment to search for the optimal combination of process parameters within different component gradients.
[0069] Step S160: Verify the optimal combination of process parameters through physical experiments to verify the accuracy of the prediction model.
[0070] It is understood that this embodiment can employ a dual-wire plasma additive manufacturing platform, which consists of two wire feeders (ET-2016), a plasma power supply (DML-VO3AD), a three-axis walking mechanism, and a gas control system. The flow rates of the ion gas and the shielding gas (Ar, 99.99% purity) are 1.8 L / min and 15 L / min, respectively, and the height of the welding torch from the sample is fixed at 10 mm. Interlayer temperature is measured using an infrared thermometer, and the next layer is deposited when the interlayer temperature cools to 100°C. Based on the above experimental process, this embodiment uses an L16(4³) orthogonal array to design 16 sets of process parameter combinations under five composition gradients (316L, 75wt%316L / 25wt%IN625, 50wt%316L / 50wt%IN625, 25wt%316L / 75wt%IN625, and IN625), resulting in a total of 80 experiments. In this context, 75wt%316L / 25wt%IN625 indicates that in a specific location (or gradient layer), the mass fraction of 316L stainless steel is 75wt%, and the remaining 25wt% is another material (IN625) used in combination with it. 50wt%316L indicates that in a specific location (or gradient layer), the mass fraction of 316L stainless steel is 50wt%, and the remaining 50wt% is another material used in combination with it. 25wt%316L indicates that in a specific location (or gradient layer), the mass fraction of 316L stainless steel is 25wt%, and the remaining 75wt% is another material used in combination with it. "wt%" stands for "mass percentage," calculated as: mass of a component / total mass of the mixture × 100%. IN625 is a high-temperature nickel-based alloy, primarily composed of iron, nickel, and cobalt, and is designed for long-term operation at temperatures above 600℃. By adjusting the wire feeding speed of 316L and IN625 welding wires in real time, a continuous gradient change in material composition can be achieved, forming a 316L / IN625 functionally graded material in which the composition and microstructure vary with position.
[0071] It is understood that the process parameters in this embodiment include wire feed speed, welding speed, and welding current, and the forming response data collected under each test group includes weld width, weld height, dilution rate, and wetting angle. Specifically, as shown... Figure 2The schematic diagram shown illustrates the sampling locations and melt pool contours for a single-layer, single-pass deposition layer. To minimize output errors, this embodiment obtains melt width (W), melt height (H), and wetting angle by taking multiple cross-sections (e.g., cross-sections 1, 2, and 3) at different locations within the single-pass, single-layer deposition layer. and ) and dilution rate ( The average value of () is used as the input parameter.
[0072] Specifically, in this embodiment, an L16(4³) orthogonal array was used to design experiments for each component gradient, and the process parameter levels are as follows:
[0073] Wire feeding speed ( ): 1.8, 2.0, 2.2, 2.4 m / min (extending to 2.6 m / min in some areas);
[0074] Welding speed ( ): 30, 40, 50, 60 cm / min;
[0075] Welding current ( ): 110, 120, 130, 140, 150 A (adjusted according to ingredients).
[0076] It is understandable that, after obtaining the process parameters and forming response data, this embodiment constructs the following quadratic polynomial regression model to describe the nonlinear relationship between the process parameters and forming response data:
[0077] ;
[0078] In the formula, Indicates the output parameters. Indicates the intercept. Denotes the coefficient of the linear term. Denotes the coefficient of the quadratic term. Represents the coefficient of the interaction term. This indicates the error term.
[0079] The quadratic polynomial regression model under each gradient can be expressed by the following formula:
[0080] ;
[0081] In the formula, y i Let represent the quadratic polynomial regression model under the i-th gradient.
[0082] like Figure 3 The image shows a comparison between the predicted melt height and the actual melt height from the quadratic polynomial regression model with output parameters. Figure 4The image shows a comparison between the predicted and actual weld width values from the quadratic polynomial regression model with the output parameters. Figure 5 The image shows a comparison between the predicted and actual dilution rates of the output parameter quadratic polynomial regression model. Figure 6 The comparison between the predicted and actual wetting angle values of the quadratic polynomial regression model for the output parameters is shown, confirming that the quadratic polynomial regression model has significant representational power for the input-output parameter mapping relationship. The input parameters were normalized using Z-score, and the model accuracy was evaluated using RMSE and R². All models showed p-values less than 0.05 and R² greater than 0.8, meeting the prediction requirements.
[0083] Understandably, after obtaining the quadratic polynomial regression model, this embodiment constructs a virtual simulation environment with component gradients and process parameters as the state space and process parameter adjustments as the action space. Specifically, the virtual simulation environment of this embodiment includes a state space, an action space, and a multi-objective reward function. The state space defines a 20-dimensional state vector containing component identifiers and process parameters for five component regions. The action space defines a 15-dimensional action vector representing the continuous adjustment of process parameters within each component region. The multi-objective reward function uses multi-gradient melt width consistency as the core reward, introducing constraint penalty terms for melt height, dilution rate, and wetting angle.
[0084] It is understandable that, such as Figure 7 As shown, the multidimensional continuous reinforcement model in this embodiment includes a state space, an action space, a multi-objective reward function, an environment model mapping module, and a dual-delay deep deterministic policy gradient model. In this embodiment, the state space, action space, multi-objective reward function, and environment model mapping module are all set within a virtual simulation environment. The preset deep deterministic policy gradient model interacts dynamically with the virtual simulation environment through an agent. Specifically, the preset deep deterministic policy gradient model in this embodiment interacts dynamically with the virtual simulation environment constructed based on a quadratic polynomial regression model through an agent, autonomously searching for the optimal combination of process parameters within five sets of discrete component gradient domains, simultaneously achieving global objective optimization of melt channel width consistency, deposition layer stability, and component control accuracy.
[0085] Specifically, the state space of this embodiment is defined by the following formula:
[0086] ;
[0087] In the formula, For the first Component identifiers for each component gradient, Then it is the first The wire feed speed, welding speed, and welding current are at different gradients.
[0088] Since the multidimensional continuous reinforcement model in this embodiment includes 5 gradients, the state space of this embodiment is a 20-dimensional continuous vector.
[0089] This embodiment employs a continuous action space construction strategy, defining the process parameter adjustment amount for each component gradient as a continuous action variable, whose output is represented as a 15-dimensional continuous vector. Therefore, the state space is defined by the following formula:
[0090] ;
[0091] In the formula, For the first The gradient interval Continuous quantities ( =0.1 corresponds to a maximum adjustment range of 10%.
[0092] The multi-objective reward function in this embodiment integrates three key optimization objectives: consistent control of gradient widths across sediment layers, ensuring continuity of the deposition process, and high-precision composition gradients. Specifically, it is expressed as the following formula:
[0093] ;
[0094] In the formula, , and These represent the penalty term weighting coefficients corresponding to melting height, dilution rate, and wetting angle, respectively. This represents the melt width error term, used to reflect the main optimization objective; , and Let represent the penalty terms corresponding to melting height, dilution rate, and wetting angle, respectively.
[0095] Specifically, As a penalty for high melt height, the weighting coefficient is applied when the melt height is in the range of 3.5~4.0mm. A value of 0 is applied; a penalty is applied if the value exceeds 3.5~4.0mm. Adjusted to 100; As a dilution penalty, if the dilution rate is less than 30%, the weighting factor is adjusted. A value of 0 is penalized when it exceeds 30%. Adjusted to 100; To compensate for the wetting angle, the wetting angle is between 55 and 75 degrees. Within the range, weighting coefficient It is 0, deviating from 55~75. When within range, Adjusted to 100.
[0096] It is understood that the dual-delay deep deterministic policy gradient model in this embodiment employs a dual-Critic network, delayed policy updates, and a target policy smoothing mechanism to address the overestimation problem in deep deterministic policy gradients. Specifically, as... Figure 8 As shown, the dual-delay preset depth deterministic policy gradient model includes a first network branch and a second network branch. The first network branch includes an online network for action policy generation (Network Actor), an online network for first action policy evaluation (Network Critic1), and an online network for second action policy evaluation (Network Critic2). The action policy generation online network is used to generate online action policies based on input data. The first and second action policy evaluation online networks evaluate the online action policies using action value functions, respectively. The second network branch includes a target network for action policy generation (Target Network Actor), a target network for first action policy evaluation (Target Network Critic1), and a target network for second action policy evaluation (Target Network Critic2). The target network for action policy generation is used to generate target action policies based on input data. The first and second action policy evaluation target networks evaluate the target action policies with added Gaussian noise using action value functions, combined with the evaluation results of the first and second action policy evaluation online networks.
[0097] In this embodiment, the policy update frequency of the action policy generation online network is lower than the value update frequency of the first action policy evaluation target network and the second action policy evaluation target network; the policy update frequency of the action policy generation target network is lower than the value update frequency of the first action policy evaluation target network and the second action policy evaluation target network. Specifically, this embodiment sets up an Actor network (policy network). The update frequency is lower than that of the Critic network (e.g., update the Actor once after every two Critic updates) to ensure that the value function converges before optimizing the policy, thus avoiding oscillations in the policy network due to unstable Q-value estimation.
[0098] Understandably, in Figure 8 The model structure shown uses a dual Critic network; this embodiment uses two independent Critic networks ( , Parallel evaluation of state-action value, taking the smaller value as the target Q-value for updating, reduces overestimation bias, as shown in the following formula:
[0099] ;
[0100] In the formula, Generated by the Actor target network, For the Critic target network, This represents the discount factor. The dual-network architecture reduces the correlation of estimation errors and improves the stability of value function evaluation by decoupling the parameter space.
[0101] Understandably, based on Figure 8 In the aforementioned model structure, the target policy smoothing process in this embodiment can be achieved by introducing Gaussian noise into the target action to enhance the Critic's adaptability to action perturbations, as shown in the formula:
[0102] ;
[0103] The strategy in this embodiment updates the trajectory through a noise injection smoothing strategy, thereby suppressing the oversensitivity of the Q function to small changes in motion.
[0104] It is understood that this embodiment is for Figure 8 The model shown is optimized by prior pool size and hyperparameters to obtain... Figure 9 The diagram illustrates the relationship between the number of training iterations and the model's convergence speed and stability within the prior pool size. Figure 10 The diagram illustrates the relationship between the number of training iterations and the model's convergence speed and stability in the target network update rate. Figure 11 The diagram showing the relationship between the number of training iterations and the model's convergence speed and stability in the discount factor, and... Figure 12 The diagram illustrates the relationship between the number of training iterations and the model's convergence speed and stability within the learning rate. Through analysis... Figure 9 , Figure 10 , Figure 11 and Figure 12 The graphs show that the number of prior action pools and hyperparameters have a significant impact on the convergence speed and stability of the model. Therefore, the prior pool size in this embodiment is set to 128, the target network update rate is 0.5, the Actor network learning rate is 0.0001, the Critic network learning rate is 0.001, the discount factor is 0.75, and the number of training iterations is 500.
[0105] It is understood that by analyzing the value distribution corresponding to different state spaces under different gradients using the method of the embodiments of this application, the following can be obtained: Figure 13 , Figure 14 , Figure 15 , Figure 16 and Figure 17 The distribution shown is as follows. (Through...) Figure 13 , Figure 14 , Figure 15 , Figure 16 and Figure 17This indicates that the distribution of high reward value states differs across different composition regions of the 316L / IN625 functionally graded material. Furthermore, the multidimensional continuous enhancement model based on the dual-delay deep deterministic policy gradient algorithm can achieve synergistic optimization of multi-composition gradient process parameters. After training, the optimal combination of process parameters output by the model is shown in Table 1 below.
[0106] Table 1
[0107] Component region Wire feed speed (m / min) Welding speed (cm / min) Welding current (A) 316L 2.19 50.70 138.83 75wt% 316L 2.32 46.56 121.35 50wt% 316L 2.28 48.58 137.15 25wt% 316L 2.23 47.45 132.15 IN625 2.12 42.42 135.42
[0108] This embodiment verifies the method by conducting experiments on the parameter combination. The results show that the predicted mean square error of the melt width is 0.027 mm, the measured mean square error of the melt width is 0.031 mm, and the prediction accuracy is 87.10%. The deposition layer morphology in each composition region is good, without defects, and the melt width is highly consistent, which verifies the effectiveness and practicality of the method in this embodiment.
[0109] As can be seen from the above, the method of this application embodiment can effectively solve the problem of optimizing the functional parameters of 316L / IN625 functionally graded materials in dual-wire plasma arc additive manufacturing, and achieve process parameter optimization across composition gradients.
[0110] Reference Figure 18 This application provides a process parameter optimization system for functionally graded materials, the system comprising:
[0111] The acquisition module is used to obtain the process parameters and forming response data corresponding to different composition gradients through orthogonal experiments;
[0112] The first construction module is used to construct a quadratic polynomial regression model based on the process parameters and the forming response data, wherein the quadratic polynomial regression model is used to characterize the mapping relationship between process parameters and forming response data within different composition gradients.
[0113] The second building module is used to create a virtual simulation environment using a quadratic polynomial regression model.
[0114] The third building module is used to construct a multidimensional continuous reinforcement model based on the dual-delay deep deterministic policy gradient algorithm.
[0115] The virtual simulation module is used to dynamically interact with the virtual simulation environment through an intelligent agent based on a multidimensional continuous reinforcement model in order to search for the optimal combination of process parameters within different composition gradients.
[0116] The verification module is used to verify the optimal combination of process parameters through physical experiments, thereby verifying the accuracy of the prediction model.
[0117] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0118] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program, which is implemented when executed by a processor. Figure 1 The method shown.
[0119] It is understood that the content of the above method embodiments is applicable to this medium embodiment. The specific functions implemented in this medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical coding feature maps; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing process parameters of functionally graded materials, characterized in that, The method includes the following steps: Orthogonal experiments were conducted to obtain the process parameters and forming response data corresponding to different composition gradients. Based on the process parameters and the forming response data, a quadratic polynomial regression model is constructed. The quadratic polynomial regression model is used to characterize the mapping relationship between process parameters and forming response data within different composition gradients. The quadratic polynomial regression model is used as the virtual simulation environment; A multidimensional continuous reinforcement model is constructed based on the dual-delay deep deterministic policy gradient algorithm; Based on the multidimensional continuous reinforcement model, the agent dynamically interacts with the virtual simulation environment to search for the optimal combination of process parameters within different component gradients. The optimal combination of process parameters was verified through physical experiments to validate the accuracy of the prediction model.
2. The process parameter optimization method according to claim 1, characterized in that, The composition gradient includes five composition regions: 316L, 75wt%316L / 25wt%IN625, 50wt%316L / 50wt%IN625, 25wt%316L / 75wt%IN625, and IN625.
3. The process parameter optimization method according to claim 1, characterized in that, The process parameters include wire feed speed, welding speed, and welding current.
4. The process parameter optimization method according to claim 1, characterized in that, The forming response data includes melt width, melt height, dilution rate, and wetting angle.
5. The process parameter optimization method according to claim 1, characterized in that, The quadratic polynomial regression model is as follows: ; In the formula, Indicates the output parameters. Indicates the intercept. Denotes the coefficient of the linear term. Denotes the coefficient of the quadratic term. Represents the coefficient of the interaction term. This indicates the error term.
6. The process parameter optimization method according to claim 1, characterized in that, The multidimensional continuous reinforcement model includes a state space, an action space, a multi-objective reward function, an environment model mapping module, and a dual-delay deep deterministic policy gradient model; the state space, action space, multi-objective reward function, and environment model mapping module are all set in the virtual simulation environment; the dual-delay deep deterministic policy gradient model interacts dynamically with the virtual simulation environment through an agent.
7. The process parameter optimization method according to claim 6, characterized in that, The multi-objective reward function is as follows: ; In the formula, , and These represent the penalty term weighting coefficients corresponding to melting height, dilution rate, and wetting angle, respectively. , and Let represent the penalty terms corresponding to melt height, dilution rate, and wetting angle, respectively; This represents the melt width error term, used to reflect the main optimization objective.
8. The process parameter optimization method according to claim 6, characterized in that, The dual-delay deep deterministic strategy gradient model includes a first network branch and a second network branch. The first network branch includes an online network for action policy generation, an online network for first action policy evaluation, and an online network for second action policy evaluation; the online network for action policy generation is used to generate online action policies based on input data; the online network for first action policy evaluation and the online network for second action policy evaluation evaluate the online action policies using action value functions, respectively. The second network branch includes an action policy generation target network, a first action policy evaluation target network, and a second action policy evaluation target network. The action policy generation target network is used to generate a target action policy based on input data. The first action policy evaluation target network and the second action policy evaluation target network evaluate the target action policy after adding Gaussian noise by using an action value function and combining the evaluation results of the first action policy evaluation online network and the second action policy evaluation online network. The policy update frequency of the action policy generation online network is lower than the value update frequency of the first action policy evaluation target network and the second action policy evaluation target network. The policy update frequency of the target network for generating the action policy is lower than the value update frequency of the target network for evaluating the first action policy and the target network for evaluating the second action policy.
9. A process parameter optimization system for functionally graded materials, characterized in that, The system includes: The acquisition module is used to obtain the process parameters and forming response data corresponding to different composition gradients through orthogonal experiments; The first construction module is used to construct a quadratic polynomial regression model based on the process parameters and the forming response data. The quadratic polynomial regression model is used to characterize the mapping relationship between process parameters and forming response data within different composition gradients. The second construction module is used to use the quadratic polynomial regression model as a virtual simulation environment. The third building module is used to construct a multidimensional continuous reinforcement model based on the dual-delay deep deterministic policy gradient algorithm. The virtual simulation module is used to dynamically interact with the virtual simulation environment through an intelligent agent based on the multidimensional continuous reinforcement model, so as to search for the optimal combination of process parameters within different component gradients. The verification module is used to verify the optimal combination of process parameters through physical experiments, thereby verifying the accuracy of the prediction model.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.
Citation Information
Cited By
A method for production quality control of medical mytilus mucus nasal spray
CN122243307A