Intelligent characterization method for hot working performance of metal

By constructing an Arrhenius high-temperature constitutive model and using deep reinforcement learning methods, combined with a deep neural network and an Actor-Critic architecture, the problem of dynamic evaluation of multiple coupled parameters in metal hot working was solved, realizing intelligent quantitative characterization and precise forming of metal hot working performance.

CN121168271APending Publication Date: 2025-12-19CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511372386.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve dynamic evaluation and real-time control of multiple coupled parameters during metal hot working. Traditional hot working diagram methods exhibit strong static discreteness, while machine learning methods rely on a large amount of labeled data and have high computational costs, making them difficult to adapt to different materials or process conditions.

Method used

Data was obtained through hot compression experiments, an Arrhenius high-temperature constitutive model and a three-dimensional activation energy map were constructed, a deep neural network was introduced to establish a virtual world model, and a deep reinforcement learning and deep deterministic policy gradient algorithm were used to construct an agent based on the Actor-Critic architecture to achieve continuous dynamic optimization of hot processing parameters.

Benefits of technology

It achieves intelligent quantitative characterization of metal hot working properties, reduces computation and training costs, is applicable to the precise forming of various metals, provides globally optimal process paths, and has high optimization accuracy and strong reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168271A_ABST
    Figure CN121168271A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent characterization method for metal hot working performance, and belongs to the technical field of intelligent manufacturing, and the method specifically comprises the following steps: obtaining metal thermal deformation data, and then constructing a research data set; establishing an Arrhenius high-temperature constitutive model, generating a three-dimensional activation energy diagram, and constructing a three-dimensional hot working diagram at the same time; constructing a virtual world model training environment, and defining a state, an action and a reward function in a hot working process based on a reinforcement learning framework; a deep reinforcement learning method and a deep deterministic strategy gradient algorithm are adopted to construct an Actor-Critic architecture agent, and continuous dynamic optimization of process parameters is realized through interaction between the Actor-Critic architecture agent and a world model; and embedding the trained intelligent agent into an actual hot working flow, outputting an optimal process path, and overlapping and fusing the optimal process path with the three-dimensional activation energy-hot working diagram to realize intelligent quantitative characterization of the metal hot working performance. Through the scheme disclosed by the invention, the representation accuracy, adaptability and flexibility are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present disclosure relates to the technical field of intelligent manufacturing, in particular to an intelligent characterization method of metal hot working performance. BACKGROUND

[0002] At present, in the field of aerospace and high-end equipment manufacturing, the precise control and performance optimization of the hot working process of metal components are directly related to the microstructure performance and service reliability of the formed parts. During hot deformation, the high-temperature rheological behavior of metals is strongly coupled by multiple parameters such as deformation temperature, strain rate and strain. Improper process parameters can easily cause flow instability, insufficient dynamic recrystallization or cracking defects, resulting in degradation of product performance or even scrap. In order to effectively avoid the above defects and achieve precise control of the processing process, it is crucial to develop an efficient and accurate hot working performance characterization method.

[0003] The hot working map is an effective tool for characterizing the hot working performance of metal materials, which is composed of power dissipation map and flow instability map, and can be used to reveal the safe deformation region corresponding to different microstructure evolution mechanisms and identify the flow instability region that needs to be avoided. Patent No. CN118010523A discloses a method for determining the hot working parameter interval of TiAl alloy, which is based on the true stress-strain curve data obtained by hot compression test, calculates the strain rate sensitivity index, dissipation efficiency factor and material flow instability parameter of TiAl alloy after isothermal hot simulation compression, obtains the power dissipation map and flow instability region, and superimposes to obtain the hot working map of TiAl alloy. In this way, the hot working process parameters are adjusted to improve the rheological behavior of TiAl alloy, and the processability and comprehensive performance of TiAl alloy are improved. However, this method mainly relies on the hot working map for static and discrete characterization, and can only reflect the processing characteristics under a specific strain, which is difficult to realize dynamic evaluation and real-time control of multi-parameter coupling in continuous deformation process, and lacks modeling and prediction ability for nonlinear and high-dimensional process states.

[0004] The numerical simulation and machine learning-based hot working performance characterization method is gradually applied, aiming to solve the above-mentioned deficiencies of the traditional processing map method in dynamicity and nonlinear modeling capability. The paper "Enhancing Constitutive Description and Workability Characterization of Mg Alloy During Hot Deformation Using Machine Learning-Based Arrhenius-type Model" proposes a new method of Arrhenius model based on machine learning, which is used to improve the accuracy and generalization ability of constitutive description and workability characterization of magnesium alloy during hot deformation. This method combines the Arrhenius model and particle swarm optimization-artificial neural network to improve the prediction accuracy of flow behavior description and enhance its generalization ability. However, the model construction process of this method depends on a large number of pre-defined labeled data, which leads to its only applicability to specific processes, thereby limiting its application range. The paper "Processing Parameters Optimization in Hot Forging of AISI 4340 Steel Using Instability Map and Reinforcement Learning" uses DMM model and Prasad instability criterion to construct instability map, combined with finite element analysis and Q-learning reinforcement learning algorithm to determine the optimal process parameters. However, it highly depends on high-fidelity finite element simulation environment, with high computational cost and low efficiency, which makes it difficult to support large-scale interactive training required by reinforcement learning; in addition, the physical-based simulation model has weak generalization ability and insufficient adaptability to different materials or process conditions.

[0005] It can be seen that it is urgent to construct an intelligent characterization method of metal hot working performance that can overcome the limitations of the above methods, that is, it can break through the static discreteness of the traditional processing map, avoid the dependence of machine learning on a large number of labeled data and the high computational cost of numerical simulation, have dynamic response and intelligent decision-making ability, and realize the precise control of the forming quality of metal parts. SUMMARY

[0006] Therefore, the embodiments of the present disclosure provide an intelligent characterization method of metal hot working performance, at least partially solving the problems of low precision, narrow application range, and poor flexibility of the existing metal hot working performance characterization method.

[0007] The embodiments of the present disclosure provide an intelligent characterization method of metal hot working performance, comprising:

[0008] Step 1, obtain metal hot deformation data through thermal compression test, interpolate and pretreat the data, and construct a research data set;

[0009] Step 2, construct an Arrhenius high-temperature constitutive model based on the research data set, calculate the activation energy value under different strain conditions, and generate a three-dimensional activation energy graph, and construct a three-dimensional hot processing map based on dynamic material model analysis and rheological instability criterion;

[0010] Step 3, introduce a deep neural network to establish a virtual world model training environment, and define the state, action and reward function in the hot processing process based on a reinforcement learning framework;

[0011] Step 4, use deep reinforcement learning method and deep deterministic policy gradient algorithm to construct an agent based on Actor-Critic architecture, and realize continuous dynamic optimization of hot processing parameters through interaction between the agent and the virtual world model;

[0012] Step 5, embed the trained agent into the actual hot processing process, output the optimal hot processing path, and superimpose the path on the three-dimensional activation energy and three-dimensional hot processing map to obtain intelligent quantitative characterization of metal hot working performance.

[0013] According to a specific implementation mode of the embodiment of the present disclosure, the metal hot deformation data includes stress-strain data under different deformation temperatures and strain rates;

[0014] The pretreatment includes normalization, and the expression of normalization is

[0015]

[0016] wherein, represents a sample value in the original data, represents the minimum value of the feature in the original data, represents the maximum value of the feature in the original data.

[0017] According to a specific implementation mode of the embodiment of the present disclosure, the step 2 specifically includes:

[0018] Step 2.1, construct an Arrhenius high-temperature constitutive model based on the research data set

[0019]

[0020] wherein, represents the true stress, represents the deformation temperature, represents the strain rate, , is a material constant, is the activation energy of thermal deformation, is the ideal gas constant, is the stress multiplication coefficient;

[0021] Step 2.2, based on the Arrhenius high-temperature constitutive model, the activation energy value under different strain conditions is calculated, and then the deformation temperature T, the logarithm of strain rate and the strain amount are plotted in a three-dimensional coordinate system composed of the three factors, to obtain a three-dimensional activation energy diagram of the metal;

[0022] Step 2.3, in the dynamic material model analysis, the power dissipation coefficient

[0023]

[0024] wherein, represents the initial strain, is the strain element, represents the strain rate sensitivity coefficient.

[0025] Step 2.4, the rheological instability criterion is calculated as

[0026]

[0027] Step 2.5, based on the stress-strain data of the metal, the relationship between the true stress and the logarithm of strain rate is fitted by a cubic polynomial, and the expression is as follows:

[0028]

[0029]

[0030] wherein, are constant term, first-order term coefficient, second-order term coefficient and third-order term coefficient, respectively;

[0031] Step 2.6, according to the expression, the strain rate sensitivity index m is calculated, and then the power dissipation coefficient and the rheological instability parameter are obtained, and then in the three-dimensional coordinate system composed of the deformation temperature T, the logarithm of strain rate and the strain amount , a three-dimensional power dissipation diagram and a three-dimensional rheological instability diagram are plotted respectively, and the two diagrams are superimposed, to obtain a three-dimensional hot working diagram of the metal.

[0032] According to a specific implementation manner of the embodiment of the present disclosure, the virtual world model comprises two parts of a dynamic model and a reward model;

[0033] The dynamic model predicts the next state based on the current state and action, enabling the agent to learn the state transition dynamics of the environment to facilitate the decision-making process;

[0034] The reward model estimates the immediate reward obtained by the agent in a given state-action, assisting the agent in evaluating the effects of potential actions and optimizing the policy;

[0035] The dynamic model adopts a deep neural network architecture to establish a nonlinear mapping relationship between metal hot processing parameters and performance indicators. The dynamic model takes deformation temperature T, strain rate and strain ε as input, and outputs η, and During the training process, the dynamic model generates prediction results through forward propagation, calculates the gradient of each layer through back propagation, updates the weights using the gradient descent method, and loops until the termination condition is met.

[0036] According to a specific implementation mode of the embodiment of the present disclosure, the expression of the reinforcement learning framework is

[0037]

[0038] wherein, represents the state, represents the action, represents the reward, represents the state transition probability function, represents the discount factor, in the process of metal hot deformation, the state is deformation temperature T, strain rate and strain ε, power dissipation coefficient η, rheological instability parameter , and hot deformation activation energy ; the action is the adjustment of deformation temperature T, strain rate ;

[0039] The reward function is

[0040]

[0041]

[0042]

[0043]

[0044] wherein, is the immediate reward obtained by the agent from the environment at time step , which is obtained by adding , and ; , and respectively represent the rewards associated with , and ; , , , , , are respectively the maximum, minimum values of , , ; denotes the restriction of the variable d to the interval , ; denotes the restriction of the variable e to the interval , ; denotes the restriction of the total reward to the interval , .

[0045] According to a specific implementation manner of an embodiment of the present disclosure, the agent comprises a policy function, a value function and an update algorithm;

[0046] The policy function selects an action according to a current state s and is constantly optimized by a policy gradient method to maximize a cumulative reward, the cumulative reward is the sum of reward values from a simulation starting time to an ending time, wherein denotes an immediate reward at a time step , is a discount factor;

[0047] The value function evaluates an expected return under a given state-action pair, and helps the agent to judge which state or action is more beneficial, when the system is in a state s, an action a is selected and an action is taken according to the policy π, an expected value of a long-term cumulative reward generated thereby is represented by an action value function:

[0048]

[0049] wherein, denotes an expectation under the policy , denotes a value that has been observed for the current state and the action ;

[0050] An optimal form of the action value function represents a maximum expected value of a cumulative reward that can be obtained by following an optimal policy , wherein, is an action performed in a state . instant reward, is an action transferred to the next state, represents the action selected in the subsequent state to maximize .

[0051] According to an implementation of an embodiment of the present disclosure, the step of constructing an agent based on an Actor-Critic architecture using a deep reinforcement learning method and a deep deterministic policy gradient algorithm includes:

[0052] The agent first generates an action through an online policy network, interacts with the environment, and collects experience tuples of states, actions, rewards, and next states, which are stored in an experience replay memory pool;

[0053] During training, the algorithm randomly samples a small batch of experience data, calculates the target Q value through an online value network, and updates the parameters of the online value network to minimize the prediction error;

[0054] The online policy network adjusts the online policy parameters according to the gradient direction provided by the online value network to maximize the cumulative reward,

[0055] The parameters of the target policy network and the target value network are gradually synchronized with the parameter changes of the online network through a soft update mechanism to ensure training stability;

[0056] The above process is iteratively executed until a preset maximum training round is reached, or until the average cumulative reward no longer significantly improves.

[0057] According to an implementation of an embodiment of the present disclosure, the step of gradually synchronizing the parameter changes of the online network through a soft update mechanism includes:

[0058] A loss function is constructed as follows:

[0059]

[0060] wherein, contains the online value network predicts the Q value based on the current state and the action , is the parameter of the online value network; the target value network predicts the Q value based on the next state and the target action , is the parameter of the target value network; the target policy network selects the target action based on the next state , parameters of the target policy network; immediate reward; discount factor;

[0061] updating the parameters of the target policy network and the target value network using a smooth soft update method

[0062]

[0063] wherein, soft update rate, parameters of the online policy network.

[0064] According to a specific implementation manner of the embodiments of the present disclosure, the expression of the parameter update of the online policy network and the online value network is:

[0065]

[0066]

[0067] wherein, and learning rate, , denotes the gradient of the Q value with respect to the parameter , denotes the gradient of the loss function with respect to the parameter .

[0068] According to a specific implementation manner of the embodiments of the present disclosure, the steps of learning to achieve continuous dynamic optimization of hot working process parameters through interaction of the agent with the virtual world model include:

[0069] The online policy network generates a deterministic action according to the current state, and the online value network evaluates the value of the state and the action;

[0070] The agent executes after injecting exploration noise in the generated action, interacts with the virtual world model, collects experience transfer tuples (s, a, r, aμ, s') and stores them in the experience replay buffer Dμ;

[0071] The virtual world model generates synthetic trajectory data by simulating the dynamics of the environment and stores them in the world model buffer DM;

[0072] When optimizing the policy, a batch of data is sampled from the buffer, the loss is calculated using the sampled data, and the parameters of the online policy network and the online value network are updated;

[0073] Through a soft update mechanism, a mixing coefficient τ is used to slowly synchronize the parameters of the target policy network and the target value network;

[0074] The above steps are iteratively performed to achieve continuous dynamic optimization of the process parameters.

[0075] The intelligent characterization scheme of the metal hot working performance in the embodiments of the present disclosure includes: step 1, obtaining metal hot deformation data through hot compression testing, and performing interpolation and preprocessing on the data to construct a research data set; step 2, constructing an Arrhenius high-temperature constitutive model based on the research data set, calculating the activation energy value under different strain conditions and generating a three-dimensional activation energy map, and constructing a three-dimensional hot working diagram based on dynamic material model analysis and rheological instability criterion; step 3, introducing a deep neural network to establish a virtual world model training environment, and defining the state, action and reward function in the hot working process based on a reinforcement learning framework; step 4, using a deep reinforcement learning method and a deep deterministic policy gradient algorithm to construct an intelligent agent based on an Actor-Critic architecture, and realizing continuous dynamic optimization of the hot working process parameters through the interaction between the intelligent agent and the virtual world model; step 5, embedding the trained intelligent agent into the actual hot working process, outputting the optimal hot working process path, and superimposing the path on the three-dimensional activation energy and the three-dimensional hot working diagram to obtain the intelligent quantitative characterization of the metal hot working performance.

[0076] The present disclosure has the following advantages: through the scheme of the present disclosure, the dynamic optimization of the process parameters is realized based on a deep reinforcement learning framework, the hot working parameters can be continuously adjusted according to the real-time state (deformation temperature, strain rate, strain, etc.) in the deformation process, the scheme is suitable for the intelligent characterization of the hot working performance of various metals, provides a globally optimal process path for precise forming, and has the characteristics of high optimization precision and strong reliability; the deep reinforcement learning framework used is an unsupervised learning mode, does not need to rely on predefined label data, and the intelligent agent generates training samples autonomously through interaction with the virtual world model, significantly reduces the dependence on physical experiments or high-fidelity simulations, and effectively reduces the data acquisition cost; a world model based on a deep neural network is proposed, which is used to efficiently predict material hot working performance indicators such as power dissipation coefficient and rheological instability parameter. The model establishes a nonlinear mapping between metal process parameters and performance indicators through supervised learning, not only significantly reduces the calculation and training cost, but also solves the simulation efficiency bottleneck of traditional finite element methods in complex multi-scale coupling problems, and provides a new virtual environment construction method for metal hot working process optimization; the DDPG framework is used to fuse the value function and the policy gradient advantage, and the parameter adjustment amount in the continuous action space is directly output through the policy network, breaking through the limitation of traditional value function methods that can only handle discrete actions, and effectively solving the optimization problem of high-dimensional continuous state space. BRIEF DESCRIPTION OF DRAWINGS

[0077] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and the drawings can be obtained by those of ordinary skill in the art without creative effort based on the drawings.

[0078] Figure 1 A flowchart of a method for intelligent characterization of metal hot working performance provided by the embodiments of the present disclosure is shown in the figure.

[0079] Figure 2 A three-dimensional activation energy diagram provided by the embodiments of the present disclosure is shown in the figure.

[0080] Figure 3 A three-dimensional hot working diagram provided by the embodiments of the present disclosure is shown in the figure.

[0081] Figure 4 A DRL-world model framework provided by the embodiments of the present disclosure is shown in the figure.

[0082] Figure 5 A deep neural network structure and training flow provided by the embodiments of the present disclosure is shown in the figure.

[0083] Figure 6 A DDPG algorithm framework provided by the embodiments of the present disclosure is shown in the figure.

[0084] Figure 7 A hot working performance intelligent characterization result provided by the embodiments of the present disclosure is shown in the figure. DETAILED DESCRIPTION

[0085] The embodiments of the present disclosure will be described in detail below with reference to the drawings.

[0086] The embodiments of the present disclosure will be described in detail below with reference to the drawings.

[0087] It is to be understood that the embodiments described hereinbelow within the scope of the appended claims. It will be apparent to one of ordinary skill in the art that aspects described herein can be implemented in a wide variety of forms, and that any specific structure and / or function described herein is merely illustrative. An aspect described herein can be implemented alone or in combination with any other aspect(s). For example, an apparatus can be implemented using any number of the aspects set forth herein. Additionally, an apparatus can be implemented using other structure and / or functionality.

[0088] It is also to be understood that the diagrams provided in the following embodiments are only schematic and that the dimensions of the various components are not necessarily to scale with one another but may have been drawn in a schematic manner for the sake of legibility only. In particular, the shapes and dimensions of the various components can be varied in accordance with a specific design requirement and / or constraints.

[0089] Furthermore, in the following description, numerous specific details are set forth in order to provide a thorough understanding of the examples. However, it will be apparent to one of ordinary skill in the art that the aspects described herein can be practiced without these specific details.

[0090] The embodiments of the present disclosure provide a method for intelligent characterization of metal hot working performance, which can be applied to the metal working process in the fields of aerospace and high-end equipment manufacturing.

[0091] Referring to Figure 1 A flowchart of a method for intelligent characterization of metal hot working performance is provided in the embodiments of the present disclosure. As shown in Figure 1 The method mainly includes the following steps:

[0092] Step 1: Obtain metal hot deformation data through hot compression test, and perform interpolation and pretreatment on the data to construct a research data set;

[0093] In specific implementation, the metal hot deformation data includes stress-strain data under different deformation temperatures and strain rates, and the data is derived from the magnesium / aluminum dissimilar metal cold spraying manufacturing process test and hot compression test conducted in the laboratory. In the cold spraying test, an AZ31 magnesium alloy hot-rolled plate with a size of 140 mm × 400 mm × 15 mm is used as the substrate, and domestic Al6061 alloy powder with a powder particle size range of 10-90 μm is used. N2 is selected as the process gas. The PSC-1000 type high-pressure cold spraying equipment is used to deposit the Al6061 alloy powder on the magnesium alloy hot-rolled plate to prepare an Al6061 coating with a thickness of 2 mm.

[0094] Hot compression tests were performed on cylindrical AZ31 / 6061 dissimilar metal composite samples with a size of Ф8 mm x 12 mm using a Gleeble-3500 thermal simulation testing machine. The test conditions were as follows: a deformation of 0.7, a deformation temperature of 523 K, 573 K, 623 K, and 673 K, and a strain rate of , , and . After obtaining the stress-strain data under different deformation temperatures and strain rates through the hot compression test, interpolation processing was performed on the deformation temperature (523-673 K) and strain rate , respectively. Finally, a research data set containing 70,000 groups of data was established. At the same time, in the process of model construction, to ensure the objectivity of the evaluation, the original data was randomly divided into a training set (70%, 49,000) and a test set (30%, 21,000). The data needs to be preprocessed before training, and the data is normalized, which can eliminate the dimensional differences between different data and accelerate the convergence of the model. The normalization calculation formula is as follows:

[0095]

[0096] wherein represents the sample value in the original data, represents the minimum value of the feature in the original data, represents the maximum value of the feature in the original data.

[0097] Step 2, based on the research data set, an Arrhenius high-temperature constitutive model is constructed, the activation energy value under different strain conditions is calculated, and a three-dimensional activation energy graph is generated, and a three-dimensional hot processing map is constructed based on dynamic material model analysis and rheological instability criterion;

[0098] In specific implementation, the activation energy under different strain conditions is calculated according to the Arrhenius high-temperature constitutive model equation . In the equation, represents the true stress, represents the deformation temperature, represents the strain rate, , is a material constant, is the thermal deformation activation energy, is the ideal gas constant, is the stress multiplication coefficient. According to the calculation result, a three-dimensional activation energy graph as shown in Figure 2 is drawn.

[0099] In the DMM analysis, the power dissipation coefficient The calculation formula is:

[0100]

[0101] wherein, represents the initial strain, usually takes a very small value; is the strain element; represents the strain rate sensitivity coefficient.

[0102] The rheological instability criterion is , wherein m is the strain rate sensitivity coefficient, which can be expressed by formula under specific strain and deformation temperature conditions.

[0103] Based on the stress-strain data of the metal, the relationship between the true stress and the logarithm of the strain rate is fitted by a cubic polynomial, and the expression is as follows:

[0104]

[0105]

[0106] wherein are constant term, first-order term coefficient, second-order term coefficient and third-order term coefficient, respectively.

[0107] The strain rate sensitivity index m is calculated according to the expression, and then the values of the power dissipation coefficient and the rheological instability parameter are obtained. Then in the three-dimensional coordinate system composed of the deformation temperature , the logarithm of the strain rate and the strain , the three-dimensional power dissipation diagram and the three-dimensional rheological instability diagram are drawn respectively, and the two are superimposed to draw the three-dimensional hot working diagram of the metal as shown in Figure 3 .

[0108] Step 3, introduce a deep neural network to establish a virtual world model training environment, and define the state, action and reward function in the hot working process based on a reinforcement learning framework;

[0109] In specific implementation, the reinforcement learning basic framework is modeled by a Markov decision process, which can be expressed by the following formula: . Wherein, represents the state, which is a general description of the environment at a certain time, and the whole constitutes the state space. In this embodiment, the state is the deformation temperature T, the strain rate and the strain ε, the power dissipation coefficient η, the rheological instability parameter , the thermal activation energy Q act ; Action, which is the decision made by the agent based on the current state, is the set of all possible actions. In the process of hot deformation of magnesium / aluminum dissimilar metals, the deformation temperature and strain rate need to be accurately and continuously controlled, so the action space is continuous. Reward, State transition probability function, Discount factor.

[0110] It is usually assumed that the reward is a function of the current state , the current action , and the next state , and the reward function is denoted as . In the process of hot deformation of metals, the unstable region where is less than 0 should be avoided. In addition, when is greater than 0, higher act and lower Q act usually represent better machinability. Based on the above evaluation indicators, the reward function is established as:

[0111]

[0112] where is the immediate reward obtained by the agent from the environment at time step , which is obtained by adding , and ; , is the maximum and minimum of ; represents that the total reward is limited to the interval[ , ]; , and represent the rewards related to , Q and , which can be obtained by the following formula:

[0113]

[0114]

[0115]

[0116] where , , , are the maximum and minimum of Q and ; represents that the variable d is limited to the interval[ , ]; represents restricting the variable e to the interval [ , ].

[0117] To address the problem of low efficiency and high cost of training agents in real environments, this patent proposes a world model based on deep neural networks as an environment simulator for DRL, providing an efficient and controllable training environment and generating simulated experience for agents. The developed world model framework is shown in Figure 4 . The world model mainly consists of two parts: dynamic model and reward model. The model means as follows:

[0118] Dynamic model: This model predicts the next state based on the current state and action, enabling agents to learn the state transition dynamics of the environment, thus facilitating the decision-making process. Reward model: This model estimates the immediate reward an agent will receive for a given state-action pair, helping the agent evaluate the effects of potential actions and optimize its strategy.

[0119] The dynamic model uses a deep neural network architecture to establish a nonlinear mapping relationship between material thermal processing parameters and performance indicators. In the model, deformation temperature T, strain rate and strain ε are used as inputs, and corresponding η, and Q act are used as outputs. The specific process of building the neural network is as follows:

[0120] To balance the model complexity and generalization ability, the network architecture design uses two hidden layers. Through experimental verification, the number of neurons in each layer is determined to be 64. Among them, ReLU activation function is used between the input layer and the hidden layer, and between the hidden layers. To prevent overfitting, 10-fold cross-validation is used. Training mechanism: forward propagation generates predictions, backward propagation calculates the gradient of each layer, and gradient descent is used to update the weights of each layer, and this cycle continues until the termination condition. The deep neural network framework and training process are shown in Figure 5 .

[0121] The present embodiment uses Python language and Pytorch library to establish a neural network model: the training batch size is set to 32, the learning rate is 0.001, the optimizer is selected as Adam adaptive gradient update algorithm, and the loss function is Mean Squared Error (MSE) to measure the prediction deviation. The training period is set to 10000 rounds, the loss value is output every 10 rounds to monitor the training progress, the model checkpoint is saved every 100 rounds and the test set performance is evaluated. Batch gradient descent is implemented on the training set to optimize network weights, and the average loss value is calculated regularly on the test set. The training error and validation error are recorded respectively to form a complete trajectory of model performance evolution.

[0122] The present embodiment of the disclosure uses correlation coefficient R, MSE, Average Relative Error (AARE) as the performance evaluation index of the developed model, and the calculation formula is as follows:

[0123]

[0124]

[0125]

[0126] wherein, is the predicted value of the i-th point, is the target value of the i-th point, and are the average values of the sets and respectively.

[0127] Step 4, using deep reinforcement learning method and deep deterministic policy gradient algorithm, constructing an agent based on Actor-Critic architecture, realizing continuous dynamic optimization of hot working process parameters through interaction learning between agent and virtual world model;

[0128] In deep reinforcement learning (DRL), an agent is an entity that can perceive the environment and take actions based on its state. The goal of the agent is to maximize the cumulative reward through interaction with the environment. The process of agent interacting with the environment usually involves three components: policy function, value function and update algorithm.

[0129] The policy function is the rule or function that the agent uses to select actions based on the current state, which maps the state to the action and determines the behavior of the agent at each time. The policy function used in the present embodiment is a deterministic policy, denoted as , which maps the state As input, directly output the action. That is, based on the input strain, deformation temperature, and strain rate, it directly outputs the changes in deformation temperature and strain rate.

[0130] The policy function selects an action based on the current state *s* and continuously optimizes it using the policy gradient method to maximize the cumulative reward. (Cumulative reward) This is the sum of the reward values ​​from the start to the end of the simulation. Indicates time step Instant rewards It is the discount factor, which is 0.99 in this example.

[0131] The value function evaluates the expected reward for a given state-action pair, helping the agent determine which states or actions are more advantageous. When the system is in state s, choosing action a and following policy π, the expected long-term cumulative reward can be represented by the action value function:

[0132]

[0133] in Indicating in strategy The expectations below This indicates that the current state has been observed. and actions The value of .

[0134] Its optimal form The representative follows the optimal strategy The maximum expected cumulative reward that can be obtained. It is a state Next action Instant rewards It is to perform an action The next state after that, Indicates the subsequent state Choose the one that maximizes The action.

[0135] The Deep Deterministic Policy Gradient (DDPG) algorithm is a DRL algorithm oriented towards continuous action spaces. Its core advantage lies in its ability to effectively handle complex decision-making problems involving high-dimensional continuous states and action spaces. In the optimization of hot deformation processes for magnesium / aluminum dissimilar metals, the agent generates actions through an online policy network and collects empirical data into a memory pool. Subsequently, it samples and trains an online value network to minimize the Q-value error and uses its gradient to update the online policy network. The target network ensures training stability through soft updates, ultimately achieving policy optimization. Figure 6The framework diagram of the DDPG algorithm.

[0136] The Actor-Critic architecture uses two deep neural networks: the online policy network μ(s, θ μ ) and the online value network Q(s, a, θ Q ) to learn the deterministic policy function and the optimal action value function , where θ μ and θ Q are the parameters of the online policy network and the online value network. Specifically, the online policy network directly outputs a deterministic action based on the input state, optimizes the policy by gradient ascent, and maximizes the Q value evaluated by the Critic; the online value network outputs the Q value based on the input state and the action output by the online policy network, and minimizes the estimation error of the Q value using the TD algorithm. Through this architecture, DDPG can effectively learn and continuously optimize the policy in complex continuous control tasks, improving the decision-making ability of the agent. The specific construction process of the online policy network and the online value network is as follows:

[0137] The online policy network structure adopts a three-layer fully connected architecture to realize the mapping of state to action. The input layer receives 3-dimensional state observations (ε / T / ), the first hidden layer contains 16 neurons, the second hidden layer contains 32 neurons, and the output layer generates a 2-dimensional action vector (ΔT / Δ ), which is constrained to the interval by tanh, and is linearly mapped to the target range in actual deployment.

[0138] The online value network structure adopts a dual-branch fusion design to realize state-action value evaluation. The state branch inputs 3-dimensional state, which is mapped to 16-dimensional features by the first fully connected layer, and after ReLU activation, it is compressed to 32-dimensional core features by the second fully connected layer. The action branch inputs 2-dimensional action, which is mapped to a 32-dimensional vector by an independent fully connected layer, and is aligned with the state feature dimension. The 32-dimensional outputs of the two branches are fused by element-wise addition, and after ReLU activation, they are transmitted to the output layer. Finally, the fusion features are converted to scalar Q value by the output fully connected layer.

[0139] The target network can alleviate the instability in the policy update process, reduce the uncertainty in the Q value estimation, and thus improve the stability of the training and speed up the convergence. Therefore, a loss function is constructed as follows:

[0140]

[0141] which contains the online value network , which predicts the Q value based on the current state and action , are parameters of the online value network; target value network based on the next state and the target action predicts the Q value, are parameters of the target policy network; target policy network based on the next state selects the target action , are parameters of the target policy network.

[0142] Instead of the constantly changing online policy network and online value network, the target policy network and target value network are used, at which time the value of the target term becomes stable, and the loss function can be calculated.

[0143] Although the structure of the target network is the same as that of the online network, its weights use a smooth soft update method, which helps to prevent overestimation or oscillation that may occur when directly copying the main network weights. The parameters of the target network are updated by the following expression:

[0144]

[0145] wherein, is a constant between 0.001 and 0.005, called the soft update rate. In this paper, since the scene is a complex task of material deformation, the input is a high-dimensional space of deformation temperature and strain rate, so 0.1 is taken to ensure stability.

[0146] The policy network maximizes the expected Q value by gradient ascent, and the gradient formula of is as follows:

[0147]

[0148] Thus, the parameter update method of the online policy network and the online value network can be obtained as follows:

[0149]

[0150]

[0151] wherein, and are learning rates, takes 0.000001, takes 0.00000245, represents the loss function of the parameter .

[0152] In this embodiment, the online value network can evaluate the current strategy, provide guidance for online policy network update, and improve learning stability and efficiency.

[0153] The complete implementation process of the agent decision evolution is as follows:

[0154] First, initialize the online policy network μ (s, θ μ ), Q (s, a, θ Q ) and the world model M (s, a, θ M ); initialize the target network μ' (s, θ μ' ) and Q' (s, a, θ Q' ) by copying the parameters from μ and Q; at the same time, initialize the experience replay buffer D μ and the world model replay buffer D M to be empty;

[0155] In the exploration phase, the agent performs an action a t = μ (s, θ μ ) + (σ) from the initial state s0, injects exploration noise, observes the immediate reward r t and the next state s t+1 ; the experience transition tuple (s, a, r, a μ , s') is stored in the experience replay buffer D μ to eliminate sample correlation. The framework enhances the environment dynamic simulation capability through the world model, which generates synthetic trajectories by jointly optimizing state transition and reward prediction tasks, and stores them in the world model buffer D M .

[0156] In the policy optimization phase, a small batch of data (s, a, r, s') is sampled from D μ , and μ' and Q' are used to update θ Q by minimizing the time difference error; then the policy is improved by using gradient ascent on θ Q , thereby updating θ μ ; finally, the target network is updated with a mixing coefficient to stabilize training. Through the iterative cycle of exploration, experience storage, world model optimization and policy update, the decision-making ability of the agent in complex continuous control scenarios is gradually improved.

[0157] Based on the above framework, Python is used as the programming language to build the DRL model, and the training parameters are set as follows: the total training period is N=3200, the maximum number of steps in a single training period is 1000, the training batch size is 16, the training iteration number is 10, and the added action noise scale is set to σ=0.3.

[0158] Step 5, embedding the trained agent into the actual hot working process, outputting the optimal hot working process path, and superimposing the path on the three-dimensional activation energy and three-dimensional hot working diagram to obtain the intelligent quantitative characterization of the metal hot working performance.

[0159] In specific implementation, the trained DRL model can be used for intelligent planning of the magnesium / aluminum dissimilar metal hot deformation process: the agent generates a deformation temperature-strain rate combination through DDPG policy iteration, predicts the activation energy, dissipation coefficient and instability parameter in combination with the world model, traverses the 0.1-0.7 interval at a 0.01 strain step, constructs a complete process parameter optimization sequence, and obtains the optimal hot working process path. Superimposing it on the three-dimensional activation energy-hot working diagram obtains the intelligent quantitative characterization result of the magnesium / aluminum dissimilar metal hot working performance as shown in FIG. 6. Figure 7 As can be seen, the regions where the orange curve intersects with the activation energy-hot working diagram all have low thermal activation energy and avoid the instability region, indicating the rationality of the DRL model prediction. Hot working the material under the parameters of the path can maximize the avoidance of unstable regions, reduce instability defects, and obtain the optimal hot working performance.

[0160] The intelligent characterization method of the metal hot working performance provided in this embodiment realizes dynamic optimization of process parameters based on a deep reinforcement learning framework, can continuously adjust the hot working parameters according to the real-time state (deformation temperature, strain rate, strain, etc.) in the deformation process, is suitable for intelligent characterization of the hot working performance of various metals, provides a globally optimal process path for precise forming, and has the characteristics of high optimization precision and strong reliability.

[0161] Based on the same inventive concept, the embodiment also provides a metal hot working performance intelligent characterization system, comprising a data acquisition and calculation processing module, a hot working window building module, an environment model and rule framework building module, an agent model training and decision generation module, and a hot working performance intelligent characterization module; wherein the data acquisition and calculation processing module is used to generate different combinations of metal hot deformation process parameters, simultaneously acquire physical experimental data, and establish a research data set; the hot working window building module is used to calculate activation energy, power dissipation coefficient and flow instability parameters based on stress-strain data, and then build an activation energy-hot working diagram to determine the metal hot working process window; the environment model and rule framework building module is used to establish a world model based on a deep neural network as an environment simulator, replace the traditional finite element simulation environment, provide an efficient and controllable training environment for the DRL model and generate simulation experience; at the same time, define state variables representing the process, controllable process action parameters and reward functions quantifying optimization objectives to constitute a DRL rule framework; the agent model training and decision generation module is used to build an agent model based on the DDPG algorithm to perceive the environment and take actions according to the state it is in; deeply integrate each link to build a closed-loop control framework of metal hot deformation process parameter autonomous optimization, establish an intelligent decision cycle of "interaction-storage-sampling-update"; the hot working performance intelligent characterization module is used to realize dynamic adjustment of metal hot deformation process parameters through real-time interaction of the agent, autonomously output the optimal metal hot working process path; superimpose the optimal process path on the three-dimensional activation energy-hot working diagram to realize intelligent quantitative characterization of the metal hot working performance.

[0162] The metal hot working performance intelligent characterization system provided by the embodiment has the same inventive concept and beneficial effects as the method described above, and details are not repeated here.

[0163] It should be understood that parts of the present disclosure can be realized by hardware, software, firmware or a combination thereof.

[0164] The above is merely specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any changes or replacements within the technical scope disclosed by the present disclosure can be easily conceived by those skilled in the art, which should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A smart characterization method for the hot working properties of metals, characterized in that, include: Step 1: Obtain metal hot deformation data through hot compression tests, and then interpolate and preprocess the data to construct the research dataset; Step 2: Construct an Arrhenius high-temperature constitutive model based on the research dataset, calculate the activation energy values ​​under different strain conditions and generate a three-dimensional activation energy map, and construct a three-dimensional thermal processing map based on dynamic material model analysis and rheological instability criteria. Step 3: Introduce a deep neural network to establish a virtual world model training environment, and define the state, action and reward functions in the hot processing process based on a reinforcement learning framework; Step 4: Using deep reinforcement learning and deep deterministic policy gradient algorithm, construct an agent based on the Actor-Critic architecture. Through the interaction and learning between the agent and the virtual world model, achieve continuous dynamic optimization of the hot processing parameters. Step 5: Embed the trained agent into the actual hot working process, output the optimal hot working process path, and superimpose the path with the three-dimensional activation energy and the three-dimensional hot working diagram to obtain an intelligent quantitative characterization of the metal hot working performance.

2. The method according to claim 1, characterized in that, The metal hot deformation data includes stress-strain data under different deformation temperatures and strain rates; The preprocessing includes normalization, and the normalization expression is as follows: ; in, Represents the sample values ​​in the original data. This represents the minimum value of this feature in the original data. This represents the maximum value of this feature in the original data.

3. The method according to claim 2, characterized in that, Step 2 specifically includes: Step 2.1: Construct the Arrhenius high-temperature constitutive model based on the research dataset. ; in, Represents the actual stress. Indicates the deformation temperature. Indicates strain rate, , For material constants, The activation energy for thermal deformation, Let be the ideal gas constant. This is the stress multiplication factor; Step 2.2: Calculate the activation energy values ​​under different strain conditions based on the Arrhenius high-temperature constitutive model, and then calculate the values ​​at the deformation temperature T and the logarithm of the strain rate. and dependent variables In the three-dimensional coordinate system, the three-dimensional activation energy diagram of the metal is plotted; Step 2.3, in the dynamic material model analysis, calculate the power dissipation coefficient. ; ; in, Indicates the initial strain. It is a strain-dependent infinitesimal element. Indicates the strain rate sensitivity coefficient; Step 2.4, calculate the rheological instability criterion as follows: ; Step 2.5, based on the stress-strain data of the metal, analyze the true stress. and the logarithm of strain rate The relationship between the elements is fitted with a cubic polynomial, yielding the following expression: ; ; in, These are the coefficients of the constant term, the linear term, the quadratic term, and the cubic term, respectively. Step 2.6: Calculate the strain rate sensitivity index m according to the expression, and then obtain the power dissipation coefficient. and rheological instability parameters The value, and then the deformation temperature T, the logarithm of the strain rate and dependent variables In the three-dimensional coordinate system, a three-dimensional power dissipation diagram and a three-dimensional rheological instability diagram are drawn respectively, and the two are superimposed to draw a three-dimensional thermal working diagram of the metal.

4. The method according to claim 3, characterized in that, The virtual world model consists of two parts: a dynamic model and a reward model. Dynamic models predict the next state based on the current state and actions, enabling agents to learn the dynamics of state transitions in the environment to facilitate the decision-making process. The reward model predicts the immediate reward that an agent will receive given a state-action, helping the agent to evaluate the effectiveness of potential actions and optimize strategies. The dynamic model employs a deep neural network architecture to establish a nonlinear mapping relationship between metal hot working parameters and performance indicators. The dynamic model uses deformation temperature T and strain rate as parameters... and strain As input, output , and During training, the dynamic model generates prediction results through forward propagation, calculates the gradients of each layer through backpropagation, updates the weights using gradient descent, and repeats the process until the termination condition is met.

5. The method according to claim 4, characterized in that, The expression for the reinforcement learning framework is: ; in, (State) represents the state. (Action) indicates an action. (Reward) means reward. This represents the state transition probability function. This represents the discount factor, which is used during the hot deformation of a metal, where the state is defined by the deformation temperature T and the strain rate. and strain Power dissipation coefficient Rheological instability parameters Activation energy of thermal deformation The action is the deformation temperature T and the strain rate. Adjustments; The reward function is: ; ; ; ; in, Is the agent in time step Instant rewards obtained from the environment, by , and Add them together to get; , and Respectively represent and , and Related rewards; , , , , , They are respectively , , The maximum and minimum values, This means restricting the variable d to the interval [ , ], This means restricting the variable e to the interval. Inside, Indicates the total reward Limited to the range Inside.

6. The method according to claim 5, characterized in that, The intelligent agent includes a policy function, a value function, and an update algorithm; The policy function selects an action based on the current state s and continuously optimizes it using the policy gradient method to maximize the cumulative reward. The sum of the reward values ​​from the start to the end of the simulation, where Indicates time step Instant rewards It is a discount factor; The value function evaluates the expected reward for a given state-action pair, helping the agent determine which states or actions are more advantageous. When the system is in state s, if action a is chosen and policy π is followed, the expected long-term cumulative reward is represented by the action value function. ; in, Indicating in strategy The expectations below This indicates that the current state has been observed. and actions The value; The optimal form of the action value function The representative follows the optimal strategy The maximum expected cumulative reward that can be obtained, where, It is a state Next action Instant rewards It is to perform an action The next state after that, Indicates the subsequent state Choose the one that maximizes The action.

7. The method according to claim 6, characterized in that, The steps for constructing an agent based on the Actor-Critic architecture using deep reinforcement learning methods and deep deterministic policy gradient algorithms include: The agent first generates actions through an online policy network, interacts with the environment, and collects experience tuples of state, action, reward and next state, storing them in the experience replay memory pool. During training, the algorithm randomly samples small batches of empirical data, calculates the target Q value through an online value network, and updates its parameters to minimize the prediction error; The online policy network adjusts the online policy parameters based on the gradient direction provided by the online value network to maximize cumulative rewards. The parameters of the target policy network and the target value network are gradually synchronized with the parameter changes of the online network through a soft update mechanism to ensure training stability; The above process is iterated until the preset maximum number of training rounds is reached, or until the average cumulative reward no longer increases significantly.

8. The method according to claim 7, characterized in that, The steps for gradually synchronizing the parameters of the target policy network and the target value network with the parameter changes of the online network through a soft update mechanism include: Construct a loss function as follows: ; This includes online value networks. Based on the current state and actions Predict the Q value. For parameters of the online value network; target value network Based on the next state and target action Predict the Q value. Parameters of the target value network; target policy network Based on the next state Select target action , These are the parameters of the target policy network; For instant rewards; Discount factor; The parameters of the target policy network and the target value network are updated using a smooth soft update method. ; in, For soft update rate, These are the parameters for the online policy network.

9. The method according to claim 8, characterized in that, The expressions for updating the parameters of the online policy network and the online value network are as follows: ; ; in, and It's the learning rate. , Indicates the Q value relative to the parameter gradient, Represents the loss function For parameters The gradient.

10. The method according to claim 9, characterized in that, The step of achieving continuous dynamic optimization of thermal processing parameters through interactive learning between an intelligent agent and a virtual world model includes: An online policy network generates deterministic actions based on the current state, and an online value network evaluates the value of the state and the actions. The agent injects exploratory noise into the generated actions before executing them, interacts with the virtual world model, collects experience transfer tuples (s, a, r, aμ, s') and stores them in the experience replay buffer Dμ; The virtual world model generates synthetic trajectory data by simulating environmental dynamics and stores it in the world model buffer DM; During policy optimization, batch data is sampled from the buffer, and the loss is calculated using the sampled data to update the parameters of the online policy network and the online value network. The parameters of the target policy network and the target value network are slowly synchronized using a soft update mechanism with mixed coefficients τ. The above steps are executed iteratively to achieve continuous dynamic optimization of process parameters.

Citation Information

Patent Citations

  • TiAl alloy and hot working parameter interval determination method and processing method thereof

    CN118010523A