Intelligent decision-making method for forging-heat treatment quality energy efficiency collaborative process parameters
By constructing a multi-objective evaluation model that integrates mechanism and data-driven approaches and using deep reinforcement learning algorithms, the problems of strong dependence on human experience and low efficiency in the optimization of forging-heat treatment process parameters are solved, and efficient and reliable decision-making for the coordinated optimization of quality and energy efficiency is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV OF TECH
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-08
AI Technical Summary
Existing methods for optimizing forging-heat treatment process parameters rely on manual experience, which is inefficient and makes it difficult to obtain a balanced optimal solution set, thus failing to effectively resolve the conflict and trade-off between quality and energy efficiency.
A multi-objective evaluation model integrating mechanism and data-driven approaches is constructed. By combining deep reinforcement learning algorithms, the forging-heat treatment process is modeled as a Markov decision process. Through multi-objective space decomposition and parallel strategy optimization, a Pareto process parameter solution set is generated.
It enables efficient searching in high-dimensional parameter space and outputs a set of solutions for quality and energy efficiency co-optimization that can be flexibly selected, thereby improving the reliability and efficiency of process optimization and ensuring scientific decision-making in engineering practice.
Smart Images

Figure CN121995877A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of heat treatment process control technology, and in particular to an intelligent decision-making method for process parameters that coordinates forging and heat treatment quality and energy efficiency. Background Technology
[0002] Forging and heat treatment are core manufacturing processes for critical load-bearing components in aerospace, automotive, and energy industries. Process parameters directly determine the geometric accuracy, microstructure, and service reliability of the parts. Meanwhile, forging heating, holding, and heat treatment processes consume a significant amount of energy, making them typical high-energy-consuming manufacturing processes. In the context of green manufacturing, how to reduce energy consumption and improve energy efficiency while ensuring quality and performance is a critical issue that urgently needs to be addressed.
[0003] Currently, the optimization of forging-heat treatment process parameters mainly relies on the following technical solutions: First, methods based on engineering experience and experimental design, where process engineers use their experience and knowledge, combined with orthogonal experiments, response surface methodology, and other methods, to search for and select optimal parameters within a limited set. Second, methods based on physical mechanism models and numerical simulations, using finite element analysis, thermodynamic calculations, and other models to predict forging deformation, temperature field, and microstructure evolution, and then determining the process window through parameter scanning or simple optimization algorithms. Third, methods based on traditional intelligent optimization algorithms or single-objective reinforcement learning, employing search techniques such as genetic algorithms and particle swarm optimization, or combining reinforcement learning with a single performance index for policy learning.
[0004] However, the aforementioned traditional approaches have limitations: they rely heavily on human experience, have low optimization efficiency, and are difficult to obtain a balanced optimal solution set. First, experience-based and experimental methods heavily depend on expert knowledge and incur high trial-and-error costs. Second, while numerical simulation methods have high accuracy, they are computationally time-consuming and cannot support large-scale optimization in high-dimensional parameter spaces. Third, traditional intelligent optimization and single-objective reinforcement learning methods cannot effectively address the inherent conflicts and trade-offs between multiple objectives such as quality and energy efficiency. They often only obtain a local optimal solution for a single objective or a subjectively weighted compromise solution, making it difficult to systematically generate a uniformly distributed Pareto optimal solution set. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent decision-making method for the process parameters of forging-heat treatment quality and energy efficiency coordination, so as to solve the limitations of the traditional scheme mentioned in the background art, which are highly dependent on human experience, have low optimization efficiency, and are difficult to obtain the equilibrium optimal solution set.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent decision-making of process parameters for forging-heat treatment with coordinated quality and energy efficiency, comprising the following steps: collecting historical production data, experimental data, and simulation data of the forging-heat treatment process; constructing a sample library containing process parameters, quality indicators, and energy efficiency indicators; constructing a multi-objective evaluation model based on the sample library, integrating a mechanism model and a data-driven model, to characterize the mapping relationship between process parameters and quality and energy efficiency indicators; modeling the forging-heat treatment process as a Markov decision process, defining a state space, action space, and multi-objective reward vector, and constructing a reinforcement learning environment with process parameters as actions and quality and energy efficiency indicators as multi-objective rewards; and based on the quality and energy efficiency indicators output by the multi-objective evaluation model... The algorithm defines a multi-objective space, generates several reference points within this space, decomposes the multi-objective optimization task into several sub-problems, and associates each reference point with at least one parameterized policy. A deep reinforcement learning algorithm is used to perform parallel sampling and policy updates for the policies corresponding to each reference point. A scalarization function based on the reference points is used to scalarize the multi-objective rewards, and policy update constraints are introduced. After each iteration, the current process parameter solutions are filtered using non-dominated relations and reference point neighborhood criteria, updating the elite process parameter archive for quality and energy efficiency synergy. When training converges or a preset termination condition is met, the Pareto process parameter solution set for the multi-objective quality and energy efficiency multi-objective is output from the elite process parameter archive to guide the formulation of the forging-heat treatment process scheme.
[0007] Optionally, the multi-objective evaluation model is constructed by weighted fusion of a mechanistic model and a data-driven model, and its calculation formula is as follows: In the formula: The results are predicted by the mechanistic model. For data-driven model output, Here, 'a' represents the fusion weighting coefficient, 'm' represents the process parameter vector, and 'm' represents the target index.
[0008] Optionally, the multi-objective reward vector includes at least one quality objective and one energy efficiency objective. The quality objective is at least one of geometric dimensions, yield strength, grain size, defect rate, microstructure deviation, or mechanical property deviation. The energy efficiency objective is the total energy consumption per unit.
[0009] Optionally, the step of generating several reference points in the multi-objective space, decomposing the multi-objective optimization task into several sub-problems, and associating each reference point with at least one parameterized strategy specifically includes: generating a set of uniformly distributed reference vectors in the multi-objective space, and normalizing the reference vectors so that the sum of the elements of each reference vector is 1; defining a scalar evaluation function for each reference vector, wherein the scalar evaluation function is obtained by calculating the weighted maximum deviation between the performance of process parameters on each objective and the ideal point, thereby transforming the multi-objective evaluation into a single-objective scalar value; and associating each reference vector with a parameterized decision strategy so that the optimization objective of the parameterized decision strategy is to minimize the expected value of its scalar function.
[0010] Optionally, the steps of employing a deep reinforcement learning algorithm to perform parallel sampling and policy updates for the policies corresponding to each reference point, using a scalarization function based on the reference point to scalarize the multi-objective rewards, and introducing policy update constraints specifically include: constructing a policy network and a state-value network for each parameterized decision policy; for each reference vector, calculating the single-step scalarized reward based on its scalarized evaluation function, and then calculating the temporal difference error and advantage function estimation; updating the policy parameters using a policy gradient method with pruning constraints, where the policy loss function for the k-th reference vector is: In the formula: θ is the strategy parameter, and k is the number of reference points. The probability ratio between the old and new strategies. Let k be the dominance function corresponding to the k-th reference point. The preset constraint coefficients are defined, and clip is the truncation function. The policy parameters are updated by averaging the policy losses of all reference vectors to obtain the total loss, and the value network parameters are updated by minimizing the prediction error of the value network.
[0011] Optionally, the step of using non-dominated relations and reference point neighborhood criteria to screen the current process parameter solutions specifically includes: using a non-dominated sorting method to determine the dominance relations between process parameter solutions, and limiting the number of solutions in each reference point neighborhood based on the reference point neighborhood, so as to maintain the uniformity and diversity of the Pareto solution set in the multi-objective space.
[0012] Optionally, when optimizing the forging-heat treatment process of aluminum alloy automotive steering knuckles, the process parameters include at least one of the following: billet heating temperature, heating and holding time, preforming parameters, die forging deformation speed, die preheating temperature, solution treatment temperature and time, and artificial aging temperature and time; when optimizing the forging-heat treatment process of aircraft landing gear forgings, the process parameters include at least one of the following: isothermal forging temperature, die forging reduction, normalizing temperature, quenching heating temperature and holding time, and tempering or graded tempering temperature and time.
[0013] On the other hand, the present invention also provides an intelligent decision-making system for forging-heat treatment quality and energy efficiency collaborative process parameters, comprising: a sample library construction module, used to collect historical production data, experimental data, and simulation data of the forging-heat treatment process, and construct a sample library containing process parameters, quality indicators, and energy efficiency indicators; an evaluation model construction module, used to construct a multi-objective evaluation model based on the sample library, which integrates a mechanism model and a data-driven model to characterize the mapping relationship between process parameters and quality and energy efficiency indicators; a reinforcement learning environment construction module, used to model the forging-heat treatment process as a Markov decision process, define a state space, action space, and multi-objective reward vector, and construct a reinforcement learning environment with process parameters as actions and quality and energy efficiency indicators as multi-objective rewards; and a strategy association module, used to associate the quality indicators and energy efficiency output by the multi-objective evaluation model. The system comprises the following modules: a multi-objective space, a reference point generated within this space, a decomposition of the multi-objective optimization task into sub-problems, and association of each reference point with at least one parameterized policy; a policy update module employing a deep reinforcement learning algorithm to perform parallel sampling and policy updates for the policies corresponding to each reference point, scalarizing the multi-objective rewards using a reference point-based scalarization function, and introducing policy update constraints; an elite solution set filtering module filtering the current process parameter solutions after each iteration using non-dominated relations and reference point neighborhood criteria, updating the elite process parameter file for quality and energy efficiency synergy; and an optimal solution set output module outputting the Pareto process parameter solution set for quality and energy efficiency multi-objectives from the elite process parameter file when training converges or a preset termination condition is met, to guide the formulation of the forging-heat treatment process scheme.
[0014] On the other hand, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described intelligent decision-making method for the synergistic process parameters of forging-heat treatment quality and energy efficiency.
[0015] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described intelligent decision-making method for the coordinated process parameters of forging-heat treatment quality and energy efficiency.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This application effectively overcomes the limitations of a single information source by constructing a multi-objective evaluation model that integrates mechanism and data-driven approaches, providing a high-fidelity virtual evaluation environment for process optimization and significantly improving the reliability of the optimization foundation. By modeling the forging-heat treatment process as a Markov decision process, a key shift from static parameter optimization to dynamic sequential decision-making is achieved, enabling the optimization strategy to fully consider the temporal and state dependencies of the production process, greatly enhancing the practical applicability of the strategy. A multi-objective space decomposition strategy is adopted, transforming the complex global optimization problem into a series of parallelizable sub-problems, which not only significantly reduces the difficulty of solving the problem but also ensures, from a mechanism perspective, that the final solution set can cover multiple trade-offs. By introducing a constrained deep reinforcement learning algorithm for parallel strategy optimization and selection, efficient search in a high-dimensional parameter space is achieved while ensuring training stability, and a uniformly distributed elite solution set that approximates the true Pareto front is dynamically maintained. The final output is a Pareto process parameter solution set that can be flexibly selected rather than a single solution, enabling engineers to make scientific decisions based on actual needs, truly realizing the efficient implementation of quality and energy efficiency synergistic optimization from theory to engineering practice. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the method steps of the present invention.
[0018] Figure 2 This is a schematic diagram of the optimization results of the compromise solution for the three objectives of this invention.
[0019] Figure 3 This is a schematic diagram of the optimization results of the compromise solution for the four objectives of this invention.
[0020] Figure 4 This is a schematic diagram of the system structure of the present invention.
[0021] In the diagram: 10 - Sample library construction module, 20 - Evaluation model construction module, 30 - Reinforcement learning environment construction module, 40 - Policy association module, 50 - Policy update module, 60 - Elite solution set selection module, 70 - Optimal solution set output module. Detailed Implementation
[0022] The present invention will now be clearly and completely described in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] Those skilled in the art will understand that, unless explicitly stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in the specification of this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0025] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0026] It should be understood that the sequence number and size of each step in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.
[0027] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0028] Please refer to Figures 1-3This invention discloses an intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters, comprising the following steps: collecting historical production data, experimental data, and simulation data of the forging-heat treatment process to construct a sample library containing process parameters, quality indicators, and energy efficiency indicators; constructing a multi-objective evaluation model based on the sample library, integrating a mechanism model and a data-driven model to characterize the mapping relationship between process parameters and quality and energy efficiency indicators; modeling the forging-heat treatment process as a Markov decision process, defining a state space, action space, and multi-objective reward vector, and constructing a reinforcement learning environment with process parameters as actions and quality and energy efficiency indicators as multi-objective rewards; and determining multi-objective rewards based on the quality and energy efficiency indicators output by the multi-objective evaluation model. In the multi-objective space, several reference points are generated to decompose the multi-objective optimization task into several sub-problems, and each reference point is associated with at least one parameterized policy. A deep reinforcement learning algorithm is used to sample and update the policies corresponding to each reference point in parallel. The multi-objective reward is scalarized using a reference point-based scalarization function, and policy update constraints are introduced. After each iteration, the current process parameter solution is screened using non-dominated relations and reference point neighborhood criteria, and the elite process parameter file for quality and energy efficiency synergy is updated. When the training converges or a preset termination condition is reached, the Pareto process parameter solution set for the multi-objective of quality and energy efficiency is output from the elite process parameter file to guide the formulation of the process scheme for the forging-heat treatment process.
[0029] This application effectively overcomes the limitations of a single information source by constructing a multi-objective evaluation model that integrates mechanism and data-driven approaches, providing a high-fidelity virtual evaluation environment for process optimization and significantly improving the reliability of the optimization foundation. By modeling the forging-heat treatment process as a Markov decision process, a key shift from static parameter optimization to dynamic sequential decision-making is achieved, enabling the optimization strategy to fully consider the temporal and state dependencies of the production process, greatly enhancing the practical applicability of the strategy. A multi-objective space decomposition strategy is adopted, transforming the complex global optimization problem into a series of parallelizable sub-problems, which not only significantly reduces the difficulty of solving the problem but also ensures, from a mechanism perspective, that the final solution set can cover multiple trade-offs. By introducing a constrained deep reinforcement learning algorithm for parallel strategy optimization and selection, efficient search in a high-dimensional parameter space is achieved while ensuring training stability, and a uniformly distributed elite solution set that approximates the true Pareto front is dynamically maintained. The final output is a Pareto process parameter solution set that can be flexibly selected rather than a single solution, enabling engineers to make scientific decisions based on actual needs, truly realizing the efficient implementation of quality and energy efficiency synergistic optimization from theory to engineering practice.
[0030] In some embodiments, the multi-objective evaluation model is constructed by weighted fusion of a mechanistic model and a data-driven model, and its calculation formula is as follows: In the formula: The results are predicted by the mechanistic model. For data-driven model output, Here, 'a' represents the fusion weighting coefficient, 'm' represents the process parameter vector, and 'm' represents the target index.
[0031] Specifically, for the multi-objective evaluation and Markov decision process modeling problem, the process parameter vector of the forging-heat treatment process is set as follows: ; In the formula: The billet heating temperature, For heating and heat preservation time, This represents the total deformation. For the press speed, This refers to the quenching temperature. For cooling rate, This is the tempering temperature. For tempering time, etc.
[0032] Define the multi-objective vector as: ; In the formula: Indicates the degree of quality loss or defect rate. Indicates the comprehensive energy consumption per unit. It can indicate penalties for substandard mechanical properties or deviations in organization, etc.
[0033] An evaluation method that integrates mechanistic and data-driven models is adopted. ; In the formula: The results are predicted by the mechanistic model. For data-driven model output, Here, 'a' represents the fusion weighting coefficient, 'm' represents the process parameter vector, and 'm' represents the target index.
[0034] This application constructs a multi-objective evaluation model that integrates a mechanistic model and a data-driven model using a weighted fusion approach. This achieves a deep integration of prior physical knowledge and data-driven principles. This fusion mechanism retains the rigor and extrapolation capabilities of the mechanistic model in describing physical processes while fully utilizing the data-driven model's ability to extract complex nonlinear relationships from historical data. This effectively addresses the shortcomings of single models in prediction accuracy and generalization performance. The introduction of weighting coefficients allows the model to dynamically adjust its dependence based on specific objective characteristics, maintaining high prediction reliability even with different materials and process routes. This establishes an accurate objective function mapping relationship for subsequent optimization processes, fundamentally improving the credibility of the entire optimization system.
[0035] In some embodiments, the multi-objective reward vector includes at least one quality objective and one energy efficiency objective. The quality objective is at least one of geometric dimensions, yield strength, grain size, defect rate, microstructure deviation, or mechanical property deviation. The energy efficiency objective is the total energy consumption per unit.
[0036] This application specifically defines the quality and energy efficiency objectives within the multi-objective reward vector. The quality objective is refined into multi-dimensional indicators such as geometric dimensions, yield strength, and grain size, enabling a comprehensive measurement of the overall quality level of the manufactured part. Meanwhile, the energy efficiency objective, defined as the comprehensive energy consumption per unit, directly corresponds to the core requirements of green manufacturing. This detailed indicator definition ensures that the reward signals in the reinforcement learning environment accurately reflect the value orientation in real production, guiding the agent to consistently pursue the dual objectives of quality improvement and energy consumption reduction when exploring the process parameter space, thus avoiding deviations in optimization direction caused by ambiguous objective definitions.
[0037] In some embodiments, the step of generating a plurality of reference points in the multi-objective space, decomposing the multi-objective optimization task into a plurality of sub-problems, and associating each reference point with at least one parameterized strategy specifically includes: generating a set of uniformly distributed reference vectors in the multi-objective space, and normalizing the reference vectors so that the sum of the elements of each reference vector is 1; defining a scalar evaluation function for each reference vector, wherein the scalar evaluation function is obtained by calculating the weighted maximum deviation between the performance of the process parameters on each objective and the ideal point, thereby transforming the multi-objective evaluation into a single-objective scalar value; and associating each reference vector with a parameterized decision strategy so that the optimization objective of the parameterized decision strategy is to minimize the expected value of its scalar function.
[0038] Specifically, the forging-heat treatment process is represented as a Markov decision process (MDP).<S,A,P,r,γ> ,in, The status includes the current process number, billet temperature characteristics, cumulative strain, etc. For an action, it corresponds to a subset of process parameters that need to be decided in the current process step; The state transition probability; For multi-objective instant reward vectors, ; This is the discount factor.
[0039] Define a discounted reward for the m-th objective: ; For the multi-reference point decomposition and scalarization problem, a set of reference vectors is constructed in the target space: ; For the k-th reference vector, a scalarization function based on Chebyshev distance is used: ; In the formula: For ideal points, typically estimated from historical best or near-optimal values, each reference vector... With a parameterized strategy Corresponding. Make .
[0040] This application significantly reduces the search difficulty in high-dimensional objective spaces by generating uniformly distributed reference vectors in a multi-objective space and establishing a scalarization function, transforming the global optimization task into a series of sub-problems centered around different preference directions. Each reference vector guides the search towards a specific region of the Pareto front, while the scalarization function transforms the multi-objective evaluation into an optimizable scalar value through weighted maximum deviation calculation, ensuring that the solution set can cover different trade-off preferences. The uniform distribution of the reference vectors ensures the uniformity of the solution set distribution in the objective space from a methodological perspective.
[0041] In some embodiments, the steps of employing a deep reinforcement learning algorithm to perform parallel sampling and policy updates for policies corresponding to each reference point, scalarizing multi-objective rewards using a reference point-based scalarization function, and introducing policy update constraints specifically include: constructing a policy network and a state-value network for each parameterized decision policy; for each reference vector, calculating the single-step scalarized reward based on its scalarized evaluation function, and then calculating the temporal difference error and advantage function estimation; updating the policy parameters using a policy gradient method with pruning constraints, wherein the policy loss function for the k-th reference vector is: In the formula: θ is the strategy parameter, and k is the number of reference points. The probability ratio between the old and new strategies. Let k be the dominance function corresponding to the k-th reference point. The preset constraint coefficients are defined, and clip is the truncation function. The policy parameters are updated by averaging the policy losses of all reference vectors to obtain the total loss, and the value network parameters are updated by minimizing the prediction error of the value network.
[0042] Specifically, for the problem of deep policy update and advantage function estimation, a parameterized policy network is adopted. With State Value Network Updated using a policy gradient framework. For the reference vector... The scalarized instantaneous return is defined as: ; Based on state value function Constructing the timing difference error: ; The advantage function is calculated using the generalized advantage estimation (GAE) form: ; Where λ is the attenuation coefficient of the dominance estimate.
[0043] Using a constrained policy gradient update approach, the policy loss function corresponding to the k-th reference vector is defined as follows: ; In the formula: θ is the strategy parameter, and k is the number of reference points. The probability ratio between the old and new strategies. Let k be the dominance function corresponding to the k-th reference point. Here, the preset constraint coefficients are defined, and clip is the truncation function; where, .
[0044] The total loss is obtained by weighting or averaging the losses corresponding to each reference vector. ; The value network parameter φ is updated by minimizing the mean square error: ; in, This is the target value for scalarized returns.
[0045] This application employs a policy gradient method with pruning constraints for parallel policy optimization. The pruning mechanism, by limiting the step size of policy updates, effectively prevents training oscillations caused by policy mutations, enabling smooth convergence of the learning process. Parallel processing of policy updates corresponding to each reference vector fully utilizes computational resources and accelerates the exploration of the Pareto front. The introduction of a value network provides a reliable advantage function estimate for policy updates by accurately estimating state values, thereby guiding the agent to learn long-term beneficial decision policies. This learning framework ensures efficient and stable policy search in multi-objective environments.
[0046] In some embodiments, the step of using non-dominated relations and reference point neighborhood criteria to screen the current process parameter solutions specifically includes: using a non-dominated sorting method to determine the dominance relations between process parameter solutions, and limiting the number of solutions in each reference point neighborhood based on the reference point neighborhood, so as to maintain the uniformity and diversity of the Pareto solution set in the multi-objective space.
[0047] This application achieves intelligent maintenance of the elite solution set by combining a dynamic selection mechanism of non-dominated sorting and a reference point neighborhood criterion. Non-dominated sorting ensures that the archived solutions are the optimal solutions in the current iteration, effectively preserving excellent individuals from the evolutionary process; while the reference point neighborhood constraint controls the number of solutions in each preference direction, preventing the solution set from becoming overly concentrated in certain regions and ensuring the diversity and breadth of the solution set distribution. This dual selection mechanism enables the algorithm to maintain a good balance between exploration and utilization, driving the solution set to continuously approach the true Pareto front while maintaining the good distribution characteristics of the solution set in the target space.
[0048] In some embodiments, when optimizing the forging-heat treatment process of aluminum alloy automotive steering knuckles, the process parameters include at least one of billet heating temperature, heating and holding time, preforming parameters, die forging deformation speed, die preheating temperature, solution treatment temperature and time, and artificial aging temperature and time; when optimizing the forging-heat treatment process of aircraft landing gear forgings, the process parameters include at least one of isothermal forging temperature, die forging reduction, normalizing temperature, quenching heating temperature and holding time, and tempering or graded tempering temperature and time.
[0049] Specifically, for the problem of non-dominated solution selection and Pareto front construction, after each training iteration, the process parameter solutions generated by the current strategy are collected. and its target value The non-dominated sorting and reference point neighborhood mechanism are used for filtering.
[0050] For any two solutions and When the following conditions are met: , and , Then it is called a solution Dominant Solution The solution set that is not dominated by other solutions is selected as the candidate Pareto front, and the number of solutions near each reference point is limited according to the neighborhood distribution of the reference point, thereby maintaining the uniformity and diversity of the solution set in the multi-objective space.
[0051] This application targets aluminum alloy automotive steering knuckles, covering the key parameters of the entire process chain from heating and forming to heat treatment. For high-strength steel forgings used in aircraft landing gear, it focuses on controlling core parameters such as forging temperature field, deformation amount, and heat treatment regime. It fully considers the different material properties and functional requirements of components, providing a universal solution for process optimization of complex forgings, and has broad industrial application prospects.
[0052] Please refer to Figure 4On the other hand, the present invention also provides an intelligent decision-making system for forging-heat treatment quality and energy efficiency collaborative process parameters, comprising: a sample library construction module, used to collect historical production data, experimental data, and simulation data of the forging-heat treatment process, and construct a sample library containing process parameters, quality indicators, and energy efficiency indicators; an evaluation model construction module, used to construct a multi-objective evaluation model based on the sample library, which integrates a mechanism model and a data-driven model to characterize the mapping relationship between process parameters and quality and energy efficiency indicators; a reinforcement learning environment construction module, used to model the forging-heat treatment process as a Markov decision process, define a state space, action space, and multi-objective reward vector, and construct a reinforcement learning environment with process parameters as actions and quality and energy efficiency indicators as multi-objective rewards; and a policy association module, used to associate the quality indicators and energy efficiency output by the multi-objective evaluation model. The system comprises the following modules: a multi-objective space, a reference point generated within this space, a decomposition of the multi-objective optimization task into sub-problems, and association of each reference point with at least one parameterized policy; a policy update module employing a deep reinforcement learning algorithm to perform parallel sampling and policy updates for the policies corresponding to each reference point, scalarizing the multi-objective rewards using a reference point-based scalarization function, and introducing policy update constraints; an elite solution set filtering module filtering the current process parameter solutions after each iteration using non-dominated relations and reference point neighborhood criteria, updating the elite process parameter file for quality and energy efficiency synergy; and an optimal solution set output module outputting the Pareto process parameter solution set for quality and energy efficiency multi-objectives from the elite process parameter file when training converges or a preset termination condition is met, to guide the formulation of the forging-heat treatment process scheme.
[0053] Specifically, the present invention will be further described below through specific embodiments.
[0054] Example 1: The test subject was a 6082 aluminum alloy steering knuckle blank for a certain vehicle model. Its original process flow included: blank blanking and induction heating, pre-forming, pre-forging, final forging, trimming and shaping, solution treatment, artificial aging, cleaning and machining, and inspection. The main parameters of the original process were: blank heating temperature approximately 470–490℃, holding time 20–30 min; final forging temperature approximately 440–460℃, deformation speed 80–110 mm / s; solution treatment temperature 535–545℃, holding time 2.0–2.5 h, water cooling; artificial aging temperature 170–180℃, holding time 6–8 h. While meeting the mechanical performance requirements, this original process had issues such as high energy consumption per unit part, large dispersion in hardness and strength, and incomplete filling and slight folding in some batches.
[0055] Data from several recent production batches were collected, including billet size and material, heating temperature and time, forging load curve, final forging temperature, solution treatment and aging process parameters, energy metering data, etc.; quality inspection data included appearance defect rate, internal defect detection results, microstructure level, Brinell hardness, yield strength and tensile strength, etc.
[0056] The following quality objectives were selected: defect rate and mechanical property deviation; hardness and strength; and energy efficiency: comprehensive energy consumption per unit. A multi-objective evaluation model was constructed by weighted fusion of mechanistic model and data-driven model.
[0057] The state vector includes the current process type, billet temperature range, cumulative equivalent strain, previous process parameters, and cumulative energy consumption. The action vector corresponds to different parameter subsets in different processes, including heating temperature, holding time, final forging target temperature, deformation rate, solution temperature and time, and aging temperature and time. Several reference vectors are preset in the multi-objective space, and the multi-objective deep reinforcement learning algorithm of this invention is used for training. After training, a set of Pareto process parameter solutions is obtained.
[0058] As shown in Tables 1 and 2, under the premise of ensuring that key performance indicators such as yield strength and tensile strength meet the product technical conditions, the optimized process of this invention reduces the comprehensive energy consumption of a unit 6082 aluminum alloy steering knuckle by about 10% to 15%, significantly reduces quality fluctuations, and significantly reduces the defect rate, thus verifying the effectiveness of the method of this invention on aluminum alloy forgings.
[0059] Table 1: Comparison of main process parameters of 6082 aluminum alloy steering knuckle.
[0060]
[0061] Table 2: Comparison of quality and energy efficiency indicators of 6082 aluminum alloy steering knuckle.
[0062]
[0063] Example 2: The test subject was a 300M steel integral forging for landing gear of a certain type of aircraft. The original process flow included: pretreatment of large-size billets, multi-directional die forging / ring forging, isothermal / temperature-controlled forging, normalizing, quenching, tempering or staged tempering, non-destructive testing and dimensional inspection, and service condition simulation evaluation. The main parameters of the original process were: preheating temperature 400-450℃, final heating temperature 1150-1200℃; isothermal forging temperature 900-930℃, total reduction 40%-50%; normalizing temperature 900-930℃; quenching heating temperature 870-890℃, oil cooling; tempering temperature 300-320℃, holding time 2.5-3.5h, with some processes using staged tempering. This original process met the current aviation specifications for strength and toughness, but it had problems such as high heating and holding energy consumption, large residual stress, and sensitivity to fluctuations in billet properties.
[0064] In this embodiment, the following main multi-objective indicators are selected: quality objective: a comprehensive quality loss function composed of internal defect indicators, peak residual tensile stress, and fatigue life safety margin; energy efficiency objective: comprehensive energy consumption per unit; and robustness objective: the probability that key performance indicators exceed tolerances under fluctuations in billet performance and equipment precision disturbances. A mechanism-data fusion approach is used to establish the objective function for multi-objective reward evaluation in a reinforcement learning environment.
[0065] The state vector includes the current process, average temperature and temperature gradient characteristics of the billet / forging, cumulative deformation, current residual stress estimation, and cumulative energy consumption; the action vector includes the isothermal forging temperature range, reduction distribution, forging end temperature, quenching heating temperature and holding time, and tempering / stage tempering temperature and time. A set of Pareto process parameter solutions is obtained by parallel invocation of the finite element thermo-mechanical-phase transformation coupling model and the fatigue life assessment model through a high-performance computing platform, combined with the multi-objective deep reinforcement learning algorithm of this invention.
[0066] As shown in Tables 3 and 4, under the premise of meeting or slightly improving the fatigue life safety margin and internal defect control level, the optimized process scheme obtained by the method of the present invention can reduce the unit comprehensive energy consumption of 300M steel landing gear forgings by about 10% to 18%, significantly reduce the peak value of residual tensile stress, reduce the sensitivity to billet and equipment fluctuations, and increase the pass rate of key performance indicators from about 96% to about 99%. This shows that the method of the present invention also has good applicability and promotion value for high-strength steel large load-bearing forgings.
[0067] Table 3: Comparison of main process parameters for 300M steel landing gear forgings.
[0068]
[0069] Table 4: Comparison of quality and energy efficiency indicators of 300M steel landing gear forgings.
[0070]
[0071] Example 3: The steps are similar to those in Example 1 and will not be repeated here. This time, yield strength and grain size are used as two quality objectives, and total energy consumption is the third objective, for a total of three optimization objectives. The optimization results of the three objectives obtained by this method are as follows: Figure 2 As shown, the optimal compromise solution was obtained. This solution yielded a yield strength of 382.82 MPa, a grain size of 20.31 μm, and an energy consumption of 0.6 kWh, demonstrating the feasibility of the proposed method.
[0072] Example 4:
[0073] The steps are similar to those in Example 2 and will not be repeated here. This time, yield strength and grain size are used as two quality targets, press energy consumption as the first energy consumption target, and gas furnace energy consumption as the second energy consumption target, for a total of four optimization targets. The optimization results of the four targets obtained by this method are as follows: Figure 3 As shown, the optimal compromise solution was obtained. This solution yielded a yield strength of 1490.3 MPa, a grain size of 46.4 micrometers, a press energy consumption of 984.9 kWh, and a gas furnace energy consumption of 1183.9 kWh. This demonstrates the feasibility of the proposed method.
[0074] On the other hand, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described intelligent decision-making method for the synergistic process parameters of forging-heat treatment quality and energy efficiency.
[0075] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described intelligent decision-making method for the coordinated process parameters of forging-heat treatment quality and energy efficiency.
[0076] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, database, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0078] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention's specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for intelligent decision-making of process parameters for synergistic quality and energy efficiency in forging and heat treatment, characterized by the following steps: include: Collect historical production data, experimental data and simulation data of the forging-heat treatment process, and build a sample library containing process parameters, quality indicators and energy efficiency indicators; Based on the aforementioned sample library, a multi-objective evaluation model is constructed that integrates the mechanism model and the data-driven model to characterize the mapping relationship between process parameters and quality and energy efficiency indicators. The forging-heat treatment process is modeled as a Markov decision process, defining the state space, action space, and multi-objective reward vector, and constructing a reinforcement learning environment with process parameters as actions and quality indicators and energy efficiency indicators as multi-objective rewards. Based on the quality and energy efficiency indicators output by the multi-objective evaluation model, a multi-objective space is determined, several reference points are generated in the multi-objective space, the multi-objective optimization task is decomposed into several sub-problems, and each reference point is associated with at least one parameterized strategy. A deep reinforcement learning algorithm is used to perform parallel sampling and policy update for the policies corresponding to each reference point. The multi-objective reward is scalarized using a reference point-based scalarization function, and policy update constraints are introduced. After each iteration, the current process parameter solutions are screened using non-dominated relations and reference point neighborhood criteria, and the elite process parameter file for quality and energy efficiency synergy is updated. When the training converges or reaches the preset termination condition, the Pareto process parameter solution set with multiple objectives of quality and energy efficiency is output from the elite process parameter file to guide the formulation of the process scheme for the forging-heat treatment process.
2. The intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters according to claim 1, characterized in that, The multi-objective evaluation model is constructed by weighted fusion of a mechanistic model and a data-driven model, and its calculation formula is as follows: ; In the formula: The results are predicted by the mechanistic model. For data-driven model output, Here, 'a' represents the fusion weighting coefficient, 'm' represents the process parameter vector, and 'm' represents the target index.
3. The intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters according to claim 1, characterized in that, The multi-objective reward vector includes at least one quality objective and one energy efficiency objective. The quality objective is at least one of geometric dimensions, yield strength, grain size, defect rate, microstructure deviation, or mechanical property deviation. The energy efficiency objective is the total energy consumption per unit.
4. The intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters according to claim 1, characterized in that, The steps of generating several reference points in the multi-objective space, decomposing the multi-objective optimization task into several sub-problems, and associating each reference point with at least one parameterized policy specifically include: In the multi-objective space, a set of uniformly distributed reference vectors are generated, and the reference vectors are normalized so that the sum of the elements of each reference vector is 1. For each reference vector, a scalar evaluation function is defined. The scalar evaluation function converts the multi-objective evaluation into a single-objective scalar value by calculating the weighted maximum deviation between the performance of the process parameters on each objective and the ideal point. Each reference vector is associated with a parameterized decision policy such that the optimization objective of the parameterized decision policy is to minimize the expected value of its scalar function.
5. The intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters according to claim 4, characterized in that, The steps of employing a deep reinforcement learning algorithm to perform parallel sampling and policy updates for policies corresponding to each reference point, using a reference point-based scalarization function to scalarize multi-objective rewards, and introducing policy update constraints specifically include: For each parameterized decision strategy, construct a policy network and a state-value network; For each reference vector, calculate the single-step scalarization return based on its scalarization evaluation function, and then calculate the time series difference error and the advantage function estimate. The policy parameters are updated using a policy gradient method with pruning constraints. The policy loss function for the k-th reference vector is: ; In the formula: θ is the strategy parameter, and k is the number of reference points. The probability ratio between the old and new strategies. Let k be the dominance function corresponding to the k-th reference point. The preset constraint coefficients are used, and clip is the truncation function; The policy parameters are updated by averaging the policy losses of all reference vectors to obtain the total loss, and the value network parameters are updated by minimizing the prediction error of the value network.
6. The intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters according to claim 1, characterized in that, The step of using non-dominated relations and the reference point neighborhood criterion to screen the current process parameter solution specifically includes: The non-dominated sorting method is used to determine the dominance relationship between process parameter solutions, and the number of solutions in each reference point neighborhood is limited based on the reference point neighborhood to maintain the uniformity and diversity of the Pareto solution set in the multi-objective space.
7. The intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters according to claim 1, characterized in that, When optimizing the forging-heat treatment process of aluminum alloy automotive steering knuckles, the process parameters include at least one of the following: billet heating temperature, heating and holding time, preforming parameters, die forging deformation speed, die preheating temperature, solution treatment temperature and time, and artificial aging temperature and time. When optimizing the forging-heat treatment process of aircraft landing gear forgings, the process parameters include at least one of the following: isothermal forging temperature, die forging reduction, normalizing temperature, quenching heating temperature and holding time, and tempering or graded tempering temperature and time.
8. A smart decision-making system for forging-heat treatment quality and energy efficiency synergy process parameters, characterized in that, include: The sample library construction module is used to collect historical production data, test data and simulation data of the forging-heat treatment process, and build a sample library containing process parameters, quality indicators and energy efficiency indicators. The evaluation model construction module is used to construct a multi-objective evaluation model that integrates the mechanism model and the data-driven model based on the sample library, so as to characterize the mapping relationship between process parameters and quality indicators and energy efficiency indicators. The reinforcement learning environment construction module is used to model the forging-heat treatment process as a Markov decision process, define the state space, action space and multi-objective reward vector, and construct a reinforcement learning environment with process parameters as actions and quality indicators and energy efficiency indicators as multi-objective rewards. The strategy association module is used to determine the multi-objective space based on the quality index and energy efficiency index output by the multi-objective evaluation model, generate several reference points in the multi-objective space, decompose the multi-objective optimization task into several sub-problems, and associate each reference point with at least one parameterized strategy. The policy update module is used to perform parallel sampling and policy update of the policy corresponding to each reference point using a deep reinforcement learning algorithm. It uses a reference point-based scalarization function to scalarize the multi-objective reward and introduces policy update constraints. The elite solution set filtering module is used to filter the current process parameter solutions after each iteration using non-dominated relations and reference point neighborhood criteria, and update the elite process parameter file for quality and energy efficiency synergy. The optimal solution set output module is used to output the Pareto process parameter solution set with multiple objectives of quality and energy efficiency from the elite process parameter file when the training converges or reaches the preset termination condition, so as to guide the formulation of the process scheme for the forging-heat treatment process.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the intelligent decision-making method for forging-heat treatment quality and energy efficiency synergy process parameters as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent decision-making method for forging-heat treatment quality-energy efficiency synergy process parameters as described in any one of claims 1 to 7.