A method and system for optimizing the manufacturing of titanium-steel composite plates based on a neural network model
By collecting interface data during the manufacturing of titanium-steel composite plates and optimizing process parameters using adaptive memory networks and dual-delay deep Q networks, the problems of insufficient prediction accuracy and unreliable optimization results in existing technologies are solved, high-precision interface bonding strength prediction and process parameter optimization are achieved, and the system's adaptability and engineering practicality are improved.
Patent Information
- Application Number
- CN202510774749.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing neural network optimization methods in the manufacturing of titanium-steel composite plates lack consideration of the material constitutive relationship and physical constraints of the manufacturing process, resulting in insufficient prediction accuracy and unreliable optimization results. It is difficult to balance the relationship between exploration and utilization, and there is a lack of dynamic feedback mechanism, which affects the engineering practicality of the optimization effect.
By collecting interface data during the manufacturing process of titanium-steel composite plates, the interface bonding features are generated by using spatiotemporal feature extraction, and the interface bonding strength is predicted by combining the adaptive memory network and the material constitutive equation library. The optimized process parameter combination is generated through the double-delay deep Q network, and the process parameters are optimized by adopting the phase change temperature point staged control and theoretical-differential gradient fusion method to dynamically adjust the optimization strategy.
It improves the accuracy and interpretability of interface bonding strength prediction, balances exploration and utilization in the process of process parameter optimization, achieves precise control of process parameters near the phase change temperature point, and enhances the system's adaptability and the engineering practicality of the optimization results.
Smart Images

Figure CN120297159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and in particular to a method and system for optimizing the manufacturing of titanium-steel composite plates based on a neural network model. Background Art
[0002] As a new type of composite material, titanium-steel clad plates face a key technical challenge in their manufacturing process: controlling the interfacial bonding strength. Existing neural network optimization methods primarily use a single deep learning model to predict interfacial bonding strength, but lack consideration of the material's constitutive relationships and the physical constraints of the manufacturing process. This results in insufficient prediction accuracy and unreliable optimization results.
[0003] Conventional process parameter optimization methods for titanium-steel composite plates typically rely on expert experience or simple deep reinforcement learning algorithms, which are unable to accurately capture the dynamic changes in interface bonding characteristics across time and space. Furthermore, when optimizing parameters near phase transition temperatures, the lack of adaptive control strategies and theoretical models can lead to performance fluctuations and localized convergence issues.
[0004] Current applications of neural networks in the optimization of titanium-steel composite plate manufacturing have failed to effectively integrate theoretical models with experimental data, and lack systematic consideration of process parameter stability. In particular, existing methods struggle to balance exploration and utilization during the optimization process, and lack dynamic feedback mechanisms, hindering the engineering practicality of the optimization results. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for optimizing the manufacturing of titanium-steel composite plates based on a neural network model, which can solve the problems in the prior art.
[0006] According to a first aspect of an embodiment of the present invention, a method for optimizing the manufacturing of titanium-steel composite plates based on a neural network model is provided, comprising: collecting interface data during the manufacturing process of titanium-steel composite plates, and generating interface bonding features through time-space dimensional feature extraction; inputting the interface bonding features into an adaptive memory network, and predicting the interface bonding strength based on a material constitutive equation library and a temperature field-stress field coupling calculation unit; based on the interface bonding strength prediction results and process parameters, generating an optimized process parameter combination through a dual-delay deep Q network structure with a priority experience replay mechanism; using the optimized process parameter combination as the initial population, and adopting a phased control method of phase change temperature points and a theoretical-differential gradient fusion method to generate an optimal process parameter solution based on performance indicators and process stability; and feeding back the deviation between the actual measured performance indicators and the predicted results to the deep Q network and the genetic algorithm to dynamically adjust the process parameter optimization strategy.
[0007] In an optional embodiment, the interface bonding characteristics are input into an adaptive memory network, and based on the material constitutive equation library and the temperature field-stress field coupling calculation unit, the interface bonding strength is predicted, including: inputting the interface bonding characteristics into the material constitutive equation library, the material constitutive equation library introduces the interface layer strengthening relationship based on the Johnson-Cook model, and calculates the stress-strain relationship of the interface bonding layer; based on the stress-strain relationship, a temperature field-stress field coupling calculation unit is used to perform a separate solution to obtain the interface layer stress distribution; the interface layer stress distribution is converted into a multi-resolution tensor field representation, and the stress gradient characteristics, stress field singular point distribution characteristics and temperature-stress covariance characteristics are extracted and mapped to the deep learning feature space; the deep learning features are input into a long short-term memory network with a physical constraint embedding layer and an interface heterogeneity adaptive attention mechanism, and the interface bonding strength prediction result is output.
[0008] In an optional embodiment, the method further includes: the interface layer strengthening relationship is:
[0009] ;in is the interface layer stress, is the matrix stress, is the grain size, is the dislocation density, and is a material constant used to calculate the stress-strain relationship of the interface bonding layer; the coupling calculation unit includes: heat conduction equation: ;in, is the material density, is the specific heat capacity; is the partial derivative of temperature with respect to time, which indicates the rate of change of temperature with time; is the temperature, For time, is the thermal conductivity, is the plastic work-to-heat conversion coefficient, is stress, is the strain rate, is the friction coefficient, is the interface pressure, is the relative slip velocity, is the gradient operator, is the divergence operator; thermal stress mapping:
[0010] ;in, is the thermal strain, is the coefficient of thermal expansion, is the temperature, is the initial temperature; the stress field equilibrium equation is: ;in, is stress, is the body force, is a divergence operator; the physical constraint embedding layer embeds the first law of thermodynamics and the continuity equation into the network structure as loss function terms; the interface heterogeneity adaptive attention mechanism calculates the spatial attention score based on the query vector and key vector of the interface position, and combines the unit state and hidden state to capture the temporal dependency of the interface features, and outputs the interface binding strength prediction result.
[0011] In an optional embodiment, based on the interface bonding strength prediction result and process parameters, an optimized process parameter combination is generated through a dual-delay deep Q network structure with a priority experience replay mechanism, including: constructing a reinforcement learning state vector using the interface bonding strength prediction value, the process parameter set of the titanium-steel composite material, the process parameter change trend, the interface bonding strength target deviation, and the process stability index; constructing an action space using the adjustment amount of the process parameter set; and constraining the maximum adjustment range of the action space based on the process parameter safety window; constructing a basic reward function using the strength reward item, the stability reward item, and the efficiency reward item, and proportionally adjusting them through a physical constraint correction item to obtain a multi-objective reward function; mapping the state vector and the action space to the feature space through an encoder respectively, and then fusing them into a dual-delay deep Q network having a main network and two target networks with different update cycles, wherein the two target networks update network parameters with different time delays; setting a priority experience replay mechanism based on a temporal difference error, wherein the temporal difference error is used to calculate a priority, and the priority is used to determine a sampling probability; selecting a process parameter adjustment action based on the output value of the dual-delay deep Q network using an ε-greedy strategy, constraining parameters that exceed the process parameter safety window, and outputting an optimized process parameter combination.
[0012] In an optional embodiment, the multi-objective reward function includes: the strength reward item is calculated based on the deviation between the interface bonding strength target value and the predicted value, the stability reward item is calculated based on the ratio of the process parameter adjustment amount to the current parameter value, and the efficiency reward item is calculated based on the ratio of the process parameter adjustment amount to the maximum adjustment amplitude; the intergranular reaction layer thickness influence coefficient is calculated based on the deviation relationship between the intergranular reaction layer thickness and the critical thickness and the optimal thickness; the temperature gradient influence coefficient is calculated based on the deviation relationship between the interface temperature gradient and the optimal temperature gradient; the thermal stress influence coefficient is calculated based on the ratio relationship between thermal stress and allowable stress; the constraint weight is determined according to the sensitivity of the influence coefficient to the process parameters, and the physical constraint correction term is constructed by the weighted sum of each influence coefficient and the corresponding constraint weight; the basic reward function is proportionally adjusted by the physical constraint correction term to obtain a multi-objective reward function; according to the characteristics of the process stage, the temperature gradient constraint weight is increased in the heating stage, the intergranular reaction layer thickness constraint weight is increased in the insulation stage, and the thermal stress constraint weight is increased in the cooling stage.
[0013] In an optional embodiment, the optimized process parameter combination is used as the initial population, and the phase change temperature point is controlled in stages and the theory-differential gradient fusion method is adopted to generate the optimal process parameter solution based on the performance index and process stability. The method includes: using the optimized process parameter combination as the initial population, obtaining the phase change temperature point of the titanium steel composite material, and calculating the current optimization stage according to the population distribution. The optimization stage includes the exploration stage, the transition stage and the refinement stage. The corresponding gradient step parameters and mutation probability parameters are set for different optimization stages. When the process parameter value is close to the phase change temperature point, the disturbance amplitude of the crossover operation is reduced and the local search accuracy is improved. In the non-dominated sorting process, the process parameter is calculated according to the phase change temperature point and the optimization stage. The theoretical gradient value of the target is confidence-weighted with the differential gradient value to obtain a fusion gradient value; the performance indicators include interface bonding strength, tensile strength and elongation; based on the population evolution process, the titanium steel interface diffusion sufficiency, the rationality of the interface layer thickness and the brittle phase content are calculated respectively, the degree of constraint violation is determined, and the degree of influence of process parameter fluctuations on performance indicators is calculated; in each generation of population evaluation, the actual fluctuation value of the process parameters is obtained, and the process stability score is calculated in combination with the degree of influence, and the process cost score is calculated based on the energy consumption value, processing time and equipment operating parameters; the process stability score, process cost score, fusion gradient value and constraint violation degree are weighted to obtain a comprehensive score, and the process parameter combination with the highest comprehensive score is selected as the optimization result.
[0014] In an optional embodiment, the theoretical gradient value and the differential gradient value are confidence-weighted to obtain a fusion gradient value, including: calculating the rate of change of element concentration with temperature and time during interface diffusion, the temperature sensitivity coefficient of the intermetallic compound phase fraction near the phase transition temperature point, and the contribution coefficient of the interface microstructure characteristic parameters to the bonding strength to obtain the corresponding theoretical gradient; selecting the top N parameter combinations with the highest scores in the current population, applying a perturbation amount, and calculating the ratio of the performance index change value to the perturbation amount to obtain a differential gradient value; determining the empirical support based on the similarity between the process parameters and the historical data, reducing the credibility of the theoretical prediction near the phase transition temperature point, and determining the confidence interval based on the historical prediction error; increasing the differential gradient weight in the exploration stage to expand the search range, increasing the theoretical gradient weight in the refinement stage, implementing differentiated fusion of the performance indicators near the phase transition temperature point, and weighting the theoretical gradient and the differential gradient according to the confidence level to obtain a fusion gradient.
[0015] A second aspect of an embodiment of the present invention provides a titanium steel composite plate manufacturing optimization system based on a neural network model, comprising:
[0016] The first unit is used to collect interface data during the manufacturing process of titanium-steel composite plates and generate interface bonding features through time-space dimension feature extraction;
[0017] The second unit is used to input the interface bonding characteristics into an adaptive memory network, and predict the interface bonding strength based on the material constitutive equation library and the temperature field-stress field coupling calculation unit;
[0018] The third unit is used to generate an optimized process parameter combination based on the interface bonding strength prediction results and process parameters through a double-delay deep Q network structure with a priority experience replay mechanism;
[0019] The fourth unit is used to use the optimized process parameter combination as the initial population, adopt the phase change temperature point stage control and theoretical-differential gradient fusion method to generate the optimal process parameter solution based on performance indicators and process stability;
[0020] The fifth unit is used to feed back the deviation between the actual measured performance indicators and the predicted results to the deep Q network and genetic algorithm to dynamically adjust the process parameter optimization strategy.
[0021] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0022] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0023] This paper improves the accuracy and interpretability of interface bond strength prediction by introducing a material constitutive equation library based on interface layer strengthening relations and a temperature-stress field coupling calculation unit, combined with an adaptive memory network with a physical constraint embedding layer. A dual-delay deep Q network with a prioritized experience replay mechanism effectively balances exploration and utilization during process parameter optimization, improving the stability of parameter adjustments.
[0024] The present invention achieves precise control of process parameters near the phase transition temperature point through a phased control strategy for the phase transition temperature point and a theory-differential gradient fusion mechanism. A comprehensive evaluation system based on process stability scores and process cost scores ensures the engineering practicality of the optimization results. Combined with a dynamic feedback mechanism, the optimization strategies of the deep Q network and genetic algorithm are continuously optimized, thereby improving the system's adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flow chart of a method for optimizing the manufacturing of titanium-steel composite plates based on a neural network model according to an embodiment of the present invention;
[0026] Figure 2 Flowchart for calculating fusion gradient values. DETAILED DESCRIPTION
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0028] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0029] Figure 1 FIG. 1 is a flow chart of a method for optimizing the manufacturing of titanium-steel composite plates based on a neural network model according to an embodiment of the present invention. Figure 1As shown, the method includes: collecting interface data in the manufacturing process of titanium steel composite plates, generating interface bonding features through time and space dimension feature extraction; inputting the interface bonding features into an adaptive memory network, and predicting the interface bonding strength based on a material constitutive equation library and a temperature field-stress field coupling calculation unit; based on the interface bonding strength prediction results and process parameters, generating an optimized process parameter combination through a double-delay deep Q network structure with a priority experience replay mechanism; using the optimized process parameter combination as the initial population, adopting a phased control of phase change temperature points and a theoretical-differential gradient fusion method, generating an optimal process parameter solution based on performance indicators and process stability; feeding back the deviation between the actual measured performance indicators and the predicted results to the deep Q network and the genetic algorithm, and dynamically adjusting the process parameter optimization strategy.
[0030] In an optional embodiment, the interface bonding characteristics are input into an adaptive memory network, and based on a material constitutive equation library and a temperature field-stress field coupling calculation unit, the interface bonding strength is predicted, including: inputting the interface bonding characteristics into the material constitutive equation library, the material constitutive equation library introduces an interface layer strengthening relationship based on the Johnson-Cook model, and calculates the stress-strain relationship of the interface bonding layer; based on the stress-strain relationship, a temperature field-stress field coupling calculation unit is used to perform a separate solution to obtain the interface layer stress distribution; the interface layer stress distribution is converted into a multi-resolution tensor field representation, and stress gradient characteristics, stress field singular point distribution characteristics, and temperature-stress covariance characteristics are extracted and mapped to a deep learning feature space; the deep learning features are input into a long short-term memory network with a physical constraint embedding layer and an interface heterogeneity adaptive attention mechanism, and the interface bonding strength prediction result is output;
[0031] Wherein, the interface layer strengthening relationship is:
[0032] ;in is the interface layer stress, is the matrix stress, is the grain size, is the dislocation density, and Material constants are used to calculate the stress-strain relationship of the interface bonding layer; the coupling calculation unit includes: heat conduction equation: ;in, is the material density, is the specific heat capacity; is the partial derivative of temperature with respect to time, which indicates the rate of change of temperature with time; is the temperature, For time, is the thermal conductivity, is the plastic work-to-heat conversion coefficient, is stress, is the strain rate, is the friction coefficient, is the interface pressure, is the relative slip velocity, is the gradient operator, is the divergence operator; thermal stress mapping:
[0033] ;in, is the thermal strain, is the coefficient of thermal expansion, is the temperature, is the initial temperature; the stress field equilibrium equation is: ; is stress, is the body force, is a divergence operator; the physical constraint embedding layer embeds the first law of thermodynamics and the continuity equation into the network structure as loss function terms; the interface heterogeneity adaptive attention mechanism calculates the spatial attention score based on the query vector and key vector of the interface position, and combines the unit state and hidden state to capture the temporal dependency of the interface features, and outputs the interface binding strength prediction result.
[0034] For example, during the manufacturing process of titanium-steel composite plates, interface data is collected to extract dynamic features such as temperature change rate, pressure fluctuation period, and diffusion rate in the temporal dimension. In the spatial dimension, static features such as interface thickness distribution, grain size variation, and element concentration gradient are extracted. These features together constitute the interface bonding feature set.
[0035] The material constitutive equation library uses the improved Johnson-Cook model and introduces the interface layer strengthening relationship based on the original model. This relationship takes into account the effects of matrix stress, grain size and dislocation density, where matrix stress is obtained through tensile testing, grain size is obtained through image analysis, and dislocation density is measured by X-ray diffraction peak width method. Material constants and Through experimental calibration, the specific operation is to conduct tensile tests under different process parameters and obtain the optimal constant value through curve fitting. For typical titanium steel composite materials, The value range is 110-130MPa·μm 0.5 , The value range is 0.15-0.25 MPa·m. During the calculation, the basic parameters of the base material are first input, including the yield strength (850 MPa), hardening coefficient (0.15), strain rate sensitivity index (0.025), and thermal softening index (1.1) of the titanium alloy, as well as the corresponding parameters of the steel (yield strength 520 MPa, hardening coefficient 0.18, strain rate sensitivity index 0.018, thermal softening index 1.05). The calculated results of the interface layer strengthening relationship are then superimposed with the results of the basic Johnson-Cook model to obtain the complete interface layer stress-strain relationship.
[0036] After obtaining the stress-strain relationship, the temperature field-stress field coupling calculation unit is used for separate solution to obtain the stress distribution of the interface layer. The calculation unit includes three main components: heat conduction equation, thermal stress mapping and stress field balance equation. During the calculation, the initial temperature field (normal temperature 25°C) and boundary conditions (such as thermal convection boundary) are first set, and then the temperature distribution of each time step is calculated according to the heat conduction equation. For a typical titanium-steel composite plate, the material density is ρtitanium = 4500kg / m³ and ρsteel = 7800kg / m³, and the specific heat capacity is Titanium = 544 J / (kg·k) and Steel = 460 J / (kg·K), thermal conductivity ktitanium = 21.9 W / (m·K) and ksteel = 50.2 W / (m·K), plastic work to heat conversion coefficient η is 0.9, and interface friction coefficient μ is 0.3. Thermal stress mapping is based on the calculated temperature field and the thermal expansion coefficient of the material (8.6×10 -6 / K, steel is 12×10 -6 / K) to calculate the resulting thermal strain. Finally, the stress distribution at the interface layer is solved using the stress field equilibrium equation, combined with boundary conditions (such as fixed constraints and loading conditions). The computational domain is divided into a 100×50×20 grid, with a time step of 0.01s, and the total computation time is the entire manufacturing process.
[0037] The obtained stress distribution in the interface layer is converted into a multi-resolution tensor field representation. The three-dimensional stress distribution data is reconstructed into a fourth-order tensor, where the first three dimensions correspond to spatial coordinates and the fourth dimension contains six independent stress components. The tensor is then decomposed into different resolution levels using a wavelet transform, typically using a four-level decomposition to capture stress characteristics at the macro, meso, and micro scales. Based on this, three key features are extracted: stress gradient features, stress field singularity distribution features, and temperature-stress covariance features. Stress gradient features are obtained by calculating the stress difference between adjacent grid points, with particular attention paid to stress gradient variations near the interface. Stress field singularity distribution features are identified by detecting sudden changes and discontinuities in the stress field using a local extreme value detection algorithm with a threshold set at 20% of the average stress value. Temperature-stress covariance features are obtained by calculating the covariance matrix between the temperature field and each component of the stress field, reflecting the correlation between temperature changes and stress responses. The extracted features are normalized and dimensionality reduced before being mapped to a 128-dimensional deep learning feature space.
[0038] Deep learning features are input into a long short-term memory network with a physical constraint embedding layer and an adaptive attention mechanism for interface heterogeneity, which outputs predictions of interface binding strength. The network architecture consists of an input layer (128 nodes), a physical constraint embedding layer (64 nodes), two LSTM layers (128 units each), an adaptive attention layer for interface heterogeneity, and an output layer (1 node). The physical constraint embedding layer constructs the first law of thermodynamics and the continuity equation as loss function terms. Specifically, this is achieved by setting violation metrics for energy conservation and mass conservation, and adding corresponding penalty terms when the network predictions violate physical laws. The penalty coefficient is set to 0.2 and fine-tuned on a validation set. The interface heterogeneity-adaptive attention mechanism calculates a spatial attention score based on the query vector and key vector of the interface location. The number of attention heads is set to 8, and the feature dimension is 16. It also combines the cell state and hidden state to capture the temporal dependencies of interface features. In an LSTM network, the cell state stores long-term memory, while the hidden state represents the output signal at the current moment. The cell state retains key information about the entire interface formation process, while the hidden state reflects the interface state at the current moment. This weighted fusion of information from these two states enables the network to simultaneously consider historical accumulation effects and current instantaneous characteristics, thereby accurately identifying the temporal evolution of interface binding strength. The network is trained with a batch size of 32, a learning rate of 0.001, the Adam optimizer, and 200 epochs. Early stopping is used to prevent overfitting; training is terminated if the validation set loss does not decrease for 10 consecutive epochs.
[0039] This technical solution successfully achieved high-precision prediction of the interface bonding strength of titanium-steel composite plates by combining multi-source data acquisition and deep learning methods. It not only took into account the basic principles of materials science and mechanics, but also introduced advanced data processing and deep learning technologies to overcome the limitations of traditional methods that rely on a large amount of experimental and empirical knowledge. The high accuracy of the prediction results can effectively guide the optimization of process parameters, significantly improve production efficiency and product quality, reduce R&D costs, and provide reliable technical support for the design and manufacture of composite materials. It has important theoretical and practical value.
[0040] In an optional embodiment, based on the interface bonding strength prediction result and process parameters, an optimized process parameter combination is generated through a dual-delay deep Q network structure with a priority experience replay mechanism, including: constructing a reinforcement learning state vector using the interface bonding strength prediction value, the process parameter set of the titanium-steel composite material, the process parameter change trend, the interface bonding strength target deviation, and the process stability index; constructing an action space using the adjustment amount of the process parameter set; and constraining the maximum adjustment range of the action space based on the process parameter safety window; constructing a basic reward function using the strength reward item, the stability reward item, and the efficiency reward item, and proportionally adjusting them through a physical constraint correction item to obtain a multi-objective reward function; mapping the state vector and the action space to the feature space through an encoder respectively, and then fusing them into a dual-delay deep Q network having a main network and two target networks with different update cycles, wherein the two target networks update network parameters with different time delays; setting a priority experience replay mechanism based on a temporal difference error, wherein the temporal difference error is used to calculate a priority, and the priority is used to determine a sampling probability; selecting a process parameter adjustment action based on the output value of the dual-delay deep Q network using an ε-greedy strategy, constraining parameters that exceed the process parameter safety window, and outputting an optimized process parameter combination.
[0041] Exemplarily, the reinforcement learning state vector contains five parts: the predicted value of interface bonding strength, the set of process parameters of titanium steel composite materials, the trend of process parameter changes, the target deviation of interface bonding strength and the process stability index. The predicted value of interface bonding strength is provided by the above prediction. The set of process parameters includes parameters such as pressing temperature, pressing pressure, holding time, cooling rate and interface cleanliness, and each parameter is normalized to the range of 0-1. For example, the pressing temperature range is 700-1000℃, which is normalized according to (actual temperature-700) / (1000-700). The trend of process parameter changes is obtained by calculating the rate of change of parameter adjustments for 5 consecutive times, reflecting the dynamic adjustment direction and amplitude of the parameters. The target deviation of interface bonding strength is calculated as the difference ratio between the predicted strength and the target strength (usually set to 350MPa). If the predicted strength is 320MPa, the deviation is (320-350) / 350=-8.57%. Process stability indicators include temperature fluctuation, pressure consistency, and interface diffusion uniformity. Temperature fluctuation is calculated as the ratio of the standard deviation of temperature to the average temperature during the manufacturing process. Pressure consistency is the coefficient of variation of pressure in different areas of the interface. Interface diffusion uniformity is characterized by the variance of the element distribution gradient. The entire state vector has a dimension of 50, and normalization is used to ensure that the value range of each dimension is consistent.
[0042] The action space consists of adjustments to a set of process parameters, including press temperature (±50°C), press pressure (±5 MPa), hold time (±2 minutes), cooling rate (±2°C / s), and interface cleanliness (±10%). Each adjustment is discretized into nine levels, ranging from -4 to +4, corresponding to different percentage adjustments. For example, a press temperature adjustment level of 2 corresponds to a 20°C increase. To ensure process safety, a process parameter safety window is defined to constrain the maximum adjustment range in the action space. This safety window is determined based on material properties and equipment limitations. For example, the press temperature must be no less than 750°C and no more than 950°C, the press pressure must be no less than 10 MPa and no more than 25 MPa, and the hold time must be no less than 5 minutes and no more than 20 minutes. If the parameter adjustment output by the optimization algorithm causes the parameter to exceed the safety window, the system automatically clamps the parameter value within the window boundary.
[0043] A multi-objective reward function is constructed. The state vector and action space are mapped to feature spaces using encoders, and then fused and fed into a dual-delay deep Q-network. The state encoder utilizes a three-layer fully connected neural network architecture, with 50 input nodes (corresponding to the state vector dimension), 128 and 64 hidden layer nodes, respectively, and 32 output nodes. The action encoder is a two-layer fully connected network with an input dimension equal to the dimension of the process parameter adjustment (typically 5), 16 hidden layer nodes, and an output dimension of 8. The encoded features are fused into a 40-dimensional vector through concatenation and fed into a dual-delay deep Q-network. This network consists of a main network and two target networks. The main network is used for current decision making, while the two target networks are used to calculate target Q values. The network architecture is a four-layer fully connected network with 128, 256, and 128 hidden layer nodes, respectively, and an output layer node equal to the size of the action space. The two target networks update their parameters with different time delays: the first target network is updated every 10 training iterations, and the second target network is updated every 20 iterations. This dual-delay architecture effectively reduces the bias in Q-value estimation.
[0044] A prioritized experience replay mechanism is set based on temporal difference error (TDE). An experience replay pool is constructed to store experience tuples generated during reinforcement learning interactions. Each tuple contains the current state, the action performed, the reward received, the next state, and a termination flag. TDE is calculated as the actual reward received plus a discount factor multiplied by the maximum Q-value of the next state, minus the estimated Q-value corresponding to the current state and action. This error reflects the difference between the prediction and the actual reward. The absolute value of the error plus a small constant (such as 0.01) is used as the priority of the experience. Experiences with higher priorities have a greater probability of being sampled. The sampling probability is calculated as the priority raised to the power of α (α is set to 0.6). Importance sampling weights are also introduced to prevent over-learning of high-priority experiences. Each training session samples a batch of 64 experiences from the TDE pool for network updates.
[0045] Based on the output of the double-delayed deep Q-network, an ε-greedy strategy is used to select process parameter adjustment actions. The initial ε value of the strategy is set to 0.9, indicating a 90% probability of random exploration and a 10% probability of selecting the action with the highest Q value. As training progresses, the ε value decreases at a rate of 0.995, reaching a minimum of no less than 0.05 to ensure that the system maintains a certain level of exploration capability. After selecting an action, constraints are applied to parameters that exceed the process parameter safety window. If a parameter exceeds the lower limit of the safety window after adjustment, it is set to the lower limit; if it exceeds the upper limit, it is set to the upper limit. The final output is the optimized process parameter combination.
[0046] This technical solution achieves intelligent optimization of process parameters for titanium-steel composite materials by integrating deep reinforcement learning with knowledge in the field of materials manufacturing. Compared with traditional parameter adjustment methods, this solution can simultaneously consider multiple objectives such as interface bonding strength, process stability, and production efficiency, significantly reducing the cost of manual trial and error and the optimization cycle. In particular, the introduction of the dual-delay Q network structure and the priority experience replay mechanism effectively improves learning efficiency and decision-making accuracy, enabling the system to quickly adapt to different production conditions and material properties. This solution provides an efficient and intelligent solution for optimizing process parameters for titanium-steel composite materials and other composite materials, with significant practical value and promotion prospects.
[0047] In an optional embodiment, the multi-objective reward function includes: the strength reward item is calculated based on the deviation between the interface bonding strength target value and the predicted value, the stability reward item is calculated based on the ratio of the process parameter adjustment amount to the current parameter value, and the efficiency reward item is calculated based on the ratio of the process parameter adjustment amount to the maximum adjustment amplitude; the intergranular reaction layer thickness influence coefficient is calculated based on the deviation relationship between the intergranular reaction layer thickness and the critical thickness and the optimal thickness; the temperature gradient influence coefficient is calculated based on the deviation relationship between the interface temperature gradient and the optimal temperature gradient; the thermal stress influence coefficient is calculated based on the ratio relationship between thermal stress and allowable stress; the constraint weight is determined according to the sensitivity of the influence coefficient to the process parameters, and the physical constraint correction term is constructed by the weighted sum of each influence coefficient and the corresponding constraint weight; the basic reward function is proportionally adjusted by the physical constraint correction term to obtain a multi-objective reward function; according to the characteristics of the process stage, the temperature gradient constraint weight is increased in the heating stage, the intergranular reaction layer thickness constraint weight is increased in the insulation stage, and the thermal stress constraint weight is increased in the cooling stage.
[0048] Exemplarily, the strength reward item is calculated based on the deviation between the target value and the predicted value of the interface bonding strength. In the specific implementation, the target interface bonding strength is set to 350MPa, and the predicted value is provided by the aforementioned interface bonding strength prediction model. When the predicted strength is lower than the target value, the reward value decreases in proportion to the predicted value and the target value; when the predicted strength is equal to the target value, the reward value reaches the maximum (set to 1.0); when the predicted strength exceeds the target value, the reward will also increase appropriately, but the increase is limited to avoid excessive optimization causing material performance redundancy. For example, when the predicted strength is 320MPa, the strength reward is calculated as 1.0-(350-320) / 350=0.914; when the predicted strength is 370MPa, the strength reward is calculated as 1.0+0.2×(370-350) / 350=1.114, where 0.2 is the gain coefficient after exceeding the target to prevent excessive pursuit of high strength and ignoring other goals.
[0049] The stability bonus is calculated based on the ratio of the process parameter adjustment to the current parameter value. This bonus is intended to encourage smooth parameter adjustments and avoid drastic parameter fluctuations that can lead to process instability. The bonus is first calculated for each process parameter as the adjustment percentage, that is, the ratio of the adjustment to the current parameter value. For example, if the current pressing temperature is 830°C and the adjustment is +15°C, the adjustment percentage is 15 / 830 = 1.81%. A weighted average of all parameter adjustment percentages is then calculated, with weights determined based on the parameter's sensitivity to interfacial bonding strength. Typically, temperature and pressure receive higher weights (0.3 each), followed by holding time and cooling rate (0.15 each), and interface cleanliness (0.1). Finally, the weighted average is applied to a decreasing function. When the average adjustment is zero, the stability bonus is maximized (1.0); as the average adjustment increases, the bonus decreases. Specifically, an exponentially decreasing function is used: when the average adjustment is 5%, the stability bonus drops to 0.6, and when it reaches 10%, it drops to 0.25.
[0050] The efficiency bonus is calculated based on the ratio of the process parameter adjustment amount to the maximum adjustment range. This primarily considers the impact of process parameter adjustments on production efficiency, encouraging improvements in production efficiency while ensuring quality. For key parameters influencing production efficiency, such as hold time and cooling rate, the ratio of the adjustment amount to the maximum allowable adjustment range is calculated. For example, if the current hold time is 12 minutes, the adjustment amount is -2 minutes, and the maximum adjustment range is ±5 minutes, the ratio is (-2) / 5 = -0.4. For adjustments that shorten the hold time or increase the cooling rate (in line with improving efficiency), the efficiency bonus increases linearly based on the ratio; for adjustments that extend the hold time or slow the cooling rate, the efficiency bonus decreases linearly. For other parameters, such as temperature and pressure, adjustments that improve efficiency will receive a small bonus, while adjustments that detract from them will receive a small penalty. Ultimately, the efficiency bonuses for each parameter are weighted and aggregated to form the efficiency bonus.
[0051] After the basic reward function is constructed, it is adjusted using a physical constraint correction term. This physical constraint correction term takes into account three key physical factors: intergranular reaction layer thickness, interface temperature gradient, and thermal stress. The intergranular reaction layer thickness influence coefficient is calculated based on the deviation relationship between intergranular reaction layer thickness and the critical and optimal thicknesses. The intergranular reaction layer thickness of titanium-steel composites significantly affects interface strength, and an optimal thickness range exists. Metallurgical analysis has determined that when the reaction layer is too thin (less than 2 μm), the interface bonding is insufficient; when the reaction layer is too thick (over 15 μm), the interface becomes brittle and strength decreases. The optimal thickness range is 5-8 μm. The influence coefficient is calculated using a piecewise function: when the thickness is within the optimal range, the coefficient is 1.0; when the thickness deviates from the optimal range but does not reach the critical value, the coefficient decreases linearly; when the thickness falls below the lower critical limit or exceeds the upper critical limit, the coefficient drops rapidly to below 0.3. For example, when the reaction layer thickness is 10 μm, the influence coefficient is 1.0-(10-8) / (15-8)×0.7=0.84.
[0052] The temperature gradient influence coefficient is calculated based on the deviation relationship between the interface temperature gradient and the optimal temperature gradient. If the temperature gradient of the titanium-steel composite interface is too large, it will lead to thermal stress concentration, while if it is too small, it will be detrimental to interface diffusion and bonding. Through thermodynamic analysis, it was determined that the optimal temperature gradient range is 20-40℃ / mm. The calculation method of the influence coefficient is similar to that of the intergranular reaction layer, using a piecewise function: when the temperature gradient is within the optimal range, the coefficient is 1.0; when it deviates from the optimal range, the coefficient gradually decreases; when the temperature gradient exceeds 80℃ / mm or is lower than 10℃ / mm, the coefficient drops below 0.2. For example, when the temperature gradient is 50℃ / mm, the influence coefficient is 1.0-(50-40) / (80-40)×0.8=0.8.
[0053] The thermal stress influence coefficient is calculated based on the ratio of thermal stress to allowable stress. Excessive thermal stress can cause interface cracking or deformation, affecting the bonding quality. According to material mechanics analysis, the upper limit of allowable thermal stress in the titanium-steel interface area is approximately 200MPa. The influence coefficient is calculated using an exponential decreasing function: when the ratio of thermal stress to allowable thermal stress is 0, the coefficient is 1.0; when the ratio is 0.5, the coefficient is 0.9; when the ratio is 0.8, the coefficient is 0.7; when the ratio is 1.0, the coefficient is 0.5; when the ratio exceeds 1.0, the coefficient drops rapidly to below 0.2. For example, when the thermal stress is 160MPa, the ratio is 160 / 200=0.8, and the influence coefficient is 0.7.
[0054] The constraint weights are determined based on the sensitivity of the influence coefficients to the process parameters, and the weighted sum of each influence coefficient and the corresponding constraint weight is used to construct the physical constraint correction term. Different process parameters have different degrees of influence on the three physical factors, so differentiated constraint weights are set. For example, the pressing temperature has a greater influence on the thickness of the intergranular reaction layer and the temperature gradient, with constraint weights of 0.4 and 0.4 respectively, and has a smaller influence on thermal stress with a weight of 0.2; the pressing pressure has a moderate influence on the thickness of the intergranular reaction layer with a weight of 0.3, a smaller influence on the temperature gradient with a weight of 0.2, and a larger influence on thermal stress with a weight of 0.5. By calculating the weighted physical constraint influence of each process parameter, the physical constraint correction term is obtained.
[0055] The final multi-objective reward function is obtained by scaling the base reward function using the physical constraint modifier. The base reward function is a weighted sum of the strength, stability, and efficiency rewards, with weights of 0.5, 0.3, and 0.2, respectively. The physical constraint modifier acts as a multiplicative factor to adjust the base reward, ensuring that the reward function conforms to physical laws. For example, when the base reward is 0.85 and the physical constraint modifier is 0.9, the final reward is 0.85 × 0.9 = 0.765.
[0056] To adapt to the characteristics of different process stages, the physical constraint weights are dynamically adjusted. During the heating phase, temperature gradients significantly affect interfacial bonding, so the temperature gradient constraint weight is increased to 1.5 times the standard value. During the holding phase, intergranular reaction layer formation is a critical process, so the intergranular reaction layer thickness constraint weight is increased to 1.5 times the standard value. During the cooling phase, thermal stress concentration is the primary risk, so the thermal stress constraint weight is increased to 1.5 times the standard value. This dynamic weight adjustment mechanism enables the reward function to accurately reflect the key objectives of different process stages, improving optimization effectiveness.
[0057] For example, during the insulation phase of a titanium-steel composite manufacturing process, the intergranular reaction layer thickness is 4 μm, the temperature gradient is 45°C / mm, and the thermal stress is 120 MPa. The calculated influence coefficient for the intergranular reaction layer thickness is 0.95, the temperature gradient influence coefficient is 0.9, and the thermal stress influence coefficient is 0.8. Considering the insulation phase, the intergranular reaction layer thickness constraint weight is increased from the standard value of 0.4 to 0.6, while the temperature gradient and thermal stress constraint weights remain at the standard values. The final calculated physical constraint correction term is 0.6 × 0.95 + 0.25 × 0.9 + 0.15 × 0.8 = 0.915. If the base reward value at this point is 0.82, the final reward value is 0.82 × 0.915 = 0.75.
[0058] This multi-objective reward function construction method achieves a scientific evaluation of the process parameter optimization process of titanium-steel composite materials by comprehensively considering the three major objectives of interface bonding strength, process stability and production efficiency, and dynamically adjusting them in combination with key physical constraints. The deep integration of materials science and reinforcement learning can not only guide the system to find the optimal process parameter combination, but also ensure that the optimization results conform to physical laws and process requirements. Compared with traditional single-objective optimization methods, this method improves material performance while taking into account production efficiency and process stability, greatly reducing the trial and error costs in the optimization process.
[0059] In an optional embodiment, the optimized process parameter combination is used as the initial population, and the phase change temperature point is controlled in stages and the theory-differential gradient fusion method is adopted to generate the optimal process parameter solution based on the performance index and process stability. The method includes: using the optimized process parameter combination as the initial population, obtaining the phase change temperature point of the titanium steel composite material, and calculating the current optimization stage according to the population distribution. The optimization stage includes the exploration stage, the transition stage and the refinement stage. The corresponding gradient step parameters and mutation probability parameters are set for different optimization stages. When the process parameter value is close to the phase change temperature point, the disturbance amplitude of the crossover operation is reduced and the local search accuracy is improved. In the non-dominated sorting process, the process parameter is calculated according to the phase change temperature point and the optimization stage. The theoretical gradient value of the target is confidence-weighted with the differential gradient value to obtain a fusion gradient value; the performance indicators include interface bonding strength, tensile strength and elongation; based on the population evolution process, the titanium steel interface diffusion sufficiency, the rationality of the interface layer thickness and the brittle phase content are calculated respectively, the degree of constraint violation is determined, and the degree of influence of process parameter fluctuations on performance indicators is calculated; in each generation of population evaluation, the actual fluctuation value of the process parameters is obtained, and the process stability score is calculated in combination with the degree of influence, and the process cost score is calculated based on the energy consumption value, processing time and equipment operating parameters; the process stability score, process cost score, fusion gradient value and constraint violation degree are weighted to obtain a comprehensive score, and the process parameter combination with the highest comprehensive score is selected as the optimization result.
[0060] Exemplarily, the process parameter combination optimized by reinforcement learning is used as the initial population, which contains multiple groups of process parameter combinations, each group including key parameters such as pressing temperature, pressing pressure, holding time, cooling rate and interface cleanliness. In order to ensure population diversity, 20 groups of parameter variants are generated around the benchmark combination. The variant generation rule is to add perturbations according to the Gaussian distribution on the basis of the benchmark parameters, and the perturbation amplitude is ±5% of the parameter value. For example, if the benchmark pressing temperature is 865°C, the temperature range of the variant is approximately distributed between 822-908°C. After the initial population is constructed, the phase transition temperature points of the titanium-steel composite material are obtained through the material database, including the α→β phase transition temperature of the titanium alloy (about 882°C), the austenite transformation temperature of the steel (about 723°C) and the activation temperature of the interface diffusion reaction (about 750°C). These temperature points serve as an important reference in the optimization process.
[0061] The optimization process is divided into three stages: exploration, transition, and refinement. The current stage is determined by calculating the degree of dispersion of the population parameters. The specific method is to calculate the coefficient of variation (the ratio of the standard deviation to the mean) of each process parameter in the population. When the coefficient of variation is greater than 0.15, it is determined to be in the exploration stage; when the coefficient of variation is between 0.08-0.15, it is determined to be in the transition stage; when the coefficient of variation is less than 0.08, it is determined to be in the refinement stage. In the exploration stage, a larger gradient step size parameter (0.1) and a higher mutation probability (0.3) are set to encourage the algorithm to search within a larger range; in the transition stage, the gradient step size parameter is reduced to 0.05 and the mutation probability is reduced to 0.15; in the refinement stage, the gradient step size parameter (0.02) and the mutation probability (0.05) are further reduced to concentrate resources on a refined search near the optimal solution.
[0062] In particular, when the process parameter values are close to the phase transition temperature point, more cautious parameter adjustment is required. For example, when the pressing temperature approaches 882°C (within the range of ±10°C), the titanium alloy undergoes an α→β phase transition, and the interface reaction mechanism changes significantly. To prevent the algorithm from crossing this sensitive area and missing the optimal solution, the perturbation amplitude of the crossover operation is reduced to half of the standard value, and the local search accuracy is improved. When the parameter value enters the range of ±1.5% of the phase transition temperature point, the parameter perturbation coefficient of the crossover operator is reduced from the standard value of 0.2 to 0.1, and 10 additional sampling points are added in this area for fine evaluation.
[0063] During the non-dominated sorting process, the theoretical gradient values of the process parameters on the performance indicators are calculated based on the phase transition temperature and the optimization stage. These values are then weighted with the differential gradient values to obtain a fused gradient value. First, based on materials science theory, the theoretical gradient values of each process parameter on the performance indicator are calculated, including the specific variation patterns of the theoretical gradient near the phase transition temperature. Simultaneously, individuals with high scores in the population are selected and their differential gradient values are calculated by applying small perturbations. Weightings are adjusted based on the optimization stage. For regions near the phase transition temperature, the confidence level of the theoretical gradient is reduced, and the weight of the local differential gradient is increased to improve adaptability. Performance evaluation includes three aspects: interface bonding strength, tensile strength, and elongation. The target value for interface bonding strength is set at 350 MPa, obtained using the previously described interface bonding strength prediction model; the target value for tensile strength is set at 620 MPa; and the target value for elongation is set at 15%. The weights of these three indicators are determined based on application requirements. Typically, interface bonding strength has the highest weight (0.5), followed by tensile strength (0.3), and elongation has the lowest weight (0.2).
[0064] Based on the population evolution process, the process constraint indicators are calculated, and the diffusion sufficiency of the titanium-steel interface is calculated. This indicator reflects the integrity of the diffusion of interface elements. The ratio of the element penetration depth to the ideal depth is estimated by the diffusion model, and the complete sufficiency is 1.0. For example, when the pressure is maintained at 865°C for 15 minutes, the diffusion depth of the interface titanium element to the steel side is about 5μm, and the ideal diffusion depth is 6μm, so the sufficiency is 0.83. Secondly, the rationality of the interface layer thickness is calculated. According to the aforementioned optimal thickness range of the interface layer (5-8μm), the degree of conformity between the actual thickness and the optimal range is evaluated. The brittle phase content is calculated again, focusing on the formation of brittle intermetallic compounds such as TiFe2. The lower the content, the better. Finally, the constraint violation is determined based on the degree of deviation of the three indicators, and the penalty function method is used to convert the constraint violation into a fitness penalty.
[0065] Calculate the impact of process parameter fluctuations on performance indicators. In actual production, process parameters are difficult to precisely control and fluctuate to a certain extent. Through sensitivity analysis, evaluate the impact of each parameter fluctuation on performance. For example, a pressing temperature fluctuation of ±5°C affects the interfacial bonding strength by approximately ±3 MPa, while a pressing pressure fluctuation of ±1 MPa affects approximately ±2 MPa. Based on the sensitivity analysis results, establish a sensitivity matrix for process parameters and performance indicators to quantify the impact weight of different parameter fluctuations.
[0066] In each generation of population evaluation, the actual fluctuation values of the process parameters are obtained, and the process stability score is calculated based on the degree of influence. The actual fluctuation values can be obtained through industrial data collection or simulation estimation. For example, the temperature control fluctuation of a typical production line is about ±3°C, and the pressure control fluctuation is about ±0.5MPa. The process stability score is calculated as a negative exponential function of the weighted sum of the fluctuation effects of each parameter. The score range is 0-1, and the higher the more stable. At the same time, the process cost score is calculated based on energy consumption values, processing time and equipment operating parameters. Energy consumption mainly considers heating energy consumption (related to temperature and time) and pressing energy consumption (related to pressure and time); processing time includes heating time, holding time and cooling time; equipment operating parameters include equipment load rate and maintenance cost factor. The process cost score is also normalized to the range of 0-1, and the higher it is, the lower the cost.
[0067] Finally, the process stability score, process cost score, fusion gradient value, and constraint violation degree are weighted together to produce a comprehensive score. The weighting coefficients are set based on actual production requirements; typically, performance indicators are weighted at 0.5, process stability at 0.3, and process cost at 0.2. The constraint violation degree is deducted from the total score as a penalty. The process parameter combination with the highest comprehensive score is selected as the final optimization result.
[0068] Actual performance indicators of titanium-steel composite plates are regularly collected, and deviations from the model's predicted values are calculated. This deviation data is used to update the deep Q network's experience replay pool, adjust the reward function weights, and revise the state transition model. Furthermore, this deviation data is used to adjust the genetic algorithm's fitness function, update the theoretical-to-differential weight ratio in gradient fusion, and calibrate the prediction model near the phase transition temperature. For example, a global parameter update is performed every 50 batches of material produced, and local fine-tuning is performed every 10 batches to ensure continuous alignment between the prediction model and the actual process.
[0069] For example, the initial process parameters of a titanium-steel composite plate production line were: pressing temperature 865°C, pressing pressure 20 MPa, holding time 15 minutes, cooling rate 4°C / s, interface cleanliness 90%, and interface bonding strength 348 MPa. After optimization using this method, the optimal process parameters were adjusted to: pressing temperature 878°C (close to but not exceeding the phase transition temperature of titanium alloy), pressing pressure 21 MPa, holding time 16 minutes, cooling rate 4.5°C / s, and interface cleanliness 92%. After optimization, the interface bonding strength increased to 356 MPa, the tensile strength reached 628 MPa, and the elongation was 16.2%. All performance indicators met or exceeded the target values. At the same time, the process stability score was 0.91 (a significant improvement compared to the initial 0.85), and the process cost score was 0.88 (slightly lower than the initial 0.9, indicating a slight cost increase but within an acceptable range). The overall score was 0.925, significantly higher than the initial solution's 0.87.
[0070] This technical solution achieves precise optimization of process parameters for titanium-steel composite materials by combining evolutionary algorithms with materials science theory. The innovation of this method lies in the introduction of phase-by-phase control of phase transition temperature points and a theoretical-differential gradient fusion mechanism, enabling the optimization process to fully consider material properties and phase transition patterns. Multi-objective balanced optimization is achieved through a comprehensive evaluation of performance indicators, process stability, and process costs. Compared to traditional methods that rely solely on experience or a single algorithm, this solution significantly improves optimization accuracy and efficiency, reduces trial-and-error costs, and ensures the engineering feasibility of the optimization results, providing a systematic and intelligent solution for process optimization of titanium-steel composite materials and other similar composite materials.
[0071] In an optional embodiment, the theoretical gradient value and the differential gradient value are confidence-weighted to obtain a fusion gradient value, including: calculating the rate of change of element concentration with temperature and time during interface diffusion, the temperature sensitivity coefficient of the intermetallic compound phase fraction near the phase transition temperature point, and the contribution coefficient of the interface microstructure characteristic parameters to the bonding strength to obtain the corresponding theoretical gradient; selecting the top N parameter combinations with the highest scores in the current population, applying a perturbation amount, and calculating the ratio of the performance index change value to the perturbation amount to obtain a differential gradient value; determining the empirical support based on the similarity between the process parameters and the historical data, reducing the credibility of the theoretical prediction near the phase transition temperature point, and determining the confidence interval based on the historical prediction error; increasing the differential gradient weight in the exploration stage to expand the search range, increasing the theoretical gradient weight in the refinement stage, implementing differentiated fusion of the performance indicators near the phase transition temperature point, and weighting the theoretical gradient and the differential gradient according to the confidence level to obtain a fusion gradient.
[0072] For example, combined Figure 2 The fusion gradient value calculation flow chart is used for explanation: For the interface diffusion process, the modified Fick's second law is used to describe the concentration distribution of elements at the interface. The relationship between the diffusion coefficient of titanium in the steel matrix and temperature follows the Arrhenius formula. In the typical process temperature range (800-900℃), the diffusion coefficient of titanium in steel is approximately 1×10-12 to 5×10-11 m² / s. Based on this, the rate of change of titanium concentration with temperature is calculated. At 850℃, the rate of increase of the titanium concentration gradient at the interface is approximately 0.033μm for every 10℃ increase in temperature. -1 The corresponding diffusion depth increases by about 0.5 μm / 15 min. Combined with the quantitative relationship between diffusion depth and bonding strength (each 1 μm increase in diffusion depth increases bonding strength by 10-15 MPa until saturation), the theoretical gradient of temperature on strength is obtained.
[0073] The temperature sensitivity coefficient of the intermetallic compound phase fraction near the phase transition temperature point was focused on changes near the α→β phase transition temperature (approximately 882°C) of titanium alloy. Through thermodynamic calculations and phase diagram analysis, the temperature sensitivity coefficients of each phase were quantitatively determined: in the range of 870-880°C, the temperature sensitivity coefficient of the TiFe phase fraction was 0.2% / °C, and in the range of 880-890°C, the temperature sensitivity coefficient of the TiFe2 phase fraction was 0.3% / °C.
[0074] The contribution coefficients of the interface microstructure characteristic parameters to the bonding strength were quantitatively determined through microstructural analysis and mechanical testing. In titanium-steel composites, the interface layer is composed of a solid solution layer, a diffusion layer, and an intermetallic compound layer. Scanning electron microscopy and transmission electron microscopy analysis, combined with nanoindentation testing, quantified the contribution coefficients of each parameter: for interface layer thicknesses in the 5-8μm range, the contribution coefficient was 25MPa / μm; for TiFe phase fractions in the 30-40% range, the contribution coefficient was 3MPa / %; and for TiFe2 phase fractions exceeding 15%, the contribution coefficient was -5MPa / %. Based on these quantitative relationships, the theoretical gradient values of each process parameter on the performance indicators were calculated.
[0075] Theoretical gradients were calculated for the different characteristics of the process parameters. The theoretical gradient of pressing temperature on interface strength is approximately 0.8 MPa / °C in the 800-870°C range, increasing to 1.2 MPa / °C in the 870-882°C range, and dropping to 0.5 MPa / °C or even becoming negative above 882°C. The theoretical gradient of pressing pressure on interface strength is approximately 1.5 MPa / MPa in the 15-25 MPa range, and decreases after exceeding 25 MPa. The theoretical gradient of holding time on interface strength exhibits a nonlinear relationship, with a large gradient (approximately 3 MPa / minute) in the initial stage (5-10 minutes), a decreasing gradient (approximately 1 MPa / minute) in the middle stage (10-20 minutes), and approaching zero in the later stages (>20 minutes). The theoretical gradient of cooling rate on interface strength is approximately -0.5 MPa / (°C / s) in the 1-5°C / s range, indicating that appropriately reducing the cooling rate can help reduce thermal stress and improve bonding strength.
[0076] Select the top N parameter combinations with the highest scores in the current population, apply a perturbation, and calculate the ratio of the performance index change value to the perturbation value to obtain the differential gradient value. The N value is usually set to 20% of the population size (for example, if the population size is 50, then N=10). For each selected parameter combination, a small perturbation is applied to a single process parameter, keeping other parameters unchanged, and then the performance change is calculated using the aforementioned interface strength prediction model. The perturbation amount is set to ±1% of the parameter standard range. For example, for a pressing temperature of 865°C, the perturbation amount is ±8.65°C. To reduce the influence of random errors, each parameter is subjected to 3 positive perturbations and 3 negative perturbations, and the average effect is taken. The differential gradient value is calculated as the ratio of the performance index change value to the perturbation amount. For example, if a temperature increase of 8.65°C results in an increase of 7.2MPa in interface strength, the differential gradient value is 7.2 / 8.65=0.83MPa / °C.
[0077] For multidimensional parameter spaces, an orthogonal design method is used to reduce the number of experiments. Each of the N high-scoring combinations is regarded as a center point, and a partial factor orthogonal experiment is constructed around it. Each process parameter only considers two levels (center point value ± disturbance amount). The main effects and interaction effects of each parameter are calculated based on the orthogonal experiment results to obtain a more accurate differential gradient value matrix. In addition, to improve computational efficiency, the gradient values of similar parameter areas are clustered and analyzed. Areas with a similarity of more than 95% share the same gradient value.
[0078] Based on the similarity between process parameters and historical data, empirical support is determined to reduce the confidence of theoretical predictions near the phase transition temperature. The confidence interval is then determined based on historical prediction errors. A historical database of process parameters and performance indicators is first constructed, containing valid records from past experiments and production. For the process parameter combination to be evaluated, the Euclidean distance to the historical data point is calculated. The K closest historical data points (K is typically set to 5) are selected and the mean and standard deviation of the prediction errors are calculated. Empirical support is calculated as a negative exponential function of the distance; smaller distances indicate higher support. For example, when the Euclidean distance is 0.1 in the standardized parameter space, the empirical support is 0.9; at a distance of 0.3, the support drops to 0.7. In particular, the confidence of theoretical predictions in the parameter region near the phase transition temperature requires adjustment. A confidence adjustment function is developed by analyzing the prediction accuracy near the phase transition temperature in historical data. For example, for the titanium alloy α→β phase transition temperature (882°C), the average error of the theoretical prediction increases by approximately 50% within a ±5°C range. Therefore, the confidence of the theoretical gradient is reduced to 0.6 times the standard region. Similarly, similar confidence adjustments are applied near the austenite transformation temperature (723°C) and interface diffusion activation temperature (750°C) of steel. Combining historical prediction errors and phase transformation region characteristics, the confidence interval of the theoretical gradient value is determined, usually expressed as ±2 times the standard deviation.
[0079] A differentiated gradient fusion strategy is employed at different stages of optimization. During the exploration phase, the differential gradient weight is set to 0.7, while the theoretical gradient weight is 0.3. During the transition phase, both weights are 0.5. During the refinement phase, the differential gradient weight is reduced to 0.3, while the theoretical gradient weight is increased to 0.7. This dynamic weight allocation strategy leverages the advantages of differential gradients in exploring new areas and the accuracy of theoretical gradients in refined optimization.
[0080] Differentiated fusion is implemented for performance indicators near the phase transition temperature point. Within the ±3% range of the phase transition temperature point, a more conservative fusion strategy is adopted, increasing the detection point density and reducing the step size. For example, near the α→β phase transition temperature of titanium alloy (882°C), the temperature step size is reduced from the standard 5°C to 2°C, and the weight of the local differential gradient is increased in the gradient fusion in this area to improve adaptability. When the differential gradient is consistent with the theoretical gradient direction, a larger learning step size is used; when the two gradient directions are inconsistent, a smaller step size is used and the detection point density is increased to avoid missing the optimal solution.
[0081] The final calculation formula for the fused gradient is the weighted sum of the theoretical gradient and the differential gradient: fused gradient = W_theory × theoretical gradient × credibility + W_difference × differential gradient, where W_theory and W_difference are the weights of the theoretical gradient and differential gradient, respectively, satisfying W_theory + W_difference = 1. Credibility is a confidence factor for the theoretical gradient adjusted based on empirical support and phase transition region.
[0082] This technical solution achieves intelligent fusion of gradient information by combining theoretical models with actual data, especially providing precise optimization guidance for complex material behaviors near phase transition temperature points; it effectively overcomes the limitations of single gradient methods and improves optimization efficiency and accuracy; the theoretical gradient provides global guidance based on material science theory, while the differential gradient captures local behaviors in actual systems that may be ignored by theory. The combination of the two achieves a balance between exploration and refinement in the optimization algorithm; in particular, the treatment of special areas such as near phase transition temperature points reflects the method's in-depth consideration of material properties, making the optimization results more in line with actual process requirements.
[0083] A second aspect of an embodiment of the present invention provides a titanium steel composite plate manufacturing optimization system based on a neural network model, comprising:
[0084] The first unit is used to collect interface data during the manufacturing process of titanium-steel composite plates and generate interface bonding features through time-space dimension feature extraction;
[0085] The second unit is used to input the interface bonding characteristics into an adaptive memory network, and predict the interface bonding strength based on the material constitutive equation library and the temperature field-stress field coupling calculation unit;
[0086] The third unit is used to generate an optimized process parameter combination based on the interface bonding strength prediction results and process parameters through a double-delay deep Q network structure with a priority experience replay mechanism;
[0087] The fourth unit is used to use the optimized process parameter combination as the initial population, adopt the phase change temperature point stage control and theoretical-differential gradient fusion method to generate the optimal process parameter solution based on performance indicators and process stability;
[0088] The fifth unit is used to feed back the deviation between the actual measured performance indicators and the predicted results to the deep Q network and genetic algorithm to dynamically adjust the process parameter optimization strategy.
[0089] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0090] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0091] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing the manufacturing of titanium-steel composite plates based on a neural network model, characterized in that: include: Collect interface data during the manufacturing process of titanium-steel composite plates, and generate interface bonding features through spatiotemporal feature extraction; The interface bonding characteristics are input into an adaptive memory network, and the interface bonding strength is predicted based on a material constitutive equation library and a temperature field-stress field coupling calculation unit; Based on the interface bonding strength prediction results and process parameters, the optimized process parameter combination is generated through a double-delay deep Q network structure with a priority experience replay mechanism; The optimized process parameter combination is used as the initial population, and the phase change temperature point is controlled in stages and the theoretical-differential gradient fusion method is adopted to generate the optimal process parameter solution based on performance indicators and process stability, including: using the optimized process parameter combination as the initial population, obtaining the phase change temperature point of the titanium steel composite material, and calculating the current optimization stage according to the population distribution. The optimization stage includes the exploration stage, the transition stage and the refinement stage. The corresponding gradient step parameters and mutation probability parameters are set for different optimization stages; when the process parameter value is close to the phase change temperature point, the disturbance amplitude of the crossover operation is reduced and the local search accuracy is improved; in the non-dominated sorting process, according to the phase change temperature point and the optimization stage, the theoretical gradient value of the process parameter to the performance indicator is calculated, and the confidence weighted difference gradient value is obtained to obtain the fusion gradient value. ; The performance indicators include interface bonding strength, tensile strength and elongation; based on the population evolution process, the titanium steel interface diffusion sufficiency, interface layer thickness rationality and brittle phase content are calculated respectively, the degree of constraint violation is determined, and the influence of process parameter fluctuations on performance indicators is calculated; in each generation of population evaluation, the actual fluctuation value of the process parameters is obtained, and the process stability score is calculated in combination with the influence degree, and the process cost score is calculated based on the energy consumption value, processing time and equipment operating parameters; the process stability score, process cost score, fusion gradient value and constraint violation degree are weightedly calculated to obtain a comprehensive score, and the process parameter combination with the highest comprehensive score is selected as the optimization result; the deviation between the actual measured performance indicators and the predicted results is fed back to the deep Q network and genetic algorithm to dynamically adjust the process parameter optimization strategy.
2. The method according to claim 1, characterized in that The interface bonding characteristics are input into an adaptive memory network, and based on a material constitutive equation library and a temperature field-stress field coupling calculation unit, the interface bonding strength is predicted, including: inputting the interface bonding characteristics into the material constitutive equation library, the material constitutive equation library introduces an interface layer strengthening relationship based on the Johnson-Cook model, and calculates the stress-strain relationship of the interface bonding layer; based on the stress-strain relationship, a temperature field-stress field coupling calculation unit is used to perform a separate solution to obtain the interface layer stress distribution; the interface layer stress distribution is converted into a multi-resolution tensor field representation, and stress gradient characteristics, stress field singular point distribution characteristics and temperature-stress covariance characteristics are extracted and mapped to a deep learning feature space; the deep learning features are input into a long short-term memory network with a physical constraint embedding layer and an interface heterogeneity adaptive attention mechanism, and the interface bonding strength prediction result is output.
3. The method according to claim 2, characterized in that The method further includes: the interface layer strengthening relationship is: ,in is the interface layer stress, is the matrix stress, is the grain size, is the dislocation density, and is the material constant, which is used to calculate the stress-strain relationship of the interface bonding layer; The coupling calculation unit includes: heat conduction equation: in, is the material density, is the specific heat capacity; is the partial derivative of temperature with respect to time, which indicates the rate of change of temperature with time; is the temperature, For time, is the thermal conductivity, is the plastic work-to-heat conversion coefficient, is stress, is the strain rate, is the friction coefficient, is the interface pressure, Relative slip velocity, is the gradient operator, is the divergence operator; thermal stress mapping: ;in, is the thermal strain, is the coefficient of thermal expansion, is the temperature, is the initial temperature; the stress field equilibrium equation is: ;in, is stress, is the body force, is the divergence operator; The physical constraint embedding layer embeds the first law of thermodynamics and the continuity equation into the network structure as loss function terms. The interface heterogeneity adaptive attention mechanism calculates the spatial attention score based on the query vector and key vector of the interface position, and combines the unit state and hidden state to capture the temporal dependency of the interface features, and outputs the interface binding strength prediction result.
4. The method according to claim 1, wherein Based on the interface bonding strength prediction results and process parameters, an optimized process parameter combination is generated using a dual-delayed deep Q-network structure with a prioritized experience replay mechanism. The method includes: constructing a reinforcement learning state vector using the interface bonding strength prediction value, the process parameter set of the titanium-steel composite material, the process parameter variation trend, the target deviation of the interface bonding strength, and the process stability index; constructing an action space using the adjustment amount of the process parameter set; and constraining the maximum adjustment range of the action space based on the process parameter safety window; constructing a basic reward function using the strength reward term, the stability reward term, and the efficiency reward term, and proportionally adjusting it using a physical constraint correction term to obtain a multi-objective reward function; mapping the state vector and the action space to the feature space via an encoder, and then fusing them into a dual-delayed deep Q-network with a main network and two target networks with different update periods, wherein the two target networks update network parameters with different time delays; setting a prioritized experience replay mechanism based on the temporal difference error, wherein the temporal difference error is used to calculate the priority, and the priority is used to determine the sampling probability; and selecting the process parameter adjustment action based on the output value of the dual-delayed deep Q-network using an ε-greedy strategy, constraining parameters that exceed the process parameter safety window, and outputting the optimized process parameter combination.
5. The method according to claim 4, characterized in that The multi-objective reward function includes: the strength reward item is calculated based on the deviation between the target value and the predicted value of the interface bonding strength, the stability reward item is calculated based on the ratio of the process parameter adjustment amount to the current parameter value, and the efficiency reward item is calculated based on the ratio of the process parameter adjustment amount to the maximum adjustment amplitude; the intergranular reaction layer thickness influence coefficient is calculated based on the deviation relationship between the intergranular reaction layer thickness and the critical thickness and the optimal thickness; the temperature gradient influence coefficient is calculated based on the deviation relationship between the interface temperature gradient and the optimal temperature gradient; the thermal stress influence coefficient is calculated based on the ratio relationship between thermal stress and allowable stress; the constraint weight is determined according to the sensitivity of the influence coefficient to the process parameters, and the weighted sum of each influence coefficient and the corresponding constraint weight is used to construct a physical constraint correction item; the basic reward function is proportionally adjusted by the physical constraint correction item to obtain a multi-objective reward function; according to the characteristics of the process stage, the temperature gradient constraint weight is increased in the heating stage, the intergranular reaction layer thickness constraint weight is increased in the insulation stage, and the thermal stress constraint weight is increased in the cooling stage.
6. The method according to claim 1, characterized in that The confidence-weighted fusion gradient value is obtained by weighting the theoretical gradient value and the differential gradient value, including: calculating the rate of change of element concentration with temperature and time during interface diffusion, the temperature sensitivity coefficient of the intermetallic compound phase fraction near the phase transition temperature point, and the contribution coefficient of the interface microstructure characteristic parameters to the bonding strength to obtain the corresponding theoretical gradient; selecting the top N parameter combinations with the highest scores in the current population, applying a perturbation amount, and calculating the ratio of the performance index change value to the perturbation amount to obtain the differential gradient value; determining the empirical support based on the similarity between the process parameters and the historical data, reducing the credibility of the theoretical prediction near the phase transition temperature point, and determining the confidence interval based on the historical prediction error; increasing the differential gradient weight in the exploration stage to expand the search range, increasing the theoretical gradient weight in the refinement stage, implementing differentiated fusion of the performance indicators near the phase transition temperature point, and weighting the theoretical gradient and the differential gradient according to the confidence level to obtain the fusion gradient.
7. A titanium steel composite plate manufacturing optimization system based on a neural network model, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to collect interface data during the manufacturing process of titanium-steel composite plates and generate interface bonding features through time-space dimension feature extraction; The second unit is used to input the interface bonding characteristics into an adaptive memory network, and predict the interface bonding strength based on the material constitutive equation library and the temperature field-stress field coupling calculation unit; The third unit is used to generate an optimized process parameter combination based on the interface bonding strength prediction results and process parameters through a double-delay deep Q network structure with a priority experience replay mechanism; The fourth unit is used to use the optimized process parameter combination as the initial population, adopt the phase change temperature point staged control and theoretical-differential gradient fusion method to generate the optimal process parameter solution based on performance indicators and process stability; The fifth unit is used to feed back the deviation between the actual measured performance indicators and the predicted results to the deep Q network and genetic algorithm to dynamically adjust the process parameter optimization strategy.
8. An electronic device, characterized in that: include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Design objective-based whole-process technological parameter integrated intelligent optimization method
CN118195820A
Distributed multi-unmanned aerial vehicle intelligent dash decision-making method for three-dimensional urban environment
CN119105523A