Gear machining parameter adaptive optimization system based on deep reinforcement learning

The gear machining parameter adaptive optimization system, which utilizes deep reinforcement learning, adjusts machining parameters in real time, resolving the conflict between gear surface integrity and blade aerodynamic performance in integrated blade disk-gear machining, thereby improving the safety and reliability of the machining process.

CN120909133AActive Publication Date: 2025-11-07NANTONG XINSIDI ELECTROMECHANICAL CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511438052.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously optimize gear surface integrity and blade aerodynamic performance in integrated bladed disk-gear machining, leading to contradictions during the machining process and making it impossible to meet the needs of both.

Method used

An adaptive optimization system for gear machining parameters based on deep reinforcement learning is adopted. Through closed-loop control of state perception, virtual twin prediction, decision reward calculation and machining parameter execution module, machining parameters are adjusted in real time. Combined with multi-source heterogeneous data and reduced-order surrogate model, the system can achieve forward-looking prediction and optimization of the machining process.

Benefits of technology

This technology enables real-time optimization of machining parameters under dynamic disturbances, synergistically improving gear surface integrity and blade aerodynamic performance. It resolves the technical contradictions that traditional methods cannot address simultaneously, and enhances the safety and reliability of the machining process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909133A_ABST
    Figure CN120909133A_ABST
Patent Text Reader

Abstract

The invention discloses a gear machining parameter adaptive optimization system based on deep reinforcement learning, and belongs to the technical field of mechanical precision machining and artificial intelligence crossing, and the system comprises a state sensing module which is used for collecting multi-source heterogeneous data in the machining process in real time, and constructing a state vector of the current state of the system; the virtual twinborn prediction module is used for receiving the state vector and the candidate processing parameter action and predicting three-dimensional physical field distribution in a future time step; and the decision return calculation module is used for calculating a scalar return value used for guiding decision according to the three-dimensional physical field distribution.The technical contradiction that in the prior art, when an integrated component is machined, the two requirements of strengthening the gear surface integrity and maintaining the blade aerodynamic performance conflict with each other and cannot be taken into account is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of mechanical precision machining and artificial intelligence, in particular to a gear machining parameter adaptive optimization system based on deep reinforcement learning. BACKGROUND

[0002] In the field of aero-engine, blisk-gear is a structure integrating blades, disk and gear in one. This structure has significant advantages in improving thrust-weight ratio and reducing structural weight, but also brings unprecedented manufacturing challenges. The integrated structure leads to the contradiction between the performance requirements of different functional areas. In order to ensure high fatigue life, the gear area requires the machined surface to have a specific residual compressive stress distribution, low roughness and high microhardness, i.e. excellent surface integrity. In order to ensure aerodynamic efficiency, the blade area requires the deformation generated during machining to be minimized to maintain the accuracy of the aerodynamic performance retention.

[0003] The traditional machining method usually adopts a fixed parameter strategy based on expert experience or offline simulation, and the optimization goal is the optimal quality of a single process or a single target. In finish machining of gears, in order to obtain ideal surface integrity, cutting parameter combinations that can introduce larger residual compressive stress are often used. Such parameters often produce huge cutting force and cutting heat, which are conducted to the blade part which is relatively weak in rigidity through the integrated structure, inevitably causing micron-level permanent plastic deformation of the blade, damaging its profile tolerance, and further leading to the decline of aerodynamic performance. This phenomenon reveals a deep technical contradiction: in the machining of integrated components, any machining strategy aimed at strengthening the performance of the gear surface may sacrifice the aerodynamic performance of the blade; especially under the influence of dynamic variables such as tool wear, the contradiction becomes more acute. The conventional method of increasing cutting load to compensate for wear will directly exacerbate the damage to the blade. The existing technology cannot actively avoid this problem, but can only passively compromise between the two conflicting goals of gear life and blade efficiency.

[0004] The technical problem to be solved by the present application is to provide a machining parameter adaptive optimization system that can overcome the above technical contradiction in the integrated machining of blisk-gear.

[0005] The above information disclosed in the background section is only used to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0007] The purpose of the present application is to provide a gear machining parameter adaptive optimization system based on deep reinforcement learning to solve the problems raised in the background.

[0008] The technical scheme of the present application is as follows:

[0009] A state perception module is configured to collect multi-source heterogeneous data in real time during the machining process and construct a state vector of the current state of the system.

[0010] A virtual twin prediction module is configured to receive the state vector and candidate machining parameter actions and predict a three-dimensional physical field distribution at a future time step.

[0011] A decision reward calculation module is configured to calculate a scalar reward value for guiding decision-making based on the three-dimensional physical field distribution.

[0012] A deep reinforcement learning decision module is configured to take the state vector as input and take the scalar reward value as a learning goal, and output an optimal machining parameter action.

[0013] A machining parameter execution module is configured to receive the optimal machining parameter action and convert it into a numerical control instruction to drive the machine tool to perform machining.

[0014] Preferably, the multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

[0015] Preferably, the core of the virtual twin prediction module is a pre-set reduced-order proxy model; and the three-dimensional physical field distribution includes a three-dimensional stress field, a three-dimensional strain field, and a three-dimensional temperature field.

[0016] Preferably, the decision reward calculation module is configured to first calculate a dimensionless stress field health index based on the three-dimensional stress field, in combination with pre-set weight hyperparameters, workpiece geometric volume, and material yield strength.

[0017] Preferably, the calculation of the stress field health index includes: quantifying the residual stress gain effect of the gear target area and the von Mises equivalent stress damage effect of the blade key area, respectively; and performing weighted summation on the gain effect and the damage effect based on the weight hyperparameters.

[0018] Preferably, the decision reward calculation module is further configured to: based on the stress field health index, introduce a hard constraint penalty term for plastic deformation, and construct a final scalar reward value.

[0019] Preferably, the introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the blade key area predicted by the virtual twin prediction module with a pre-set blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a pre-set penalty factor is applied, otherwise no penalty is applied.

[0020] Preferably, the deep reinforcement learning decision module has a deep Q network intelligent agent built-in, and the intelligent agent learns through continuous interaction with the machining environment.

[0021] Preferably, the optimal machining parameter actions include spindle speed, feed rate and depth of cut.

[0022] The present application provides a gear machining parameter adaptive optimization system based on deep reinforcement learning by improving, compared with the prior art, has the following improvements and advantages:

[0023] 1. The technical solution of the present application constructs a closed-loop control architecture integrating state perception, virtual twin prediction, decision reward calculation, deep reinforcement learning decision and machining parameter execution, successfully converts a multi-objective, multi-constrained nonlinear optimization problem into a sequence decision problem that can be efficiently solved, and solves the technical contradiction in the prior art that the two demands of strengthening gear surface integrity and maintaining blade aerodynamic performance conflict with each other and cannot be considered at the same time;

[0024] 2. The state perception module of the system can construct a more comprehensive and robust machining state image than a single signal source by fusing cutting force, cutting zone temperature, machining vibration signal and tool wear amount and other multi-source heterogeneous data. The introduction of multi-dimensional physical information greatly improves the perception accuracy of the system to the dynamic disturbance of the machining process, laying a solid foundation for subsequent accurate prediction and decision-making;

[0025] 3. The virtual twin prediction module of the system takes the pre-set reduced-order proxy model as the core, breaks through the technical bottleneck that high-fidelity physical simulation models cannot be used for real-time control due to time-consuming calculation, and can accurately predict the three-dimensional stress field, three-dimensional strain field and three-dimensional temperature field distribution inside the workpiece that cannot be seen in milliseconds. This prediction ability of the future physical state gives the entire control system foresight, enabling it to actively avoid potential machining damage, and is a key prerequisite for optimizing multiple contradictory performance targets in coordination;

[0026] 4. The decision reward calculation module of the system establishes a set of innovative evaluation paradigm containing the processing technology target; it converts the high-dimensional, complex physical field distribution into a low-dimensional scalar reward value; by calculating the dimensionless stress field health index, it first maps and unifies the macro-level conflicting processing targets, i.e. the residual stress gain effect of the gear and the equivalent stress damage effect of the blade, to a calculable and optimized micro-physical quantity, providing a clear and smooth optimization direction with physical meaning for intelligent decision-making; Furthermore, the module further introduces a hard constraint penalty term for plastic deformation; by comparing the predicted maximum equivalent plastic strain with the material safety threshold with clear physical source, it sets an absolute safety boundary for the system's learning exploration process, ensuring that in the process of seeking the optimal solution, any risk of permanent damage to the workpiece is avoided with the highest priority, significantly improving the safety and reliability of the optimization process, the deep Q network agent in the system's deep reinforcement learning decision-making module adopts a model-independent learning method; this enables the system to learn and discover good processing strategies autonomously through interaction and experience accumulation with the environment without relying on accurate analytical processing process mathematical models; this gives the system the ability to handle highly nonlinear, strongly coupled complex processing scenarios, making it highly adaptive and robust;

[0027] 5. The processing parameter execution module of the system seamlessly converts the optimal processing parameter action output by the agent, i.e. the spindle speed, feed rate and cutting depth, into instructions executable by the CNC machine; this opens up a complete path from high-level intelligent decision-making to low-level physical control, ensuring the final executability of the optimization results, making the entire adaptive optimization closed loop effectively realized. BRIEF DESCRIPTION OF DRAWINGS

[0029] The present application will be further explained in conjunction with the accompanying drawings and examples:

[0030] Figure 1 is a flow chart of the system of the present application. DETAILED DESCRIPTION

[0032] To make the purpose, technical scheme and advantages of the present application clearer and more apparent, the present application will be further described in detail below in conjunction with specific examples.

[0033] Example 1

[0034] Please refer to Figure 1 , the present application provides a gear processing parameter adaptive optimization system based on deep reinforcement learning, comprising:

[0035] A state perception module is configured to collect multi-source heterogeneous data in real time during the machining process and construct a state vector of the current state of the system.

[0036] A virtual twin prediction module is configured to receive the state vector and candidate machining parameter actions and predict a three-dimensional physical field distribution in a future time step.

[0037] A decision reward calculation module is configured to calculate a scalar reward value for guiding decision-making based on the three-dimensional physical field distribution.

[0038] A deep reinforcement learning decision-making module is configured to take the state vector as input and output an optimal machining parameter action based on the scalar reward value as a learning goal.

[0039] A machining parameter execution module is configured to receive the optimal machining parameter action and convert it into numerical control instructions to drive the machine tool to perform machining.

[0040] A gear machining parameter adaptive optimization system based on deep reinforcement learning is provided to solve the technical contradiction between gear surface integrity reinforcement and blade aerodynamic performance maintenance in integrated component machining. In this embodiment, the system realizes real-time adaptive regulation of machining parameters by constructing a closed-loop control process of perception-prediction-decision-making-execution. The system includes a state perception module, a virtual twin prediction module, a decision reward calculation module, a deep reinforcement learning decision-making module, and a machining parameter execution module.

[0041] The state perception module is configured to capture dynamic physical information of the machining process in real time and provide a basis for decision-making. In this embodiment, the module integrates a sensor network to continuously collect multi-source heterogeneous data reflecting the current machining state and integrates them into a standardized vector representation, i.e., a state vector. The state vector is the input basis for the subsequent calculation and decision-making of all modules.

[0042] The virtual twin prediction module is configured to provide a millisecond-level prediction capability for future machining results to realize forward-looking control. In this embodiment, the module receives the state vector output by the state perception module and the candidate machining parameter actions provided by the deep reinforcement learning decision-making module as joint inputs. Based on these inputs, the module can quickly predict the three-dimensional physical field distribution of the entire workpiece in a very short future time step, thereby explicitly displaying internal stress, strain, and other information that is not visible, and providing data support for quantitative evaluation of the impact of machining.

[0043] A decision reward calculation module aims to convert high-dimensional and complex physical field prediction information into low-dimensional scalar reward values that can guide the learning of the reinforcement learning agent; in this embodiment, the module calculates dimensionless scalar values from the three-dimensional physical field distribution output by the virtual twin prediction module through a set of pre-set mathematical norms containing processing technology targets; the design of the reward value changes the optimization target from the traditional final processing quality to dynamic regulation and control of the health of the physical field during processing;

[0044] A deep reinforcement learning decision module aims to autonomously learn and generate optimal processing strategies as the core decision unit of the system; in this embodiment, the module takes the state vector provided by the state perception module as the observation input of the environment, and takes the scalar reward value generated by the decision reward calculation module as the target signal for learning and optimization; through continuous interaction and trial and error with the processing environment, the module can learn the mapping relationship from the current state to the optimal processing parameter action;

[0045] A processing parameter execution module aims to convert the abstract parameter instructions output by the decision module into specific operations that can be executed by the machine tool in the physical world; in this embodiment, the module serves as an interface that connects digital decision-making with physical processing, receives the optimal processing parameter action output by the deep reinforcement learning decision module, and compiles it into standard numerical control instructions, which then drive the numerical control machine tool to complete the corresponding cutting processing action, thereby completing the entire closed-loop control;

[0046] This embodiment successfully converts the multi-objective, multi-constrained nonlinear optimization problem into a sequence decision problem that can be efficiently solved by the coordinated work of the above-mentioned five modules, combining complex processing mechanisms with efficient intelligent decision-making; the overall technical effect is that it can real-time and predictively regulate and control processing parameters, dynamically compensate for disturbances such as tool wear, and simultaneously improve the surface integrity of the gear and the profile retention of the blade, overcoming the technical contradiction in the prior art that the two cannot be considered simultaneously.

[0047] The multi-source heterogeneous data includes cutting force, cutting zone temperature, processing vibration signal, and tool wear amount;

[0048] In this embodiment, the multi-source heterogeneous data collected by the state perception module is specified; multi-source heterogeneous data refers to a collection of physical signals of different sources and types that collectively represent the processing state, and its function is to provide the system with a comprehensive and accurate portrait of the current processing environment; in this embodiment, the data set specifically includes:

[0049] Cutting force: its function is to represent the interaction strength between the tool and the workpiece, and its source is real-time collection through a multi-component force sensor installed on the spindle or workpiece clamp;

[0050] Cutting zone temperature: the role is to reflect the thermal load level of the machining area, the source is measured by embedding micro-thermocouples near the workpiece or tool, or by non-contact infrared thermal imager;

[0051] Machining vibration signal: the role is to monitor the stability of the machining process, the source is collected by installing acceleration sensors or acoustic emission sensors on key parts of the machine tool;

[0052] Tool wear: the role is to quantify the degree of tool loss, which is a key dynamic variable, and the source is obtained by periodic evaluation through online monitoring systems such as machine vision;

[0053] To ensure the accuracy and consistency of the collected data, the sensor network in this embodiment is configured as follows: the multi-component force sensor is a three-axis piezoelectric force sensor installed on the machine tool turret, used to measure the cutting force in three directions; the acceleration sensor is a MEMS accelerometer installed on the spindle box, used to monitor high-frequency vibrations during machining; the micro-thermocouple is integrated into the tool holder near the cutting area; the machine vision system for evaluating tool wear consists of an industrial camera with a ring light source fixed inside the machine tool guard, which takes a photo of the tool tip at the gap of each machining cycle, and calculates the width of the tool flank wear band through image processing algorithms;

[0054] These data are fused into a state vector:

[0055]

[0056] Where, And are the real-time measurements of cutting temperature and tool wear, respectively, and And represent the key feature scalars extracted from the real-time collected cutting force signal and vibration signal, such as the root mean square value, where the cutting force and vibration signal are processed through feature extraction, such as calculating the root mean square value in the sampling window, to obtain the corresponding scalar value, which can provide more rich and robust state information for the system than a single signal source; : the state vector of the current state of the system; : the real-time measurement of tool wear; The technical effect of this embodiment is that by introducing multi-dimensional physical signals, the perception accuracy of the system for the machining state, especially for dynamic disturbances such as tool wear, is greatly improved, laying a solid foundation for subsequent accurate prediction and decision-making.

[0057] The core of the virtual twin prediction module is a pre-set reduced-order proxy model; the three-dimensional physical field distribution includes three-dimensional stress field, three-dimensional strain field and three-dimensional temperature field;

[0058] ​In this embodiment, the internal structure and output content of the virtual twin prediction module are specified. The core of the virtual twin prediction module is a preset reduced-order proxy model. The preset reduced-order proxy model refers to a lightweight neural network model that has been pre-trained and can quickly reproduce the input-output relationship of a high-fidelity physical simulation model. The working principle is as follows: a multiphysics finite element model with huge computational load but extremely high accuracy is constructed; through model reduction techniques such as intrinsic orthogonal decomposition, the key dynamic characteristics of the high-fidelity model are extracted. These characteristics are used to train a neural network, enabling it to simulate calculations that would take hours for a high-fidelity model to complete at millisecond speeds. This preset method allows it to be embedded into the real-time control loop.

[0059] In this embodiment, the reduced-order surrogate model is specifically a fully connected feedforward neural network; the network includes an input layer, three hidden layers, and an output layer; the input layer receives data from the state vector. and candidate action vectors The features are assembled from the data; the three hidden layers have 256, 128, and 64 neurons respectively, all using modified linear units as activation functions. The output layer outputs a flattened one-dimensional vector, which can be reconstructed into discrete node values ​​of three-dimensional stress, strain, and temperature fields. The model is trained using the Adam optimizer, with an initial learning rate set to... The batch size is 64. The training dataset consists of 10,000 input-output pairs generated by high-fidelity finite element software such as ABAQUS simulation, and the simulation input machining parameters are fully randomly sampled within their process window;

[0060] The three-dimensional physical field distribution output by this module includes:

[0061] Three-dimensional stress field: characterizes the distribution of internal forces at various points inside a workpiece under load;

[0062] Three-dimensional strain field: characterizes the degree of deformation at various points inside the workpiece;

[0063] Three-dimensional temperature field: characterizes the temperature distribution inside the workpiece;

[0064] The technical advantage of this embodiment is that by adopting a reduced-order proxy model, it solves the problem that high-fidelity physical simulation cannot be used for real-time control, and realizes high-speed and high-precision prediction of multiple physical fields inside the processing process. This enables the system to make decisions based on the prediction of future stress, strain and temperature, thereby realizing predictive control, which is a key prerequisite for achieving the goals of actively avoiding damage and collaboratively optimizing performance.

[0065] Example 2

[0066] The decision reward calculation module is used to firstly calculate a non-dimensional stress field health index based on the three-dimensional stress field, in combination with preset weight hyperparameters, workpiece geometric volume and material yield strength;

[0067] The calculation of the stress field health index includes: quantifying the residual stress gain effect of the gear target area and the von Mises equivalent stress damage effect of the blade key area respectively; and performing weighted summation on the gain effect and the damage effect based on the weight hyperparameters;

[0068] In the embodiment, the calculation logic of the decision reward calculation module is specified; the calculation logic of the module is to calculate a non-dimensional stress field health index based on the three-dimensional stress field output by the virtual twin prediction module; the non-dimensional stress field health index, is a scalar index that integrates the two mutually conflicting physical targets of gear gain and blade damage into a unified evaluation system, and the role is to provide a continuous and smooth optimization target for the reinforcement learning agent;

[0069] The calculation of the index separately quantizes the effects of the gear and the blade, and performs weighted summation; to achieve this goal, the calculation formula is defined as follows:

[0070]

[0071] Among them, : The stress field health index at the moment is a non-dimensional scalar, which is calculated by the module;

[0072] The weight hyperparameters are non-dimensional and represent the relative importance of the gear surface gain and the blade damage, respectively, and are obtained by multi-objective optimization algorithm or inverse reinforcement learning optimization from successful processing cases;

[0073] In the embodiment, the values of and are determined by grid search method; we traverse in the preset interval, for example, , , with a step of 0.1, and use a set of standard offline test working conditions to evaluate the final convergence performance of the reinforcement learning agent under each set of parameters, and select the parameter combination that can fastest and most stably balance the gear residual compressive stress and the blade equivalent stress; in the embodiment, the selected values are and ;

[0074] : Yield strength of the material, which is the characteristic stress used for non-dimensionalizing the system, is obtained from standard tensile tests;

[0075] : Volume of the target machining region of the gear and volume of the key aerodynamic profile of the blade, respectively, are extracted directly from the computer-aided design model of the workpiece;

[0076] : Residual stress tensor of the tooth surface region predicted by the virtual twin module, used to quantify the residual stress gain effect;

[0077] : Tooth surface stress evaluation function, defined in this embodiment as:

[0078]

[0079] Wherein, : First principal stress; this definition makes the beneficial residual compressive stress, produce positive rewards;

[0080] : Von Mises equivalent stress of the blade region predicted by the virtual twin module, used to quantify the von Mises equivalent stress damage effect;

[0081] The technical effect of this embodiment is that by establishing This core index, the macroscopic and mutually conflicting machining targets, gear life and blade precision, are converted into calculable and optimized micro-process physical quantities; this provides the intelligent agent with a clear and physically meaningful optimization direction, making collaborative optimization possible.

[0082] Embodiment 3

[0083] The decision reward calculation module is further used to: based on the stress field health index, and introducing a hard constraint penalty term for plastic deformation, to construct the final scalar reward value;

[0084] The introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the key region of the blade predicted by the virtual twin prediction module with the preset blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty is applied;

[0085] In this embodiment, the calculation logic of the decision reward calculation module is further completed; based on the stress field health index, the module introduces a hard constraint penalty term for plastic deformation to construct the final scalar reward value; the hard constraint penalty term refers to a mechanism that gives a huge negative incentive when the system's behavior touches a certain insurmountable safety boundary, and the role is to force the agent to ensure the structural integrity of the workpiece as the highest priority in any case;

[0086] The final scalar reward value is defined as follows:

[0087]

[0088] Wherein, : The final reward value received by the agent at the moment is a dimensionless scalar, calculated by this module as the learning goal of the deep reinforcement learning decision module;

[0089] : The stress field health index calculated by the foregoing formula;

[0090] : The dimensionless plastic deformation penalty factor, the value is set to be much larger than The expected absolute value of a positive number, which comes from hyperparameter optimization, and the role is to represent a severe punishment for behavior that crosses the damage threshold;

[0091] To ensure the effectiveness of the penalty, The value of is set to be at least two orders of magnitude higher than the theoretical maximum value of Considering that After dimensionless processing, its absolute value is usually within 10, in this embodiment, Is directly set to a fixed large value, for example ;

[0092] : The maximum equivalent plastic strain of the key area of the blade predicted by the virtual twin prediction module;

[0093] : The preset plastic deformation safety threshold of the blade material; this is a parameter with a clear physical source, and the determination process is: through uniaxial tensile test of the grade of high-temperature alloy used for this part, the accurate stress-strain curve is obtained; combined with the accurate three-dimensional model of the part, finite element simulation analysis is carried out to calibrate the maximum equivalent plastic strain value of the blade surface without permanent deformation;

[0094] : The Heaviside step function, the working principle is like a switch, when the internal parameter is 1 if greater than zero, otherwise 0;

[0095] The technical effect of the embodiment is that by introducing a hard constraint penalty term based on explicit physical thresholds, an absolute safety boundary is drawn for the system's exploratory learning process; this ensures that the system will prioritize avoiding any risk of causing permanent plastic deformation of the blade in the process of seeking optimal machining parameters, greatly improving the safety and reliability of the optimization process.

[0096] Embodiment 4

[0097] The deep reinforcement learning decision module is built-in with a deep Q network agent, which learns through continuous interaction with the machining environment;

[0098] In this embodiment, the internal implementation of the deep reinforcement learning decision module is specified; the module is built-in with a deep Q network agent; the deep Q network agent is a deep reinforcement learning algorithm, which works by using a deep neural network to approximate the optimal action value function, Q function, so as to be able to handle high-dimensional state input space; the agent learns through continuous interaction with the machining environment, i.e. constantly updates its internal neural network weights according to the current state output action , obtains the reward from the environment and the next state , and uses these experience data to iteratively update the weights of its internal neural network, so that it can output the action that brings the maximum long-term reward for any given state;

[0099] In this embodiment, the deep Q network is specifically a neural network containing four fully connected layers; the input layer receives the state vector , the subsequent two hidden layers contain 128 and 64 neurons respectively, and use ReLU activation function; the output layer is a linear layer with neurons, where is the total number of actions in the discretized action space, and each neuron corresponds to the Q value of a discrete action; during training, the - greedy strategy is used for exploration, where is linearly decayed from 1.0 to 0.01. The size of the experience replay pool is 20000, the batch size during training is 128, the learning rate is , and the Adam optimizer is used for gradient descent;

[0100] The technical effect of the embodiment is that the model-independent reinforcement learning method such as the deep Q network is used, so that the system does not need to establish an accurate and analytical machining process mathematical model; the agent can independently find an excellent machining strategy through trial and error and experience accumulation like a human expert, thereby having the ability to process a highly nonlinear and strongly coupled complex machining scene, and having strong adaptability and robustness.

[0101] The optimal machining parameter action includes a spindle speed, a feed rate, and a cutting depth.

[0102] In the embodiment, the content of the optimal machining parameter action output by the deep reinforcement learning decision module is limited; the optimal machining parameter action refers to a set of machine tool control parameters that can maximize the long-term return under the current state and are decided by the agent; in the embodiment, the action is defined as a vector:

[0103]

[0104] wherein, : candidate action vector, or optimal machining parameter action;

[0105] : spindle speed, speed of controlling rotation of a tool;

[0106] : feed rate, speed of controlling movement of the tool relative to a workpiece;

[0107] : cutting depth, depth of cutting of the tool into the workpiece in a single cutting;

[0108] The three parameters are the most core and directly controllable process parameters in numerical control machining, and their combination directly determines the generation of cutting force, cutting heat, and the final machining quality.

[0109] Since the standard deep Q network algorithm is suitable for a discrete action space, the machining parameters , the spindle speed, , the feed rate, and , the cutting depth in the embodiment are continuous variables, and therefore the action space needs to be discretized, and each parameter is divided into several levels within its process allowable range; for example, the spindle speed is discretized into 5 levels within the range of [6000, 12000] rpm; the feed rate is discretized into 5 levels within the range of [500, 1500] mm / min; and the cutting depth is discretized into 4 levels within the range of [0.1, 0.5] mm; therefore, the total size of the discretized action space is ; The deep reinforcement learning decision module outputs the optimal index of the 100 combined actions, and the machining parameter execution module decodes the index into specific parameter combination instructions;

[0110] The technical effect of the embodiment is that the decision output of the system is directly corresponding to the core control parameters at the physical execution level of the machine tool; this makes the optimization results of the system be seamlessly converted into specific numerical control instructions by the machining parameter execution module, ensuring the executability of the decision, constituting a complete path from high-level decision to low-level control, and enabling the whole adaptive optimization closed loop to be realized.

[0111] To ensure the robustness of the system in actual application, the system also includes a monitoring and safety module (not detailed); the module is responsible for detecting outliers of sensor data, and triggering safety strategies when the data exceeds the pre-set reasonable physical range, such as switching to a conservative set of default machining parameters, to prevent abnormal input from leading to wrong decisions; at the same time, the output actions of the agent decision are also bounded, ensuring that they are always within the process window allowed by the machine tool.

[0112] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A deep reinforcement learning based adaptive optimization system for gear machining parameters, characterized in that, The method comprises the following steps: a state perception module for collecting multi-source heterogeneous data in real time during the machining process and constructing a state vector representing the current state of the system; a virtual twin prediction module for receiving the state vector and candidate machining parameter actions and predicting the three-dimensional physical field distribution at the future time step; a decision reward calculation module for calculating a scalar reward value for guiding decision-making based on the three-dimensional physical field distribution; a deep reinforcement learning decision module for taking the state vector as input and taking the scalar reward value as the learning goal to output the optimal machining parameter action; a machining parameter execution module for receiving the optimal machining parameter action and converting it into numerical control instructions to drive the machine tool to perform machining.

2. The system of claim 1, wherein, The multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

3. The system of claim 1, wherein, The core of the virtual twin prediction module is a pre-set reduced-order surrogate model; the three-dimensional physical field distribution includes three-dimensional stress field, three-dimensional strain field, and three-dimensional temperature field.

4. The system of claim 1, wherein, The decision reward calculation module is used to first calculate the dimensionless stress field health index based on the three-dimensional stress field, combined with pre-set weight hyperparameters, workpiece geometric volume, and material yield strength.

5. The deep reinforcement learning based gear machining parameter adaptive optimization system according to claim 4, wherein, The calculation of the stress field health index includes: quantifying the residual stress gain effect of the gear target area and the von Mises equivalent stress damage effect of the blade key area respectively; and weighting the gain effect and damage effect based on the weight hyperparameters.

6. The deep reinforcement learning based gear machining parameter adaptive optimization system of claim 4, wherein, The decision reward calculation module is further used to: based on the stress field health index, introduce a hard constraint penalty term for plastic deformation, and construct the final scalar reward value.

7. The system of claim 6, wherein the system is configured to: The introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the blade key area predicted by the virtual twin prediction module with the pre-set blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a pre-set penalty factor is applied, otherwise no penalty is applied.

8. The system of claim 1, wherein, The deep reinforcement learning decision module has a deep Q network agent built-in, which learns through continuous interaction with the machining environment.

9. The system of claim 1, wherein, The optimal machining parameter action includes spindle speed, feed rate, and cutting depth.

Citation Information

Patent Citations

  • Centrifugal impeller comprehensive optimization method based on digital twinning and reinforcement learning

    CN111241752A

  • Front edge reinforced edge shape finish machining efficiency and residual stress deformation collaborative optimization method

    CN118798012A

  • Electromagnetic-assisted nickel-based alloy laser melting deposition crack inhibition method and electromagnetic-assisted nickel-based alloy laser melting deposition crack inhibition system

    CN119407199A

  • Cutting parameter dynamic autonomous decision-making method and system based on deep reinforcement learning

    CN119511954A

  • Numerical control tool machining monitoring method

    CN119828596A