A gear machining parameter adaptive optimization system based on deep reinforcement learning

The gear machining parameter adaptive optimization system, which utilizes deep reinforcement learning, collects multi-source data in real time and predicts the three-dimensional physical field distribution. This resolves the contradiction between gear surface integrity and blade aerodynamic performance in integrated bladed disk-gear machining, thereby improving the safety and reliability of the machining process.

CN120909133BActive Publication Date: 2026-01-27NANTONG XINSIDI ELECTROMECHANICAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511438052.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-01-27
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously optimize gear surface integrity and blade aerodynamic performance in integrated bladed disk-gear machining, leading to contradictions during the machining process and failing to meet the conflicting requirements of both.

Method used

An adaptive optimization system for gear machining parameters based on deep reinforcement learning is adopted. Through a state perception module, a virtual twin prediction module, a decision reward calculation module, and a deep reinforcement learning decision module, multi-source heterogeneous data is collected in real time to predict the three-dimensional physical field distribution. The system then performs autonomous learning through a deep Q-network agent to output the optimal machining parameters and actions.

Benefits of technology

This technology enables the collaborative optimization of gear surface integrity and blade aerodynamic performance during real-time machining, resolving the technical contradictions that traditional methods cannot address simultaneously, and improving the safety and reliability of the machining process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909133B_ABST
    Figure CN120909133B_ABST
Patent Text Reader

Abstract

The application discloses a gear machining parameter adaptive optimization system based on deep reinforcement learning, belongs to the technical field of mechanical precision machining and artificial intelligence, and comprises: a state sensing module, which is used for collecting multi-source heterogeneous data in a machining process in real time and constructing a state vector of a current state of the system; a virtual twin prediction module, which is used for receiving the state vector and candidate machining parameter actions and predicting three-dimensional physical field distribution in future time steps; and a decision reward calculation module, which is used for calculating a scalar reward value for guiding decision-making according to the three-dimensional physical field distribution. The application solves the technical contradiction that the two requirements of reinforcing gear surface integrity and maintaining blade aerodynamic performance conflict with each other and cannot be considered simultaneously in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of precision machining and artificial intelligence, specifically to an adaptive optimization system for gear machining parameters based on deep reinforcement learning. Background Technology

[0002] In the field of aero-engines, the integral bladed disk-gear structure integrates blades, disk, and transmission gears into a single unit. This structure offers significant advantages in improving thrust-to-weight ratio and reducing structural weight, but it also presents unprecedented manufacturing challenges. The integrated structure leads to conflicting performance requirements among different functional areas. To ensure high fatigue life, the gear area requires a specific residual compressive stress distribution, low roughness, and high microhardness on the machined surface—that is, excellent surface integrity. Conversely, to ensure aerodynamic efficiency, the blade area requires minimal deformation during machining to maintain precise aerodynamic performance.

[0003] Traditional machining methods typically employ fixed-parameter strategies based on expert experience or offline simulation, with the optimization goal of achieving the best quality for a single process or objective. In gear finishing, to obtain ideal surface integrity, combinations of cutting parameters that introduce significant residual compressive stress are often used. These parameters often generate enormous cutting forces and heat, which are transmitted through the integrated structure to the less rigid blade section, inevitably causing micron-level permanent plastic deformation of the blade, damaging its profile and leading to a decrease in aerodynamic performance. This phenomenon reveals a profound technical contradiction: in the machining of integrated components, any machining strategy aimed at enhancing tooth surface performance may come at the expense of blade aerodynamic performance; this contradiction becomes even more acute under the influence of dynamic variables such as tool wear. The conventional practice of increasing cutting loads to compensate for wear directly exacerbates blade damage. Existing technologies cannot proactively avoid this problem and can only passively compromise between the two conflicting goals of gear life and blade efficiency.

[0004] The technical problem that this invention aims to solve is to provide a machining parameter adaptive optimization system that can overcome the above-mentioned technical contradictions in the integrated machining scenario of an integral bladed disk-gear.

[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide an adaptive optimization system for gear machining parameters based on deep reinforcement learning, so as to solve the problems mentioned in the background art.

[0007] The technical solution of the present invention includes:

[0008] The state awareness module is used to collect multi-source heterogeneous data in real time during the processing and construct it into a state vector of the current system state.

[0009] The virtual twin prediction module is used to receive the state vector and candidate processing parameters, and predict the three-dimensional physical field distribution in the future time step;

[0010] The decision reward calculation module is used to calculate scalar reward values ​​to guide decision-making based on the three-dimensional physical field distribution.

[0011] The deep reinforcement learning decision module is used to take the state vector as input and, based on the scalar reward value as the learning objective, output the optimal processing parameters and actions.

[0012] The machining parameter execution module is used to receive the optimal machining parameter actions and convert them into CNC instructions to drive the machine tool to perform machining.

[0013] Preferably, the multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

[0014] Preferably, the core of the virtual twin prediction module is a preset reduced-order proxy model; the three-dimensional physical field distribution includes a three-dimensional stress field, a three-dimensional strain field, and a three-dimensional temperature field.

[0015] Preferably, the decision reward calculation module is used to first calculate a dimensionless stress field health index based on the three-dimensional stress field and in combination with preset weight hyperparameters, workpiece geometric volume and material yield strength.

[0016] Preferably, the calculation of the stress field health index includes: quantifying the residual stress gain effect in the target area of ​​the gear and the von Mises equivalent stress damage effect in the key area of ​​the blade, respectively; and weighting and summing the gain effect and damage effect based on the weight hyperparameter.

[0017] Preferably, the decision reward calculation module is further used to: construct the final scalar reward value based on the stress field health index and by introducing a hard constraint penalty term for plastic deformation.

[0018] Preferably, the introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the blade critical region predicted by the virtual twin prediction module with a preset blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty is applied.

[0019] Preferably, the deep reinforcement learning decision module incorporates a deep Q-network agent, which learns through continuous interaction with the processing environment.

[0020] Preferably, the optimal machining parameters include spindle speed, feed rate, and depth of cut.

[0021] This invention provides an improved adaptive optimization system for gear machining parameters based on deep reinforcement learning, which has the following improvements and advantages compared with the prior art:

[0022] 1. This technical solution constructs a closed-loop control architecture that integrates state perception, virtual twin prediction, decision reward calculation, deep reinforcement learning decision-making, and processing parameter execution. This successfully transforms a multi-objective, multi-constraint nonlinear optimization problem into a sequential decision problem that can be solved efficiently. It also solves the technical contradiction in the existing technology where the two requirements of enhancing the integrity of the gear surface and maintaining the aerodynamic performance of the blades are conflicting and cannot be simultaneously met during the processing of integrated components.

[0023] 2. The system's state perception module, by integrating multi-source heterogeneous data such as cutting force, cutting zone temperature, machining vibration signals, and tool wear, can construct a more comprehensive and robust machining state profile than a single signal source. The introduction of this multi-dimensional physical information greatly improves the system's perception accuracy of dynamic disturbances in the machining process, laying a solid foundation for subsequent accurate prediction and decision-making.

[0024] 3. The system's virtual twin prediction module, with a preset reduced-order proxy model as its core, breaks through the technical bottleneck that high-fidelity physical simulation models cannot be used for real-time control due to computational time. This module can predict the invisible three-dimensional stress field, three-dimensional strain field, and three-dimensional temperature field distribution inside the workpiece with high precision at the millisecond level. This ability to predict future physical states gives the entire control system foresight, enabling it to actively avoid potential processing damage, which is a key prerequisite for synergistically optimizing multiple conflicting performance objectives.

[0025] 4. The system's decision-reward calculation module establishes an innovative evaluation paradigm that incorporates processing technology objectives. It transforms high-dimensional, complex physical field distributions into a low-dimensional scalar reward value. By calculating a dimensionless stress field health index, it for the first time maps and unifies conflicting macroscopic processing objectives—namely, the residual stress gain effect of gears and the equivalent stress damage effect of blades—to a calculable and optimizable microscopic physical quantity, providing a physically clear and smooth optimization direction for intelligent decision-making. Furthermore, this module introduces a hard constraint penalty term for plastic deformation; by comparing the predicted maximum equivalent plastic strain with the material strain having a clear physical origin… By comparing material safety thresholds, an absolute safety boundary is defined for the system's learning and exploration process. This ensures that any risk that could cause permanent damage to the workpiece is avoided with the highest priority during the search for the optimal solution, significantly improving the safety and reliability of the optimization process. The system's deep reinforcement learning decision module incorporates a deep Q-network agent and employs a model-independent learning method. This allows the system to learn autonomously and discover superior processing strategies through trial and error and experience accumulation in interaction with the environment, without relying on a precise analytical mathematical model of the machining process. This endows the system with the ability to handle highly nonlinear, strongly coupled, and complex machining scenarios, giving it strong adaptability and robustness.

[0026] 5. The system's machining parameter execution module seamlessly converts the optimal machining parameters output by the intelligent agent—namely, the core process parameters such as spindle speed, feed rate, and depth of cut—into executable instructions for the CNC machine tool. This opens up a complete pathway from high-level intelligent decision-making to low-level physical control, ensuring the ultimate executability of the optimization results and enabling the entire adaptive optimization closed loop to be effectively realized. Attached Figure Description

[0027] The present invention will be further explained below with reference to the accompanying drawings and embodiments:

[0028] Figure 1 This is a flowchart of the system of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0030] Example 1

[0031] Please see Figure 1 This invention provides an adaptive optimization system for gear machining parameters based on deep reinforcement learning, comprising:

[0032] The state awareness module is used to collect multi-source heterogeneous data in real time during the processing and construct it into a state vector of the current system state.

[0033] The virtual twin prediction module is used to receive the state vector and candidate processing parameters, and predict the three-dimensional physical field distribution in the future time step;

[0034] The decision reward calculation module is used to calculate scalar reward values ​​to guide decision-making based on the three-dimensional physical field distribution.

[0035] The deep reinforcement learning decision module is used to take the state vector as input and, based on the scalar reward value as the learning objective, output the optimal processing parameters and actions.

[0036] The machining parameter execution module is used to receive the optimal machining parameter actions and convert them into CNC instructions to drive the machine tool to perform machining.

[0037] An adaptive optimization system for gear machining parameters based on deep reinforcement learning is proposed. The aim is to resolve the technical contradiction between enhancing the surface integrity of gears and maintaining the aerodynamic performance of blades in the machining of integrated components. In this embodiment, the system achieves real-time adaptive control of machining parameters by constructing a closed-loop control process of perception-prediction-decision-execution. The system includes a state perception module, a virtual twin prediction module, a decision reward calculation module, a deep reinforcement learning decision module, and a machining parameter execution module.

[0038] The state perception module aims to capture dynamic physical information of the processing process in real time, providing a basis for system decision-making. In this embodiment, the module continuously collects multi-source heterogeneous data reflecting the current processing state through an integrated sensor network, and integrates it into a standardized vectorized representation, namely the state vector. This state vector is the input basis for all subsequent modules to perform calculations and decisions.

[0039] The virtual twin prediction module aims to provide a millisecond-level prediction capability for future processing results, thereby achieving forward-looking control. In this embodiment, the module receives a state vector output by the state perception module and candidate processing parameters provided by the deep reinforcement learning decision module as joint inputs. Based on this input, the module can quickly predict the three-dimensional physical field distribution of the entire workpiece within a very short time step in the future, thereby making invisible internal stress, strain and other information explicit, and providing data support for quantitatively evaluating the impact of processing.

[0040] The decision reward calculation module aims to transform high-dimensional, complex physical field prediction information into low-dimensional scalar reward values ​​that can guide reinforcement learning agents in their learning. In this embodiment, the module calculates dimensionless scalar values ​​based on the three-dimensional physical field distribution output by the virtual twin prediction module, using a set of preset mathematical paradigms that contain the processing technology objectives. The design of this reward value shifts the optimization objective from the traditional final processing quality to the dynamic control of the physical field health during the processing.

[0041] The deep reinforcement learning decision module aims to serve as the core decision-making unit of the system, autonomously learning and generating the optimal processing strategy. In this embodiment, the module uses the state vector provided by the state perception module as the observation input to the environment, and the scalar reward value generated by the decision reward calculation module as the target signal for learning and optimization. Through continuous interaction and trial and error with the processing environment, the module can learn the mapping relationship from the current state to the optimal processing parameter action.

[0042] The purpose of the machining parameter execution module is to transform the abstract parameter instructions output by the decision module into specific operations that the machine tool can perform in the physical world. In this embodiment, the module serves as the interface for connecting digital decision-making with physical machining. It receives the optimal machining parameter actions output by the deep reinforcement learning decision module, compiles them into standard CNC instructions, and then drives the CNC machine tool to complete the corresponding cutting actions, thereby completing the entire closed-loop control.

[0043] This embodiment combines complex machining mechanisms with efficient intelligent decision-making through the coordinated work of the above five modules, successfully transforming a multi-objective, multi-constraint nonlinear optimization problem into a sequential decision problem that can be solved efficiently. The overall technical effect is that it can control machining parameters in real time and predictively, and while dynamically compensating for disturbances such as tool wear, it can synergistically improve the surface integrity of gears and the profile retention of blades, overcoming the technical contradiction that the two cannot be balanced in the prior art.

[0044] Multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

[0045] In this embodiment, the multi-source heterogeneous data collected by the state perception module is specified. Multi-source heterogeneous data refers to a set of physical signals from different sources and of different types, but which together characterize the processing state. Its purpose is to provide the system with a comprehensive and accurate picture of the current processing environment. In this embodiment, the data set specifically includes:

[0046] Cutting force: Its function is to characterize the intensity of the interaction between the tool and the workpiece. It is obtained in real time by a multi-component force sensor installed on the spindle or workpiece fixture.

[0047] Cutting zone temperature: Its function is to reflect the thermal load level of the machining area. It is obtained by measuring through miniature thermocouples embedded near the workpiece or tool, or by non-contact infrared thermal imager.

[0048] Vibration signals during machining: Their function is to monitor the stability of the machining process. They are collected by accelerometers or acoustic emission sensors installed in key parts of the machine tool.

[0049] Tool wear: Its function is to quantify the degree of tool wear. This is a key dynamic variable, which is obtained through periodic evaluation by online monitoring systems such as machine vision.

[0050] To ensure the accuracy and consistency of the collected data, the sensor network configuration in this embodiment is as follows: the multi-component force sensor is a triaxial piezoelectric force sensor mounted on the machine tool turret, used to measure... Cutting forces in three directions; the accelerometer is a MEMS accelerometer mounted on the spindle housing to monitor high-frequency vibrations during machining; a miniature thermocouple is integrated into the tool holder near the cutting area; the machine vision system for assessing tool wear consists of an industrial camera with a ring light source fixed inside the machine tool guard, which takes pictures of the tool tip between each machining cycle and calculates the width of the flank wear band using image processing algorithms;

[0051] Combine these data into a single state vector:

[0052]

[0053] in, and These are the real-time measured values ​​of cutting temperature and tool wear, respectively. and These represent key feature scalars extracted from real-time acquired cutting force and vibration signals, such as root mean square (RMS) values. The cutting force and vibration signals undergo feature extraction, such as calculating their RMS values ​​within the sampling window, to obtain corresponding scalar values. This provides the system with richer and more robust state information than a single signal source. : The state vector of the current state of the system; The real-time measurement value of tool wear; the technical effect of this embodiment is that by introducing multi-dimensional physical signals, the system’s perception accuracy of machining status, especially dynamic disturbances such as tool wear, is greatly improved, laying a solid foundation for subsequent accurate prediction and decision-making.

[0054] The core of the virtual twin prediction module is a preset reduced-order proxy model; the three-dimensional physical field distribution includes a three-dimensional stress field, a three-dimensional strain field, and a three-dimensional temperature field.

[0055] In this embodiment, the internal structure and output content of the virtual twin prediction module are specified. The core of the virtual twin prediction module is a preset reduced-order proxy model. The preset reduced-order proxy model refers to a lightweight neural network model that has been pre-trained and can quickly reproduce the input-output relationship of a high-fidelity physical simulation model. The working principle is as follows: a multiphysics finite element model with huge computational load but extremely high accuracy is constructed; through model reduction techniques such as intrinsic orthogonal decomposition, the key dynamic characteristics of the high-fidelity model are extracted. These characteristics are used to train a neural network, enabling it to simulate calculations that would take hours for a high-fidelity model to complete at millisecond speeds. This preset method allows it to be embedded into the real-time control loop.

[0056] In this embodiment, the reduced-order surrogate model is specifically a fully connected feedforward neural network; the network includes an input layer, three hidden layers, and an output layer; the input layer receives data from the state vector. and candidate action vectors The features are assembled from the data; the three hidden layers have 256, 128, and 64 neurons respectively, all using modified linear units as activation functions. The output layer outputs a flattened one-dimensional vector, which can be reconstructed into discrete node values ​​of three-dimensional stress, strain, and temperature fields. The model is trained using the Adam optimizer, with an initial learning rate set to... The batch size is 64. The training dataset consists of 10,000 input-output pairs generated by high-fidelity finite element software such as ABAQUS simulation, and the simulation input machining parameters are fully randomly sampled within their process window;

[0057] The three-dimensional physical field distribution output by this module includes:

[0058] Three-dimensional stress field: characterizes the distribution of internal forces at various points inside a workpiece under load;

[0059] Three-dimensional strain field: characterizes the degree of deformation at various points inside the workpiece;

[0060] Three-dimensional temperature field: characterizes the temperature distribution inside the workpiece;

[0061] The technical advantage of this embodiment is that by adopting a reduced-order proxy model, it solves the problem that high-fidelity physical simulation cannot be used for real-time control, and realizes high-speed and high-precision prediction of multiple physical fields inside the processing process. This enables the system to make decisions based on the prediction of future stress, strain and temperature, thereby realizing predictive control, which is a key prerequisite for achieving the goals of actively avoiding damage and collaboratively optimizing performance.

[0062] Example 2

[0063] The decision return calculation module is used to first calculate the dimensionless stress field health index based on the three-dimensional stress field, combined with preset weight hyperparameters, workpiece geometry and material yield strength.

[0064] The calculation of the stress field health index includes: quantifying the residual stress gain effect in the target area of ​​the gear and the von Mises equivalent stress damage effect in the key area of ​​the blade, respectively; and weighting and summing the gain effect and damage effect based on the weight hyperparameter.

[0065] In this embodiment, the calculation logic of the decision return calculation module is specified; the calculation logic of this module is to calculate a dimensionless stress field health index based on the three-dimensional stress field output by the virtual twin prediction module; the dimensionless stress field health index, This refers to the scalar index that integrates the two conflicting physical objectives of gear gain and blade damage into a unified evaluation system. Its function is to provide continuous and smooth optimization objectives for reinforcement learning agents.

[0066] The calculation of this index quantifies the effects of the two key regions, gears and blades, separately and then performs a weighted summation; to achieve this goal, the calculation formula is defined as follows:

[0067]

[0068] in, : The stress field health index at time t is a dimensionless scalar, calculated by this module;

[0069] The dimensionless weighted hyperparameters represent the relative importance of tooth surface gain and blade damage, respectively. They are obtained through multi-objective optimization algorithms or by inverse reinforcement learning optimization from successful machining cases.

[0070] In this embodiment, a grid search method is used to determine... and The value; we are in a preset range, for example, , The test iterates through the system in steps of 0.1 and uses a set of standard offline test cases to evaluate each test case. The final convergence performance of the reinforcement learning agent under the given parameters is determined by selecting the parameter combination that most quickly and stably balances the residual compressive stress of the gear and the equivalent stress of the blade. In this embodiment, the selected value is... and ;

[0071] The yield strength of a material is a dimensionless characteristic stress that is used as a reference physical quantity to treat the system. It is obtained through standard material tensile tests.

[0072] These are the volumes of the target machining area of ​​the gear and the key aerodynamic profile of the blade, respectively, which are extracted directly from the computer-aided design model of the workpiece.

[0073] The residual stress tensor of the tooth surface region predicted by the virtual twin module is used to quantify the residual stress gain effect.

[0074] The tooth surface stress evaluation function, in this embodiment, is defined as:

[0075]

[0076] in, First principal stress; this definition allows for favorable residual compressive stress. Generate positive rewards;

[0077] The von Mises equivalent stress in the blade region, predicted by the virtual twin module, is used to quantify the von Mises equivalent stress damage effect.

[0078] The technical advantage of this embodiment lies in establishing This core indicator transforms macroscopic and conflicting processing objectives, such as gear life and blade precision, into calculable and optimizable microscopic process physical quantities. This provides intelligent agents with clear and physically meaningful optimization directions, making collaborative optimization possible.

[0079] Example 3

[0080] The decision reward calculation module is further used to: construct the final scalar reward value based on the stress field health index and by introducing a hard constraint penalty term for plastic deformation;

[0081] The introduction of hard constraint penalty terms includes: comparing the maximum equivalent plastic strain of the blade critical region predicted by the virtual twin prediction module with a preset blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty is applied.

[0082] In this embodiment, the calculation logic of the decision reward calculation module has been further improved. Based on the stress field health index, the module introduces a hard constraint penalty term for plastic deformation to construct the final scalar reward value. The hard constraint penalty term refers to a mechanism that provides a huge negative incentive when the behavior of the system touches a certain insurmountable safety boundary. Its function is to force the agent to prioritize the structural integrity of the workpiece under any circumstances.

[0083] The final scalar return value The calculation formula is defined as follows:

[0084]

[0085] in, : The final reward value received by the agent at each moment is a dimensionless scalar, calculated by this module, and serves as the learning target of the deep reinforcement learning decision module.

[0086] The stress field health index is calculated using the aforementioned formula.

[0087] The dimensionless plastic deformation penalty factor is set to a value much greater than... A positive absolute value of expectation comes from hyperparameter optimization and represents a severe penalty for behavior that crosses the damage threshold.

[0088] To ensure the effectiveness of punishment, The value is set to be at least greater than The theoretical maximum value is two orders of magnitude higher. Considering After dimensionless processing, its absolute value is usually less than 10. In this embodiment, It is directly set to a fixed, large value, for example ;

[0089] : The maximum equivalent plastic strain in the critical region of the blade predicted by the virtual twin prediction module;

[0090] : The preset safety threshold for plastic deformation of the blade material; this is a parameter with a clear physical source, and the determination process is as follows: by conducting a uniaxial tensile test on the high-temperature alloy used in the part, its accurate stress-strain curve is obtained; combined with the accurate three-dimensional model of the part, finite element simulation analysis is performed to calibrate the maximum equivalent plastic strain value of the blade surface without permanent deformation.

[0091] The Heaviside step function works like a switch; when the internal parameters... The function value is 1 when the value is greater than zero, and 0 otherwise.

[0092] The technical advantage of this embodiment is that by introducing a hard constraint penalty term based on a clear physical threshold, an absolute safety boundary is defined for the system's exploration and learning process. This ensures that the system will avoid any risks that may cause permanent plastic deformation of the blades with the highest priority when seeking the optimal processing parameters, which greatly improves the safety and reliability of the optimization process.

[0093] Example 4

[0094] The deep reinforcement learning decision-making module incorporates a deep Q-network agent, which learns through continuous interaction with the processing environment.

[0095] In this embodiment, the internal implementation of the deep reinforcement learning decision-making module is specified. This module incorporates a deep Q-network agent. The deep Q-network agent is a deep reinforcement learning algorithm that uses a deep neural network to approximate the optimal action-value function, the Q-function, thereby enabling it to handle high-dimensional state input spaces. This agent learns through continuous interaction with the processing environment, that is, it continuously adjusts its behavior based on the current state. Output Action To reap rewards from the environment and the next state It uses this empirical data to iteratively update the weights of its internal neural network, enabling it to output actions that yield the greatest long-term rewards for any given state.

[0096] In this embodiment, the deep Q-network is specifically a neural network containing four fully connected layers; the input layer receives the state vector. The subsequent two hidden layers contain 128 and 64 neurons respectively, and employ the ReLU activation function; the output layer has... A linear layer of neurons, wherein It represents the total number of actions in the discretized action space, with each neuron corresponding to the Q-value of a discrete action; during training, it uses... - Explore using a greedy strategy, where The learning rate decreases linearly from 1.0 to 0.01. The experience replay pool size is 20,000, the batch size during training is 128, and the learning rate is... And use the Adam optimizer for gradient descent;

[0097] The technical advantage of this embodiment is that by using model-agnostic reinforcement learning methods such as deep Q-networks, the system does not need to establish a precise, analytical mathematical model of the processing procedure. The agent can autonomously discover excellent processing strategies through trial and error and experience accumulation, just like human experts. This gives it the ability to handle highly nonlinear and strongly coupled complex processing scenarios, and it has strong adaptability and robustness.

[0098] Optimal machining parameters include spindle speed, feed rate, and depth of cut;

[0099] In this embodiment, the content of the optimal machining parameter action output by the deep reinforcement learning decision module is defined; the optimal machining parameter action refers to a set of machine tool control parameters decided by the agent in the current state that maximizes long-term rewards; in this embodiment, this action is defined as a vector:

[0100]

[0101] in, Candidate action vector, or optimal processing parameter action;

[0102] Spindle speed: controls the speed at which the tool rotates;

[0103] Feed rate: controls the speed at which the tool moves relative to the workpiece;

[0104] : Depth of cut, which controls the depth to which the tool penetrates the workpiece in a single cut;

[0105] These three parameters are the most critical and directly controllable process parameters in CNC machining. Their combination directly determines the cutting force, the generation of cutting heat, and the final machining quality.

[0106] Since the standard depth Q-network algorithm is applicable to discrete action spaces, while the processing parameters in this embodiment... Spindle speed, feed rate and Since the cutting depth is a continuous variable, the motion space needs to be discretized. We divide each parameter into several levels within its process-allowed range; for example, the spindle speed... Discretized into 5 levels within the range of [6000, 12000] rpm; feed rate It is discretized into 5 levels within the range of [500, 1500] mm / min; cutting depth It is discretized into 4 levels within the range of [0.1, 0.5] mm; therefore, the total size of the discrete action space is... The deep reinforcement learning decision-making module outputs the optimal index among these 100 combined actions, and the processing parameter execution module then decodes this index into specific actions. Parameter combination instructions;

[0107] The technical advantage of this embodiment lies in directly corresponding the system's decision output with the core control parameters at the machine tool's physical execution level. This allows the system's optimization results to be seamlessly converted into specific CNC instructions by the machining parameter execution module, ensuring the executability of the decisions and forming a complete pathway from high-level decision-making to low-level control, thus enabling the entire adaptive optimization closed loop to be realized.

[0108] To ensure the robustness of the system in practical applications, this system also includes a monitoring and safety module (not detailed). This module is responsible for detecting abnormal values ​​in sensor data and triggering safety strategies when the data exceeds the preset reasonable physical range, such as switching to a set of conservative default processing parameters to prevent abnormal input from leading to incorrect decisions. At the same time, it also imposes boundary restrictions on the output actions of the agent's decisions to ensure that they always remain within the process window allowed by the machine tool.

[0109] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An adaptive optimization system for gear machining parameters based on deep reinforcement learning, characterized in that, include: The state awareness module is used to collect multi-source heterogeneous data in real time during the processing and construct it into a state vector of the current system state. The virtual twin prediction module is used to receive the state vector and candidate processing parameters, and predict the three-dimensional physical field distribution in the future time step; the three-dimensional physical field distribution includes the three-dimensional stress field, the three-dimensional strain field and the three-dimensional temperature field. The decision reward calculation module is used to calculate scalar reward values ​​to guide decision-making based on the three-dimensional physical field distribution. The deep reinforcement learning decision module is used to take the state vector as input and, based on the scalar reward value as the learning objective, output the optimal processing parameters and actions. The machining parameter execution module is used to receive the optimal machining parameter actions and convert them into CNC instructions to drive the machine tool to perform machining. The decision return calculation module is used to first calculate the dimensionless stress field health index based on the three-dimensional stress field, combined with preset weight hyperparameters, workpiece geometric volume and material yield strength. The calculation of the stress field health index includes: quantifying the residual stress gain effect in the target area of ​​the gear and the von Mises equivalent stress damage effect in the key area of ​​the blade, respectively; and weighting and summing the gain effect and damage effect based on the weight hyperparameter. The decision reward calculation module is further used to: construct the final scalar reward value based on the stress field health index and by introducing a hard constraint penalty term for plastic deformation; The introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the blade critical region predicted by the virtual twin prediction module with a preset blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty is applied.

2. The gear machining parameter adaptive optimization system based on deep reinforcement learning according to claim 1, characterized in that, The multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

3. The gear machining parameter adaptive optimization system based on deep reinforcement learning according to claim 1, characterized in that, The core of the virtual twin prediction module is a preset downgraded proxy model.

4. The gear machining parameter adaptive optimization system based on deep reinforcement learning according to claim 1, characterized in that, The deep reinforcement learning decision module incorporates a deep Q-network agent, which learns through continuous interaction with the processing environment.

5. The gear machining parameter adaptive optimization system based on deep reinforcement learning according to claim 1, characterized in that, The optimal machining parameters include spindle speed, feed rate, and depth of cut.

Citation Information

Patent Citations

  • Long flexible blade monitoring method and system based on multi-source data fusion

    CN119982384A

  • Assembly workshop production line construction system based on digital twinning

    CN120493592A