Gear-blade co-optimization adaptive machining parameter system

By constructing a gear-blade collaborative optimization machining parameter adaptive system, multi-source data is collected in real time, the three-dimensional physical field distribution is predicted, the scalar reward value is calculated, and the optimal machining parameters are output. This solves the conflict between gear surface integrity and blade aerodynamic performance in integrated bladed disk-gear machining and achieves efficient multi-objective optimization.

CN120974865BActive Publication Date: 2026-01-23NANTONG XINSIDI ELECTROMECHANICAL CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511505049.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-23
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously optimize gear surface integrity and blade aerodynamic performance in integrated bladed disk-gear machining, resulting in a conflict between gear life and blade efficiency, which cannot be balanced.

Method used

An adaptive processing parameter system is constructed by employing a state-aware module, a virtual twin prediction module, a decision reward calculation module, and a deep reinforcement learning decision module. It collects multi-source heterogeneous data in real time, predicts the three-dimensional physical field distribution, calculates the scalar reward value, and outputs the optimal processing parameter action.

Benefits of technology

It achieves synergistic optimization of gear surface integrity and blade aerodynamic performance, overcomes the contradiction that traditional technologies cannot balance, improves the accuracy and safety of dynamic disturbance perception in the machining process, and ensures the reliability and feasibility of the optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974865B_ABST
    Figure CN120974865B_ABST
Patent Text Reader

Abstract

The gear-vane collaborative optimization machining parameter adaptive system of the present application belongs to the cross technical field of mechanical precision machining and artificial intelligence, comprising: a state perception module for real-time acquisition of multi-source heterogeneous data in the machining process and construction of a state vector of the current state of the system; a virtual twin prediction module for receiving the state vector and candidate machining parameter actions and predicting the three-dimensional physical field distribution in the future time step; a decision reward calculation module for calculating a scalar reward value for guiding decision-making according to the three-dimensional physical field distribution, the present application solves the technical contradiction that the two requirements of strengthening gear surface integrity and maintaining vane aerodynamic performance degree conflict with each other and cannot be considered in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of precision machining and artificial intelligence, specifically to an adaptive machining parameter system for gear-blade collaborative optimization. Background Technology

[0002] In the field of aero-engines, the integral bladed disk-gear structure integrates blades, disk, and transmission gears into a single unit. This structure offers significant advantages in improving thrust-to-weight ratio and reducing structural weight, but it also presents unprecedented manufacturing challenges. The integrated structure leads to conflicting performance requirements among different functional areas. To ensure high fatigue life, the gear area requires a specific residual compressive stress distribution, low roughness, and high microhardness on the machined surface—that is, excellent surface integrity. Conversely, to ensure aerodynamic efficiency, the blade area requires minimal deformation during machining to maintain precise aerodynamic performance.

[0003] Traditional machining methods typically employ fixed-parameter strategies based on expert experience or offline simulation, with the optimization goal of achieving the best quality for a single process or objective. In gear finishing, to obtain ideal surface integrity, combinations of cutting parameters that introduce significant residual compressive stress are often used. These parameters often generate enormous cutting forces and heat, which are transmitted through the integrated structure to the less rigid blade section, inevitably causing micron-level permanent plastic deformation of the blade, damaging its profile and leading to a decrease in aerodynamic performance. This phenomenon reveals a profound technical contradiction: in the machining of integrated components, any machining strategy aimed at enhancing tooth surface performance may come at the cost of sacrificing the retention of blade aerodynamic performance; this contradiction becomes even more acute under the influence of dynamic variables such as tool wear. The conventional practice of increasing cutting loads to compensate for wear directly exacerbates blade damage. Existing technologies cannot proactively avoid this problem and can only passively compromise between the two conflicting goals of gear life and blade efficiency.

[0004] The technical problem that this invention aims to solve is to provide a machining parameter adaptive optimization system that can overcome the above-mentioned technical contradictions in the integrated machining scenario of an integral bladed disk-gear.

[0005] The information disclosed in the background section above is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide an adaptive machining parameter system for gear-blade collaborative optimization, in order to solve the problems mentioned in the background art; specifically, the technical solution of this invention includes:

[0007] The state awareness module is used to collect multi-source heterogeneous data in real time during the processing and construct it into a state vector of the current system state.

[0008] The virtual twin prediction module is used to receive the state vector and candidate processing parameters, and predict the three-dimensional physical field distribution in the future time step;

[0009] The decision reward calculation module is used to calculate scalar reward values ​​to guide decision-making based on the three-dimensional physical field distribution.

[0010] The deep reinforcement learning decision module is used to take the state vector as input and, based on the scalar reward value as the learning objective, output the optimal processing parameters and actions.

[0011] The machining parameter execution module is used to receive the optimal machining parameter actions and convert them into CNC instructions to drive the machine tool to perform machining.

[0012] Preferably, the multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

[0013] Preferably, the core of the virtual twin prediction module is a preset reduced-order proxy model; the three-dimensional physical field distribution includes a three-dimensional stress field, a three-dimensional strain field, and a three-dimensional temperature field.

[0014] Preferably, the decision reward calculation module is used to first calculate a dimensionless stress field health index based on the three-dimensional stress field and in combination with preset weight hyperparameters, workpiece geometric volume and material yield strength.

[0015] Preferably, the calculation of the stress field health index includes: quantifying the residual stress gain effect in the target area of ​​the gear and the von Mises equivalent stress damage effect in the key area of ​​the blade, respectively; and weighting and summing the gain effect and damage effect based on the weight hyperparameter.

[0016] Preferably, the decision reward calculation module is further used to: construct the final scalar reward value based on the stress field health index and by introducing a hard constraint penalty term for plastic deformation.

[0017] Preferably, the introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the blade critical region predicted by the virtual twin prediction module with a preset blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty is applied.

[0018] Preferably, the deep reinforcement learning decision module incorporates a deep Q-network agent, which learns through continuous interaction with the processing environment.

[0019] Preferably, the optimal machining parameters include spindle speed, feed rate, and depth of cut.

[0020] This invention provides an improved adaptive system for gear-blade collaborative optimization of machining parameters, which, compared with the prior art, has the following improvements and advantages:

[0021] 1. This technical solution constructs a closed-loop control architecture that integrates state perception, virtual twin prediction, decision reward calculation, deep reinforcement learning decision-making, and processing parameter execution. This successfully transforms a multi-objective, multi-constraint nonlinear optimization problem into a sequential decision problem that can be solved efficiently. It also solves the technical contradiction in the existing technology where the two requirements of enhancing the integrity of the gear surface and maintaining the aerodynamic performance of the blades are conflicting and cannot be simultaneously met during the processing of integrated components.

[0022] 2. The system's state perception module, by integrating multi-source heterogeneous data such as cutting force, cutting zone temperature, machining vibration signals, and tool wear, can construct a more comprehensive and robust machining state profile than a single signal source. The introduction of this multi-dimensional physical information greatly improves the system's perception accuracy of dynamic disturbances in the machining process, laying a solid foundation for subsequent accurate prediction and decision-making.

[0023] 3. The system's virtual twin prediction module, with a preset reduced-order proxy model as its core, breaks through the technical bottleneck that high-fidelity physical simulation models cannot be used for real-time control due to computational time. This module can predict the invisible three-dimensional stress field, three-dimensional strain field, and three-dimensional temperature field distribution inside the workpiece with high precision at the millisecond level. This ability to predict future physical states gives the entire control system foresight, enabling it to actively avoid potential processing damage, which is a key prerequisite for synergistically optimizing multiple conflicting performance objectives.

[0024] 4. The system's decision reward calculation module establishes an innovative evaluation paradigm that incorporates processing objectives. It transforms the high-dimensional, complex physical field distribution into a low-dimensional scalar reward value. By calculating a dimensionless stress field health index, it maps and unifies conflicting macroscopic processing objectives—namely, the residual stress gain effect of gears and the equivalent stress damage effect of blades—onto a calculable and optimizable microscopic physical quantity, providing a physically clear and smooth optimization direction for intelligent decision-making. Furthermore, this module introduces a hard constraint penalty term for plastic deformation. By comparing the predicted maximum equivalent plastic strain with a material safety threshold with a clear physical source, it defines an absolute safety boundary for the system's learning and exploration process. This ensures that any risk leading to permanent damage to the workpiece is avoided with the highest priority during the search for the optimal solution, significantly improving the safety and reliability of the optimization process.

[0025] 5. The system's deep reinforcement learning decision module incorporates a deep Q-network agent and employs a model-independent learning method. This allows the system to learn autonomously and discover superior processing strategies through trial and error and experience accumulation in interaction with the environment, without relying on precise analytical mathematical models of the processing process. This endows the system with the ability to handle highly nonlinear, strongly coupled, and complex processing scenarios, giving it strong adaptability and robustness.

[0026] 6. The system's machining parameter execution module seamlessly converts the optimal machining parameters output by the intelligent agent—namely, the core process parameters such as spindle speed, feed rate, and depth of cut—into executable instructions for the CNC machine tool. This opens up a complete pathway from high-level intelligent decision-making to low-level physical control, ensuring the ultimate executability of the optimization results and enabling the entire adaptive optimization closed loop to be effectively realized. Attached Figure Description

[0027] The present invention will be further explained below with reference to the accompanying drawings and embodiments:

[0028] Figure 1 This is a flowchart of the system of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. Example 1

[0030] Please see Figure 1 This invention provides a gear-blade collaborative optimization machining parameter adaptive system, comprising:

[0031] The state awareness module is used to collect multi-source heterogeneous data in real time during the processing and construct it into a state vector of the current system state.

[0032] The virtual twin prediction module is used to receive the state vector and candidate processing parameters, and predict the three-dimensional physical field distribution in the future time step;

[0033] The decision reward calculation module is used to calculate scalar reward values ​​to guide decision-making based on the three-dimensional physical field distribution.

[0034] The deep reinforcement learning decision module is used to take the state vector as input and, based on the scalar reward value as the learning objective, output the optimal processing parameters and actions.

[0035] The machining parameter execution module is used to receive the optimal machining parameter actions and convert them into CNC instructions to drive the machine tool to perform machining.

[0036] An adaptive machining parameter system for gear-blade collaborative optimization aims to resolve the technical contradiction between enhancing gear surface integrity and maintaining blade aerodynamic performance in integrated component machining. In this embodiment, the system achieves real-time adaptive control of machining parameters by constructing a closed-loop control process of perception-prediction-decision-execution. The system includes a state perception module, a virtual twin prediction module, a decision reward calculation module, a deep reinforcement learning decision module, and a machining parameter execution module.

[0037] The state perception module aims to capture dynamic physical information of the processing process in real time, providing a basis for system decision-making. In this embodiment, the module continuously collects multi-source heterogeneous data reflecting the current processing state through an integrated sensor network, and integrates it into a standardized vectorized representation, namely the state vector. This state vector is the input basis for all subsequent modules to perform calculations and decisions.

[0038] The virtual twin prediction module aims to provide a millisecond-level prediction capability for future processing results, thereby achieving forward-looking control. In this embodiment, the module receives a state vector output by the state perception module and candidate processing parameters provided by the deep reinforcement learning decision module as joint inputs. Based on this input, the module can quickly predict the three-dimensional physical field distribution of the entire workpiece within a very short time step in the future, thereby making invisible internal stress, strain and other information explicit, and providing data support for quantitatively evaluating the impact of processing.

[0039] The decision reward calculation module aims to transform high-dimensional, complex physical field prediction information into low-dimensional scalar reward values ​​that can guide reinforcement learning agents in their learning. In this embodiment, the module calculates dimensionless scalar values ​​based on the three-dimensional physical field distribution output by the virtual twin prediction module, using a set of preset mathematical paradigms that contain the processing technology objectives. The design of this reward value shifts the optimization objective from the traditional final processing quality to the dynamic control of the physical field health during the processing.

[0040] The deep reinforcement learning decision module aims to serve as the core decision-making unit of the system, autonomously learning and generating the optimal processing strategy. In this embodiment, the module uses the state vector provided by the state perception module as the observation input to the environment, and the scalar reward value generated by the decision reward calculation module as the target signal for learning and optimization. Through continuous interaction and trial and error with the processing environment, the module can learn the mapping relationship from the current state to the optimal processing parameter action.

[0041] The purpose of the machining parameter execution module is to transform the abstract parameter instructions output by the decision module into specific operations that the machine tool can perform in the physical world. In this embodiment, the module serves as the interface for connecting digital decision-making with physical machining. It receives the optimal machining parameter actions output by the deep reinforcement learning decision module, compiles them into standard CNC instructions, and then drives the CNC machine tool to complete the corresponding cutting actions, thereby completing the entire closed-loop control.

[0042] This embodiment combines complex machining mechanisms with efficient intelligent decision-making through the coordinated work of the above five modules, successfully transforming a multi-objective, multi-constraint nonlinear optimization problem into a sequential decision problem that can be solved efficiently. The overall technical effect is that it can control machining parameters in real time and predictively, and while dynamically compensating for disturbances such as tool wear, it can synergistically improve the surface integrity of gears and the profile retention of blades, overcoming the technical contradiction that the two cannot be balanced in the prior art. Example 2

[0043] Multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

[0044] In this embodiment, the multi-source heterogeneous data collected by the state perception module is specified. Multi-source heterogeneous data refers to a set of physical signals from different sources and of different types, but which together characterize the processing state. Its purpose is to provide the system with a comprehensive and accurate picture of the current processing environment. In this embodiment, the data set specifically includes:

[0045] Cutting force: Its function is to characterize the intensity of the interaction between the tool and the workpiece. It is obtained in real time by a multi-component force sensor installed on the spindle or workpiece fixture.

[0046] Cutting zone temperature: Its function is to reflect the thermal load level of the machining area. It is obtained by measuring through miniature thermocouples embedded near the workpiece or tool, or by non-contact infrared thermal imager.

[0047] Vibration signals during machining: Their function is to monitor the stability of the machining process. They are collected by accelerometers or acoustic emission sensors installed in key parts of the machine tool.

[0048] Tool wear: Its function is to quantify the degree of tool wear. This is a key dynamic variable, which is obtained through periodic evaluation by online monitoring systems such as machine vision.

[0049] To ensure the accuracy and consistency of the collected data, the sensor network configuration in this embodiment is as follows: the multi-component force sensor is a triaxial piezoelectric force sensor mounted on the machine tool turret, used to measure... Cutting forces in three directions; the accelerometer is a MEMS accelerometer mounted on the spindle housing to monitor high-frequency vibrations during machining; a miniature thermocouple is integrated into the tool holder near the cutting area; the machine vision system for assessing tool wear consists of an industrial camera with a ring light source fixed inside the machine tool guard, which takes pictures of the tool tip between each machining cycle and calculates the width of the flank wear band using image processing algorithms;

[0050] Considering the latency of periodic machine vision measurements, this module further integrates a tool wear prediction sub-model based on a recurrent neural network. This sub-model is trained using historical wear data and continuous cutting force and vibration signal features, enabling it to estimate tool wear in real time and continuously during the intervals between two direct visual measurements. This estimate will be fused and updated with periodic measurements, thereby providing a more rapid and delay-free wear state representation for the state vector, significantly improving the system's ability to anticipate and adapt to situations such as sudden tool breakage.

[0051] Combine these data into a single state vector:

[0052]

[0053] in, : The state vector of the current state of the system; Real-time measurement of tool wear; This refers to the cutting temperature. and These represent key feature scalars extracted from real-time acquired cutting force and vibration signals, such as root mean square (RMS) values. The cutting force and vibration signals undergo feature extraction, such as calculating their RMS values ​​within the sampling window, to obtain corresponding scalar values. This provides the system with richer and more robust state information than a single signal source.

[0054] To ensure the comprehensiveness of the state vector, feature extraction of cutting force and vibration signals goes beyond the root mean square value in the time domain. Frequency domain features are also introduced, such as calculating the spectral entropy and dominant frequency amplitude using Fast Fourier Transform (FFT) to accurately identify unstable states like chatter during machining. To ensure the effectiveness of data fusion, the system specifies the synchronization mechanism and sampling frequency for each sensor. Force and vibration signals are synchronously acquired at a high frequency of at least 10kHz via hardware triggering, while slowly varying signals such as temperature and wear are updated at a lower frequency. Through timestamp alignment and interpolation algorithms, these signals are ultimately fused into a unified state vector. ;

[0055] This embodiment greatly improves the system's perception accuracy of machining status, especially dynamic disturbances such as tool wear, by introducing multi-dimensional physical signals, laying a solid foundation for subsequent accurate prediction and decision-making. Example 3

[0056] The core of the virtual twin prediction module is a preset reduced-order proxy model; the three-dimensional physical field distribution includes a three-dimensional stress field, a three-dimensional strain field, and a three-dimensional temperature field.

[0057] In this embodiment, the internal structure and output content of the virtual twin prediction module are specified. The core of the virtual twin prediction module is a preset reduced-order proxy model. The preset reduced-order proxy model refers to a lightweight neural network model that has been pre-trained and can quickly reproduce the input-output relationship of a high-fidelity physical simulation model. The working principle is as follows: a multiphysics finite element model with huge computational load but extremely high accuracy is constructed; through model reduction techniques such as intrinsic orthogonal decomposition, the key dynamic characteristics of the high-fidelity model are extracted. These characteristics are used to train a neural network, enabling it to simulate calculations that would take hours for a high-fidelity model to complete at millisecond speeds. This preset method allows it to be embedded into the real-time control loop.

[0058] In this embodiment, the reduced-order surrogate model is specifically a fully connected feedforward neural network; the network includes an input layer, three hidden layers, and an output layer; the input layer receives data from the state vector. and candidate action vectors The features are assembled from the data; the three hidden layers have 256, 128, and 64 neurons respectively, all using modified linear units as activation functions. The output layer outputs a flattened one-dimensional vector, which can be reconstructed into discrete node values ​​of three-dimensional stress, strain, and temperature fields. The model is trained using the Adam optimizer, with an initial learning rate set to... The batch size is 64. The training dataset consists of 10,000 input-output pairs generated by high-fidelity finite element software such as ABAQUS simulation, and the simulation input machining parameters are fully randomly sampled within their process window;

[0059] The three-dimensional physical field distribution output by this module includes:

[0060] Three-dimensional stress field: characterizes the distribution of internal forces at various points inside a workpiece under load;

[0061] Three-dimensional strain field: characterizes the degree of deformation at various points inside the workpiece;

[0062] Three-dimensional temperature field: characterizes the temperature distribution inside the workpiece;

[0063] This embodiment solves the problem that high-fidelity physical simulation cannot be used for real-time control by adopting a reduced-order proxy model, and realizes high-speed and high-precision prediction of multiple physical fields inside the processing process. This enables the system to make decisions based on the prediction of future stress, strain and temperature, thereby realizing predictive control, which is a key prerequisite for achieving the goals of actively avoiding damage and collaboratively optimizing performance.

[0064] To address the model drift issue that may arise in pre-trained models when faced with changes in actual operating conditions, this system introduces an online calibration and uncertainty assessment mechanism. This mechanism periodically compares the key physical quantities predicted by the twin model with the actual measurements from the state-aware module, calculating the prediction residuals. These residual vectors serve as an additional dimension, along with the original state vector. These parameters are input into the deep reinforcement learning decision-making module. This allows the agent to not only perceive the current physical state but also dynamically perceive the prediction accuracy and confidence level of the virtual twin model. As a result, when the model fails, it can autonomously adjust its decision-making strategy and select more robust processing parameters, ensuring the robustness of the entire closed-loop system in real and complex environments. Example 4

[0065] The decision return calculation module is used to first calculate the dimensionless stress field health index based on the three-dimensional stress field, combined with preset weight hyperparameters, workpiece geometry and material yield strength.

[0066] The calculation of the stress field health index includes: quantifying the residual stress gain effect in the target area of ​​the gear and the von Mises equivalent stress damage effect in the key area of ​​the blade, respectively; and weighting and summing the gain effect and damage effect based on the weight hyperparameter.

[0067] In this embodiment, the calculation logic of the decision return calculation module is specified; the calculation logic of this module is to calculate a dimensionless stress field health index based on the three-dimensional stress field output by the virtual twin prediction module; the dimensionless stress field health index, This refers to the scalar index that integrates the two conflicting physical objectives of gear gain and blade damage into a unified evaluation system. Its function is to provide continuous and smooth optimization objectives for reinforcement learning agents.

[0068] The calculation of this index quantifies the effects of the two key regions, gears and blades, separately and then performs a weighted summation. To achieve this goal, the calculation formula is defined as follows:

[0069]

[0070] in, : The stress field health index at time t is a dimensionless scalar, calculated by this module;

[0071] The dimensionless weighted hyperparameters represent the relative importance of tooth surface gain and blade damage, respectively. They are obtained through multi-objective optimization algorithms or by inverse reinforcement learning optimization from successful machining cases.

[0072] In this embodiment, a grid search method is used to determine... and The value; we are in a preset range, for example, , The test is performed with a step size of 0.1, and each test case is evaluated using a set of standard offline test conditions. The final convergence performance of the reinforcement learning agent under the given parameters is determined by selecting the parameter combination that most quickly and stably balances the residual compressive stress of the gear and the equivalent stress of the blade. In this embodiment, the selected value is... and ;

[0073] The yield strength of a material is a dimensionless characteristic stress that is used as a reference physical quantity to treat the system. It is obtained through standard material tensile tests.

[0074] These are the volumes of the target machining area of ​​the gear and the key aerodynamic profile of the blade, respectively, which are extracted directly from the computer-aided design model of the workpiece.

[0075] The residual stress tensor of the tooth surface region predicted by the virtual twin module is used to quantify the residual stress gain effect.

[0076] The tooth surface stress evaluation function, in this embodiment, is defined as:

[0077]

[0078] in, First principal stress; this definition allows for favorable residual compressive stress. Generate positive rewards;

[0079] The von Mises equivalent stress in the blade region, predicted by the virtual twin module, is used to quantify the von Mises equivalent stress damage effect.

[0080] This embodiment establishes... This core indicator transforms macroscopic and conflicting processing objectives, such as gear life and blade precision, into calculable and optimizable microscopic process physical quantities. This provides intelligent agents with clear and physically meaningful optimization directions, making collaborative optimization possible. Example 5

[0081] The decision reward calculation module is further used to: construct the final scalar reward value based on the stress field health index and by introducing a hard constraint penalty term for plastic deformation;

[0082] The introduction of hard constraint penalty terms includes: comparing the maximum equivalent plastic strain of the blade critical region predicted by the virtual twin prediction module with a preset blade material plastic deformation safety threshold; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty is applied.

[0083] In this embodiment, the calculation logic of the decision reward calculation module has been further improved. Based on the stress field health index, the module introduces a hard constraint penalty term for plastic deformation to construct the final scalar reward value. The hard constraint penalty term refers to a mechanism that provides a huge negative incentive when the behavior of the system touches a certain insurmountable safety boundary. Its function is to force the agent to prioritize the structural integrity of the workpiece under any circumstances.

[0084] The final scalar return value The calculation formula is defined as follows:

[0085]

[0086] in, : The final reward value received by the agent at each moment is a dimensionless scalar, calculated by this module, and serves as the learning target of the deep reinforcement learning decision module.

[0087] The stress field health index is calculated using the aforementioned formula.

[0088] The dimensionless plastic deformation penalty factor is set to a value much greater than... A positive absolute value of expectation comes from hyperparameter optimization and represents a severe penalty for behavior that crosses the damage threshold.

[0089] To ensure the effectiveness of punishment, The value is set to be at least greater than The theoretical maximum value is two orders of magnitude higher. Considering After dimensionless processing, its absolute value is usually less than 10. In this embodiment, It is directly set to a fixed, large value, for example ;

[0090] : The maximum equivalent plastic strain in the critical region of the blade predicted by the virtual twin prediction module;

[0091] : The preset safety threshold for plastic deformation of the blade material; this is a parameter with a clear physical source, and the determination process is as follows: by conducting a uniaxial tensile test on the high-temperature alloy used in the part, its accurate stress-strain curve is obtained; combined with the accurate three-dimensional model of the part, finite element simulation analysis is performed to calibrate the maximum equivalent plastic strain value of the blade surface without permanent deformation.

[0092] The Heaviside step function works like a switch; when the internal parameters... The function value is 1 when the value is greater than zero, and 0 otherwise.

[0093] This embodiment introduces a hard constraint penalty term based on a clear physical threshold, which defines an absolute safety boundary for the system's exploration and learning process. This ensures that the system will avoid any risks that may cause permanent plastic deformation of the blades with the highest priority when seeking the optimal processing parameters, greatly improving the safety and reliability of the optimization process. Example 6

[0094] The deep reinforcement learning decision-making module incorporates a deep Q-network agent, which learns through continuous interaction with the processing environment.

[0095] In this embodiment, the internal implementation of the deep reinforcement learning decision-making module is specified. This module incorporates a deep Q-network agent. The deep Q-network agent is a deep reinforcement learning algorithm that uses a deep neural network to approximate the optimal action-value function, the Q-function, thereby enabling it to handle high-dimensional state input spaces. This agent learns through continuous interaction with the processing environment, that is, it continuously adjusts its behavior based on the current state. Output Action To reap rewards from the environment and the next state It uses this empirical data to iteratively update the weights of its internal neural network, enabling it to produce actions that yield the greatest long-term reward for any given state output.

[0096] In this embodiment, the deep Q-network is specifically a neural network containing four fully connected layers; the input layer receives the state vector. The subsequent two hidden layers contain 128 and 64 neurons respectively, and employ the ReLU activation function; the output layer has... A linear layer of neurons, wherein It represents the total number of actions in the discretized action space, with each neuron corresponding to the Q-value of a discrete action; during training, it uses... - Use a greedy strategy to explore, where The learning rate decreases linearly from 1.0 to 0.01. The experience replay pool size is 20,000, the batch size during training is 128, and the learning rate is... And use the Adam optimizer for gradient descent;

[0097] This embodiment employs model-agnostic reinforcement learning methods such as deep Q-networks, which eliminates the need for the system to establish precise, analytical mathematical models of the processing procedures. The agent can autonomously discover superior processing strategies through trial and error and experience accumulation, just like human experts. This enables it to handle highly nonlinear and strongly coupled complex processing scenarios, exhibiting strong adaptability and robustness. Example 7

[0098] Optimal machining parameters include spindle speed, feed rate, and depth of cut;

[0099] In this embodiment, the content of the optimal machining parameter action output by the deep reinforcement learning decision module is defined; the optimal machining parameter action refers to a set of machine tool control parameters decided by the agent in the current state that maximizes long-term rewards; in this embodiment, this action is defined as a vector:

[0100]

[0101] in, Candidate action vector, or optimal processing parameter action;

[0102] Spindle speed: controls the speed at which the tool rotates;

[0103] Feed rate: controls the speed at which the tool moves relative to the workpiece;

[0104] : Depth of cut, which controls the depth to which the tool penetrates the workpiece in a single cut;

[0105] These three parameters are the most critical and directly controllable process parameters in CNC machining. Their combination directly determines the cutting force, the generation of cutting heat, and the final machining quality.

[0106] Since the standard depth Q-network algorithm is applicable to discrete action spaces, while the processing parameters in this embodiment... Spindle speed, feed rate and Since the cutting depth is a continuous variable, the motion space needs to be discretized. We divide each parameter into several levels within its process-allowed range; for example, the spindle speed... Discretized into 5 levels within the range of [6000, 12000] rpm; feed rate It is discretized into 5 levels within the range of [500, 1500] mm / min; cutting depth It is discretized into 4 levels within the range of [0.1, 0.5] mm; therefore, the total size of the discrete action space is... The deep reinforcement learning decision-making module outputs the optimal index among these 100 combined actions, and the processing parameter execution module then decodes this index into specific actions. Parameter combination instructions;

[0107] This embodiment directly correlates the system's decision output with the core control parameters at the machine tool's physical execution level. This allows the system's optimization results to be seamlessly converted into specific CNC instructions by the machining parameter execution module, ensuring the executability of the decisions and forming a complete pathway from high-level decision-making to low-level control, thus enabling the entire adaptive optimization closed loop to be realized.

[0108] To ensure the robustness of the system in practical applications, this system also includes a monitoring and safety module. This module is responsible for detecting abnormal values ​​in sensor data and triggering safety strategies when the data exceeds a preset reasonable physical range. For example, it may switch to a set of conservative default processing parameters to prevent abnormal input from leading to incorrect decisions. At the same time, it also imposes boundary restrictions on the output actions of the agent's decisions to ensure that they always remain within the process window allowed by the machine tool.

[0109] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A gear-blade collaborative optimization machining parameter adaptive system, characterized in that, include: The state awareness module is used to collect multi-source heterogeneous data in real time during the processing and construct it into a state vector of the current system state. The virtual twin prediction module is used to receive the state vector and candidate processing parameters, and predict the three-dimensional physical field distribution in the future time step; The decision reward calculation module is used to calculate scalar reward values ​​to guide decision-making based on the three-dimensional physical field distribution. The deep reinforcement learning decision module is used to take the state vector as input and, based on the scalar reward value as the learning objective, output the optimal processing parameters and actions. The machining parameter execution module is used to receive the optimal machining parameter actions and convert them into CNC instructions to drive the machine tool to perform machining. The decision return calculation module is used to first calculate the dimensionless stress field health index based on the three-dimensional stress field, combined with preset weight hyperparameters, workpiece geometric volume and material yield strength. The calculation of the stress field health index includes: quantifying the residual stress gain effect in the target area of ​​the gear and the von Mises equivalent stress damage effect in the key area of ​​the blade, respectively; The gain effect and damage effect are then summed using a weighted hyperparameter. The decision reward calculation module is further used to: construct the final scalar reward value based on the stress field health index and by introducing a hard constraint penalty term for plastic deformation.

2. The gear-blade collaborative optimization machining parameter adaptive system according to claim 1, characterized in that, The multi-source heterogeneous data includes cutting force, cutting zone temperature, machining vibration signal, and tool wear.

3. The gear-blade collaborative optimization machining parameter adaptive system according to claim 1, characterized in that, The core of the virtual twin prediction module is a preset reduced-order proxy model; the three-dimensional physical field distribution includes a three-dimensional stress field, a three-dimensional strain field, and a three-dimensional temperature field.

4. The gear-blade collaborative optimization machining parameter adaptive system according to claim 1, characterized in that, The introduction of the hard constraint penalty term includes: comparing the maximum equivalent plastic strain of the critical region of the blade predicted by the virtual twin prediction module with a preset safety threshold for plastic deformation of the blade material; when the maximum equivalent plastic strain is greater than the safety threshold, a preset penalty factor is applied, otherwise no penalty factor is applied.

5. The gear-blade collaborative optimization machining parameter adaptive system according to claim 1, characterized in that, The deep reinforcement learning decision module incorporates a deep Q-network agent, which learns through continuous interaction with the processing environment.

6. The gear-blade collaborative optimization machining parameter adaptive system according to claim 1, characterized in that, The optimal machining parameters include spindle speed, feed rate, and depth of cut.

Citation Information

Patent Citations

  • Dynamic deformation distribution optimization method for large-specification bar rolling integrated with digital twinning

    CN120306402A