Numerical control machining parameter optimization method and device based on deep reinforcement learning
Through deep reinforcement learning, the optimization of CNC machining parameters is solved, and the problem of parameter selection depends on experience in traditional methods is realized, and automated and intelligent optimization under complex conditions is achieved, processing quality and efficiency is improved, and energy consumption and labor costs are reduced.
Patent Information
- Application Number
- CN202510391733.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
AI Technical Summary
The selection of traditional CNC machining parameters depends on experience or heuristic algorithms, making it difficult to achieve global optimization under complex machining conditions, and lacks real-time feedback mechanisms, resulting in unstable processing quality and limited efficiency improvement.
The CNC machining parameter optimization method based on deep reinforcement learning is adopted. By obtaining real-time data, using particle swarm optimization algorithm and deep reinforcement learning model, the cutting speed, feed speed and cutting depth are dynamically adjusted, the optimization objective function is constructed, and the parameters are automatically optimized.
It improves the degree of automation and optimization efficiency of the processing process, and can adaptively optimize processing parameters under different materials and complex working conditions, ensure processing quality and efficiency, reduce manual intervention, reduce energy consumption and production costs.
Smart Images

Figure CN120353190A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to numerical control machining, and in particular to a method and device for optimizing numerical control machining parameters based on deep reinforcement learning. Background Art
[0002] Numerical control machining refers to a processing method for machining parts on a numerically controlled machine tool. The processing procedures of numerically controlled machine tools and traditional machine tools are generally the same, but obvious changes have also occurred. It is a mechanical processing method that uses digital information to control the displacement of parts and tools. It is an effective way to solve problems such as variable part varieties, small batch sizes, complex shapes, high precision, and to achieve high-efficiency and automated processing.
[0003] During the numerical control machining process, cutting force, material removal rate, and tool life are key indicators for measuring machining efficiency and effectiveness. However, the selection of traditional machining parameters mainly relies on experience or heuristic algorithms. This method is difficult to achieve the global optimal selection of parameters under complex machining conditions. Relying on manual debugging or empirical formulas, it is difficult to quickly adapt to different machining tasks and complex machining environments. Under changing machining conditions, such as the influence of factors such as different materials, tool wear, and temperature fluctuations, traditional optimization methods are difficult to adjust machining parameters in real time, easily resulting in unstable machining quality, lacking an automated optimization ability based on a real-time feedback mechanism, and limiting the improvement space of machining efficiency and product quality. Summary of the Invention
[0004] The purpose of the present invention is to at least solve one of the deficiencies of the prior art, and provide a method and device for optimizing numerical control machining parameters based on deep reinforcement learning.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] Specifically, a method for optimizing numerical control machining parameters based on deep reinforcement learning is proposed, including the following:
[0007] Obtain real-time data of numerical control machining, and perform analog-to-digital conversion and preprocessing on the real-time data to obtain preprocessed data;
[0008] Initialize machining parameters, where the machining parameters include cutting speed v c , feed rate f, and cutting depth a p ;
[0009] Determine the optimization goal, and construct an optimization objective function according to the optimization goal. The optimization goal includes simultaneously satisfying minimizing machining time T, minimizing surface roughness Q s , and reducing energy consumption C during the cutting process;
[0010] Calculate the current preliminary combination of machining parameters [v c , f, a p using the particle swarm optimization algorithm;
[0011] Input the combination of machining parameters [v c , f, a p into the pre-trained deep reinforcement learning model, and output the optimal combination of machining parameters through the action of the pre-trained deep reinforcement learning model
[0012] Use the optimal combination of machining parameters to perform numerical control machining control.
[0013] Furthermore, specifically, construct the optimization objective function according to the optimization objective as follows:
[0014] F(x) = α·Q s + β·T + γ·C;
[0015] where α, β, γ are pre-determined objective weights, which depend on the priority of the machining task.
[0016] Furthermore, specifically, calculate the current preliminary combination of machining parameters [v c , f, a p using the particle swarm optimization algorithm, including:
[0017] Record the position of the particle as the machining parameters [v c , f, a p , and the velocity of the particle is the moving direction and amplitude of the particle in the search space. The global optimal solution g best is the best position of all current particles in history, and the local optimal solution P best represents the best position of a certain particle itself in history.
[0018] For each particle i, update its velocity and position based on the following formula:
[0019] v i,d (t + 1) = ω·v i,d (t) + c1·r1·(p best,i,d - x i,d (t)) + c2·r2·(g best,d - x i,d (t)),
[0020] v i,d (t): The velocity of particle i in dimension d, which includes cutting speed, feed speed, and cutting depth;
[0021] w: Inertia weight, which controls the balance of particle movement;
[0022] c1, c2: Learning factors, representing the degree to which particles learn from themselves and the global optimal solution;
[0023] r1, r2: Random numbers within the range of [0, 1], used to introduce randomness;
[0024] p best,i,d : The historical best position of particle i;
[0025] g best,d : The global best position;
[0026] Update position:
[0027] x i,d (t + 1) = x i,d (t) + v i,d (t + 1),
[0028] Compare the fitness value F(x) of the current particle with its historical best value p best , and update the historical best position of the particle, where the fitness function F(x) is designed based on the optimization objective function J:
[0029] F(x) = -J(x) = -(w1·T(x) + w2·Q(x) + w3·C(x));
[0030] Compare the P of all particles best , and update the global best position g best ;
[0031] If the maximum number of iterations T max is reached, or the change in the global optimal solution g best tends to be stable, that is, the corresponding error is less than the threshold ∈, then terminate the optimization, otherwise perform the next round of iteration.
[0032] After multiple rounds of iteration, the particle swarm converges to the global optimal solution g best , and obtain the optimized combination of machining parameters [v c , f, a p .
[0033] Furthermore, specifically, the value of ω is 1.0, and the values of c1 and c2 are 2.0.
[0034] Furthermore, specifically, the reward function of the pre-trained deep reinforcement learning model is
[0035] R(s, a) = α 1 ·Q s + β 1 ·T + γ 1 ·C;
[0036] Among them, s is the state space, which consists of the machining environment of the CNC lathe, namely the cutting speed, feed rate, and cutting depth features. The action space a is the adjustment of machining parameters, that is, the amplitude of increasing or decreasing the cutting speed, feed rate, and cutting depth;
[0037] The pre-trained deep reinforcement learning model includes an input layer: the state vector S, a hidden layer: a multi-layer fully connected network containing the ReLU activation function, and an output layer: the action value Q(S,A), corresponding to the expected return for each action.
[0038] Furthermore, specifically, the optimal machining parameter combination is output through the action of the pre-trained deep reinforcement learning model including,
[0039] Initialize a replay pool of a fixed size to store experience samples (s t , a t , r t , s t+1 ), randomly initialize the weights of the behavior network and the target network, and load the current preliminary machining parameter combination [v c , f, a p as a benchmark;
[0040] In the sampling stage, use the current state s t , and use the behavior network to select the action a t :
[0041]
[0042] ∈: exploration rate, adopt the ∈-greedy strategy,
[0043] According to the action a t Adjust the machining parameters, observe the next state s t+1 and the reward r t ,
[0044] Store the sample (s t , a t , r t , s t+1 ) into the experience replay pool, randomly sample a small batch of experiences (s i , a i , r i , s i+1 ) from the replay pool,
[0045] The target value is calculated based on the target network:
[0046]
[0047] where γ is the discount factor, measuring the importance of future rewards,
[0048] The network update uses the mean squared error (MSE) to optimize the behavior network:
[0049]
[0050] where N represents the batch size,
[0051] Every fixed number of steps, the weights of the behavior network are copied to the target network. If the maximum number of iterations is reached or the reward converges, the training stops, and finally the optimal processing parameters after training are output:
[0052] The present invention also proposes a device for optimizing numerical control processing parameters based on deep reinforcement learning, including the following:
[0053] A data acquisition module, which is used to obtain real-time data of numerical control processing, and perform analog-to-digital conversion and preprocessing on the real-time data to obtain preprocessed data;
[0054] An initialization module, which is used to initialize processing parameters, and the processing parameters include cutting speed v c , feed speed f, and cutting depth a p ;
[0055] An optimization model module, which is used to determine an optimization goal and construct an optimization objective function according to the optimization goal. The optimization goal includes simultaneously satisfying minimizing the processing time T, minimizing the surface roughness Q s , reducing the energy consumption C during the cutting process; calculating the current preliminary combination of processing parameters [v c , f, a p based on the particle swarm optimization algorithm;
[0056] A deep reinforcement learning module, which is used to input the combination of processing parameters [v c , f, a p into a pre-trained deep reinforcement learning model, and output the optimal combination of processing parameters through the action of the pre-trained deep reinforcement learning model
[0057] A decision-making module, which is used to perform numerical control processing control with the optimal combination of processing parameters .
[0058] The beneficial effects of the present invention are:
[0059] The present invention proposes a numerical control machining parameter optimization method and device based on deep reinforcement learning. Through the deep reinforcement learning algorithm, it automatically optimizes the machining parameters of a numerically controlled lathe, reduces manual intervention, and improves the automation degree and optimization efficiency of the machining process. It can dynamically adjust the optimization strategy according to the real-time machining state, adaptively optimize the machining parameters under complex working conditions such as different materials, tool wear, and temperature changes, ensure the machining quality and efficiency, and continuously adjust and optimize the machining parameters based on real-time feedback information (such as cutting force, temperature, etc.) through the reinforcement learning algorithm, so as to realize the intelligent and automatic optimization of the machining process of the numerically controlled lathe. Real-time machining data is collected through sensors and combined with the deep reinforcement learning model to dynamically adjust the process parameters of the lathe to ensure that the machining process is always in the best state. Brief Description of the Drawings
[0060] By describing the embodiments shown in the accompanying drawings in detail, the above and other features of the present disclosure will become more obvious. The same reference numerals in the drawings of the present disclosure represent the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0061] Figure 1 It shows a flowchart of the numerical control machining parameter optimization method based on deep reinforcement learning of the present invention;
[0062] Figure 2 It shows a schematic diagram of an embodiment of the sensor arrangement when the present invention is applied;
[0063] Figure 3 It shows a flowchart of the present invention for outputting the optimal machining parameter combination through the action of a pre-trained deep reinforcement learning model. Detailed Embodiments
[0064] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention in combination with the embodiments and the drawings to fully understand the purpose, solution and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The same reference numerals used throughout the drawings indicate the same or similar parts.
[0065] Embodiment 1, referring to Figure 1 , the present invention proposes a numerical control machining parameter optimization method based on deep reinforcement learning, including the following:
[0066] Obtain real-time data of numerical control machining (real-time data is collected through multiple sensors installed on the lathe, see Figure 2 , Figure 2It is a typical layout example of a sensor, where T1 - T9 are multiple arranged sensors, and the real - time data is subjected to analog - to - digital conversion and pre - processing to obtain pre - processed data;
[0067] Among them, the real - time data and the corresponding sensors are as follows:
[0068] Cutting force sensor: Constraint parameter combination (to prevent overload).
[0069] Infrared temperature sensor: To prevent overheating and deformation, trigger a cooling strategy (such as reducing the cutting speed).
[0070] Power sensor: Calculate the energy consumption E.
[0071] Surface roughness sensor: Obtain the surface roughness Q of the workpiece s .
[0072] Cutting speed v c , feed rate f, and cutting depth a p are sourced from the feedback of the machine tool system;
[0073] Initialize the machining parameters, and the machining parameters include cutting speed v c , feed rate f, and cutting depth a p ;
[0074] Among them,
[0075] Cutting speed (v c ): The relative movement speed between the tool and the workpiece per unit time.
[0076] Feed rate (f): The distance that the tool moves per revolution or per minute.
[0077] Cutting depth (a p ): The depth at which the tool cuts into the workpiece.
[0078] Determine the optimization objectives, and construct an optimization objective function according to the optimization objectives. The optimization objectives include simultaneously satisfying minimizing the machining time T, minimizing the surface roughness Q s , and reducing the energy consumption C during the cutting process;
[0079] Based on the particle swarm optimization algorithm, calculate the current preliminary machining parameter combination [v c , f, a p ;
[0080] Input the machining parameter combination [v c , f, a p into the pre - trained deep reinforcement learning model, and through the action of the pre - trained deep reinforcement learning model, output the optimal machining parameter combination
[0081] With the optimal combination of machining parameters Carry out numerical control machining control.
[0082] In this Embodiment 1, through the deep reinforcement learning algorithm, the machining parameters of the numerically controlled lathe are automatically optimized, manual intervention is reduced, and the automation degree and optimization efficiency of the machining process are improved. It can dynamically adjust the optimization strategy according to the real-time machining state, and adaptively optimize the machining parameters under complex working conditions such as different materials, tool wear, and temperature changes to ensure machining quality and efficiency. Based on real-time feedback information (such as cutting force, temperature, etc.), the machining parameters are continuously adjusted and optimized through the reinforcement learning algorithm, so as to realize the intelligent and automatic optimization of the numerically controlled lathe machining process. Real-time machining data is collected through sensors, and combined with the deep reinforcement learning model, the process parameters of the lathe are dynamically adjusted to ensure that the machining process is always in the best state.
[0083] Through deep reinforcement learning technology, it is possible to intelligently optimize the processing parameters of a CNC lathe based on real-time processing data (such as cutting force, temperature, vibration, etc.). This intelligent automatic optimization can reduce manual intervention, improve the efficiency of parameter adjustment, and thus significantly shorten the processing time. By providing real-time feedback on the processing status, the present invention can dynamically adjust the processing parameters during the processing to ensure that the lathe always operates under the optimal process parameters, avoiding efficiency losses caused by improper parameter selection. By using a deep reinforcement learning model through learning and optimization, not only the processing time is concerned, but also important quality indicators such as processing accuracy and surface roughness can be optimized through a reward function, thereby improving the processing quality and reducing the scrap rate. Through the deep reinforcement learning algorithm, the system can adapt to complex working conditions such as different materials and tool wear, achieving the best processing quality without manual adjustment. By automatically optimizing the parameters, the dependence on manual operation is reduced, and the labor cost is reduced. High-quality processing can be completed in a shorter time, reducing energy consumption and production costs. Through continuous learning and optimization by deep reinforcement learning, the present invention can automatically adjust the processing parameters according to the real-time working conditions, avoiding the time required for manual intervention. Through rapid iteration and optimization, the production efficiency is improved. By optimizing the cutting parameters through deep reinforcement learning, cutting operations that consume excessive energy are avoided. Reasonable feed speed and cutting depth can minimize the tool load and reduce unnecessary energy consumption. The parameter optimization methods of traditional CNC lathes often rely on experience or static rules and cannot quickly respond to changing working conditions. The present invention uses deep reinforcement learning, which can not only automatically adjust and optimize the strategy under different materials and tool wear states, but also learn new process patterns through real-time feedback, enhancing the adaptive ability of the system. With the accumulation of processing data, the deep reinforcement learning model can continuously self-optimize and adapt to more complex processing conditions, ensuring processing stability and quality. By automatically optimizing the processing parameters through deep reinforcement learning, the technical threshold for operators is reduced. The operator only needs to monitor the equipment and processing status without relying on complex experience or manual adjustment. It can adjust the parameters in real time and effectively reduce the equipment load, thereby reducing the failure rate and maintenance requirements of the machine and lowering the maintenance cost of the equipment.
[0084] As a preferred embodiment of the present invention, specifically, an optimization objective function is constructed according to the optimization objective as follows:
[0085] F(x) = α·Q s +β·T + γ·C;
[0086] Where α, β, and γ are predetermined objective weights, depending on the priority of the processing task.
[0087] Considering that the particle swarm optimization algorithm (PSO) searches for the optimal solution by simulating the foraging behavior of a bird flock, as a preferred embodiment of the present invention, specifically, based on the particle swarm optimization algorithm, the current preliminary combination of processing parameters [vc , f, a p , including
[0088] Record the position of the particle as the machining parameter [v c , f, a p , the velocity of the particle is the moving direction and amplitude of the particle within the search space, and the global optimal solution g best is the best position in the history of all current particles, and the local optimal solution P best represents the best position in the history of a certain particle itself.
[0089] For each particle i, update its velocity and position based on the following formula:
[0090] v i,d (t + 1) = ω·v i,d (t) + c1·r1·(p best,i,d - x i,d (t)) + c2·r2·(g best,d - x i,d (t)),
[0091] v i,d (t): The velocity of particle i in dimension d, which includes cutting speed, feed speed, and cutting depth;
[0092] w: Inertia weight, controlling the balance of particle movement;
[0093] c1, c2: Learning factors, indicating the degree to which the particle learns from its own and the global optimal solutions;
[0094] r1, r2: Random numbers within the range [0, 1], used to introduce randomness;
[0095] p best,i,d : The historical best position of particle i;
[0096] g best,d : The global best position;
[0097] Update the position:
[0098] x i,d (t + 1) = x i,d (t) + v i,d (t + 1),
[0099] Compare the fitness value F(x) of the current particle with its historical best value p best , and update the historical optimal position of the particle, where the fitness function F(x) is designed based on the optimization objective function J:
[0100] F(x) = -J(x) = -(w1·T(x) + w2·Q(x) + w3·C(x));
[0101] Compare P of all particles best , update the global optimal position g best ;
[0102] If the maximum number of iterations T is reached max , or the change of the global optimal solution g best tends to be stable, that is, the corresponding error is less than the threshold ∈, then terminate the optimization, otherwise perform the next round of iteration.
[0103] After multiple rounds of iteration, the particle swarm converges to the global optimal solution g best , and obtain the optimized processing parameter combination [v c , f, a p .
[0104] As a preferred embodiment of the present invention, specifically, the value of ω is 1.0, and the values of c1 and c2 are 2.0.
[0105] Refer to Figure 3 , as a preferred embodiment of the present invention, specifically, the reward function of the pre-trained deep reinforcement learning model is
[0106] R(s, a) = α 1 ·Q s +β 1 ·T + γ 1 ·C;
[0107] Where s is the state space, which is composed of the machining environment of the CNC lathe, namely the cutting speed, feed speed, and cutting depth features, and the action space a is the adjustment of the machining parameters, that is, the amplitude of increasing or decreasing the cutting speed, feed speed, and cutting depth;
[0108] The pre-trained deep reinforcement learning model includes an input layer: the state vector S, a hidden layer: a multi-layer fully connected network containing a ReLU activation function, and an output layer: the action value Q(S, A), which corresponds to the expected return for each action.
[0109] As a preferred embodiment of the present invention, specifically, the optimal machining parameter combination is output through the action of the pre-trained deep reinforcement learning model including
[0110] Initialize a replay pool of a fixed size to store experience samples (s t , a t , r t , s t+1 ), randomly initialize the weights of the behavior network and the target network, and load the current preliminary machining parameter combination [v c , f, a p as a benchmark;
[0111] In the sampling stage, the current state s is utilized t , and the action network is used to select the action at:
[0112]
[0113] ∈: Exploration rate, adopting the ∈-greedy strategy,
[0114] According to the action a t , the machining parameters are adjusted, and the next state s is observed t+1 and the reward r t ,
[0115] The sample (s t , a t , r t , s t+1 ) is stored in the experience replay pool, and a small batch of experiences (s i , a i , r i , s i+1 ) are randomly sampled from the replay pool,
[0116] The target value is calculated based on the target network:
[0117]
[0118] where γ is the discount factor, measuring the importance of future rewards,
[0119] The network is updated using the mean squared error MSE to optimize the action network:
[0120]
[0121] where N represents the batch size,
[0122] Every fixed number of steps, the weights of the action network are copied to the target network. If the maximum number of iterations is reached or the reward converges, the training stops, and finally the optimal machining parameters after training are output:
[0123] The present invention also proposes a device for optimizing CNC machining parameters based on deep reinforcement learning, including the following:
[0124] A data acquisition module, configured to obtain real-time data of CNC machining, and perform analog-to-digital conversion and preprocessing on the real-time data to obtain preprocessed data;
[0125] An initialization module, configured to initialize machining parameters, where the machining parameters include the cutting speed v c , the feed rate f, and the cutting depth a p ;
[0126] Optimization model module, used to determine the optimization objectives and construct an optimization objective function according to the optimization objectives. The optimization objectives include simultaneously satisfying minimizing the processing time T, minimizing the surface roughness Q s , and reducing the energy consumption C during the cutting process; calculating the current preliminary combination of machining parameters [v c , f, a p ;
[0127] Deep reinforcement learning module, used to input the combination of machining parameters [v c , f, a p into the pre-trained deep reinforcement learning model, and output the optimal combination of machining parameters through the action of the pre-trained deep reinforcement learning model
[0128] Decision-making module, used to perform numerical control machining control with the optimal combination of machining parameters
[0129] In this Embodiment 2, the data acquisition module is responsible for collecting data in real time from various sensors in the CNC lathe and performing preliminary processing on this data for subsequent optimization and learning. Different types of sensors connected to the lathe (such as cutting force sensors, temperature sensors, etc.) are used to obtain relevant data in real time. The sensor data is collected in real time, the analog signals are converted into digital signals, and the collected data is preprocessed, including data denoising, time series analysis, normalization, feature extraction, etc. The optimization model module establishes and executes an optimization model based on the input machining process data, and preliminarily calculates the machining parameters suitable for the current working conditions. This module is usually implemented based on particle swarm optimization. The objective function is designed according to the key objectives in the machining process (such as machining time, surface quality, tool wear, etc.), and this is used as the direction for optimization. The constraint conditions are defined according to the actual machining limitations (such as cutting force range, machining accuracy, equipment limitations, etc.) to ensure that the optimization results meet the actual production conditions. Particle swarm optimization is used to perform preliminary parameter optimization according to the objective function and constraint conditions to ensure that the deep reinforcement learning model can converge to the appropriate machining parameters more quickly. The deep reinforcement learning module is the core module. It uses the deep reinforcement learning algorithm to continuously adjust and optimize the machining parameters through interaction with the CNC lathe, and automatically searches for the best strategy. The deep Q-network is used for modeling. The input is the real-time state in the machining process (such as cutting force, temperature, etc.), and the output is the optimized machining parameters. Based on the core algorithm of reinforcement learning, Q-learning, the system guides the learning process through a reward mechanism and gradually optimizes the machining strategy. A reasonable reward function is designed to give feedback according to the machining effect (such as machining time, surface quality, tool wear, etc.) to drive the deep reinforcement learning model to learn the best strategy. In reinforcement learning, the historical experience (state, action, reward, next state) is stored in the experience pool for subsequent training to enhance the learning efficiency of the model. The decision selection module selects the final machining parameters according to the optimization results output by the deep reinforcement learning model and the current machining environment, and adjusts the operation of the CNC lathe through the control system. The decision is made by combining the output of the deep reinforcement learning module and the suggestions of the optimization model. It can select the optimal strategy or combine the outputs of both to generate the final machining parameters. According to the real-time collected data and machining feedback, the decision module dynamically adjusts the selection strategy to ensure the stability and efficiency of the machining process. The final decision-making machining parameters are transmitted to the CNC lathe control system to perform corresponding machining operations (such as adjusting cutting speed, feed rate, cutting depth).
[0130] 1) The data acquisition module is responsible for collecting and preprocessing the process data in real time;
[0131] 2) The optimization model module provides preliminary machining parameter suggestions based on traditional optimization methods;
[0132] 3) The deep reinforcement learning module continuously adjusts and optimizes parameters according to the processing feedback;
[0133] 4) The decision-making module makes a final decision based on the output of the deep reinforcement learning model and controls the numerical control lathe for processing.
[0134] In addition, in each embodiment of the present invention, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module.
[0135] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or system, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc., that can carry the computer program code.
[0136] Although the description of the present invention has been quite detailed and particularly describes several of the above-mentioned embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be regarded as providing a broad interpretation of these claims in light of the prior art by reference to the appended claims, so as to effectively cover the intended scope of the present invention. In addition, the present invention is described above with embodiments foreseeable by the inventor for the purpose of providing a useful description, and those non-substantive modifications to the present invention that are not currently foreseeable can still represent equivalent modifications of the present invention.
[0137] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above-mentioned implementation manners. As long as it achieves the technical effects of the present invention by the same means, it should fall within the protection scope of the present invention. Within the protection scope of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.
Claims
1. A numerical control machining parameter optimization method based on deep reinforcement learning, characterized in that, including the following: Obtain the real-time data of numerical control machining, and perform analog-to-digital conversion and preprocessing on the real-time data to obtain preprocessed data; Initialize the machining parameters, where the machining parameters include cutting speed v c , feed rate f, and cutting depth a p ; Determine the optimization objectives and construct an optimization objective function according to the optimization objectives. The optimization objectives include simultaneously satisfying minimizing the processing time T, minimizing the surface roughness Q s , and reducing the energy consumption C during the cutting process; Calculate the current preliminary combination of machining parameters [v c , f, a p using the particle swarm optimization algorithm; Input the processing parameter combination [v c , f, a p into the pre-trained deep reinforcement learning model, and the optimal processing parameter combination is output through the action of the pre-trained deep reinforcement learning model With the optimal combination of machining parameters Perform numerical control machining control.
2. The method for optimizing numerical control machining parameters based on deep reinforcement learning according to claim 1, wherein Specifically, construct an optimization objective function according to the optimization objective as follows, J = α·Q s + β·T + γ·C; where α, β, γ are predetermined objective weights, depending on the priority of the machining task.
3. The numerical control machining parameter optimization method based on deep reinforcement learning according to claim 1, characterized in that Specifically, calculate the current preliminary combination of machining parameters [v c , f, a p , including Record the position of the particle as the processing parameters [v c , f, a p , the velocity of the particle is the moving direction and amplitude of the particle in the search space, and the global optimal solution g best is the best position in the history of all current particles, and the local optimal solution P best represents the best position in the history of a certain particle itself. For each particle i, update its velocity and position based on the following formula: v i,d (t + 1) = ω·v i,d (t) + c1·r1·(p best,i,d -x i,d (t)) + c2·r2·(g best,d -x i,d (t)), v i,d (t): The velocity of particle i in dimension d, which includes cutting velocity, feed velocity, and cutting depth; w: inertia weight, controlling the balance of particle movement; c1, c2: learning factors, indicating the degree to which the particle learns its own and the global optimal solution; r1, r2: random numbers in the range of [0, 1], used to introduce randomness; p best,i,d : The historical best position of particle i in dimension d; g best,d : Global best position in dimension d; Update the position: x i,d (t + 1)= x i,d (t)+ v i,d (t + 1), x: the position of the current particle; T(x): machining time; Q(x): surface roughness; C(x): cutting power consumption; w1, w2, w3: weight coefficients, reflecting the priority of each objective; Compare the fitness value F(x) of the current particle with its historical best value p best , and update the historical optimal position of the particle, where the fitness function F(x) is designed based on the optimization objective function J: F(x) = -J(x) = -(w1·T(x) + w2·Q(x) + w3·C(x)); Compare P of all particles best , update the global optimal position g best ; If the maximum number of iterations T is reached max , or the change in the global optimal solution g best tends to be stable, i.e., the corresponding error is less than the threshold ∈, then the optimization is terminated; otherwise, the next iteration is performed. After multiple rounds of iteration, the particle swarm converges to the global optimal solution g best , and an optimized combination of machining parameters [v c , f, a p is obtained.
4. The method for optimizing numerical control machining parameters based on deep reinforcement learning according to claim 3, characterized in that Specifically, the value of ω is 1.0, and the values of c1 and c2 are 2.
0.
5. The optimization method of numerical control machining parameters based on deep reinforcement learning according to claim 1, characterized in that, Specifically, the reward function of the pre-trained deep reinforcement learning model is, R(s,a) = α 1 ·Q s + β 1 ·T + γ 1 ·C; where s is the state space, which consists of the machining environment of the CNC lathe, i.e., cutting speed, feed rate, and cutting depth characteristics, and the action space a is the adjustment of machining parameters, i.e., the amplitude of increasing or decreasing the cutting speed, feed rate, and cutting depth, and α 1 , β 1 , and γ 1 are the updated values of α, β, and γ, respectively; The pre-trained deep reinforcement learning model includes an input layer: state vector S, a hidden layer: a multi-layer fully connected network containing a ReLU activation function, and an output layer: action value Q(S, A), corresponding to the expected return for each action.
6. The optimization method for numerical control machining parameters based on deep reinforcement learning according to claim 1, characterized in that Specifically, the optimal combination of processing parameters is output by the action of a pre-trained deep reinforcement learning model including Initialize a replay pool of a fixed size for storing experience samples (s t , a t , r t , s t+1 ), randomly initialize the weights of the behavior network and the target network, and load the current preliminary combination of processing parameters [v c , f, a p as a benchmark; In the sampling phase, the current state s is utilized t , and the action network is used to select an action a t : ∈: exploration rate, adopting an ∈-greedy strategy, According to action a t Adjust the processing parameters and observe the next state s t+1 and reward r t , Store the sample (s t , a t , r t , s t+1 ) in the experience replay pool and randomly sample a mini-batch of experiences (s i , a i , r i , s i+1 ) from the replay pool. The target value is calculated based on the target network: where γ is the discount factor, measuring the importance of future rewards, Q target : Output of the target network; s i+1 : The next state of the i-th sample, i.e., the state to which it transfers after i performing the action a i ; Target network Q target Predict the Q-values for all possible actions α′ of the next state s i+1 and take the maximum value, where α′ is the action variable representing all possible candidate actions in state s i+1 ; y i : The target Q value of the i-th sample, calculated by the target network: The network is updated using the mean square error MSE to optimize the behavior network: where N represents the batch size, s i ,a i : The state and action in the i-th sample, sampled from the experience replay pool; Q behavior : The predicted Q-value of the behavior network for state s i and action a i ; At fixed intervals of steps, copy the weights of the behavior network to the target network. Stop training if the maximum number of iterations is reached or the reward converges. Finally, output the optimal processing parameters after training:
7. Device for optimizing NC machining parameters based on deep reinforcement learning, characterized in that, including the following: A data acquisition module, used to obtain the real-time data of numerical control machining, and perform analog-to-digital conversion and preprocessing on the real-time data to obtain preprocessed data; Initialization module, used to initialize machining parameters, the machining parameters including cutting speed v c , feed rate f and cutting depth a p ; Optimization model module, which is used to determine the optimization objectives and construct an optimization objective function according to the optimization objectives. The optimization objectives include simultaneously satisfying the minimization of the machining time T, the minimization of the surface roughness Q s , reducing the energy consumption C during the cutting process; calculating the current preliminary combination of machining parameters [v c , f, a p ; Deep reinforcement learning module, used to input the combination of processing parameters [v c , f, a p into the pre-trained deep reinforcement learning model, and output the optimal combination of processing parameters through the action of the pre-trained deep reinforcement learning model A decision-making module for performing numerical control machining control with an optimal combination of machining parameters
Citation Information
Cited By
Adaptive machining parameter optimization method and system for numerical control machine tool
CN122044070A