Machine tool thermal error compensation method and device based on reinforcement learning

Through a reinforcement learning-based method, a mathematical model of machine tool temperature and thermal error was established, which solved the accuracy and robustness problems of machine tool thermal error in complex environments, achieved efficient thermal error compensation, adapted to various processing conditions, and improved processing accuracy and production efficiency.

CN120671286APending Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510602485.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When dealing with machine tool thermal errors, especially complex and nonlinear thermal error problems, existing technologies suffer from insufficient accuracy and poor robustness. They also lack versatility and online update capabilities, making it difficult to adapt to dynamic changes in the machine tool operating environment.

Method used

A reinforcement learning-based method is adopted to obtain machine tool temperature data and thermal error data. K-means clustering and time series grouping are used to establish a mathematical model between machine tool temperature and thermal error. Online training and strategy update are performed through reinforcement learning algorithm to achieve real-time compensation of thermal error.

Benefits of technology

The prediction accuracy and robustness of the machine tool thermal error model are improved, adapting to a variety of machine tool structures and processing conditions, reducing manual intervention, lowering modeling complexity and cost, and improving the quality consistency and production efficiency of processed products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671286A_ABST
    Figure CN120671286A_ABST
Patent Text Reader

Abstract

The invention relates to a machine tool thermal error compensation method and device based on reinforcement learning. The method comprises the following steps: obtaining overall temperature data of a machine tool; performing grouping optimization on the temperature data through k-means clustering to obtain processed temperature data; obtaining thermal error data of a linear shaft and a rotating shaft of the machine tool at each preset time node; grouping the processed temperature data and thermal error data by using a time sequence to obtain a grouping result; a grouping result is processed based on reinforcement learning, and a mathematical model between the machine tool temperature and the thermal error is established; real-time machine tool temperature is obtained and input into the mathematical model to obtain machine tool thermal errors, and machine tool thermal error compensation is carried out based on the machine tool thermal errors. And through online training and strategy updating of the reinforcement learning algorithm, the machine tool can keep high-precision performance in a complex machining environment. The method can adapt to various machine tool structures and machining conditions, independent modeling for each device is not needed, the universality of the modeling process is improved, and the application range of the modeling process is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to reinforcement learning, and in particular to a method and device for compensating thermal errors of machine tools based on reinforcement learning. Background Art

[0002] In high-precision machining scenarios, heat accumulation during machine tool operation can cause thermal deformation of components, leading to machining errors. Traditional methods based on physical models or empirical data fitting have large errors when dealing with complex, nonlinear thermal error problems. In addition, as machine tool structures become more complex, these methods are unable to fully capture the nonlinear relationships between multiple variables, resulting in insufficient accuracy and robustness in model predictions. Current thermal error compensation methods typically rely on offline modeling, lack online update capabilities, and are difficult to adapt to dynamic changes in the machine tool operating environment (such as different machining tasks, ambient temperature fluctuations, etc.). Existing methods are often designed for specific machine tools or machining conditions and lack versatility. Once the machining conditions change, re-modeling is required, resulting in long modeling cycles, low efficiency, and poor adaptability. Summary of the Invention

[0003] The purpose of the present invention is to solve at least one of the deficiencies of the prior art and to provide a method and device for machine tool thermal error compensation based on reinforcement learning.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions: Specifically, a machine tool thermal error compensation method based on reinforcement learning is proposed, which includes the following: Get the overall temperature data of the machine tool, such as Figure 3 Arrange temperature sensors at various locations; Performing grouping optimization on the temperature data by k-means clustering to obtain processed temperature data; Obtain thermal error data of the machine tool linear axis and rotary axis at each preset time node; The processed temperature data and thermal error data are grouped using time series to obtain grouping results; Processing the grouping results based on reinforcement learning to establish a mathematical model between machine tool temperature and thermal error; The real-time temperature of the machine tool is obtained, the real-time temperature of the machine tool is input into the mathematical model to obtain the thermal error of the machine tool, and the thermal error compensation of the machine tool is performed based on the thermal error of the machine tool.

[0005] Furthermore, specifically, the overall temperature data of the machine tool is obtained through the typical layout of the temperature sensors of the machine tool.

[0006] Furthermore, specifically, the thermal error data of the linear axis and the rotary axis of the machine tool at each preset time node are measured multiple times by means of a laser interferometer.

[0007] Furthermore, specifically, the processed temperature data and thermal error data are grouped using time series to obtain grouping results, including: Considering the next state thermal error in the state space Depends only on the current state , so we have, , The number of states n and the granularity of the range division need to be adjusted according to the system characteristics and data volume. Record the state sequence of the system at each time step: , Count each pair of states Number of transfers ,in: Indicates the slave state Transfer to state The number of times, Construct the state transfer matrix P, P is a A matrix where each element Indicates the slave state Transfer to state Probability of: , is the status The total number of transfers.

[0008] Furthermore, the method also includes considering that when there are few transfer samples, the number of state transfers may be zero, resulting in a sparse state transfer matrix P. To avoid overfitting, the state transfer matrix P is smoothed to obtain the updated state transfer matrix. The elements of the updated state transfer matrix are It is expressed as follows: , in is the smoothing parameter.

[0009] Furthermore, specifically, the grouping results are processed based on reinforcement learning to establish a mathematical model between the machine tool temperature and the thermal error, including: Find the steady-state distribution based on power iteration and use uniform distribution to initialize the probability distribution , , Recursive formula: , Iteration stop condition: When hour, is the set constant, stops the iteration, Each state corresponds to a thermal error value ,in, Indicates thermal errors in the x, y, and z directions; Predicted expected thermal error: ; in, :state Corresponding thermal error value : In time step , the system is in state The steady-state probability of (obtained through the Markov chain).

[0010] : predicted thermal error at time step t+k.

[0011] The mathematical model obtained by Markov process processing the expected value of the predicted thermal error is as follows: .

[0012] The present invention also proposes a device for machine tool thermal error compensation based on reinforcement learning, comprising the following: Temperature data acquisition module, used to obtain the overall temperature data of the machine tool; A clustering module, configured to group and optimize the temperature data by k-means clustering to obtain processed temperature data; A thermal error data acquisition module is used to acquire thermal error data of the linear axis and rotary axis of the machine tool at each preset time node; A grouping module is used to group the processed temperature data and thermal error data using time series to obtain grouping results; A mathematical model building module, configured to process the grouping results based on reinforcement learning to establish a mathematical model between the machine tool temperature and the thermal error; The thermal error compensation module is used to obtain the real-time temperature of the machine tool, input the real-time temperature of the machine tool into the mathematical model to obtain the thermal error of the machine tool, and perform thermal error compensation based on the thermal error of the machine tool.

[0013] The beneficial effects of the present invention are: The present invention proposes a method and device for machine tool thermal error compensation based on reinforcement learning, which uses the self-learning ability of reinforcement learning to establish an efficient mapping relationship between thermal errors and multivariable working conditions, thereby improving the prediction accuracy and robustness of the model. Through online training and strategy updating of the reinforcement learning algorithm, the machine tool can maintain high-precision performance in a complex processing environment. It can adapt to a variety of machine tool structures and processing conditions, without the need for separate modeling for each device, thereby improving the versatility and application scope of the modeling process. Through the autonomous optimization process of reinforcement learning, the dependence on manual intervention is reduced, the modeling process is simplified, the complexity and implementation cost are reduced, and a convenient solution is provided for industrial applications. Through precise thermal error compensation, processing errors are reduced, the quality consistency of processed products is improved, while the scrap rate is reduced, production efficiency is improved, and the stringent requirements of high-end manufacturing on machine tool performance are met. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other features of the present disclosure will become more apparent through a detailed description of the embodiments shown in conjunction with the accompanying drawings. The same reference numerals in the drawings of the present disclosure represent the same or similar elements. Obviously, the drawings described below are only some embodiments of the present disclosure. It is possible for a person skilled in the art to derive other drawings based on these drawings without inventive effort. In the drawings: Figure 1 FIG2 is a flow chart of a method for compensating thermal errors of machine tools based on reinforcement learning according to the present invention; Figure 2 Schematic diagram showing the method of measuring thermal error data of a linear axis and a rotary axis of a machine tool by a laser interferometer according to the present invention; Figure 3 Shown is a schematic diagram of an embodiment of the arrangement of sensors when the present invention is applied. DETAILED DESCRIPTION

[0015] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments and features in the embodiments of this application can be combined with each other unless there is a conflict. The same reference numerals used throughout the drawings indicate the same or similar parts.

[0016] Example 1, with reference to Figure 1 The present invention proposes a machine tool thermal error compensation method based on reinforcement learning, comprising the following steps: Get the overall temperature data of the machine tool, such as Figure 3 As shown (wherein each numerical label represents an example position where a sensor is arranged); Performing grouping optimization on the temperature data by k-means clustering to obtain processed temperature data; Obtain thermal error data of the machine tool linear axis and rotary axis at each preset time node; The processed temperature data and thermal error data are grouped using time series to obtain grouping results; Processing the grouping results based on reinforcement learning to establish a mathematical model between machine tool temperature and thermal error; The real-time temperature of the machine tool is obtained, the real-time temperature of the machine tool is input into the mathematical model to obtain the thermal error of the machine tool, and the thermal error compensation of the machine tool is performed based on the thermal error of the machine tool.

[0017] Leveraging the self-learning capabilities of reinforcement learning, the complex relationship between machine tool thermal errors and operating conditions can be automatically identified from large amounts of real-time data, thereby establishing a more accurate thermal error model. Temperature sensors collect overall machine tool temperature data, ensuring comprehensive and accurate data during the thermal error modeling process. Thermal error measurements using laser interferometers provide high-precision thermal error data reference for the model, ensuring the reliability of the modeling results. K-means clustering is used to optimize the grouping of temperature data, effectively reducing data redundancy, highlighting key features, and improving modeling efficiency. Time series group analysis of temperature and thermal error data fully considers dynamic factors, laying the foundation for subsequent model development. A mathematical model between machine tool temperature and thermal error accurately reflects the complex nonlinear relationship between the two. The model calculates thermal error using real-time temperature data, demonstrating strong practicality and reliability. The use of real-time temperature data enables dynamic adjustment of thermal error compensation to adapt to the changing needs of the machine tool under varying machining tasks and environmental conditions. Leveraging the self-learning capabilities of reinforcement learning, the complex nonlinear relationship between machine tool thermal error and operating conditions can be accurately captured, significantly reducing machining accuracy errors caused by temperature fluctuations. The system's dynamic adaptability supports online learning, enabling continuous model optimization based on real-time collected operating data. This autonomous optimization process, enhanced through learning, reduces reliance on manual commissioning and intervention, significantly reducing equipment commissioning and maintenance costs.

[0018] As a preferred embodiment of the present invention, specifically, the overall temperature data of the machine tool is acquired through a typical layout of the temperature sensors of the machine tool.

[0019] In this preferred embodiment, Figure 3 The typical layout of temperature sensors is shown: including the spindle, worktable, guide rails, etc.

[0020] Reference Figure 2 As a preferred embodiment of the present invention, specifically, the thermal error data of the linear axis and the rotary axis of the machine tool at each preset time node are measured by a laser interferometer in a multiple measurement manner.

[0021] In this preferred embodiment, the thermal error data of the linear axis and the rotary axis of the machine tool at each preset time node can be accurately measured by a laser interferometer.

[0022] As a preferred embodiment of the present invention, specifically, the processed temperature data and thermal error data are grouped using time series to obtain grouping results, including: Considering the next state thermal error in the state space Depends only on the current state , so we have, , The number of states n and the granularity of the range division need to be adjusted according to the system characteristics and data volume. Record the state sequence of the system at each time step: , Count each pair of states Number of transfers ,in: Indicates the slave state Transfer to state The number of times, Among them, the state sequence of the time step Each Corresponding to a state

[0023] Data is collected every 5 minutes. and It is the discretized joint state, which represents the temperature-thermal error combination of the machine tool at a certain moment.

[0024] The specific form is: , Construct the state transfer matrix P, P is a A matrix where each element Indicates the slave state Transfer to state Probability of: , is the status The total number of transfers.

[0025] As a preferred embodiment of the present invention, the method further includes, considering that when there are few transfer samples, the number of state transfers may be zero, resulting in a sparse state transfer matrix P. To avoid overfitting, the state transfer matrix P is smoothed to obtain the updated state transfer matrix. The elements of the updated state transfer matrix are It is expressed as follows: , in is the smoothing parameter.

[0026] As a preferred embodiment of the present invention, specifically, the grouping results are processed based on reinforcement learning to establish a mathematical model between the machine tool temperature and the thermal error, including: Find the steady-state distribution based on power iteration and use uniform distribution to initialize the probability distribution , , Recursive formula: , Iteration stop condition: When hour, is the set constant, stops the iteration, Each state corresponds to a thermal error value ,in, Indicates thermal errors in the x, y, and z directions; Predicted expected thermal error: ; in, :state Corresponding thermal error value : In time step , the system is in state The steady-state probability of (obtained through the Markov chain).

[0027] : predicted thermal error at time step t+k.

[0028] The mathematical model obtained by Markov process processing the expected value of the predicted thermal error is as follows: .

[0029] In this preferred embodiment, to characterize the environment's feedback reward to the agent, the Markov reward process expands the above MP from < S , P > to < S , P , R , γ >. Here, R represents the reward function, and γ represents the reward discount factor.

[0030] : State space, which is composed of temperature data collected by sensors at key positions of the machine tool (such as spindle, guide rail, etc.).

[0031] : State transition probability, temperature-thermal error state transition probability based on time series statistics.

[0032] : Reward function, the improvement of machining accuracy after compensation (error reduction).

[0033] , Return is the total expected return starting from time step t. The discounted return is as follows: , : The reward obtained in step t is the improvement of machining accuracy after one error compensation.

[0034] State Value Function (s) is in the strategy Next, the expected return from state s.

[0035] , : The state of the system at time t (temperature distribution) Markov decision process The immediate reward (Reward, R) of the Markov decision is represented by < S , A , P , R , γ > A represents a finite set of actions, at which point the immediate reward becomes , : The action to be taken at this time (select a thermal error compensation scheme).

[0036] Strategy It is used to describe the way in which the agent takes actions based on the environment observation. and actions A to action probability distribution The mapping, Indicates the probability of taking action a in state s. In this system, the strategy is the rule for the machine tool to choose which compensation method to use under a certain temperature state.

[0037] , The expected return is expressed as: , : Trajectory, that is, a complete state-action-reward sequence, a processing task process : Cumulative reward over the entire trajectory (overall compensation effect).

[0038] Represents a given initial state distribution With the strategy π, the probability of a T-step trajectory τ occurring in the Markov decision process, that is, the cumulative reward (overall compensation effect) over the entire trajectory, is as follows: , : Initial state distribution (probability of initial temperature state).

[0039] : Transition probability, indicating the transition from Take action The probability of (determined by the strategy).

[0040] Reinforcement learning optimization problems use optimization methods to improve strategies to maximize expected returns. Expressed as , : Make the expected total return The largest optimal strategy Given a policy π, the value function V(s), that is, the expected return under a given state, is to find the most effective sequence of thermal error compensation behaviors in the system. It can be expressed as: , In an MDP, given an action, there is an action-value function, which is the expected return based on the state and action as follows:

[0041] :Indicates that the state Take action The sum of the expected returns after the initial value is calculated. It is used to compare which compensatory action has greater "long-term value" in a certain state.

[0042] The state space of reinforcement learning consists of the machine tool temperature, the action space represents possible compensation actions, and the reward function is designed based on the machining accuracy after compensation. This module continuously optimizes the thermal error compensation strategy through interactive learning with the environment.

[0043] According to the above definition, we can get the state value function expression under the strategy: .

[0044] :About action The expectation of π represents the weighted sum of the probability distributions given by the strategy π for all possible actions.

[0045] : The action value function under the strategy π, which represents the action value under the state Next action , and then continue to act according to the policy π, the expected cumulative reward. The expected effect of taking a compensatory action in the current state.

[0046] In Example 2, the present invention further proposes a device for compensating machine tool thermal errors based on reinforcement learning, comprising the following: Temperature data acquisition module, used to obtain the overall temperature data of the machine tool; A clustering module, configured to group and optimize the temperature data by k-means clustering to obtain processed temperature data; A thermal error data acquisition module is used to acquire thermal error data of the linear axis and rotary axis of the machine tool at each preset time node; A grouping module is used to group the processed temperature data and thermal error data using time series to obtain grouping results; A mathematical model building module, configured to process the grouping results based on reinforcement learning to establish a mathematical model between the machine tool temperature and the thermal error; The thermal error compensation module is used to obtain the real-time temperature of the machine tool, input the real-time temperature of the machine tool into the mathematical model to obtain the thermal error of the machine tool, and perform thermal error compensation based on the thermal error of the machine tool.

[0047] In this second embodiment, a temperature data acquisition module monitors the temperature of the machine tool in real time. Temperature data acquisition is achieved through sensors installed in key locations on the machine tool. The sensors transmit the temperature data to the data processing module via a data interface. After the data acquisition module transmits the raw data, the clustering module and the grouping module filter the data for noise, remove outliers, and normalize it, extracting feature data related to thermal errors. These features serve as input to the reinforcement learning model. A thermal error compensation strategy is trained using a reinforcement learning algorithm. The state space of the reinforcement learning algorithm consists of the machine tool's temperature, the action space represents possible compensation actions, and the reward function is designed based on the post-compensation machining accuracy. This module continuously optimizes the thermal error compensation strategy through interactive learning with the environment. Based on the optimal compensation strategy output by the reinforcement learning model, the compensation decision and control module adjusts the machine tool's control instructions, optimizing the machine tool's motion trajectory or machining parameters in real time to reduce the impact of thermal errors. A self-learning module is added to collect new data continuously generated during the machine tool's machining process and optimize the reinforcement learning model based on this new data feedback, allowing the model to gradually adapt to new machining tasks and working environments.

[0048] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.

[0049] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or system that can carry the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0050] Although the present invention has been described in considerable detail and with particularity with respect to several described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but rather should be construed as providing a broad possible interpretation of these claims in view of the prior art by reference to the appended claims, thereby effectively encompassing the intended scope of the invention. In addition, the invention has been described above in terms of embodiments foreseen by the inventors for the purpose of providing a useful description, and those insubstantial modifications of the invention that are not currently foreseen may still represent equivalent modifications of the invention.

[0051] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. As long as the technical effects of the present invention are achieved by the same means, they shall fall within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods may be made.

Claims

1. A machine tool thermal error compensation method based on reinforcement learning, characterized in that: These include: Obtain overall temperature data of the machine tool; Performing grouping optimization on the temperature data by k-means clustering to obtain processed temperature data; Obtain the thermal error data of the machine tool linear axis and rotary axis at each preset time node, recorded as R; The processed temperature data and thermal error data are grouped using time series to obtain grouping results; Processing the grouping results based on reinforcement learning to establish a mathematical model between the machine tool temperature and R; The real-time temperature of the machine tool is obtained, the real-time temperature of the machine tool is input into the mathematical model to obtain the thermal error of the machine tool, and the thermal error compensation of the machine tool is performed based on the thermal error of the machine tool.

2. The machine tool thermal error compensation method based on reinforcement learning according to claim 1, characterized in that: Specifically, the overall temperature data of the machine tool is obtained through the typical layout of the temperature sensors of the machine tool.

3. The machine tool thermal error compensation method based on reinforcement learning according to claim 1, characterized in that: Specifically, the thermal error data of the linear axis and the rotary axis of the machine tool at each preset time node are measured multiple times by using a laser interferometer.

4. The method for compensating thermal errors of machine tools based on reinforcement learning according to claim 3, characterized in that: Specifically, the processed temperature data and thermal error data are grouped using time series to obtain grouping results, including: Considering the next state thermal error in the state space Depends only on the current state , so we have, , The number of states n and the granularity of the range division need to be adjusted according to the system characteristics and data volume. Record the state sequence of the system at each time step: , Count each pair of states Number of transfers ,in: Indicates the slave state Transfer to state The number of times, Construct the state transfer matrix P, P is a A matrix where each element Indicates the slave state Transfer to state Probability of: , is the status The total number of transfers.

5. The machine tool thermal error compensation method based on reinforcement learning according to claim 4, characterized in that: The method further includes taking into account that when there are few transfer samples, the number of state transfers may be zero, resulting in a sparse state transfer matrix P. To avoid overfitting, the state transfer matrix P is smoothed to obtain the updated state transfer matrix. The elements of the updated state transfer matrix are It is expressed as follows: , in is the smoothing parameter.

6. A machine tool thermal error compensation method based on reinforcement learning according to claim 4 or 5, characterized in that: Specifically, the grouping results are processed based on reinforcement learning to establish a mathematical model between the machine tool temperature and thermal error, including: Find the steady-state distribution based on power iteration and use uniform distribution to initialize the probability distribution , , Recursive formula: , Iteration stop condition: When hour, is the set constant, stops the iteration, Each state corresponds to a thermal error value ,in, Indicates thermal errors in the x, y, and z directions; Predicted expected thermal error: ; in, :state The corresponding thermal error value; : In time step , the system is in state The steady-state probability of is obtained through the Markov chain. : predicted thermal error at time step t+k; The mathematical model obtained by Markov process processing the expected value of the predicted thermal error is as follows: 。 7. A device for compensating thermal errors of machine tools based on reinforcement learning, characterized in that: These include: Temperature data acquisition module, used to obtain the overall temperature data of the machine tool; A clustering module, configured to group and optimize the temperature data by k-means clustering to obtain processed temperature data; A thermal error data acquisition module is used to acquire thermal error data of the linear axis and rotary axis of the machine tool at each preset time node; A grouping module is used to group the processed temperature data and thermal error data using time series to obtain grouping results; A mathematical model building module, configured to process the grouping results based on reinforcement learning to establish a mathematical model between the machine tool temperature and the thermal error; The thermal error compensation module is used to obtain the real-time temperature of the machine tool, input the real-time temperature of the machine tool into the mathematical model to obtain the thermal error of the machine tool, and perform thermal error compensation based on the thermal error of the machine tool.

Citation Information

Patent Citations

  • Machine tool thermal error compensation method and system

    CN117572820A

  • Battery analysis method

    CN118962455A

  • Permanent magnet motor energy storage and optimized operation method and system based on electricity price

    CN119514816A