Numerical control machining cutter axis vector motion trajectory planning method based on Q-learning

By applying the reinforcement learning method based on Q-learning in five-axis CNC machining, the problems of low efficiency and poor adaptability of tool posture smoothness planning are solved, and the global optimal tool axis vector motion trajectory planning is achieved, which improves processing quality and equipment reliability.

CN119937462AActive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202411988563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In the existing five-axis CNC machining, the tool posture smoothness planning method has problems such as low efficiency, poor adaptability and large calculation volume, making it difficult to achieve global optimization in complex machining scenarios.

Method used

Using a reinforcement learning method based on Q-learning, through C-space expression technology and grid division, the knife axis vector is compressed from three-dimensional space to two-dimensional polar coordinate space, and a reinforcement learning model is constructed, state space and action space are defined, reward functions are designed, and the Q-learning algorithm is used to search for the optimal knife axis vector motion trajectory.

Benefits of technology

It realizes efficient exploration of the optimal tool axis vector trajectory in complex machining scenarios, avoids the defects of the traditional heuristic algorithms that are prone to local optimality, ensures the global optimality of tool axis vector planning, and improves processing quality and equipment operation reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937462A_ABST
    Figure CN119937462A_ABST
Patent Text Reader

Abstract

The invention discloses a Q-learning-based numerical control machining cutter axis vector motion trajectory planning method, which comprises the following steps: compressing a cutter axis vector from three dimensions to two dimensions through a C space expression technology, carrying out meshing division, and distinguishing a feasible domain and a prohibited domain based on non-interference, non-singularity and other machining requirements; a Q-learning algorithm is combined with a reward feedback mechanism, an optimal cutter shaft path is searched in a feasible region, and the fairness of the path and the stability of machine tool movement are gradually improved through state and action space definition and reward function design; the reward function comprehensively considers factors such as path safety, cutter kinematics characteristics and cutter shaft vector variation amplitude so as to strengthen optimization planning of a cutter shaft track; and finally, realizing global optimization of the cutter shaft vector path through Q value iterative updating. According to the method, the problems of calculation complexity and local optimization of tool attitude optimization under complex working conditions are effectively solved, and the adaptability and efficiency of five-axis machining path planning are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of numerical control machining, and in particular to a method for planning a vector motion trajectory of a numerical control machining tool axis based on Q-learning. Background Art

[0002] In five-axis CNC machining, the smoothness of the tool posture directly affects the stability and quality of the machining process. Smooth tool posture changes can not only avoid machine tool overload and jitter caused by sudden posture changes, but also significantly improve the quality of the machined surface, thereby improving overall machining efficiency and accuracy. Especially in the process of complex surface machining, it is crucial to maintain a smooth change in tool posture, which can ensure the best contact between the tool and the workpiece and avoid degradation of machining quality or equipment wear caused by drastic changes in posture.

[0003] At present, there are two main methods for tool posture smoothness planning: an adjustment method based on preset rules and a global planning method based on the feasible domain.

[0004] For example, Chinese patent document with publication number CN107335847A discloses a processing method for cutting efficiency constraining tool posture; Chinese patent document with publication number CN115857429A discloses a method for generating a smooth path for a five-axis machine tool tool.

[0005] The adjustment method based on preset rules pre-generates the tool posture according to certain constraints, meets the conditions of no interference and no singularity, and completes the posture adjustment by combining local optimization. This method is simple and intuitive, but it cannot achieve global optimization, and requires multiple adjustments for different workpieces, which makes it inefficient and poorly adaptable.

[0006] The global planning method based on the feasible domain is based on the feasible domain of tool contact points without interference and singular tool angles, and realizes global smoothing optimization within the feasible domain. In theory, this method has strong universality, and additional optimization objectives (such as cutting force or dynamic characteristics optimization) can be selected according to needs without subsequent adjustments. However, its large amount of calculation leads to a long planning time, which is subject to certain limitations in practical applications.

[0007] In order to solve performance problems, heuristic algorithms are widely used in such optimization problems. However, heuristic algorithms often face limitations such as local optimal traps, dependence on domain knowledge, and poor adaptability when solving complex optimization problems, making them difficult to expand in high-dimensional and uncertain environments.

[0008] In contrast, reinforcement learning effectively balances exploration and utilization through a reward mechanism, has stronger adaptability, and does not rely on explicit domain knowledge, so it can respond more flexibly to dynamically changing environments. This makes reinforcement learning show strong robustness and global search capabilities in complex optimization problems, becoming an effective alternative to traditional heuristic algorithms. Summary of the invention

[0009] In order to solve the problems in the background technology, the present invention provides a CNC machining tool axis vector motion trajectory planning method based on Q-learning, which can efficiently explore the optimal trajectory of the tool axis vector, thereby achieving global optimization in complex machining scenarios, and providing a new solution for tool posture smoothness planning in five-axis CNC machining.

[0010] A Q-learning-based CNC machining tool axis vector motion trajectory planning method comprises the following steps:

[0011] S1. Through C-space expression technology, the tool axis vector is compressed from three-dimensional space to two-dimensional polar coordinate space, and then gridded;

[0012] S2, judging the state of the grid after gridding, and dividing it into feasible domain and forbidden domain;

[0013] S3. Construct a reinforcement learning model; in which the state space is defined as the grid set of the feasible domain of the tool contact points, the action space is defined as the set of possible actions of the agent from the current grid point to all grid points in the next feasible domain, and the reward function comprehensively considers the safety of the path, the kinematic characteristics of the tool, and the amplitude of the tool axis vector change;

[0014] S4. The Q-learning algorithm is combined with value iteration update to search for the optimal tool axis vector motion trajectory through interaction with the environment and reward feedback.

[0015] In step S1, the tool axis vector is compressed from the three-dimensional space to the two-dimensional polar coordinate space through the C space expression technology, specifically:

[0016] During the machining process, the tool axis vector is represented by a three-dimensional vector (i, j, k) in the workpiece coordinate system. When the tool axis vector is mapped to polar coordinates, only two angle parameters (α, β) are used to represent it, where:

[0017]

[0018] In the formula, α represents the direction angle of the tool axis vector on the horizontal plane, and β represents the angle between the tool axis vector and the vertical direction.

[0019] In step S1, the specific process of gridding is as follows:

[0020] Discretize the two-dimensional polar coordinate space (α, β) according to a certain resolution and divide α into n α intervals, divide β into n β intervals; each grid center point corresponds to a discrete knife axis vector (α i ,β i ), these grid points constitute the state space of the tool posture.

[0021] In step S2, for each grid point in the state space, it is determined whether it satisfies the constraints of no interference and no singularity, and marked, and the grid points that satisfy the constraints of no interference and no singularity are divided into feasible domains, and the remaining grid points are divided into prohibited domains.

[0022] In step S3, when defining the state space, the feasible domain grid set of the knife contact point is used as the state space, where the feasible domain grid of each knife contact point is described by two discretized variables α and β, which represent the set of inclination angles and azimuth angles respectively; the state space S is defined as the grid pair of the feasible domains of adjacent knife contacts, expressed as:

[0023]

[0024] Among them, k represents the current knife contact number, S k represents the kth state, S k It is expressed as:

[0025] S k ={(α k ,β k ),(α k+1 ,β k+1 )}

[0026] in, is the feasible domain mesh corresponding to the kth knife contact point, (α k+1 ,β k+1 ) is the feasible domain mesh corresponding to the k+1th knife contact point.

[0027] In step S3, the action space is defined as the set of possible actions of the agent from the current grid point to all grid points in the next feasible domain:

[0028] A k ={(α k+1 ,β k+1 )∈ feasible domain}.

[0029] In step S3, the reward function is defined as follows:

[0030]

[0031] Among them, r t(s,a) represents the immediate reward obtained by transferring to the next state s after taking action a at time t. The negative of the Euclidean distance between adjacent feasible region grid points is used as the reward value.

[0032] The specific process of step S4 is:

[0033] Initialization parameters: Randomly initialize the Q value table Q(S,A); the learning rate α determines the impact of new information on the Q value update, and its value is 0.1; the discount factor γ determines the importance of future rewards, and its value is 0.9;

[0034] The agent's starting state selection: Randomly select the feasible grid pair of the k-1th and kth knife contact points as the starting point S k ;

[0035] Action exploration strategy: In order to balance exploration and utilization, the ε-greedy exploration strategy is adopted. k In this case, the agent selects actions according to the following rules

[0036]

[0037] Among them, ε takes the value of 0.1, indicating the exploration probability; A k represents the set of all optional actions in the current state, Q(S k ,A') represents state S k The Q value of action A';

[0038] State transfer and reward calculation: according to the selected action a, transfer to the next state S' and obtain the immediate reward r(S,a);

[0039] Q value table update:

[0040]

[0041] Where Q(s,a) is the Q value of taking action a in the current state s; r is the immediate reward obtained by taking action a in the current state s; γ is the discount factor that determines the weight of future rewards; is the maximum Q value of all possible actions in the next state s′; α is the learning rate, which determines the influence of new information on the Q value update;

[0042] Path generation: Starting from the starting state, the intelligent agent gradually selects actions according to the current Q-value table to generate a continuous grid point sequence between the tool contacts; with the iterative update of the Q-value table, the accuracy and optimization effect of path planning are gradually improved, and finally the optimal tool axis vector path that meets the requirements of smoothness and safety is obtained; when the update amplitude of the Q-value table is lower than the predetermined threshold, the iteration is stopped and the optimal path result is output.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] (1) The present invention significantly reduces the complexity of the path planning problem through C-space compression and grid discretization technology.

[0045] (2) The present invention adopts a Q-learning-based reinforcement learning path planning algorithm, which effectively avoids the defect of traditional heuristic algorithms that are prone to fall into local optimality by balancing exploration and utilization, and ensures the global optimality of tool axis vector planning.

[0046] (3) The present invention comprehensively considers path smoothness, safety and kinematic characteristics of machine tools, and ensures the smoothness and stability of tool movement by designing corresponding reward functions, thereby improving processing quality and equipment operation reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flow chart of path planning using the Q-learning algorithm in the present invention.

[0048] Figure 2 It is a two-dimensional discrete grid representation of the feasible domain.

[0049] Figure 3 This is a graph showing the change in total reward during Qleaning training.

[0050] Figure 4 Schematic diagram of tool axis vector path planning in four feasible domains. DETAILED DESCRIPTION

[0051] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.

[0052] The present invention compresses the tool axis vector from three dimensions to two dimensions through C-space expression technology, and divides the two-dimensional space into grids. Based on processing requirements such as no interference and no singularity, the grid is state-judged to distinguish between feasible domains and forbidden domains. The complexity of the path planning problem is greatly reduced through C-space compression and grid discretization. During the planning process, the Q-learning algorithm searches for the optimal tool axis path in the feasible domain and combines the reward feedback of the current grid state to strengthen the mapping relationship between the environmental state and the path planning. The algorithm gradually improves the smoothness and stability of the tool vector path during iterative learning, ensuring that the path planning meets the smoothness, thereby achieving the optimal tool axis trajectory planning under complex working conditions. The specific technical solutions adopted by the present invention are as follows:

[0053] S1, C space compression and grid division

[0054] S1.1. C-space representation of three-dimensional tool axis vector

[0055] In the process of machine tool processing, the tool axis vector can be expressed by a three-dimensional vector (i, j, k) in the workpiece coordinate system. In order to simplify the calculation, the tool axis vector is mapped to the polar coordinate form, which can be expressed by only two angle parameters (α, β), where:

[0056]

[0057] In the formula, α represents the direction angle of the tool axis vector on the horizontal plane, and β represents the angle between the tool axis vector and the vertical direction.

[0058] S1.2. Grid Division

[0059] Discretize the two-dimensional polar coordinate space (α, β) according to a certain resolution and divide α into n α intervals, divide β into n β The two-dimensional grid is numbered starting from the lower left corner, using the row-by-row numbering rule, that is, numbering from left to right and from bottom to top. i,j represents the jth grid in the feasible domain of the i-th tool contact point. Each grid can be represented by a grid center point, which corresponds to a discrete tool axis vector (α i ,β i ), these grid center points constitute the state space of the tool posture.

[0060] S1.3. Distinguishing between feasible domain and forbidden domain

[0061] For each mesh, determine whether it satisfies the constraints of no interference and no singularity, based on the following assumptions: If the tool axis vectors represented by the four vertices of a mesh all satisfy the conditions of no interference and no singularity, then the mesh is feasible and marked as the feasible domain, such as Figure 2 shown.

[0062] S2. Creation of reinforcement learning model

[0063] S2.1. State space definition

[0064] The core of reinforcement learning is that the agent gradually learns the rules of the environment and optimizes the decision by interacting with the environment. In the tool axis vector path planning task, the goal of the agent is to plan a tool axis vector path for a set of continuous tool contacts, while the environment corresponds to the entire machining process. Suppose the set of continuous n tool contacts is CL = {CL1, CL2, …, CL n}, where each knife contact CL i =(x i ,y i ,z i ) represents the position coordinates of the tool center. Corresponding to each tool contact point CLi , the feasible domain of its tool posture can be defined as CO i , represented by a discretized grid. Therefore, the state space can be represented as the set of feasible regions of the knife contact points CO = {CO1, CO2, …, CO n}.

[0065] In order to facilitate rapid learning and decision-making of the intelligent agent and reduce the complexity of the path planning problem, each feasible domain CO i Meshing is performed to transform the state space into a finite and discrete form. In addition, to introduce acceleration constraints, the state input s is defined as the feasible region pair (CO) of two adjacent knife contacts. i-1 ,CO i ), that is, s=(CO i-1 ,CO i ). Through this design, the intelligent agent can comprehensively consider the dynamic constraints between the current tool contact point and the adjacent tool contact points, thereby achieving optimal planning of the tool axis vector path.

[0066] In this embodiment, the feasible domain grid set of the knife contact point is used as the state space, where the feasible domain grid of each knife contact point is described by two discretized variables α and β, which represent the set of inclination angle and azimuth angle respectively. It can be defined as a grid pair of the feasible region of adjacent knife contacts, expressed as:

[0067]

[0068] Among them, k represents the current knife contact number, S k represents the kth state, S k It is expressed as:

[0069] S k ={(α k ,β k ),(α k+1 ,β k+1 )}

[0070] in, is the feasible domain mesh corresponding to the kth knife contact point, (α k+1 ,β k+1 ) is the feasible domain mesh corresponding to the k+1th knife contact point.

[0071] S2.2 Action Space Definition

[0072] The action space is defined as the set of all possible grids in the next feasible domain that the agent can choose after selecting a grid in the current feasible domain. Specifically, the action space describes the target grids that the agent can reach from the current state (i.e., the current grid), and these target grids belong to the next feasible domain.

[0073] Through this design, the agent can explore different paths in each decision-making step, thereby achieving optimal planning of the knife axis vector. The action space can be expressed as:

[0074]

[0075] Where CO = {CO1, CO2, …, CO n} represents the set of all feasible domains, g i-1 and g i represents the feasible region pair (CO) of two adjacent knife contacts corresponding to any state input s i-1 ,CO i ) in the grid points, g i+1 Represents the next feasible region CO i+1 The grid points in .

[0076] In this embodiment, the action space is defined as the set of possible actions of the agent from the current grid point to the next grid point:

[0077] A k ={(α k+1 ,β k+1 )∈ feasible domain}

[0078] S2.3. Reward function design

[0079] In the reinforcement learning model, the reward function is a core component that directly determines the learning direction of the agent and the final optimization effect. The design of the reward function needs to comprehensively consider the constraints and optimization goals to ensure the safety, feasibility and smoothness of path planning.

[0080] S2.3.1. Ensure path safety

[0081] In CNC machining, once the path enters the forbidden domain, it may cause tool collision or singularity problems, causing uncontrollable or even irreversible effects on the machining process. Therefore, the planned tool axis vector trajectory must always be within the feasible domain.

[0082] The reward function is designed as follows:

[0083] When a path enters a forbidden domain, a larger negative reward is given:

[0084] R 域 =-100

[0085] Through this negative reward design, the reinforcement learning model can effectively avoid dangerous paths and ensure the safety of the tool axis trajectory.

[0086] S2.3.2. Satisfy the kinematic characteristics of each axis of the machine tool

[0087] The feasible domain of the tool axis vector is defined in the workpiece coordinate system, but its mapping from the workpiece coordinate system to the machine tool coordinate system involves nonlinear transformation. This makes the smoothness in the workpiece coordinate system unable to fully reflect the smoothness of the rotation angle in the machine tool coordinate system. Therefore, it is necessary to constrain the rotation angular velocity and acceleration of each axis in the machine tool coordinate system in the reward function. For the tool axis vector in the workpiece coordinate system selected in the current state s and the tool axis vector selected in the next action a According to the inverse kinematics IK, it is converted into the tool axis vector in the machine tool coordinates So the central difference method gives the expressions of angular velocity and angular acceleration of corners A and C:

[0088]

[0089] Among them, f is the feed rate, which is generally a constant value during the machining process; L is the distance between adjacent tool contacts. In order to meet the performance requirements of the machine tool, the rotational angular velocity and angular acceleration need to be limited to the following range:

[0090]

[0091] Wherein, * indicates A or C axis.

[0092] The reward function is designed as follows:

[0093] If the next action causes the angular velocity or angular acceleration to exceed the predetermined limit range, a larger negative reward will be given:

[0094] R 动 =-100

[0095] Through this design, the intelligent agent can give priority to tool axis vectors that meet kinematic constraints during path planning, thereby avoiding machine tool overload or unstable motion caused by exceeding the limit.

[0096] S2.3.3. Ensure the smoothness of the tool axis vector trajectory

[0097] The change amplitude of the tool axis vector between adjacent tool positions is an important indicator for evaluating the smoothness of the path. Too large a change amplitude may cause sudden changes in the trajectory, affecting the processing quality and efficiency. Therefore, in order to ensure the smoothness of the processing path, it is necessary to minimize the change of the tool axis vector between adjacent tool positions.

[0098] The present invention uses the negative value of the Euclidean distance between adjacent feasible domain grid points as an indicator to measure the change amplitude of the knife axis vector. When the Euclidean distance between two adjacent knife position points is small (i.e., the change amplitude of the knife axis vector is small), a higher positive reward can be given to encourage the agent to choose a smooth path. On the contrary, if the Euclidean distance is large, a larger negative reward is given to avoid sudden changes in the path. Suppose the coordinates of the grid points of two adjacent feasible domains are g i =(x i ,y i ) and g i+1 =(x i+1 ,y i+1 ).

[0099] The reward function is designed as follows:

[0100] The reward value is the negative of the Euclidean distance between adjacent feasible region grid points:

[0101]

[0102] S2.3.4. Summary of reward function definition

[0103]

[0104] S3, use Q-learning algorithm for path planning, such as Figure 1 shown.

[0105] S3.1. Initialization parameters

[0106] The Q value table Q(S,A) is randomly initialized to 0 for ease of calculation; the learning rate α determines the impact of new information on the Q value update and is set to 0.1; the discount factor γ determines the importance of future rewards and is set to 0.9.

[0107] S3.2. Decision-making process of intelligent agents

[0108] S3.2.1. Starting state selection

[0109] Randomly select the feasible mesh pair of the k-1th and kth knife contact points as the starting point S k .

[0110] S3.2.2 Action Exploration Strategy

[0111] In order to balance exploration and utilization, the ε-greedy exploration strategy is adopted. k In this case, the agent selects actions according to the following rules:

[0112]

[0113] Among them, ε takes the value of 0.1, indicating the exploration probability; A k represents the set of all optional actions in the current state, Q(S k ,A') represents state S k The Q value of action A'.

[0114] S3.2.3. State transfer and reward calculation

[0115] According to the selected action a, it transfers to the next state S' and obtains the immediate reward r(S,a).

[0116] S3.2.4. Q value table update

[0117] In the process of reinforcement learning, the agent continuously interacts with the environment and continuously updates the Q-value table to optimize path planning decisions. The core of Q-learning is to learn how to select actions and iteratively adjust the Q-value table through the following update formula:

[0118]

[0119] Where Q(s,a) is the Q value of taking action a in the current state s; r is the immediate reward obtained by taking action a in the current state s; γ is the discount factor that determines the weight of future rewards; is the maximum Q value of all possible actions in the next state s′; α is the learning rate, which determines the influence of new information on the Q value update.

[0120] S3.2.5 Action selection

[0121] In the process of updating the Q value, the agent needs to balance the relationship between exploration and exploitation. Exploitation means that the agent always chooses the action with the largest current Q value, and optimizes the path based entirely on existing experience. The learning efficiency is high, but it is easy to fall into the local optimum. Exploration means that the agent always randomly selects actions, which has low learning efficiency but can ensure the global optimum.

[0122] In order to balance exploration and exploitation, many strategies have been proposed.

[0123] S3.2.6. Iterate and update the Q value table until convergence

[0124] Q-learning optimizes the strategy by repeatedly updating the Q-value table until the change in Q-value meets the convergence condition. In each iteration, the agent selects an action based on the current state, obtains feedback through interaction with the environment, and continuously corrects the value to optimize the path planning. As the iteration proceeds, the accuracy and stability of the path planning gradually improve. Ultimately, the agent will find the optimal tool axis trajectory that meets the optimization goal within the feasible domain.

[0125] To verify the effect of the present invention, four consecutive feasible domain grids were selected for path planning in this example to simulate the complex tool posture planning task in five-axis CNC machining. In order to more clearly verify the tool axis vector path optimization effect based on the Q-learning algorithm, the following experimental conditions and parameter settings were designed.

[0126] (1) Experimental conditions and parameter settings

[0127] Meshing

[0128] The two-dimensional polar coordinate space (α, β) after the tool axis vector is compressed by C space is divided into 10×10 grids, and the center point of each grid corresponds to a discrete tool axis vector. The four feasible domains define the grid range under the requirements of non-interference and non-singular processing, and mark the feasible grid points inside them. The intelligent agent plans the optimal trajectory of the tool from the initial position to the target position in these feasible domains.

[0129] Reinforcement learning parameter setting

[0130] The number of iterations is 2000, and the learning rate α=0.1 ensures that the agent can quickly absorb new environmental information. The discount factor γ=0.9 gives a higher weight to future rewards to ensure the global optimality of the path. The action exploration strategy selects the ε-greedy exploration strategy with an initial value of ε=1. As the number of iterations increases, it exponentially decays to 0.1 to balance exploration and utilization.

[0131] (2) Experimental results and analysis

[0132] Figure 3 This is the curve of the total reward value changing with the number of iterations during the Q-learning exploration process. The final result is as follows Figure 4 As shown in Figure 1, the optimal path of the tool between the four grids is shown. The optimal path is [36,36,36,66,36], and the maximum total reward value is -6.

[0133] Experiments show that in the initial 700 iterations, the agent is in the exploration stage, resulting in large fluctuations in the reward value. As the number of iterations increases, the agent gradually reduces the rotation of negative paths, and the reward value increases steadily. It tends to stabilize after 1,000 iterations, indicating that the agent has found the global optimal path.

[0134] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for planning the vector motion trajectory of a CNC machining tool axis based on Q-learning, characterized in that: The following steps are involved: S1. Through C-space expression technology, the tool axis vector is compressed from three-dimensional space to two-dimensional polar coordinate space, and then gridded; S2, perform state discrimination on the state space after gridding, and divide it into feasible domain and forbidden domain; S3. Construct a reinforcement learning model; in which the state space is defined as the grid set of the feasible domain of the tool contact points, the action space is defined as the set of possible actions of the agent from the current grid point to all grid points in the next feasible domain, and the reward function comprehensively considers the safety of the path, the kinematic characteristics of the tool, and the amplitude of the tool axis vector change; S4. The Q-learning algorithm is combined with value iteration update to search for the optimal tool axis vector motion trajectory through interaction with the environment and reward feedback.

2. The method for planning the vector motion trajectory of a CNC tool axis based on Q-learning according to claim 1, characterized in that: In step S1, the tool axis vector is compressed from the three-dimensional space to the two-dimensional polar coordinate space through the C space expression technology, specifically: During the machining process, the tool axis vector is represented by a three-dimensional vector (i, j, k) in the workpiece coordinate system. When the tool axis vector is mapped to polar coordinates, only two angle parameters (α, β) are used to represent it, where: In the formula, α represents the direction angle of the tool axis vector on the horizontal plane, and β represents the angle between the tool axis vector and the vertical direction.

3. The Q-learning-based CNC machining tool axis vector motion trajectory planning method according to claim 2 is characterized in that: In step S1, the specific process of gridding is as follows: Discretize the two-dimensional polar coordinate space (α, β) according to a certain resolution and divide α into n α intervals, divide β into n β intervals; each grid center point corresponds to a discrete knife axis vector (α i ,β i ), these grid points constitute the state space of the tool posture.

4. The method for planning the vector motion trajectory of a CNC tool axis based on Q-learning according to claim 1, characterized in that: In step S2, for each grid point in the state space, it is determined whether it satisfies the constraints of no interference and no singularity, and marked, and the grid points that satisfy the constraints of no interference and no singularity are divided into feasible domains, and the remaining grid points are divided into prohibited domains.

5. The Q-learning-based CNC machining tool axis vector motion trajectory planning method according to claim 1, characterized in that: In step S3, when defining the state space, the feasible domain grid set of the knife contact point is used as the state space, wherein the feasible domain grid of each knife contact point is described by two discretized variables α and β, which represent the set of inclination angle and azimuth angle respectively; the state space It is defined as a pair of grids of the feasible region of adjacent knife contacts, expressed as: Among them, k represents the current knife contact number, S k represents the kth state, S k It is expressed as: S k ={(α k ,β k ),(α k+1 ,β k+1 )} in, is the feasible domain mesh corresponding to the kth knife contact point, (α k+1 ,β k+1 ) is the feasible domain mesh corresponding to the k+1th knife contact point.

6. The method for planning the vector motion trajectory of a CNC tool axis based on Q-learning according to claim 5, characterized in that: In step S3, the action space is defined as the set of possible actions of the agent from the current grid point to all grid points in the next feasible domain: A k ={(α k+1 ,β k+1 )∈ feasible domain}.

7. The method for planning the vector motion trajectory of a CNC tool axis based on Q-learning according to claim 5, characterized in that: In step S3, the reward function is defined as follows: Among them, r t (s,a) represents the immediate reward obtained by transferring to the next state s after taking action a at time t. The negative of the Euclidean distance between adjacent feasible region grid points is used as the reward value.

8. The Q-learning-based CNC machining tool axis vector motion trajectory planning method according to claim 1, characterized in that: The specific process of step S4 is: Initialization parameters: Randomly initialize the Q value table Q(S,A); the learning rate α determines the impact of new information on the Q value update, and its value is 0.1; the discount factor γ determines the importance of future rewards, and its value is 0.9; The agent's starting state selection: Randomly select the feasible grid pair of the k-1th and kth knife contact points as the starting point S k ; Action exploration strategy: In order to balance exploration and utilization, the ε-greedy exploration strategy is adopted. k In this case, the agent selects actions according to the following rules Among them, ε is taken as 0.1, which indicates the exploration probability; A k represents the set of all optional actions in the current state, Q(S k ,A') represents state S k The Q value of action A'; State transfer and reward calculation: according to the selected action a, transfer to the next state S' and obtain the immediate reward r(S,a); Q value table update: Where Q(s,a) is the Q value of taking action a in the current state s; r is the immediate reward obtained by taking action a in the current state s; γ is the discount factor that determines the weight of future rewards; is the maximum Q value of all possible actions in the next state s′; α is the learning rate, which determines the influence of new information on the Q value update; Path generation: Starting from the starting state, the intelligent agent gradually selects actions according to the current Q-value table to generate a continuous grid point sequence between the tool contacts; with the iterative update of the Q-value table, the accuracy and optimization effect of path planning are gradually improved, and finally the optimal tool axis vector path that meets the requirements of smoothness and safety is obtained; when the update amplitude of the Q-value table is lower than the predetermined threshold, the iteration is stopped and the optimal path result is output.

Citation Information

Patent Citations

  • Processing method of cutter orientation under cutting ability constraint

    CN107335847A

  • Five-axis machine tool cutter smoothing path generation method

    CN115857429A

  • Method and device for controlling a manipulator

    CN101898358A

  • Grid free-form surface toroidal cutter path planning method based on improved Butterfly subdivision

    CN105739432A

  • Autonomous navigation mobile platform for construction site

    CN114555894A