Machine learning device, control device, processing system, and machine learning method

Through machine learning device learning the correction amount of workpiece model, the problem of shape error during the processing process is solved, and high-precision processing effect is achieved.

CN112947308BActive Publication Date: 2025-06-10FANUC LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202011428167.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-10
Filing Date
2020-12-07
Publication Date
2025-06-10
Estimated Expiration
2040-12-07

AI Technical Summary

Technical Problem

During the robot processing process, there is often an error between the processing results after the workpiece is modeled and the target shape, resulting in a reduction in processing accuracy.

Method used

The correction amount is learned through the machine learning device, and the state observation unit observes the machine tool processing state data and shape error data. The learning unit associates the correction amount with the error for learning, and automatically calculates the most appropriate correction amount to reduce the error.

Benefits of technology

It realizes the rapid determination of the most appropriate correction amount in various processing states, simplifies the operation process, and improves the processing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112947308B_ABST
    Figure CN112947308B_ABST
Patent Text Reader

Abstract

The present invention provides a machine learning device, a control device, a processing system, and a machine learning method, which can reduce the error generated between the processed workpiece and the target shape when processing the workpiece based on a workpiece model that models the target shape of the workpiece. The machine learning device learns a correction amount that corrects the workpiece model so that the shape of the workpiece processed based on the workpiece model that models the workpiece is consistent with the target shape. The machine learning device includes: a state observation unit that observes, as state variables representing the current state of the environment for processing the workpiece, the processing state data of a machine tool that processes the workpiece and the measurement data of the error between the shape of the workpiece processed by the machine tool based on the workpiece model and the target shape; and a learning unit that uses the state variables and learns the correction amount in association with the error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning device, a control device, a machining system, and a machine learning method for learning a correction amount for a workpiece model. Background Art

[0002] There is known a machine learning device that learns the actions of a robot (for example, Japanese Unexamined Patent Application Publication No. 2017-064910). When machining a workpiece based on a workpiece model that models the target shape of the workpiece, an error sometimes occurs between the machined workpiece and the target shape. Therefore, a technique capable of reducing such an error is required. Summary of the Invention

[0003] In one aspect of the present disclosure, a machine learning device learns a correction amount for correcting a workpiece model that models a workpiece so that the shape of the workpiece machined based on the workpiece model matches the target shape. The machine learning device includes: a state observation unit that observes machining state data of a machine tool that machines a workpiece and measurement data of an error between the shape of the workpiece machined by the machine tool based on the workpiece model and the target shape as state variables representing the current state of the environment in which the workpiece is machined; and a learning unit that uses the state variables and learns by associating the correction amount with the error.

[0004] In another aspect of the present disclosure, a machine learning method learns a correction amount for correcting a workpiece model that models a workpiece so that the shape of the workpiece machined based on the workpiece model matches the target shape. In the machine learning method, machining state data of a machine tool that machines a workpiece and measurement data of an error between the shape of the workpiece machined by the machine tool based on the workpiece model and the target shape are observed as state variables representing the current state of the environment in which the workpiece is machined, and the state variables are used to learn by associating the correction amount with the error.

[0005] According to the present disclosure, it is possible to automatically obtain the correction amount of the workpiece model that is most appropriate for reducing the error using the learning result of the learning unit. If the correction amount can be automatically obtained, the most appropriate correction amount can be quickly determined based on the machining state data. Therefore, the operation of obtaining the correction amount in various machining states can be greatly simplified. And since the correction amount is learned based on a large data set, it is possible to accurately obtain the correction amount that is most appropriate for reducing the error. Brief Description of the Drawings

[0006] Figure 1 is a block diagram of a machine learning device according to an embodiment.

[0007] Figure 2 is a perspective view of a machine tool according to an embodiment.

[0008] Figure 3 shows an example of a workpiece produced using the Figure 2 machine tool shown.

[0009] Figure 4 shows a Figure 3 workpiece model that models the target shape of the workpiece shown.

[0010] Figure 5 is a diagram for explaining the error between the Figure 4 workpiece model shown and the workpiece after machining, and shows a case where the error is a protrusion error caused by the workpiece protruding relative to the workpiece model.

[0011] Figure 6 is a diagram for explaining the error between the Figure 4 workpiece model shown and the workpiece after machining, and shows a case where the error is a depression error caused by the workpiece being recessed relative to the workpiece model.

[0012] Figure 7 is a block diagram of a machine learning device according to another embodiment.

[0013] Figure 8 shows an Figure 7 example of the process of a learning cycle executed by the machine learning device shown.

[0014] Figure 9 schematically represents a model of a neuron.

[0015] Figure 10 schematically represents a model of a multi-layer neural network.

[0016] Figure 11 is a block diagram of a machine learning device according to yet another embodiment.

[0017] Figure 12 is a block diagram of a machining system according to one embodiment.

[0018] Figure 13 shows another Figure 2 example of a workpiece produced using the machine tool shown. DETAILED DESCRIPTION

[0019] Embodiments of the present disclosure will be described below with reference to the drawings. In the various embodiments described below, the same reference numerals are assigned to the same elements and repeated descriptions are omitted. First, reference is made to Figure 1 to describe a machine learning device 10 according to one embodiment. The machine learning device 10 is a device for learning a correction amount C such that a machine tool 100 ( Figure 2The workpiece model WM is corrected in such a way that the shape of the workpiece W machined based on the workpiece model WM that models the workpiece W is consistent with a predetermined target shape.

[0020] The following refers to Figure 2 A machine tool 100 according to an embodiment will be described. The machine tool 100 includes: a base stage 102, a translation mechanism 104, a support stage 106, a swing mechanism 108, a swing member 110, a rotation mechanism 112, a workpiece stage 114, a headstock 116, a tool 118, and a spindle movement mechanism 120.

[0021] The base stage 102 has a base plate 122 and a pivot portion 124. The base plate 122 is a substantially rectangular flat member and is disposed above the translation mechanism 104. The pivot portion 124 is integrally formed on the base plate 122 so as to protrude upward from the upper surface 122a of the base plate 122.

[0022] The translation mechanism 104 moves the base stage 102 in the x-axis direction and the y-axis direction in the machine coordinate system CM. Specifically, the translation mechanism 104 includes: an x-axis ball screw mechanism that moves the base stage 102 in the x-axis direction, a y-axis ball screw mechanism that moves the base stage 102 in the y-axis direction, a servo motor that drives the x-axis ball screw mechanism, and a servo motor that drives the y-axis ball screw mechanism (all not shown).

[0023] The support stage 106 is fixed to the base stage 102. Specifically, the support stage 106 has a base portion 126 and a motor housing portion 128. The base portion 126 is a substantially prismatic hollow member and is fixed to the upper surface 122a so as to protrude upward from the upper surface 122a of the base plate 122. The motor housing portion 128 is a substantially semi-circular hollow member and is integrally formed at the upper end portion of the base portion 126. The swing mechanism 108 includes a servo motor and the like and is disposed inside the base portion 126 and the motor housing portion 128. The swing mechanism 108 rotates the swing member 110 about the axis A1.

[0024] The swing member 110 is rotatably supported by the support table 106 and the pivot portion 124. Specifically, the swing member 110 has: a pair of holding portions 130 and 132 disposed opposite to each other in the x-axis direction of the mechanical coordinate system CM, and a motor housing portion 134 fixed to the holding portions 130 and 132. The holding portion 130 is mechanically connected to the swing movement mechanism 108 (specifically, the output shaft of the servo motor). On the other hand, the holding portion 132 is pivotally supported by the pivot portion 124 via a support shaft (not shown). The motor housing portion 134 is a substantially cylindrical hollow member, and is disposed between the holding portions 130 and 132 and integrally formed on the holding portions 130 and 132.

[0025] The rotation movement mechanism 112 includes a servo motor or the like and is disposed inside the motor housing portion 134. The rotation movement mechanism 112 rotates the workpiece stage 114 about the axis A2. The axis A2 is orthogonal to the axis A1 and rotates about the axis A1 together with the swing member 110. The workpiece stage 114 is a substantially disk-shaped member, and the workpiece W is placed above it via a jig (not shown). The workpiece stage 114 is mechanically connected to the rotation movement mechanism 112 (specifically, the output shaft of the servo motor).

[0026] The spindle head 116 is provided so as to be movable in the z-axis direction, and a tool 118 can be detachably attached to its front end. The spindle head 116 rotates the tool 118 about the axis A3 and processes the workpiece W placed on the workpiece stage 114 using the rotating tool 118. The axis A3 is orthogonal to the axis A1. The spindle movement mechanism 120 has, for example, a ball screw mechanism that reciprocates the spindle head 116 in the z-axis direction and a servo motor (both not shown) that drives the ball screw mechanism, and can move the spindle head 116 in the z-axis direction of the mechanical coordinate system CM.

[0027] A mechanical coordinate system CM is set on the machine tool 100. This mechanical coordinate system CM is fixed in a three-dimensional space and is a control coordinate system used as a reference when controlling the operation of the machine tool 100. In the present embodiment, the mechanical coordinate system CM is set such that its x-axis is parallel to the rotation axis A1 of the swing member 110 and its z-axis is parallel to the vertical direction.

[0028] The machine tool 100 relatively moves the tool 118 relative to the workpiece W placed on the workpiece stage 114 in five-axis directions by the translation movement mechanism 104, the swing movement mechanism 108, the rotation movement mechanism 112, and the spindle movement mechanism 120. Therefore, the translation movement mechanism 104, the swing movement mechanism 108, the rotation movement mechanism 112, and the spindle movement mechanism 120 constitute a movement mechanism 136 that relatively moves the tool 118 and the workpiece W.

[0029] The machine tool 100 operates according to the machining program MP, relatively moves the tool 118 and the workpiece W through the moving mechanism 136, and machines the workpiece substrate with the tool 118 rotationally driven by the spindle headstock 116 to form the workpiece W. Figure 3 An example of the workpiece W machined by the machine tool 100 is shown.

[0030] Here, when generating the machining program MP, first, the operator uses a drafting device such as CAD to generate a workpiece model WM1 that models the target shape of the workpiece W as a product. Figure 4 An example of the workpiece model WM1 is shown. In the three-dimensional virtual space where the model is generated by the drafting device, a model coordinate system CW is set, and the surface model SM1 constituting the workpiece model WM1 is defined by model points or model lines set in the model coordinate system CW.

[0031] Next, the operator inputs the generated workpiece model WM1 into a program generation device such as CAM, and the program generation device generates the machining program MP1 based on the workpiece model WM1. The machine tool 100 operates according to the machining program MP1 to machine the workpiece substrate, and as a result, the workpiece W is formed.

[0032] At this time, an error will occur between the shape of the actually formed workpiece W and the target shape of the workpiece W (i.e., the workpiece model WM1). As one of the countermeasures to eliminate this error, the operator operates the drafting device to manually correct the workpiece model WM1, and the program generation device regenerates the machining program MP based on the corrected workpiece model.

[0033] The machine learning device 10 of the present embodiment automatically learns the correction amount C for correcting the workpiece model WM1 in order to eliminate the error. The machine learning device 10 can be composed of a computer having a processor (CPU, GPU, etc.) and a memory (ROM, RAM, etc.) or software such as a learning algorithm.

[0034] As shown in Figure 1 The machine learning device 10 includes a state observation unit 12 and a learning unit 14. The state observation unit 12 observes the machining state data CD of the machine tool 100 and the measurement data of the error δ between the shape of the workpiece W after machining by the machine tool 100 based on the workpiece model WM as state variables SV representing the current state of the environment for machining the workpiece W.

[0035] The machining state data CD is data of parameters that can affect the machining accuracy of the machine tool 100, and includes, for example, at least one of the following: the dimensional error E of the machine tool 100, the temperature T1 of the machine tool 100, the ambient temperature T2 around the machine tool 100, the heat quantity Q of the machine tool 100, the power consumption P of the machine tool 100, the thermal displacement ξ of the machine tool 100, and the operation parameter OP of the machine tool 100.

[0036] The dimensional error E includes, for example, the offset E1 between the axis A1 and the axis A2. Here, it is designed that: in terms of the design dimensions, the rotation axis A1 of the swing member 110 is orthogonal to the rotation axis A2 of the workpiece stage 114, but in the actual machine tool 100, it is possible that: the axis A1 and the axis A2 are not orthogonal but offset from each other. Such an offset E1 may cause a reduction in the machining accuracy of the machine tool 100. The offset E1 can be measured in advance using an offset measuring device and digitized as a vector (distance and direction) in the machine coordinate system CM.

[0037] In addition, the dimensional error E may include: the inclination angle E2 of the axis A1 with respect to the x-axis of the machine coordinate system CM, the inclination angle E3 of the axis A3 with respect to the z-axis of the machine coordinate system CM, the inclination angle E4 of the actual movement path of the base stage 102 with respect to the x-axis or y-axis, etc. These inclination angles E2, E3, and E4 can also be measured using an offset measuring device and digitized as a vector (angle and inclination direction) in the machine coordinate system CM.

[0038] The temperature T1 of the machine tool 100 is the temperature of the components of the machine tool 100 (i.e., the base stage 102, the translational movement mechanism 104, the support stage 106, the swing movement mechanism 108, the swing member 110, the rotational movement mechanism 112, the workpiece stage 114, the spindle headstock 116, the tool 118, the spindle movement mechanism 120). The temperature T1 of the machine tool 100 can be measured during machining or after machining is completed by a first temperature sensor provided in the components of the machine tool 100.

[0039] For example, the first temperature sensor is installed on components prone to thermal displacement such as the x-axis or y-axis ball screw shaft of the translational movement mechanism 104 of the machine tool 100, the output shaft of the servo motor of the swing movement mechanism 108 or the rotational movement mechanism 112, or the ball screw shaft of the spindle movement mechanism 120 of the machine tool 100, and measures the temperature T1 of the component during machining or after machining is completed. The ambient temperature T2 can be measured using a second temperature sensor provided outside the machine tool 100. The second temperature sensor measures the ambient temperature (i.e., the atmospheric temperature) T2 before machining, during machining, or after machining of the machine tool 100.

[0040] The heat quantity Q represents the heat accumulated in the components (such as ball screw shafts) of the machine tool 100 during the machining process. As an example, the above-mentioned first temperature sensor measures the temperature T1_1 before the machine tool 100 performs machining, and then measures the temperature T1_2 at a predetermined moment (or at the end of machining) during the machining process of the machine tool 100. The heat quantity Q can be obtained by the formula Q = B×ΔT based on the difference ΔT (=T1_2 - T1_1) between the temperature T1_1 and the temperature T1_2 and the heat capacity B of the components of the machine tool 100. In addition, the heat quantity Q can also be measured using a calorimeter provided in the machine tool 100.

[0041] The power consumption P is, for example, the power consumed (or input to the machine tool 100) by the machine tool 100 from the start of machining to the end of machining. Specifically, the power (or current or voltage) input to all the servo motors and spindle motors provided in the machine tool 100 can be measured using a wattmeter (or ammeter or voltmeter), and the power consumption P can be measured based on the measured value. As an alternative, the power consumption P can also be the power consumption of each of the multiple (a total of five in this embodiment) servo motors and one or more (one in this embodiment) spindle motors provided in the machine tool 100.

[0042] The thermal displacement amount ξ represents the displacement amount (such as thermal expansion) of the components (such as ball screw shafts) of the machine tool 100 due to the heat generated during the machining process. As an example, the thermal displacement amount ξ can be estimated by calculation by introducing the above-mentioned heat quantity Q into a known experimental formula. As another example, the thermal displacement amount ξ can also be actually measured using a displacement measuring device (displacement meter, linear scale, etc.) during or after the machining process of the machine tool 100.

[0043] The operation parameter OP includes at least one of the following, namely: the acceleration α of the moving mechanism 136 (specifically, the translational moving mechanism 104, the swing moving mechanism 108, the rotational moving mechanism 112, or the spindle moving mechanism 120), the time constant τ that determines the time required for the acceleration or deceleration of the moving mechanism 136, the control gain G that determines the response speed of the control of the moving mechanism 136, and the inertia moment M of the moving mechanism 136.

[0044] For example, as the operation parameter OP, the acceleration α, the time constant τ, the control gain G, and the inertia moment M of the servo motors of the translational moving mechanism 104, the swing moving mechanism 108, the rotational moving mechanism 112, and the spindle moving mechanism 120 can be respectively obtained. In addition, as the acceleration α, the accelerations of the base stage 102 moved by the translational moving mechanism 104 in the x-axis and y-axis directions can also be obtained. The operation parameter OP can be preset by the operator and specified in the machining program MP.

[0045] The error δ can be measured using a measuring device such as a three-dimensional scanner or a three-dimensional measuring instrument with a stereo camera. Specifically, the measuring device measures the shape of the workpiece W after being machined by the machine tool 100. Next, based on the measurement result of the measuring device and the dimensional information of the target shape (workpiece model WM1), the error δ between the shape of the workpiece W and the target shape is measured. In addition, it can also be configured such that the measuring device receives the input of the workpiece model WM1 and calculates the error δ between the shape of the actually measured workpiece W and the shape of the workpiece model WM1.

[0046] The following will refer to Figures 4 - 6 to explain the error δ. Figure 4 It shows the region F where an error δ occurs between the shape of the machined workpiece W measured by the measuring device and the target shape (workpiece model WM1) in the workpiece model WM1. For example, the region F is Figure 5 as shown, a region where the surface SW of the machined workpiece W protrudes outward with respect to the surface model SM1 of the workpiece model WM1 corresponding to the surface SW. Or, the region F is Figure 6 as shown, a region where the surface SW of the machined workpiece W is recessed inward with respect to the surface model SM1 of the workpiece model WM1 corresponding to the surface SW.

[0047] As an example, the error δ includes a plurality of measurement points Pm (m = 1, 2, 3, ···) preset on the workpiece model WM1 and a plurality of errors δm between the corresponding plurality of measurement points Pm' on the machined workpiece W. At this time, the measuring device measures the shape at the plurality of measurement points Pm' on the machined workpiece W. As another example, the error δ can also be the maximum value δmax among the plurality of errors δm, the cumulative value δS = Σδm of the plurality of errors δm, or the average value δA = (Σδm) / m of the plurality of errors δm.

[0048] As still another example, the error δ can also be the volume δV of the region F between the surface SW and the surface model SM1 (that is, the integral value of the error in the region F). At this time, the measuring device can also generate a machined workpiece model MM that models the machined workpiece W based on the measured value of the shape of the machined workpiece W. Based on the machined workpiece model MM and the workpiece model WM1, the volume δV can be obtained. The state observation unit 12 observes the above-mentioned machining state data CD and the measurement data of the error δ as state variables SV.

[0049] The learning unit 14 learns the correction amount C of the workpiece model WM1 according to any learning algorithm collectively referred to as machine learning. Specifically, when the error δ between the shape of the workpiece W machined by the machine tool 100 according to the machining program MP1 and the target shape (workpiece model WM1) is measured, the drawing device corrects the workpiece model WM1 with the correction amount C to generate a new workpiece model WM2. In addition, the correction amount C is represented as a vector (magnitude and direction) in the model coordinate system CW.

[0050] Moreover, the program generation device generates a machining program MP2 based on the workpiece model WM2, and the machine tool 100 machines the workpiece base material according to the machining program MP2 to form the workpiece W. The measuring device measures the measurement data of the error δ between the shape of the machined workpiece W and the target shape. Whenever the correction of the workpiece model WM1 and the attempt to machine based on the corrected workpiece model WM2 are repeatedly executed, the state observation unit 12 observes the state variable SV, and the learning unit 14 repeatedly performs learning based on the data set including the state variable SV.

[0051] By repeating this learning cycle, the learning unit 14 can automatically identify the features that imply the correlation between the correction amount C and the error δ. Although the correlation between the correction amount C and the error δ is actually unknown at the start of the learning algorithm, the learning unit 14 gradually identifies the features and interprets the correlation as the learning progresses.

[0052] When the interpretation of the correlation between the correction amount C and the error δ reaches a level of reliability to a certain extent, the learning results repeatedly output by the learning unit 14 can be used for action selection (i.e., decision-making) on how much the workpiece model WM1 should be corrected to reduce the error δ when machining the workpiece W in the current state.

[0053] As described above, the machine learning device 10 uses the state variable SV (machining state data CD, measurement data δ) observed by the state observation unit 12, and the learning unit 14 learns the correction amount C of the workpiece model WM1 according to the machine learning algorithm. According to the machine learning device 10, the correction amount C that is most appropriate for reducing the error δ can be automatically obtained using the learning results of the learning unit 14.

[0054] If the correction amount C can be automatically obtained, the most appropriate correction amount C can be quickly determined based on the machining state data CD. Therefore, the operation of obtaining the correction amount C in various machining states can be greatly simplified. And since the correction amount C is learned based on a large data set, the correction amount C that is most appropriate for reducing the error δ can be obtained with high precision.

[0055] In addition, the state observation unit 12 can also observe identification information (such as program name, program identification number, etc.) for identifying the machining program MP and use it as the state variable SV. When the machine learning device 10 is constituted by a computer, the processor of the computer performs arithmetic processing to realize the functions of the above-mentioned state observation unit 12 and learning unit 14. On the other hand, when the machine learning device 10 is constituted by software, the machine learning device 10 realizes the functions of the above-mentioned state observation unit 12 and learning unit 14 by executing the computer program contained in the software in resources such as a processor.

[0056] In the machine learning device 10, the learning algorithm executed by the learning unit 14 is not particularly limited. For example, supervised learning, unsupervised learning, reinforcement learning, or neural network, etc. can be adopted as well-known learning algorithms in machine learning. Figure 7 Fig. shows a configuration of the machine learning device 10 in which the learning unit 14 that executes reinforcement learning as an example of the learning algorithm is provided.

[0057] Reinforcement learning is a method as follows: observing the current state (i.e., input) of the environment where the learning object exists and executing a predetermined action (i.e., output) in the current state, and repeatedly executing the cycle of giving some rewards to this action by trial and error, and learning the measure (the correction amount C in this embodiment) that maximizes the total reward as the optimal solution.

[0058] In Figure 7 In the shown machine learning device 10, the learning unit 14 includes: a reward calculation unit 16 that calculates a reward R associated with the error δ; and a function update unit 18 that updates a function EQ representing the value of the correction amount C using the reward R. The learning unit 14 learns the correction amount C by repeatedly updating the function EQ by the function update unit 18.

[0059] An example of the algorithm of the reinforcement learning executed by the learning unit 14 will be described below. The algorithm of this example is the well-known Q-learning, that is, using the state s of the agent and the action a that the agent can select in the state s as independent variables to learn the function EQ(s, a) representing the value of the action when the action a is selected in the state s.

[0060] Selecting the action a with the highest value function EQ in state s is the optimal solution. Q-learning starts in a state where the correlation between state s and action a is unknown, and repeats trial and error of selecting various actions a in any state s, thereby repeatedly updating the value function EQ and approaching the optimal solution. The structure here is as follows: As a result of selecting action a in state s, when the environment (i.e., state s) changes, a reward (i.e., the weight of action a) r corresponding to this change can be obtained, and by guiding the learning to select an action a that can obtain a higher reward R, the value function EQ can be made to approach the optimal solution in a shorter time.

[0061] The update formula of the value function EQ can generally be expressed by the following formula (1).

[0062]

[0063] In formula (1), s t and a t are the state and action at time t respectively. The state changes to s t due to action a t+1 . r t+1 is the reward obtained by the state changing from s t to s t+1 . The term maxQ represents Q when the action a that becomes the maximum value Q (considered at time t) is taken at time t + 1. α and γ are the learning coefficient and the discount rate respectively, and can be arbitrarily set according to 0 < α ≤ 1, 0 < γ ≤ 1.

[0064] When Q-learning is executed in the learning unit 14, the state variable SV observed by the state observation unit 12 corresponds to the state s in the update formula, and the action (i.e., the correction amount C) of how much the workpiece model WM1 should be corrected when machining the workpiece W in the current state corresponds to the action a in the update formula. In addition, the reward R calculated by the reward calculation unit 16 corresponds to the reward r in the update formula. The function update unit 18 repeatedly updates the function EQ representing the value of the correction amount C when machining the workpiece W in the current state through Q-learning using the reward R.

[0065] For example, regarding the reward R calculated by the reward calculation unit 16, it is a positive (plus) reward R when the error δ is smaller than the predetermined threshold ΔTh1, and a negative (minus) reward R when the error δ is equal to or greater than the threshold ΔTh1. The absolute values of the positive and negative rewards R can be the same or different from each other.

[0066] In addition, the reward calculation unit 16 can obtain a reward R that varies according to the magnitude of the error δ. For example, the reward calculation unit 16 can set the reward R = +5 when 0 ≤ δ < δth2 (< δth1), set the reward R = +2 when δth2 ≤ δ < δth3 (< δth1), and set the reward R = +1 when δth3 ≤ δ < δth1.

[0067] On the other hand, the reward calculation unit 16 can set the reward R = -1 when δth1 ≤ δ < δth4, set the reward R = -2 when δth4 < δ ≤ δth5, and set the reward R = -5 when δth5 < δ. That is, at this time, the smaller the error δ, the larger the value of the reward R calculated by the reward calculation unit 16. By obtaining such a conditionally weighted reward R, Q-learning can converge to the optimal solution in a shorter time.

[0068] In addition, the reward calculation unit 16 can obtain a reward R that varies according to the difference in the processing state data CD. For example, the reward calculation unit 16 sets a positive reward R with a larger value when the error δ is smaller than the threshold value δth1 and the control gain G included in the operation parameter OP of the processing state data CD is within a predetermined allowable range. In addition, the reward calculation unit 16 sets a positive reward R with a larger value when the error δ is smaller than the threshold value δth1 and the time constant included in the operation parameter OP is within a predetermined allowable range. At this time, it is possible to promote the learning of the correction amount C to reduce the error δ under the condition of speeding up the operation of the moving mechanism 136 of the machine tool 100.

[0069] The function update unit 18 can have an action value table, in which the state variable SV and the reward R are organized in association with the action value (for example, a numerical value) represented by the function EQ. At this time, the behavior of the function update unit 18 to update the function EQ has the same meaning as the behavior of the function update unit 18 to update the action value table.

[0070] At the start of Q-learning, the correlation between the current state of the environment and the correction amount C is unknown. Therefore, various state variables SV and rewards R are prepared in the action value table in association with the values of the action values (function EQ) set as random. In addition, if the reward calculation unit 16 obtains the error δ, it can directly calculate the corresponding reward R and record the calculated value of the reward R in the action value table.

[0071] When performing Q-learning using the reward R corresponding to the error δ, the learning is guided in the direction of selecting an action (i.e., the correction amount C) that can obtain a higher reward R. And, according to the state of the environment (i.e., the state variable SV) that changes as a result of performing the selected action in the current state, the value of the action value (function EQ) related to the action performed in the current state is rewritten, and the action value table is updated.

[0072] By repeating this update, the value of the action shown in the action value table (function EQ) is rewritten as a larger value appropriate for the action (correction amount C). This gradually clarifies the current state of the previously unknown environment (error δ) and its correlation with the action (correction amount C).

[0073] Refer to the following Figure 8 for Figure 7 an example of the learning process of the machine learning device 10 shown. The process shown starts when measuring the error δ between the shape of the workpiece W machined by the machine tool 100 according to the machining program MP1 and the target shape (workpiece model WM1). Figure 8 the process shown.

[0074] In step S1, the function update unit 18 refers to the action value table at that moment and selects the correction amount C as the action to be performed in the current state. For example, the function update unit 18 obtains the workpiece model WM1 from the drawing device and obtains the measurement data of the latest measured error δ.

[0075] Furthermore, the function update unit 18 determines the region F on the workpiece model WM1 based on the measurement data of the error δ ( Figure 4 ). And the function update unit 18 randomly selects the correction amount C for correcting the components (model points, model lines, surface model SM1) of the workpiece model WM1 existing in the region F.

[0076] Here, it can also be configured such that the function update unit 18 randomly selects the correction amount C under a predetermined condition that restricts the magnitude and direction of the correction amount C. For example, when an error δ as shown in Figure 5 occurs in the region F, the function update unit 18 can select the direction of the correction amount C for correcting the surface model SM1 as the direction D1 opposite to the direction in which the error δ (protrusion error) occurs (that is, on the opposite side of the surface SW with respect to the surface model SM1 in Figure 5 ). On the other hand, when an error δ (recess error) as shown in Figure 6 occurs in the region F, the function update unit 18 can select the direction of the correction amount C as the direction D2 opposite to the direction in which the error δ occurs.

[0077] In addition, the function update unit 18 can also select the magnitude of the correction amount C, that is, |C|, within a numerical range determined based on the error δ. For example, when setting the maximum value of the error δ in the region F as δmax, this numerical range can be set as 0 < |C| ≤ δmax. In addition, the function update unit 18 can also select the position where the workpiece model WM1 is corrected with the correction amount C as the position of the component (such as a model point) of the workpiece model WM1 where an error δ of a predetermined magnitude (for example, the maximum value δmax) occurs.

[0078] In step S2, the function update unit 18 takes in the state variable SV. Specifically, when the function update unit 18 selects the correction amount C in step S1, the drawing device corrects the components (model points, model lines, surface model SM1) of the workpiece model WM1 by the correction amount C in the model coordinate system CW to generate the workpiece model WM2. Then, the program generation device generates the machining program MP2 based on the workpiece model WM2, and the machine tool 100 machines the workpiece W according to the machining program MP2. The measuring device measures the error δ between the shape of the machined workpiece W and the target shape (workpiece model WM1).

[0079] In this step S2, the state observation unit 12 observes the machining state data CD when the machine tool 100 machines the workpiece W according to the machining program MP2 and the measurement data of the error δ between the shape of the machined workpiece W and the target shape as the state variable SV. The function update unit 18 takes in the state variable SV observed by the state observation unit 12.

[0080] In step S3, the function update unit 18 determines whether the error δ taken in in the latest step S2 is greater than or equal to the threshold value δth1. When δ≥δth1, the function update unit 18 determines "yes" and proceeds to step S5, and when δ<δth1, the function update unit 18 determines "no" and proceeds to step S4.

[0081] In step S4, the reward calculation unit 16 obtains a positive reward R. At this time, the reward calculation unit 16 can also obtain the reward R that varies according to the magnitude of the error δ as described above (specifically, the smaller the error δ, the larger the value of the reward R). The reward calculation unit 16 applies the obtained positive reward R to the update formula of the function EQ. By setting the reward R to be larger as the error δ becomes smaller in this way, the learning of the learning unit 14 is guided in the direction of selecting actions that reduce the error δ.

[0082] In step S5, the reward calculation unit 16 obtains a negative reward R and applies it to the update formula of the function EQ. At this time, the reward calculation unit 16 can also obtain the negative reward R whose absolute value becomes larger as the error δ becomes larger as described above. In addition, in this step S4, the reward calculation unit 16 can also set the reward R to 0 instead of the negative reward R and apply it to the update formula of the function EQ.

[0083] In step S6, the function update unit 18 updates the action value table (function EQ) using the state variable SV and the reward R in the current state. In this way, the learning unit 14 repeatedly updates the action value table and learns the correction amount C by repeating steps S1 to S6.

[0084] When performing the above-mentioned reinforcement learning, for example, a neural network can be used instead of Q-learning. Figure 9A model schematically representing a neuron. Figure 10 Schematically represents combining Figure 9 The model of a three - layer neural network formed by combining the neurons shown. The neural network can be composed of, for example, a processor and a memory that simulate the model of a neuron.

[0085] Figure 9 The neuron shown outputs a result y with respect to a plurality of inputs x (for example, inputs x1 to x3 in the figure). And each input x (x1, x2, x3) is multiplied by a weight w (w1, w2, w3) respectively. The relationship between the input x and the result y can be represented by the following formula (2). In addition, the input x, the result y, and the weight w are all vectors. And in formula (2), θ is the bias, and f k Is an activation function.

[0086]

[0087] Figure 10 The three - layer neural network shown inputs a plurality of inputs x (for example, inputs x1 to input x3 in the figure) from the left and outputs a result y (for example, results y1 to y3 in the figure) from the right. In the illustrated example, the inputs x1, x2, x3 are multiplied by the corresponding weights (collectively denoted as ω1), and all the inputs x1, x2, x3 are input to the three neurons N11, N12, N13.

[0088] In Figure 10 The outputs of the neurons N11 to N13 are collectively denoted as Z1. Z1 can be regarded as a feature vector that extracts the feature quantity of the input vector. In the illustrated example, the feature vector Z1 is multiplied by the corresponding weights (collectively denoted as ω2), and all the feature vectors Z1 are input to the two neurons N21, N22. The feature vector Z1 represents the feature between the weight w1 and the weight w2.

[0089] Figure 10 The outputs of the neurons N21 to N22 are collectively denoted as Z2. Z2 can be regarded as a feature vector that extracts the feature quantity of the feature vector Z1. In the illustrated example, the feature vector Z2 is multiplied by the corresponding weights (collectively denoted as ω3), and all the feature vectors Z2 are input to the three neurons N31, N32, N33. The feature vector Z2 represents the feature between the weight ω2 and the weight ω3. Finally, the neurons N31 to N33 output the results y1 to y3 respectively.

[0090] In the machine learning device 10, with the state variable SV as the input x, the learning unit 14 performs operations based on the multi-layer structure of the above neural network, so as to be able to output the correction amount C (result y). In addition, the operation modes of the neural network include a learning mode and a value prediction mode. For example, in the learning mode, a learning data set is used to learn the weights ω, and the learned weights ω can be used to judge the value of actions in the value prediction mode. In addition, in the value prediction mode, detection, classification, derivation, etc. can also be performed.

[0091] The structure of the above machine learning device 10 can be expressed as a machine learning method (or software) executed by a processor of a computer. In this machine learning method, the processor observes the processing state data CD of the machine tool 100 and the measurement data of the error δ between the shape of the workpiece W machined by the machine tool 100 based on the workpiece model WM and the target shape as the state variable SV representing the current state of the environment for machining the workpiece W, and learns the correction amount C in association with the state variable SV and the error δ.

[0092] Figure 11 Another mode of the machine learning device 10 is shown. The machine learning device 10 further includes a decision-making unit 20. The decision-making unit 20 outputs the output value of the correction amount C based on the learning result (action value table) of the learning unit 14. When the decision-making unit 20 outputs the output value C, the state (error δ) of the environment 140 for machining the workpiece W will change accordingly.

[0093] That is, the decision-making unit 20 outputs the output value C to a drawing device, and the drawing device corrects the components (model points, model lines, surface model SM1) of the workpiece model WM1 according to the output value C in the model coordinate system CW to generate a workpiece model WM2. And the program generation device generates a machining program MP2 based on the workpiece model WM2, and the machine tool 100 machines the workpiece W according to the machining program MP2. The measuring device measures the error δ between the shape of the machined workpiece W and the target shape, and the state observation unit 12 observes the state variable SV with this error δ as the measurement data in the next learning cycle.

[0094] The learning unit 14 updates, for example, a value function EQ (i.e., an action value table) using the changed state variable SV to learn the correction amount C. The decision-making unit 20 outputs the most appropriate output value C based on the learned correction amount C and according to the state variable SV. The machine learning device 10 repeats this cycle to learn the correction amount C and gradually improves the reliability of the correction amount C.

[0095] Adopt Figure 11The illustrated machine learning device 10 can change the state of the environment 140 according to the output of the decision-making unit 20. In addition, it is possible to obtain, in an external device, a function equivalent to that of the decision-making unit used in the machine learning device 10 to reflect the learning result of the learning unit 14 in the environment 140.

[0096] Next, a processing system 150 according to an embodiment will be described with reference to Figure 12 The processing system 150 includes a machine tool 100, a drawing device 152, a program generation device 154, a measurement device 156, a sensor 158, and a control device 160. The drawing device 152 is a device (such as CAD) that can generate a workpiece model WM as described above, and can be composed of a computer or software having a processor and a memory.

[0097] As described above, the program generation device 154 is a device (such as CAM) that can generate a machining program MP based on the workpiece model WM, and can be composed of a computer or software having a processor and a memory. In addition, the drawing device 152 and the program generation device 154 may be integrated into a computer-aided design device, which is a single computer having a processor and a memory. As described above, the measurement device 156 is a three-dimensional scanner or a three-dimensional measuring instrument having a stereo camera, etc., and measures the measurement error δ and sends the measurement data of the error δ to the control device 160.

[0098] The sensor 158 is used to measure the dimensional error E, temperature T1, ambient temperature T2, heat quantity Q, power consumption P, and thermal displacement amount ξ in the machining state data CD, and includes the above-mentioned offset measuring device, temperature sensor, calorimeter, power meter (voltmeter or ammeter), and displacement measuring device. The sensor 158 measures the dimensional error E, temperature T1, ambient temperature T2, heat quantity Q, power consumption P, and thermal displacement amount ξ and sends them to the control device 160 as the machining state data CD.

[0099] The control device 160 has a processor (such as CPU, GPU, etc.) 162 and a memory (such as ROM, RAM, etc.) 164. The processor 162 is communicably connected to the memory 164 via a bus 166, and communicates with the memory 164 and performs various operations. The control device 160 is communicably connected to the machine tool 100 (specifically, the moving mechanism 136), the drawing device 152, the program generation device 154, the measurement device 156, and the sensor 158, and controls the operations of these components.

[0100] In this embodiment, the machine learning device 10 is installed in the control device 160, and the processor 162 functions as the above-mentioned state observation unit 12, learning unit 14 (reward calculation unit 16 and function update unit 18), and decision-making unit 20. In addition, the processor 162 acquires the processing state data CD and the measurement data δ. Specifically, the processor 162 acquires the dimensional error E, temperature T1, ambient temperature T2, heat quantity Q, power consumption P, and thermal displacement amount ξ as the processing state data CD from the sensor 158.

[0101] Moreover, the processor 162 acquires the operation parameter OP as the processing state data CD. For example, the operation parameter OP (acceleration α, time constant τ, control gain G, inertia moment M) is preset by the operator and stored in the memory 164. The processor 162 reads and obtains the operation parameter OP from the memory 164. And, the processor 162 acquires the measurement data of the error δ from the measurement device 156. Thus, in this embodiment, the processor 162 functions as the state data acquisition unit 168 that acquires the processing state data CD and the measurement data of the error δ.

[0102] The processor 162 functions as the machine learning device 10 and can automatically learn the correction amount δ in cooperation with the machine tool 100, drawing device 152, program generation device 154, measurement device 156, and sensor 158. For example, the processor 162 can learn the most appropriate correction amount δ by executing Figure 8 the learning process shown.

[0103] In addition, the processing system 150 may further include a workpiece handling robot (not shown). The workpiece handling robot places the workpiece substrate stored at a predetermined location on the workpiece stage 114 of the machine tool 100, and takes out the workpiece W from the workpiece stage 114 after processing the workpiece substrate. And, the workpiece handling robot places the processed workpiece W in the measurement device 156, and after the measurement device 156 measures the shape and error δ of the workpiece W, takes out the workpiece W from the measurement device 156.

[0104] The processor 162 controls the workpiece handling robot to perform the above-mentioned loading and unloading of the workpiece W. With this structure, the processor 162 can automatically execute, for example, Figure 8 the machine learning process shown without manual operation by the operator.

[0105] On the other hand, the operator may also manually perform at least one process in the machine learning process. For example, the operator can operate the drawing device to manually generate the workpiece model WM, or operate the program generation device to manually generate the machining program MP.

[0106] In addition, in the above-described embodiment, for ease of understanding, the case where the region F where the error δ is generated is one has been described. However, in practice, the error δ may sometimes be generated in a plurality of regions Fi (i = 1, 2, 3, ···). At this time, the machine learning device 10 executes the above-described machine learning method for each region Fi. For example, in the case of the machine learning device 10 shown in Figure 7 the machine learning device 10 sequentially executes for each region Fi Figure 8 the process shown. Thereby, the most appropriate correction amount C can be learned in each region Fi.

[0107] In addition, in the above-described embodiment, for ease of understanding, the workpiece W having a simple shape as shown in Figure 2 has been described as an example, but the shape of the workpiece is not limited. For example, the machine learning device 10 can also execute the above-described machine learning method for the workpiece W2 as shown in Figure 13 to learn the most appropriate correction amount C. Figure 13 The workpiece W2 shown is an impeller for a fluid device such as a compressor, and has a base WA and a wing portion WB that extends curvilinearly outward from the base WA. The workpiece W2 can be machined using the machine tool 100.

[0108] In addition, the machine tool 100 is not limited to the above structure and can be of any type. For example, the machine tool 100 is not limited to cutting with the tool 118. For example, it may also be provided with a laser processing head and use a laser beam emitted from the laser processing head to machine the workpiece W.

[0109] Alternatively, instead of the above-described moving mechanism 136, a vertical multi-joint type, a horizontal multi-joint type, or a parallel link type robot may be applied as a moving mechanism for relatively moving the tool 118 (or the laser processing head) with respect to the workpiece W. At this time, the robot has a drive unit that rotationally drives the tool 118, and the machine tool 100 moves the tool 118 relative to the workpiece W using the robot and machines the workpiece W with the tool 118.

[0110] Alternatively, in Figure 12 the embodiment shown, at least one of the drawing device 152 and the program generation device 154 may be integrated as software in the control device 160. The present disclosure has been described above through embodiments, but the above-described embodiments do not limit the invention of the claims. In addition, the state observation unit 12 may also observe the correction amount C as the state variable SV. At this time, a correction amount acquisition unit for acquiring the correction amount C may be provided.

Claims

1. A machine learning device that learns a correction amount for correcting a workpiece model that models a target shape of a workpiece as a CAD model so that the shape of the workpiece machined based on the workpiece model coincides with the target shape. The machine learning device is characterized in that it includes a processor that learns the correction amount by repeatedly executing the following learning cycle: Randomly select a correction amount for correcting the initial workpiece model. Observe, as state variables, the machining state data of the machine tool when machining the workpiece with the workpiece model obtained by correcting the initial workpiece model with the selected correction amount and the measurement data of the error between the shape of the workpiece machined by the machine tool based on the corrected workpiece model and the target shape. Use the state variables and learn by associating the correction amount with the error.

2. The machine learning device according to claim 1. It is characterized in that The machining state data includes at least one of the dimensional error of the machine tool, the temperature of the machine tool, the ambient temperature around the machine tool, the heat of the machine tool, the power consumption of the machine tool, the thermal displacement amount of the machine tool, and the motion parameters of the machine tool.

3. The machine learning device according to claim 2. It is characterized in that The machine tool has: A tool for machining the workpiece; and A moving mechanism for relatively moving the tool and the workpiece. The motion parameters include at least one of the acceleration of the moving mechanism, the time constant for determining the time required for the moving mechanism to accelerate or decelerate, the control gain for determining the response speed of the control of the moving mechanism, and the inertia moment of the moving mechanism.

4. The machine learning device according to any one of claims 1 to 3. It is characterized in that The processor calculates a reward associated with the error and uses the reward to update a function representing the value of the correction amount.

5. The machine learning device according to claim 4. It is characterized in that The processor calculates different rewards according to the magnitude of the error.

6. The machine learning device according to any one of claims 1 to 3. It is characterized in that The processor outputs an output value of the correction amount based on the learning result, and observes the state variables using, as the measurement data in the next learning cycle, the error between the shape of the workpiece machined by the machine tool based on the workpiece model obtained by correcting the initial workpiece model according to the output value and the target shape.

7. A control device that controls a machine tool. The control device is characterized in that it includes: The machine learning device according to any one of claims 1 to 3; and A state data acquisition unit that acquires the machining state data and the measurement data.

8. A machining system It is characterized in that It includes: A machine tool for machining a workpiece; A measuring device for measuring the error between the shape of the workpiece machined by the machine tool and the target shape of the workpiece set in advance; and The control device according to claim 7.

9. A machine learning method that learns a correction amount for correcting a workpiece model that models the target shape of a workpiece as a CAD model, such that the shape of the workpiece machined based on the workpiece model is consistent with the target shape. The machine learning method is characterized in that: The processor learns the correction amount by repeatedly executing the following learning cycle: Randomly select a correction amount for correcting the initial workpiece model; Use the machining state data of the machine tool when machining the workpiece with the workpiece model obtained by correcting the initial workpiece model with the selected correction amount and the measurement data of the error between the shape of the workpiece machined by the machine tool based on the corrected workpiece model and the target shape as state variables for observation; and Learn using the state variables and associating the correction amount with the error.

Citation Information

Patent Citations

  • Machine learning device for learning taking-out operation of workpiece, robot system, and machine learning method

    JP2017064910A

  • Method and device for producing a master die tool

    US20110130854A1

  • Controller-equipped machining apparatus having machining time measurement function and on-machine measurement function

    US20170031328A1

  • Machine tool for generating optimum acceleration / deceleration

    US20170090459A1

  • Machine learning device and method for optimizing frequency of tool compensation of machine tool, and machine tool having the machine learning device

    US20170091667A1