Machine learning device, classification device, and control device
By extracting features from video and runtime data using machine learning devices and generating a learned model using semi-supervised learning, the problem of large training data requirements in existing technologies is solved, and high-precision learning model generation and judgment are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-07
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies require a large amount of training data to generate a high-precision, fully learned model, which leads to problems with the effort and time spent on data collection.
By extracting features from video and runtime data using machine learning devices, generating a completed learning model using semi-supervised learning methods, and assigning labels to unlabeled data using the correlation features between video and runtime data, a completed learning model with high judgment accuracy is constructed.
Even with limited training data, it is possible to generate highly accurate learned models, simplifying the data collection process and improving decision-making efficiency.
Smart Images

Figure CN117296067B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to machine learning devices, classification devices, and control devices. Background Technology
[0002] In factories and other production sites, in order to improve operational efficiency, the work actions of operators on the production line are analyzed to improve production equipment and work content.
[0003] For example, the following technology has been proposed: Operators wear sensors and lasers, and a video camera is used to film the operator. From the tracking data measured by the sensors, lasers, and the video camera, tact unit operations and minimum unit operations are extracted. Feature vectors are calculated based on the extracted minimum unit operations. The calculated feature vectors are then used to analyze and process the operator's actions, detecting errors, changes in the operator's posture, and the actual operation time. See, for example, Patent Document 1.
[0004] Existing technical documents
[0005] Patent documents
[0006] Patent Document 1: Japanese Patent Application Publication No. 2006-209468 Summary of the Invention
[0007] The problem that the invention aims to solve
[0008] Preparation and other tasks vary depending on the factory or machine. Therefore, in order to generate a learned model or classifier that makes high-precision judgments, training data needs to be prepared according to the site or machine.
[0009] However, in order to perform generative learning to complete deep learning models or classifiers, a large amount of training data is required, which presents the problem of spending effort and time collecting training data.
[0010] Therefore, it is expected that a learning-complete model with high decision accuracy can be generated even with less training data.
[0011] Methods for solving problems
[0012] (1) One aspect of the machine learning apparatus of this disclosure includes: a video data feature extraction unit that extracts features representing the actions of an operator from video data containing an operator captured by at least one camera; a training data extraction unit that, when extracting features representing a pre-registered specific action from the features extracted by the video data feature extraction unit, labels the extracted work content represented by the specific action to the corresponding video data, and extracts training data consisting of input data of the labeled video data and label data of the work content; and an operation data feature extraction unit that extracts features associated with the extracted training data from the operation data of industrial machinery. The system comprises: a feature quantity; a labeling benchmark creation unit, which creates a labeling benchmark for video data and a labeling benchmark for running data based on the feature quantity of the extracted video data and the feature quantity of the running data from the training data; a labeling unit, which labels the unlabeled video data and the running data based on the labeling benchmark for video data and the labeling benchmark for running data; and a learning unit, which performs machine learning using the training data and generates a learned model, wherein the training data includes the video data labeled by the labeling unit, and the learned model is used to classify the tasks performed by the operator based on the input video data.
[0013] (2) One aspect of the classification apparatus of the present disclosure includes: a learning completion model generated by the machine learning apparatus of (1); an input unit that inputs video data containing an operator captured by at least one camera; and a job determination unit that determines the operator's job from the video data input by the input unit based on the learning completion model.
[0014] (3) One aspect of the control device of this disclosure has the classification device of (2).
[0015] Invention Effects
[0016] According to one method, a learning-completed model with high judgment accuracy can be generated even with less training data. Attached Figure Description
[0017] Figure 1 This is a functional block diagram illustrating an example of the functional structure of a control system in one implementation.
[0018] Figure 2 This is an example of a video captured by a camera.
[0019] Figure 3 This is a graph representing an example of a time series of tasks.
[0020] Figure 4AThis is an example of a gesture used to indicate a tool change.
[0021] Figure 4B This is an example of a gesture used to indicate a tool change.
[0022] Figure 5 This is a diagram illustrating an example of the preparation time width used for product A.
[0023] Figure 6A This is a graph representing an example of the training data for each job.
[0024] Figure 6B This is a graph representing an example of the training data for each job.
[0025] Figure 7 This is a diagram representing an example of data stored in the decision result storage unit.
[0026] Figure 8 This is a flowchart illustrating the classification process of the classification device during the application phase. Detailed Implementation
[0027] Hereinafter, an embodiment of the present disclosure will be described using the accompanying drawings. Here, a machine tool is illustrated as industrial machinery. Furthermore, the present invention can also be applied to industrial robots, service robots, forging machines, and injection molding machines, among other industrial machinery.
[0028] <One Implementation Method>
[0029] Figure 1 This is a functional block diagram illustrating an example of the functional structure of a control system in one implementation. For example... Figure 1 As shown, the control system 1 includes: a machine tool 10, a camera 20, a classification device 30, and a machine learning device 40.
[0030] The machine tool 10, the camera 20 that captures video (moving images) at a predetermined frame rate, the classification device 30, and the machine learning device 40 can be directly connected to each other via a connection interface not shown. Alternatively, the machine tool 10, camera 20, classification device 30, and machine learning device 40 can also be interconnected via a network not shown, such as a LAN (Local Area Network) or the Internet. In this case, the machine tool 10, camera 20, classification device 30, and machine learning device 40 have a communication unit not shown for communicating with each other via this connection. Furthermore, as described later, the control device 110 included in the machine tool 10 may also include the classification device 30 and the machine learning device 40.
[0031] Machine tool 10 is a machine tool well known to those skilled in the art, and includes a control device 110. Machine tool 10 operates according to action commands from control device 110.
[0032] The control device 110, for example, is a numerical control device known to those skilled in the art, that generates action commands based on control information and sends the generated action commands to the machine tool 10. Thus, the control device 110 controls the actions of the machine tool 10.
[0033] Specifically, the control device 110 is a device that controls the machine tool 10 to perform prescribed machining operations. A machining program describing the actions of the machine tool 10 is provided to the control device 110. The control device 110 generates motion commands based on the provided machining program and sends these motion commands to the machine tool 10, thereby controlling the motors of the machine tool 10. These motion commands include movement commands for each axis, rotation commands for the motor driving the spindle, etc. Thus, the prescribed machining operations of the machine tool 10 are executed.
[0034] Furthermore, the control device 110 sends the machine tool 10's operating data, including action commands, door opening and closing, motor torque values, etc., to the sorting device 30. Also, the control device 110 can append time information indicating the time when the action command, door opening and closing, motor torque values, etc., were measured, to the operating data and output it to the sorting device 30, based on the clock signal of the clock (not shown) included in the control device 110. Additionally, the control device 110 can, for example, check the time of the clock (not shown) included in the camera 20 (described later) at predetermined time intervals and synchronize them.
[0035] Furthermore, if the machine tool 10 is a robot or the like, the control device 110 may be a robot control device or the like.
[0036] Furthermore, the control device 110 can be used to control various types of machinery, including but not limited to machine tools 10 and robots. Industrial machinery includes machine tools, industrial robots, service robots, forging machines, and injection molding machines.
[0037] In this embodiment, a numerical control device is exemplified as the control device 110.
[0038] The camera 20, such as a digital camera like a surveillance camera, is installed in the factory where the machine tool 10 is located. The camera 20 outputs the frame images captured at a predetermined frame rate as video to the classification device 30. In addition, the camera 20 obtains the time when each frame image is captured based on the clock signal of the clock (not shown) included in the camera 20, appends the time information indicating the time of acquisition to the video, and outputs it to the classification device 30.
[0039] Figure 2 This is an example diagram showing a video captured by camera 20. Figure 2 The text indicates that the video was captured by camera 20 from three machine tools 10.
[0040] In addition, Figure 1 The machine tool 10 is equipped with one camera 20, but it can also be equipped with two or more cameras 20 to film the machine tool 10.
[0041] The classification device 30 can also acquire video footage captured by the camera 20 with added time information during the application phase. Additionally, the classification device 30 can acquire operating data of the machine tool 10 from the control device 110, for example. By inputting the acquired video into a learning model provided by the machine learning device 40 (described later), the classification device 30 can determine (classify) the work performed by the operator on the machine tool 10 as captured in the video.
[0042] In addition, the classification device 30 can also send the obtained machine tool 10 operating data and the video data captured by the camera 20 to the machine learning device 40 described later.
[0043] Furthermore, during the application phase, the classification device 30 can receive the completed learning model generated by the machine learning device 40 (described later), input video data captured by the camera 20 into the received completed learning model, thereby determining (classifying) the operator's work content within the video, and displaying the determination result on a display unit (not shown) such as an LCD screen included in the control device 110. Moreover, if the determination result is incorrect, the classification device 30 obtains the correct work content input into the control device 110, outputs the input data of the determined video data and the label data of the obtained correct work content to the machine learning device 40, thereby causing the machine learning device 40 to update the completed learning model.
[0044] Before describing the classification device 30, the machine learning used to generate the learned model will be explained.
[0045] <Machine Learning Device 40>
[0046] The machine learning device 40 acquires video data captured by the camera 20, and, as described later, extracts feature quantities representing the operator's actions from the acquired video data.
[0047] When the machine learning device 40 extracts the feature quantity of the gesture of the specific action of the operator as described below from the extracted feature quantity, it labels the work content represented by the extracted gesture to the corresponding video data, and extracts the training data of the input data of the labeled video data and the label data of the work content.
[0048] In addition, the machine learning device 40 acquires the operating data of the machine tool 10, and extracts the feature quantities associated with the training data from the acquired operating data as described below. As described below, based on the extracted feature quantities of the training data and the operating data, it creates a labeling benchmark for video data and a labeling benchmark for operating data, respectively.
[0049] The machine learning device 40 labels at least the unlabeled video data based on the labeling benchmarks used for video data and the labeling benchmarks used for running data.
[0050] The machine learning device 40 performs supervised learning using training data containing newly labeled video data to build the learning completion model, which will be described later.
[0051] Thus, the machine learning device 40 can build a high-accuracy learning model even with less training data, and can provide the built learning model to the classification device 30.
[0052] The machine learning device 40 will be described in detail.
[0053] like Figure 1 As shown, the machine learning device 40 includes: a video data acquisition unit 401, a video data feature extraction unit 402, a training data extraction unit 403, a running data acquisition unit 404, a running data feature extraction unit 405, a label assignment benchmark creation unit 406, a label assignment unit 407, a learning unit 408, and a storage unit 409.
[0054] Storage unit 409, such as ROM (Read Only Memory) or HDD (Hard Disk Drive), stores system programs and machine learning applications to be executed by the processor (not shown) included in the machine learning device 40. Additionally, storage unit 409 stores the learning completed model generated by learning unit 408 (described later) and has a training data extraction signature storage unit 4091.
[0055] The training data extraction signature storage unit 4091 stores, for example, the job content assigned as a label and the specific action (gesture, etc.) representing each job content.
[0056] Figure 3 This is a graph representing an example of a time series of tasks.
[0057] like Figure 3 As shown, the job content includes "standby", "preparation for product A", "cleaning", and "tool replacement". Alternatively, the job content can also be defined as "preparation for product B", etc., freely defined by the operator or other users.
[0058] Figure 4A and Figure 4B This is an example of a gesture used to indicate a tool change. Figure 4A This shows an example of a gesture indicating the start of a tool change. Figure 4B This shows an example of a gesture indicating the end of a tool change.
[0059] That is, at the beginning and end of each task, the operator performs pre-registered gestures (specific actions) towards the camera 20, indicating which task is being performed. This allows the video data to be labeled with the task content represented by the gestures, which can then be used as training data.
[0060] Therefore, it is easy to collect training data for generating a learning completion model that determines the content of a task based on video data.
[0061] The video data acquisition unit 401 acquires video data including the operator captured by the camera 20 via the classification device 30.
[0062] The video data feature extraction unit 402 extracts features representing the operator's actions from the acquired video data.
[0063] Specifically, the video data feature extraction unit 402 uses known methods (e.g., Kosuke Kanno, Kenta Oku, Kyoji Kawagoe, "Multidimensional Time Series Data Motion Extraction and Classification Methods", DEIM Forum 2016 G4-5; or Shohei Uezono, Satoshi Ono, "Magic Feature Extraction of LSTM Autoencoder Series Data", Research Association of Artificial Intelligence Materials, SIG-KBS-B802-01, 2018) to extract time-series features of the operator's body (fingers, arms, legs, etc.) joints and angles from video data with added time information. Additionally, the video data feature extraction unit 402 obtains the body joint coordinates extracted from the video data and statistical features (average, peak, etc.) extracted from the time-series angle data. In addition, the video data feature extraction unit 402 judges subtle movements and extracts feature quantities (numerical control (NC) operation or operation within the machine tool 10, whether it is a subtle movement constituting preparation, the moment when preparation should be carried out, etc.) based on the coordinates and angles of the body joints.
[0064] Alternatively, the video data feature extraction unit 402 can also use the machine tool 10's operating data obtained by the operating data acquisition unit (described later) to extract features representing the operator's actions.
[0065] For example, the video data feature extraction unit 402 can determine the opening and closing time of the machine tool 10's door based on the operating data. When there are locations in the video data that change at the same time, that location is identified as the door of the machine tool 10. Thus, when the operator enters the door at that location with their upper body, the video data feature extraction unit 402 can extract features such as "operation inside the machine tool". In addition, the video data feature extraction unit 402 can also extract features such as "operating with the automatic tool changer (ATF)" when the operator's upper body is facing upwards towards the machine tool 10, and features such as "operating with the worktable" when the operator's upper body is facing downwards towards the machine tool 10.
[0066] Furthermore, the video data feature extraction unit 402 can identify the area where the operator pressed a button in the video data during the time the operator performed a NC operation from the running data. Therefore, when a hand is placed on that area, the video data feature extraction unit 402 can extract a feature such as "NC operation".
[0067] When the training data extraction unit 403 extracts the feature quantity representing a pre-registered specific action, i.e. a gesture, from the feature quantity extracted by the video data feature quantity extraction unit 402, it marks the job content represented by the extracted gesture onto the corresponding video data, and extracts the input data of the marked video data and the label data of the job content as training data.
[0068] The operation data acquisition unit 404 acquires the operation data of the machine tool 10 via the classification device 30.
[0069] The running data feature extraction unit 405 extracts features associated with the training data extracted from the running data of the machine tool 10.
[0070] Specifically, the runtime feature extraction unit 405 calculates the time width, for example, based on the start and end times of the training data extracted by the training data extraction unit 403, or calculates the average time width for each labeled job.
[0071] Figure 5 This is a diagram illustrating an example of the preparation time width used for product A.
[0072] like Figure 5As shown, the operation data feature extraction unit 405 cuts out the time based on the preparation time width of product A calculated simultaneously with the overlap, and extracts the following information from the operation data of machine tool 10 as feature quantities of operation data according to the cut-out time: signal information of machine tool 10 associated with the job to be classified (tool abnormality value calculated based on chuck opening and closing, motor torque, etc.), information on whether the same processing is performed (processing program, manufacturing number), and the operation status of machine tool 10 and the change of operation status (e.g., changes in job content such as recovery from tool change alarm to normal operation).
[0073] In addition, the video data feature extraction unit 402 can also, in the same way as the running data feature extraction unit 405, cut out video data based on the time width of each operation calculated simultaneously with the overlap, and extract the feature values of the video data at that time.
[0074] The labeling benchmark production unit 406 produces labeling benchmarks for video data and labeling benchmarks for running data based on the feature quantities of the video data extracted from the training data and the feature quantities of the running data.
[0075] Figure 6A as well as Figure 6B This is a graph representing an example of the training data for each job. In Figure 6A The data in the middle represents the training data when the job is being prepared (running status is stopped). Figure 6B The data in the middle represents the training data when the tool is changed (running in an alarm state). Additionally, Figure 6A The chuck opening / closing signal changes from "closed" to "open" at the start of operation within the machine tool and from "open" to "closed" at the end of operation. On the other hand, Figure 6B The chuck opening / closing signal remains "closed" during operation within the machine tool. Additionally, the operating status includes: running, (network) disconnected, emergency stop, temporary stop, manual operation, and preheating operation.
[0076] right Figure 6A and Figure 6B For each job of the training data shown, the labeling benchmark generation unit 406 calculates, for example, the average value of the features extracted by the video data feature extraction unit 402. For each job of the training data, the labeling benchmark generation unit 406 calculates the Mahalanobis distance between the calculated average value of the features and the features of the unlabeled video data, and uses this distance as the labeling benchmark for the video data.
[0077] In addition, the label standard production department is assigned 406 pairs. Figure 6A as well as Figure 6BFor each job in the training data shown, the average value of the features extracted by the running data feature extraction unit 405 is calculated. For each job in the training data, the labeling benchmark creation unit 406 calculates the Mahalanobis distance between the calculated average value of the features and the features of the unlabeled running data, and uses this distance as the labeling benchmark for the running data.
[0078] In addition, the labeling benchmark production unit 406 calculates the distance, but is not limited to this. For example, machine learning (e.g., decision tree algorithms such as CART (Classification and Regression Tree)) can be performed to create a classifier to calculate the probability of classifying a task as which job.
[0079] The labeling unit 407 labels the unlabeled video data and the running data based on the labeling benchmarks used for video data and the labeling benchmarks used for running data through semi-supervised learning.
[0080] Specifically, the tag assignment unit 407 calculates a weighted distance (reference) based on the distance (reference) between the video data and the running data calculated by the tag assignment reference production unit 406 and (Equation 1). If the calculated distance is less than a certain distance, the untagged video data and running data are tagged. Furthermore, the weighting coefficients "0.8" and "0.2" are examples, and any value can be used.
[0081] (Equation 1)
[0082] The distance used was calculated using two reference points: distance calculated from video data × 0.8 + distance calculated from runtime data × 0.2.
[0083] In addition, the tag assignment unit 407 can also determine which tag to assign to the untagged video data and the running data.
[0084] Thus, the labeling unit 407 is able to label complex tasks with less training data, and can label even if there are some different parts in the video data, since the feature amount of the running data is roughly the same.
[0085] Furthermore, when the probability calculated by the tag assignment reference production unit 406 is used as the reference for video data and running data, the tag assignment unit 407 can also calculate a weighted probability based on the probability (reference) of video data and running data and (Equation 2), and when the calculated probability is less than a certain probability, the untagged video data and running data can be tagged.
[0086] (Equation 2)
[0087] The probability using two benchmarks = probability calculated from video data × 0.8 + probability calculated from runtime data × 0.2
[0088] Alternatively, the labeling benchmark (e.g., the average value of a feature) can be updated sequentially using labeled video data and running data. For example, the labeling benchmark creation unit 406 can recalculate the benchmark based on the training data and labeled data using a known method (e.g., guided co-training), and the labeling unit 407 can label the data based on this benchmark.
[0089] Learning unit 408 uses training data containing video data labeled by unit 407 to perform machine learning (e.g., gradient boosting, neural networks, etc.) to construct learning completion model 361, which is used to classify the tasks performed by the operator based on the input video data.
[0090] Furthermore, the learning unit 408 provides the classification device 30 with the constructed learning completion model 361.
[0091] Furthermore, after the learning unit 408 provides the learning completed model 361 to the classification device 30, when it obtains training data consisting of new video data input data and job content label data from the classification device 30, it can use the obtained training data to perform machine learning again and update the learning completed model 361.
[0092] In addition, the Learning Department 408 can conduct learning online, in batches, or in small batches.
[0093] Online learning is a method of supervised learning that occurs immediately whenever training data is obtained from the classification device 30. Batch learning, on the other hand, involves collecting multiple sets of training data corresponding to each iteration during repeated acquisitions of training data from the classification device 30, and using all collected training data for supervised learning. Mini-batch learning, an intermediate method between online and batch learning, involves supervised learning that occurs only when training data has accumulated to a certain extent.
[0094] The above describes the machine learning used to generate the learning-completed model 361 of the classification device 30.
[0095] Next, the sorting device 30 in the application stage will be described.
[0096] <Classification device 30 in the application stage>
[0097] like Figure 1As shown, the classification device 30 in the application stage is configured to include the following parts: input unit 301, job judgment unit 302, judgment result writing unit 303, judgment result correction reading unit 304, judgment data parsing unit 305, and storage unit 306.
[0098] In addition, in order to achieve Figure 1 The sorting device 30 includes an arithmetic processing unit (not shown) such as a CPU (Central Processing Unit) to control the operation of the function blocks. Additionally, the sorting device 30 includes auxiliary storage devices (not shown) such as ROM and HDD for storing various control programs, and a main storage device (not shown) such as RAM for storing data temporarily needed when the arithmetic processing unit executes programs.
[0099] Furthermore, in the sorting device 30, the arithmetic processing unit reads the operating system and application software from the auxiliary storage device, expands the read OS and application software in the main storage device, and simultaneously performs arithmetic processing based on these OS and application software. Based on the processing results, the sorting device 30 controls each piece of hardware. Thus, [the following is achieved]... Figure 1 The processing is performed by the function blocks. That is, the sorting device 30 can be implemented through hardware and software cooperation.
[0100] The storage unit 306 can be a ROM, HDD, etc., and can also have a learning completion model 361 and a judgment result storage unit 362 together with various control programs.
[0101] The determination result storage unit 362 stores the determination result of the video data determined by the operation determination unit 302 (described later) and the feature quantity in the determined video data.
[0102] The input unit 301 inputs video data containing the operator captured by the camera 20.
[0103] The job determination unit 302 determines the operator's job from the video data input by the input unit 301 based on the learning completion model 361.
[0104] The judgment result writing unit 303 displays the judgment result of the work judgment unit 302 on the display unit (not shown) of the control device 110.
[0105] Therefore, operators can determine the quality of the classification performed by the learning model 361.
[0106] When the judgment result correction input unit 304 finds an error, it obtains the correct operation content input by the operator through the input unit (not shown) such as the keyboard and touch panel included in the control device 110, and outputs the input data of the judged video data and the label data of the obtained correct operation content to the machine learning device 40, so that the machine learning device 40 updates the learning completed model 361.
[0107] Furthermore, the reading of results and the writing of corrections can be performed using, for example, the storage unit 306 or the PMC area described later.
[0108] The judgment data analysis unit 305 detects whether there are any abnormalities in the feature quantities based on the judgment result stored in the judgment result storage unit 362 and the feature quantities in the judged video data.
[0109] Specifically, the data parsing unit 305 uses, for example, a known unsupervised anomaly detection algorithm such as the k-nearest neighbor method. Figure 7 As shown, the system detects whether there are any abnormalities in the feature quantities from the data of the same results in the judgment result storage unit 362.
[0110] Alternatively, for example, if there are 100 video data points identified as being in preparation for a task, and the characteristic quantity of the chuck opening / closing signal from ON to OFF is detected in 99 of these video data points, but no such characteristic quantity is detected in one video data point, the data analysis unit 305 can classify the undetected video data point as an anomaly. Furthermore, this single video data point classified as an anomaly can also be classified as an anomaly using known unsupervised anomaly detection algorithms such as the k-nearest neighbor method described above.
[0111] Furthermore, the data analysis unit 305 can also display the detection results on the display unit (not shown) of the control device 110 when it detects abnormalities in characteristic quantities such as chuck opening and closing signals.
[0112] In this way, the judgment data analysis unit 305 can detect errors in the judgment result and abnormalities in the operation content by detecting abnormalities in characteristic quantities such as the chuck opening and closing signal.
[0113] <Classification process of classification device 30 in the application stage>
[0114] Next, the operation of the classification process of the classification device 30 in this embodiment will be explained.
[0115] Figure 8 This is a flowchart illustrating the classification process of the classification device 30 during the application phase. The process shown here is repeatedly executed during the input of video data from the camera 20.
[0116] In step S11, video data containing the operator captured by camera 20 is input.
[0117] In step S12, the job determination unit 302 inputs the video data input in step S11 into the learning completion model 361 to determine the operator's job.
[0118] In step S13, the determination result writing unit 303 displays the determination result of step S12 on the display unit (not shown) of the control device 110.
[0119] In step S14, the determination result correction input unit 304 determines whether the correct work content input by the operator via the input unit (not shown) of the control device 110 has been obtained. If the correct work content has been obtained, the process proceeds to step S15. On the other hand, if the correct work content has not been obtained, the process proceeds to step S16.
[0120] In step S15, the judgment result correction input unit 304 outputs the input data of the judged video data and the label data of the obtained correct job content to the machine learning device 40, so that the machine learning device 40 updates the learning completed model 361.
[0121] In step S16, the determination data parsing unit 305 detects whether there are any abnormalities in the feature quantities based on the determination result stored in the determination result storage unit 362 and the feature quantities in the determined video data.
[0122] Based on the above, in one embodiment, the machine learning device 40 establishes a benchmark as follows: Training data consisting of input data (video data) and label data for each task, obtained by the operator indicating the start and end of the task via gestures towards the camera 20 during each task, and also using machine tool 10 operation data collected at the same time as the training data, is used to label both the video data and the operation data. Based on this benchmark, the machine learning device 40, through semi-supervised learning, labels the unlabeled video data from both the perspectives of the video data and the operation data. Thus, the machine learning device 40 can generate a highly accurate learning model even with limited training data, and can easily generate the learning model 361 with minimal burden on the operator without requiring new equipment.
[0123] In addition, the classification device 30 can use the learned model 361 to easily classify and identify complex and arbitrary tasks performed by the operator from the camera 20.
[0124] The above describes one embodiment, but the classification device 30 and the machine learning device 40 are not limited to the above embodiment, and also include variations and improvements within the scope of achieving the purpose.
[0125] <Variation Example 1>
[0126] In the above embodiments, the machine learning device 40 is exemplified as a device different from the control device 110 and the classification device 30, but the control device 110 or the classification device 30 may also have some or all of the functions of the machine learning device 40.
[0127] <Variation Example 2>
[0128] In addition, for example, in the above embodiment, the sorting device 30 is shown to be a different device from the control device 110, but the control device 110 may also have some or all of the functions of the sorting device 30.
[0129] Alternatively, the server may have some or all of the following components of the sorting device 30: input unit 301, job judgment unit 302, judgment result writing unit 303, judgment result correction reading unit 304, judgment data parsing unit 305, and storage unit 306. Furthermore, the functions of the sorting device 30 can be implemented in the cloud using virtual server functionality or similar methods.
[0130] Furthermore, the classification device 30 can also be a distributed processing system in which the functions of the classification device 30 are appropriately distributed to multiple servers.
[0131] <Variation Example 3>
[0132] Furthermore, as in the embodiment described above, the learning unit 408 uses training data consisting of input video data and label data of the job content to perform machine learning and construct a learning completion model 361 that classifies the jobs performed by the operator based on the input video data, but it is not limited to this. For example, when the job determination unit 302 of the classification device 30 can obtain both video data and operation data, the learning unit 408 can use training data consisting of input video data and operation data and label data of the job content to perform machine learning and construct a learning completion model 361 that classifies the jobs performed by the operator based on the input video data and operation data.
[0133] Furthermore, the functions included in the classification device 30 and the machine learning device 40 in one embodiment can be implemented separately by hardware, software, or a combination thereof. Here, implementation by software means implementation by reading and executing a program by a computer.
[0134] The various structural components included in the classification device 30 and the machine learning device 40 can be implemented using hardware, software, or a combination thereof, including electronic circuits. In the case of software implementation, the program constituting the software is installed on a computer. Furthermore, these programs can be distributed to users either by recording them on removable media or by downloading them to a user's computer via a network. In the case of hardware implementation, for example, integrated circuits (ICs) such as ASICs (Application Specific Integrated Circuits), gate arrays, FPGAs (Field Programmable Gate Arrays), and CPLDs (Complex Programmable Logic Devices) can constitute part or all of the functionality of the various components included in the aforementioned devices.
[0135] Programs can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., floppy disks, magnetic tapes, hard disks), optical-magnetic recording media (e.g., optical discs), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash memory ROMs, and RAM). Alternatively, programs can also be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transient computer-readable media can provide programs to a computer via wired communication paths such as wires and optical fibers, or via wireless communication paths.
[0136] Furthermore, the steps describing a program recorded in a recording medium naturally include processes performed in that order in a time sequence, as well as processes that are not necessarily performed in a time sequence, and processes that are performed in parallel or individually.
[0137] In other words, the machine learning apparatus, classification apparatus, and control apparatus of this disclosure can be implemented in various ways with the following structures.
[0138] (1) The machine learning apparatus 40 of this disclosure includes: a video data feature extraction unit 402, which extracts features representing the operator's actions from video data containing the operator captured by at least one camera 20; a training data extraction unit 403, which, when extracting features representing pre-registered specific actions from the features extracted by the video data feature extraction unit 402, labels the work content represented by the extracted specific actions onto the corresponding video data, and extracts training data consisting of the input data of the labeled video data and the label data of the work content; and a running data feature extraction unit 405, which extracts features associated with the extracted training data from the running data of the machine tool 10. The system comprises: a feature-based labeling unit 406, which creates labeling benchmarks for video data and for running data based on the feature quantities of the extracted training data and the feature quantities of the running data; a labeling unit 407, which labels the unlabeled video data and running data based on the labeling benchmarks for video data and the labeling benchmarks for running data; and a learning unit 408, which performs machine learning using the training data and generates a learned model 361, wherein the training data includes video data labeled by the labeling unit 407, and the learned model 361 is used to classify the work performed by the operator based on the input video data.
[0139] According to the machine learning device 40, even with less training data, it is possible to generate a learning model with high judgment accuracy.
[0140] (2) In the machine learning device 40 described in (1), the video data feature extraction unit 402 may extract feature quantities representing the operator's actions within the machine tool 10 or feature quantities representing the operator's operation on the machine tool 10 from the video data and the machine tool 10 change time in the video data and the running data.
[0141] Thus, the machine learning device 40 can generate a learning completion model 361 that can determine the operator's work content in more detail.
[0142] (3) In the machine learning apparatus 40 described in (1) or (2), the labeling reference production unit 406 calculates the average value of the features extracted by the video data feature extraction unit 402 for each job of the training data, and uses the distance between the calculated average value of the features and the features of the unlabeled video data as the labeling reference for the video data for each job of the training data, and calculates the average value of the features extracted by the running data feature extraction unit 405 for each job of the training data, and uses the distance between the calculated average value of the features and the features of the unlabeled running data as the labeling reference for the running data for each job of the training data.
[0143] Thus, the machine learning device 40 is able to label complex tasks with less training data.
[0144] (4) The classification device 30 of this disclosure includes: a learning completion model 361, which is generated by the machine learning device 40 according to any one of (1) to (3); an input unit 301, which inputs video data containing the operator captured by at least one camera 20; and a job determination unit 302, which determines the operator's job from the video data input by the input unit 301 according to the learning completion model 361.
[0145] According to the classification device 30, the operator's work can be determined with high accuracy from the video data.
[0146] (5) In the classification device 30 described in (4), the classification device 30 may also include: a judgment result writing unit 303, which causes the control device 110 of the machine tool 10 to display the judgment result of the job judgment unit 302; and a judgment result correction reading unit 304, which, in the event of an error in the judgment result, obtains the correct job content input in the control device 110, outputs the input data of the judged video data and the label data of the obtained correct job content to the machine learning device 40, so that the machine learning device 40 updates the learning completion model 361.
[0147] Thus, the operator can determine the quality of the classification performed by the learning completion model 361, and the classification device 30 can update the learning completion model 361 by receiving correct work content from the operator.
[0148] (6) In the classification device 30 described in (4) or (5), the classification device 30 may also have: a judgment data parsing unit 305, which detects whether there is any abnormality in the feature quantity based on the judgment result of the job judgment unit 302 and the feature quantity in the judged video data.
[0149] Therefore, the sorting device 30 can detect errors in the judgment result and abnormalities in the operation content by detecting anomalies.
[0150] (7) Alternatively, the classification device 30 may have the machine learning device 40 described in any one of (1) to (3).
[0151] Thus, the sorting device 30 can achieve the same effect as any of (1) to (6) above.
[0152] (8) The control device 110 of this disclosure has: the sorting device 30 as described in any one of (4) to (7).
[0153] According to the control device 110, the same effect as any of (1) to (7) above can be obtained.
[0154] Symbol Explanation
[0155] 1 Control System
[0156] 10 machine tools
[0157] 110 Control Device
[0158] 20 cameras
[0159] 30. Sorting device
[0160] 301 Input Section
[0161] 302 Work Judgment Department
[0162] 303 Judgment Result Write-in Department
[0163] 304 Decision Result Correction Input Section
[0164] 305 Judgment Data Analysis Department
[0165] 361 Learning Complete Model
[0166] 362 Judgment Result Storage Department
[0167] 40 machine learning devices
[0168] 401 Video Data Acquisition Department
[0169] 402 Video Data Feature Extraction Department
[0170] 403 Training Data Extraction Department
[0171] 404 Operational Data Acquisition Department
[0172] 405 Operational Data Feature Extraction Department
[0173] 406 Label Standard Production Department
[0174] 407 Label Assignment Department
[0175] 408 Study Department
[0176] 4091 Training data extraction, signature storage unit.
Claims
1. A machine learning device, characterized in that, have: The video data feature extraction unit extracts features representing the operator's actions from video data containing the operator captured by at least one camera. The training data extraction unit, when extracting the feature quantity representing a pre-registered specific action from the feature quantity extracted by the video data feature quantity extraction unit, marks the job content represented by the extracted specific action onto the corresponding video data, and extracts training data consisting of the input data of the marked video data and the label data of the job content; The operational data feature extraction unit extracts features associated with the extracted training data from the operational data of industrial machinery. The labeling benchmark production unit produces labeling benchmarks for video data and labeling benchmarks for running data based on the feature quantities of the video data extracted from the training data and the feature quantities of the running data. The tag assignment unit tags the untagged video data and the running data based on the tag assignment criteria used for the video data and the tag assignment criteria used for the running data; as well as The learning department uses the training data to perform machine learning and generate a learning-completed model, wherein the training data includes the video data labeled by the department, and the learning-completed model is used to classify the tasks performed by the operators based on the input video data.
2. The machine learning apparatus according to claim 1, characterized in that, The video data feature extraction unit extracts feature quantities representing the operator's actions within the industrial machinery or feature quantities representing the operator's operation of the industrial machinery from the video data and operation data based on the change time of the industrial machinery in the video data and operation data.
3. The machine learning apparatus according to claim 1 or 2, characterized in that, The labeling reference production unit calculates the average value of the features extracted by the video data feature extraction unit for each job of the training data. For each job of the training data, the distance between the calculated average value of the features and the features of the unlabeled video data is used as the labeling reference for the video data. The unit also calculates the average value of the features extracted by the running data feature extraction unit for each job of the training data. For each job of the training data, the distance between the calculated average value of the features and the features of the unlabeled running data is used as the labeling reference for the running data.
4. A sorting device, characterized in that, have: The learning completes the model, which is generated by the machine learning apparatus according to any one of claims 1 to 3; The input unit receives video data containing the operator captured by at least one camera. as well as The task determination unit determines the operator's task from the video data input by the input unit based on the learning completion model.
5. The sorting device according to claim 4, characterized in that, The sorting device has: The determination result writing unit causes the control device controlling the industrial machinery to display the determination result of the operation determination unit; and The judgment result correction input unit, in the event of an error in the judgment result, obtains the correct operation content input in the control device, and outputs the input data of the judged video data and the label data of the obtained correct operation content to the machine learning device, so that the machine learning device updates the learning completion model.
6. The sorting device according to claim 4 or 5, characterized in that, The classification device includes a judgment data analysis unit, which detects whether the feature quantity is abnormal based on the judgment result of the job judgment unit and the feature quantity in the determined video data.
7. The sorting device according to claim 4 or 5, characterized in that, The classification device has the machine learning device according to any one of claims 1 to 3.
8. A control device, characterized in that, have: The sorting device according to any one of claims 4 to 7.
Citation Information
Patent Citations
Work operation analysis device, work operation analysis method and work operation analysis program
JP2006209468A
Method and system for human motion identification based on hybrid cooperative training
CN106778796A