How to build a pre-trained model

JP7927952B2Active Publication Date: 2026-10-01KAWASAKI JUKOGYO KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025130023
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2026-10-01
Estimated Expiration
2041-09-06

Smart Images

  • Figure 0007927952000001
    Figure 0007927952000001
  • Figure 0007927952000002
    Figure 0007927952000002
  • Figure 0007927952000003
    Figure 0007927952000003
Patent Text Reader

Abstract

To efficiently obtain a high-performance trained model.SOLUTION: A construction method of a trained model includes six steps. In a first step, a computer collects data for performing machine learning for operation of a control object machine executed by a human. In a second step, the computer evaluates collection data which is the collected data, and re-collects the data when it does not satisfy a predetermined evaluation criterion. In a third step, the computer selects training data from the collection data that satisfies the evaluation criterion. In a fourth step, the computer evaluates the training data, and re-selects the training data when it does not satisfy the predetermined evaluation criterion. In a fifth step, the computer constructs the trained model, by machine learning using the training data which satisfies the evaluation criterion. In a sixth step, the computer evaluates the trained model, and re-trains the trained model when it does not satisfy the predetermined evaluation criterion.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to construction of a trained model by machine learning.

Background Art

[0002] Conventionally, there has been known a system that controls the motion of a robot or the like by using machine learning that repeatedly learns from collected data to automatically find laws and rules, and realizes functions similar to the learning ability that humans naturally have. Patent Document 1 discloses this type of system.

[0003] In the motion prediction system of Patent Document 1, a motion prediction model is constructed by causing the motion prediction model to perform machine learning on data obtained when an operator manually remotely operates a robot arm to perform work. The robot arm is automatically operated based on the output of the motion prediction model.

Prior Art Literature

Patent Literature

[0004]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0005] For example, it may be necessary to evaluate the performance of the motion prediction model (trained model) constructed in Patent Document 1. However, the autonomous motion of a robot arm is a result that comprehensively reflects multiple aspects such as the quality of remote work performed by a user during training, the quality of acquired training data, the quality of the constructed trained model, and the quality of automatic motion based on the output of the trained model. In addition, each quality often mutually influences other qualities. Therefore, it has been difficult to evaluate each quality independently. As a result, even if there is a need to improve the trained model, it is unclear what should be improved, and inefficient trial and error has been required.

[0006] This disclosure is made in light of the above circumstances, and its purpose is to efficiently obtain high-performance pre-trained models. [Means for solving the problem]

[0007] The problems that this disclosure aims to solve are as described above, and next we will explain the means and effects of solving these problems.

[0008] From the perspective of this disclosure, the following method for building a trained model is provided. Specifically, the method for building a trained model includes a first step, a second step, a third step, a fourth step, and a fifth step. In the first step, a computer collects data for machine learning human operation of a robot. In the second step, the collected data is evaluated based on a first evaluation criterion, and if the first evaluation criterion is not met, the computer recollects the data. In the third step, the computer selects training data from the collected data that meets the first evaluation criterion. In the fourth step, the training data is evaluated based on a second evaluation criterion, and if the second evaluation criterion is not met, the computer re-selects the training data. In the fifth step, the computer builds a trained model by machine learning using the training data that meets the second evaluation criterion. The selection in the third step is performed by learning the collected data. Structure This is done by a selection model, which is a machine learning model that has been built up. The aforementioned selection model is constructed at an earlier stage than the construction of the pre-trained model.

[0009] In this way, by proceeding with the process of building a pre-trained model step by step and performing evaluations at each stage, it becomes easier to narrow down the cause of any problems found at any stage. Therefore, the pre-trained model can be built smoothly. [Effects of the Invention]

[0010] According to this disclosure, high-performance pre-trained models can be obtained efficiently. [Brief explanation of the drawing]

[0011] [Figure 1] A block diagram showing the configuration of a robot operation system according to one embodiment of this disclosure. [Figure 2] A conceptual diagram explaining the operation information. [Figure 3] An example of a series of tasks performed by a robot, and a diagram illustrating each task state. [Figure 4] A diagram illustrating an example of information presented to the user during the selection of training data. [Figure 5] A diagram illustrating an example of how training data is selected from collected data. [Figure 6] This diagram shows an example of how to display the operation when partially removing the robot's movements based on the output of a trained model. [Figure 7] This diagram shows an example of what happens when the robot's movement, based on the output of a trained model, stalls midway through. [Figure 8] This diagram shows an example of how the display looks when training data is added to resolve stagnation in the robot's movement. [Figure 9] A schematic diagram illustrating the validation of the learning elements of a pre-trained model. [Figure 10] A flowchart illustrating the workflow from building a pre-trained model to operating the robot. [Modes for carrying out the invention]

[0012] Next, the disclosed embodiments will be described with reference to the drawings. First, a robot operation system 100 using a learning model constructed by the method of this embodiment will be briefly described with reference to Figure 1, etc. Figure 1 is a schematic diagram illustrating the robot operation system 100. Figure 2 is a conceptual diagram illustrating operation information. Figure 3 is a diagram showing an example of a series of tasks performed by the robot 11 and each task state.

[0013] A robot operation system 100 is a system that constructs a learning model 31 and operates a robot system 1 based on an output of the learning model 31. As a result of the robot system 1 being operated by the learning model 31, a robot 11 autonomously performs work.

[0014] The work performed by the robot 11 is arbitrary; examples of conceivable work include welding, assembly, processing, handling, painting, cleaning, and polishing.

[0015] As shown in Fig. 1, the robot system 1 includes a robot control device 10, a robot (controlled machine) 11, and an operating device 12. Each device is connected to one another via a wired or wireless network, and can exchange signals (data) with each other.

[0016] The robot control device 10 is configured by a known computer. The robot control device 10 includes an arithmetic processing unit such as a microcontroller, a CPU, an MPU, a PLC, a DSP, an ASIC, or an FPGA, a robot storage unit such as a ROM, a RAM, or an HDD, and a communication unit capable of communicating with an external device. Control applications and the like for controlling an arm unit or the like are stored in the robot storage unit.

[0017] The robot control device 10 can switch the operation mode of the robot 11 between a manual operation mode and an autonomous operation mode.

[0018] In the manual operation mode, a user manually operates the operating device 12 described below to cause the robot 11 to operate.

[0019] In the autonomous operation mode, the robot 11 automatically operates based on a result of machine learning performed in advance on the operation of the robot 11 by manual operation.

[0020] Robot 11 is configured, for example, as a vertical articulated robot with 6 degrees of freedom. Robot 11 includes an arm attached to a base. The arm has multiple joints. Each joint is provided with an actuator (e.g., an electric motor) not shown in the diagram, for driving the arm around that joint. An end effector is attached to the tip of the arm, depending on the task to be performed.

[0021] The arm and end effector of the robot 11 operate based on motion commands for operating the robot 11. These motion commands include, for example, commands for linear velocity and commands for angular velocity.

[0022] The robot 11 is equipped with sensors for detecting the robot's movements and surrounding environment. Specifically, the robot 11 is equipped with a motion sensor 11a, a force sensor 11b, and a camera 11c.

[0023] The motion sensor 11a is composed of, for example, an encoder. The motion sensor 11a is provided at each joint of the arm of the robot 11 and detects the rotation angle or angular velocity of each joint.

[0024] The force sensor 11b detects the force acting on each joint of the robot arm, or on the end effector attached to the tip of the arm, when the robot 11 is in motion. The force sensor 11b may also be configured to detect moment in addition to or instead of force.

[0025] Camera 11c detects images of the workpiece 81 (the progress of the work being performed on the workpiece 81). To detect the progress of the work, a sound sensor for detecting sound and / or a vibration sensor for detecting vibration may be provided in place of or in addition to camera 11c. Furthermore, the robot 11 or the like may be equipped with sensors that collect distance information, such as a laser scan sensor or an infrared scan sensor.

[0026] The data detected by the motion sensor 11a is motion data indicating the movement of the robot 11, while the data detected by the force sensor 11b and the camera 11c is ambient environment data indicating the state of the environment surrounding the robot 11. This ambient environment data is a state value indicating the progress of the robot 11's work at the time the sensor detects the data. The data detected by the motion sensor 11a, the force sensor 11b, and the camera 11c are collected as state information by the management device 20, which will be described later.

[0027] The operating device 12 is a component operated by the user to operate the robot 11. The operating device 12 varies depending on the task, but for example, it may be a lever operated by the user's hand or a pedal operated by the user's foot. The operating device 12 may be configured as a remote control device located in a location physically separate from the robot 11.

[0028] The operating device 12 is equipped with an operating force detection sensor 13. The operating force detection sensor 13 detects the user operating force, which is the force applied by the user to the operating device 12. If the operating device 12 is configured to be moved in various directions, the user operating force may be a value that includes the direction and magnitude of the force, for example, a vector. The user operating force may be detected not only as the force applied by the user, but also as the acceleration linked to the force (i.e., the value obtained by dividing the force applied by the user by the mass of the operating device 12).

[0029] In this embodiment, the user operating force detected by the operating force detection sensor 13 includes, for example, the force and velocity components in the x-axis (force x and velocity x) and the force and velocity components in the y-axis (force y and velocity y) in the coordinate system of the robot 11, as shown in Figure 2. The data relating to the user operating force detected by the operating force detection sensor 13 is collected as operation information by the management device 20.

[0030] As shown in Figure 1, the robot operation system 100 includes a learning model 31. In the robot operation system 100, for example, a learning model 31 used to cause the robot 11 to perform a series of operations, such as inserting a workpiece 81 into a recess 82 of a component, can be constructed using machine learning.

[0031] Specifically, the user operates the operating device 12 to move the robot 11, for example, as shown below. That is, in operation OA shown in Figure 3, with the robot 11 holding the workpiece, the workpiece 81 is positioned above the member and brought close to the surface of the member. In operation OB, the workpiece 81 is moved in that position until it comes into contact with the surface of the member. In operation OC, the workpiece 81 is moved toward the position of the recess 82. During the movement of the workpiece 81, the state in which the workpiece 81 is in contact with the surface of the member is maintained. In operation OD, the end of the workpiece 81 comes into contact with the inner wall of the recess 82. In operation OE, the workpiece 81 is inserted into the recess 82.

[0032] In this way, the user operates the robot 11 so that it operates in the order of operation OA to operation OE. By learning the relationship between the state information and the user's operating force during this process, the robot operation system 100 can construct a learning model 31 that enables the robot 11 to operate autonomously in the order of operation OA to operation OE.

[0033] As shown in Figure 1, the robot operation system 100 of this embodiment includes a robot system 1 as well as a management device 20.

[0034] The management device 20 is, for example, composed of a known computer and includes an arithmetic processing unit such as a microcontroller, CPU, MPU, PLC, DSP, ASIC, or FPGA, a robot storage unit such as ROM, RAM, or HDD, and a communication unit capable of communicating with an external device.

[0035] The robot system 1 and the control device 20 are connected to each other via a wired or wireless network and can exchange signals (data). The control device 20 may be composed of physically identical hardware to the robot control device 10 provided in the robot system 1.

[0036] The management device 20 includes a data acquisition unit 21, a training data selection unit 22, a model construction unit 23, and an operation data recording unit 24.

[0037] The data acquisition unit 21 collects data from the robot system 1. The data collected by the data acquisition unit 21 from the robot system 1 includes, as described above, state information indicating the surrounding environment data of the robot 11 and operation information reflecting the user's operating force corresponding to the surrounding environment data of the robot 11. Hereinafter, the data collected by the data acquisition unit 21 may be referred to as collected data.

[0038] The collected data is a time-series data of state information and operation information obtained when a user continuously operates the control device 12 to have the robot 11 perform a certain task (or part of a task). That is, the data collection unit 21 collects each of the state information and each of the operation information in relation to time. For example, one collected data is obtained when a user continuously operates the control device 12 to have the robot 11 perform a series of tasks including the five operations OA to OE described in Figure 3 once. The state information and operation information include measured values ​​based on detection values ​​obtained from the camera 11c and the operating force detection sensor 13, etc.

[0039] The data collection unit 21 has an evaluation function that determines whether or not the collected data meets predetermined evaluation criteria.

[0040] The conditions required for the collected data are arbitrarily determined, taking into consideration the characteristics of the learning model 31, the content of the preprocessing performed on the collected data, etc. In this embodiment, the conditions that the collected data must satisfy are that the collected data has a length of a predetermined number or more in the time series, and that data indicating a predetermined operation appears at the beginning and end of the time series.

[0041] The data collection unit 21 determines the length of the collected data, etc., according to the predetermined evaluation criteria described above. This determination is not a complex one using, for example, machine learning, but rather a rule-based one. This simplifies the processing and improves the real-time accuracy of the determination.

[0042] The data collection unit 21 may calculate a value indicating the degree to which the collected data is suitable for use as training data, and present the calculation result to the user as the effectiveness score of the collected data. The effectiveness score can be calculated, for example, as follows: A score is defined in advance for each of the conditions that the collected data must satisfy. The data collection unit 21 determines whether or not the collected data satisfies each condition, and calculates the sum of the scores corresponding to the conditions that are satisfied as the effectiveness score. Based on the effectiveness score calculation result, it may be determined whether or not the collected data satisfies the evaluation criteria.

[0043] The effectiveness level can be presented to the user, for example, by displaying a numerical value of the effectiveness level on an output device such as a display (not shown) provided by the management device 20. Instead of displaying a numerical value on the display, the effectiveness level may be represented, for example, by the color of the displayed graphic.

[0044] If the collected data does not meet the required conditions, or if the effectiveness value is unsatisfactory, the data collection unit 21 can suggest to the user how to improve the data. This suggestion can be achieved, for example, by displaying a message on the display. The message content is arbitrary, but could be a text message such as, "The device operation time is too short." The suggestion of improvement methods is not limited to text messages; it can also be done through other outputs such as icons, audio, or video.

[0045] By repeatedly performing the operation while referring to these suggestions for improvement, users can easily become proficient in the operation, even if they are initially unfamiliar with it, and operate the operating device 12 in a way that yields good data for use as training data.

[0046] If the collected data meets the predetermined evaluation criteria, the data collection unit 21 outputs the collected data to the training data selection unit 22.

[0047] The training data selection unit 22 selects the collected data input from the data collection unit 21 to obtain training data.

[0048] The management device 20 includes an input device (not shown in the figure). The input device consists of, for example, a keyboard, mouse, touch panel, etc.

[0049] The user indicates via the input device whether or not to use the collected data, previously gathered by the data acquisition unit 21 when the operating device 12 was operated, as training data for machine learning. As a result, a selected portion of the collected data is adopted as training data.

[0050] In this embodiment, the selection of training data from the collected data is determined by user input. This allows the collected data used for training to be limited to those that match the user's intentions. However, the training data selection unit 22 may also automatically select the collected data.

[0051] For example, the collected data can be automatically selected as follows. The training data selection unit 22 includes a machine learning model built using machine learning to evaluate the collected data. This machine learning model is built separately from the training model 31, and at an earlier stage than the training model 31. Hereinafter, the machine learning model built in the training data selection unit 22 may be referred to as the selection model 41.

[0052] The selection model 41 learns from the collected data gathered by the data collection unit 21. During the training phase of the selection model 41, the collected data is classified into multiple groups using an appropriate clustering method. If the collected data includes multiple work states, it is divided and classified according to each work state. Work states will be described later.

[0053] Clustering is a technique that learns distributional patterns from a large amount of data and automatically obtains multiple clusters, which are groups of data with similar characteristics. Clustering can be performed using well-known clustering methods such as neural networks (NNs), K-Means, and self-organizing maps. The number of clusters into which the work states included in the collected data are classified can be determined as appropriate. Classification may also be performed using automatic classification methods other than clustering.

[0054] Let's explain the work states. In this embodiment, for example, the collected data related to a series of operations collected by the data acquisition unit 21 is classified according to the user operation (reference operation) corresponding to the work state. For example, as shown in Figure 3, when the robot 11 is made to perform a series of operations to place the workpiece 81 into the recess 82, the work states can be classified into four states: air, contact, insertion, and completion.

[0055] Working state SA (Air) is when the robot 11 is holding the workpiece 81 and positioning it above the recess 82. Working state SB (Contact) is when the robot 11 is bringing the workpiece 81 it is holding into contact with the surface where the recess 82 is formed. Working state SC (Insertion) is when the robot 11 is inserting the workpiece 81 it is holding into the recess 82. Working state SD (Completed) is when the workpiece 81 it is holding into the recess 82 is fully inserted.

[0056] Thus, the four work states classify the series of operations performed by the robot 11 into individual processes. When the robot 11's operations proceed correctly, the work states transition in the following order: work state SA (in the air), work state SB (contact), work state SC (insertion), and work state SD (completion).

[0057] The data that the selection model 41 learns can be, for example, a combination of any one work state and the next work state associated with that work state (i.e., the next work state to transition to), and at least one set of state information and the user operation force associated with this state information. This allows the selection model 41 to learn the order of work states and the order of the corresponding operation forces. The machine learning of the selection model 41 can also be described as data clustering.

[0058] The above-mentioned work states SA, SB, SC, and SD are representative examples, and in reality, a large number of different work states can exist. Let's consider a case where the operator makes the robot 11 perform the same task several times, and work state SA1 corresponding to one set of state information and operating force, work state SA2 corresponding to another set of state information and operating force, and work state SA3 corresponding to yet another set of state information and operating force are collected. Due to variations in the operator's actions and the circumstances, these work states SA1, SA2, and SA3 are, strictly speaking, different from each other. However, since work states SA1, SA2, and SA3 have common characteristics, they will be classified into the same cluster (cluster of work states SA).

[0059] As described above, the sorting model 41 uses machine learning to reflect the time order of the output of the operating force. Simply put, the sorting model 41 learns at least one set of state information and operating force combinations corresponding to each of the work states SA, SB, SC, and SD, and also learns the work order, such as work state SA being followed by work state SB. This allows the sorting model 41 to perform classification that reflects the time-series information of the operating force. In other words, it can reflect each operating force associated with each work state in the order of the work.

[0060] As described above, this state information is sensor information detected by the motion sensor 11a, force sensor 11b, and camera 11c (for example, working conditions such as position, velocity, force, moment, and video). This state information may also include information calculated based on the sensor information (for example, values ​​showing the change in sensor information over time from the past to the present).

[0061] In the inference phase, the training model 41 can estimate and output a criterion operation corresponding to the state information associated with the time-series information of the input collected data. Instead of the estimated criterion operation, the selection model 41 may output information about the cluster to which the criterion operation belongs. In this case, the selection model 41 can output the similarity between the operation information of the input collected data and the estimated criterion operation. The similarity can be defined, for example, using a known Euclidean distance. The similarity output by the selection model 41 can be used as an evaluation value for evaluating the collected data.

[0062] As explained in Figure 3, in a series of operations performed by the robot 11 represented by a single set of collected data, the operation state transitions sequentially. Taking this into consideration, the similarity output by the selection model 41 is performed for the operation information within each predetermined time range, as shown in Figure 4. Specifically, the training data selection unit 22 assigns a label (corresponding information) indicating the cluster to which the reference operation output by the selection model 41 belongs to the collected data within a predetermined time range of the collected data if the evaluation value output by the selection model 41 is above a predetermined threshold. On the other hand, if the evaluation value output by the selection model 41 falls below the predetermined threshold, the training data selection unit 22 does not assign a label to that time range.

[0063] Figure 4 shows an example where the operation information within a predetermined time range in the collected data was similar to the standard operation corresponding to work state SA, and therefore the operation information was labeled with the numerical value "1". Similarly, the time ranges labeled with the numerical values ​​"2", "3", and "4" in Figure 4 indicate that the operation information within those time ranges is similar to the standard operation corresponding to work states SB, SC, and SD. Figure 4 also shows the time ranges that were not labeled.

[0064] For time ranges in the collected data where no labels are assigned, the situation differs significantly from any of the work states learned by the selection model 41, and it may be inappropriate to use this collected data as training data. Therefore, the training data selection unit 22 does not adopt this time range as training data. On the other hand, for time ranges in the collected data where labels are assigned, the training data selection unit 22 adopts them as training data.

[0065] The labels output by the selection model 41 do not necessarily have to be used for the automatic decision of whether or not to adopt the collected data as training data. For example, the labels may be presented as reference information when the user manually decides whether or not to adopt the collected data as training data. Figure 4 shows an example of a screen presented to the user. In this example, the operation information (e.g., operation force) included in the collected data is visually displayed in the form of a graph, and data portions where time-series information is continuous and assigned the same numerical label are displayed as a single block. This makes it easier for the user to decide whether or not to adopt the collected data as training data.

[0066] In this example, clustering is employed in the selection model 41. Therefore, it becomes easy to extract only the data from a specific time range (block) in which the operations are effective, from the collected data obtained by the user performing a series of operations, and use it as training data. Figure 5 shows an example of data used as training data. In the example in Figure 5, one of the five collected data sets is entirely used as training data, while for the remaining three, only a portion of the multiple labeled blocks is used as training data.

[0067] The training data selection unit 22 has a function to evaluate the selected training data. The evaluation criteria are arbitrary. For example, the evaluation may check whether the number of training data is sufficient, or whether there is insufficient training data corresponding to some work states.

[0068] If the training data meets the evaluation criteria, the training data selection unit 22 outputs the training data to the model building unit 23.

[0069] The model building unit 23 constructs a learning model 31 to be used in the robot system 1 using machine learning (e.g., supervised learning). Hereafter, a learning model that has completed training may be referred to as a trained model.

[0070] The model building unit 23 uses the training data output from the training data selection unit 22 to build the learning model 31. As described above, the training data corresponds to the collected data that has been selected. Therefore, the training data, like the collected data, includes at least ambient environment data (i.e., state information) that reflects the working state of the robot 11, and user operation force (i.e., operation information) associated with the ambient environment data.

[0071] The learning model 31 is a neural network with a general configuration, for example, having an input layer, a hidden layer, and an output layer. Each layer has multiple units that mimic brain cells. The hidden layer is placed between the input layer and the output layer and is composed of an appropriate number of intermediate units. State information (training data) input to the learning model 31 in the model building unit 23 flows through the input layer, the hidden layer, and the output layer in that order. The number of hidden layers can be determined as appropriate. However, the format of the learning model 31 is arbitrary and not limited to this.

[0072] In this learning model 31, the data input to the input layer is state information that reflects the surrounding environment data described above. The data output by the output layer is the estimated result of the detection value of the operating force detection sensor 13. This essentially represents the estimated user operating force. Therefore, the data output by the output layer indicates the user operation estimated by the learning model 31.

[0073] Each input unit and each intermediate unit are connected by a path through which information flows, and each intermediate unit and each output unit are connected by a path through which information flows. In each path, the influence (weight) that the information from the upstream unit has on the information from the downstream unit is set.

[0074] During the training phase of the learning model 31, the model building unit 23 inputs state information into the learning model 31 and compares the operating force output from the learning model 31 with the operating force performed by the user. The model building unit 23 updates the learning model 31 by updating the weights, for example, using a known algorithm such as backpropagation, so that the error obtained from this comparison becomes smaller.

[0075] Since the learning model 31 is not limited to a neural network, updating the learning model 31 is not limited to backpropagation. For example, the learning model 31 can also be updated using the well-known algorithm SOM (Self-organizing maps). Learning is achieved by continuously performing such processing.

[0076] The model building unit 23 has a function to evaluate the constructed learning model 31. The criteria for this evaluation are arbitrary. For example, the evaluation may be performed on whether the learning model 31 outputs the expected user operation force when specific state information is input. The evaluation of working time and power consumption may also be performed by simulation using a 3D model of the robot 11 or the like.

[0077] Regarding the evaluation of the learning model 31, the model building unit 23 can be configured to also present the training data that forms the basis of the inference output by the learning model 31 when it operates the learning model 31 in the inference phase. This makes it easier for humans or other personnel to evaluate the learning model 31.

[0078] For example, consider the case where the learning model 31 is constructed not by a neural network, but by clustering, similar to the selection model 41 described above. In the training phase, each training data point is plotted as a point in a multidimensional feature space. Each training data point is assigned identification information that makes it possible to uniquely identify that training data point.

[0079] Once all the training data is plotted, multiple clusters are determined using the appropriate clustering method described above, similar to the selection model 41. Next, the model building unit 23 finds data that represents each cluster. Hereafter, this data may be referred to as a node. A node can be, for example, data that corresponds to the centroid of each cluster in a multidimensional space. Once nodes have been found for all clusters, the training phase is complete.

[0080] In the inference phase of the learning model 31, state information at a certain point in time is input to the learning model 31. The learning model 31 finds one or more nodes that have features similar to the given state information. Similarity can be defined, for example, using the Euclidean distance.

[0081] When a node with features similar to the state information is found, the learning model 31 determines the user operation force (in other words, the detected value of the operation force detection sensor 13) contained in the data of that node. If multiple nodes similar to the state information are detected, the user operation forces of the multiple nodes are appropriately combined. The learning model 31 outputs the obtained user operation force as the estimated operation force described above. At this time, the learning model 31 outputs identification information that identifies the training data corresponding to that node. This makes it possible to identify the training data that served as the basis for the learning model 31's inference.

[0082] For example, suppose in the above simulation, it is found that some of the actions of the robot 11 based on the learning model 31 are undesirable, and the user wants to prevent those actions from being performed. Figure 6 shows an example of the simulation screen displayed on the control device 20's display. The simulation screen displays a timeline, and labels corresponding to the aforementioned work states, which classify the series of actions performed by the robot 11 in the simulation, are displayed along the time axis. The user operates the appropriate input device to select the time range (in other words, the work state) corresponding to the action to be deleted, using the above blocks as units. Then, the display highlights the portion (block) of the training data corresponding to the work state of the action that has been specified for deletion. The method of highlighting is arbitrary. As a result, the user can intuitively grasp the magnitude of the impact of deletion and decide whether or not to actually delete the action.

[0083] When a user instructs the model to delete an action, the plot of training data corresponding to the portion to be deleted is removed from the clustering result, which is the learning result of the learning model 31. If the learning model 31 then operates in the inference phase, it will output an operational force based on different training data instead of the operational force that caused the undesirable action. In this way, the learning result of the learning model 31 can be partially modified.

[0084] For example, suppose that in the above simulation, a situation occurs where the robot 11's operation based on the learning model 31 stalls. To give a specific example, in the operation OD in Figure 3, where the end of the workpiece 81 is to be brought into contact with the inner wall of the recess 82, the position of the workpiece 81 does not coincide with the position of the recess 82, so it is not possible to make contact with the inner wall of the recess 82, and as a result the robot 11 stops without performing any further operations. Figure 7 shows an example of a simulation result where the robot 11's operation stalls.

[0085] Suppose the user wants to add one or more new training data to resolve this situation. This training data performs a series of actions, similar to the other training data. However, the new training data includes, before the aforementioned action OD, a movement in which the workpiece 81 is in contact with the surface near the recess 82 and moved slightly in various directions in the horizontal plane so that the center of the workpiece 81 aligns with the center of the recess 82.

[0086] When a user instructs the system to add new training data, a plot of that training data is added to the clustering results in the multidimensional space described above. This corresponds to additional learning being performed in the learning model 31.

[0087] After further training of the learning model 31, the user instructs the system to run the simulation again under the same conditions as above. In this simulation, the robot 11 performs new actions that were not present in the previous simulation, and as a result, the robot successfully completes a series of actions. Figure 8 shows an example of the simulation results in this case. On the simulation results display screen, the user operates the input device to select the part of the timeline that corresponds to the newly performed action. If the corresponding part of the newly added training data is highlighted in response to this selection, the user can determine that the stagnation in the simulation has been resolved due to the contribution of the newly added training data.

[0088] In this configuration, the user is presented with the consequences of deleting or adding content to the learning model 31. As a result, the user can edit the learning content of the learning model 31 with confidence.

[0089] If the learning model 31 is determined to meet the evaluation criteria, the model building unit 23 outputs information to that effect to the operation data recording unit 24.

[0090] The motion data recording unit 24 transmits the output of the learning model 31 to the robot system 1 to make the robot 11 operate autonomously, and records the motion data. This motion data is used, for example, to verify the autonomous operation of the robot 11. As will be described in more detail later, the motion data recording unit 24 can also be used in subsequent actual operation scenarios.

[0091] The motion data recording unit 24 receives state information indicating the surrounding environment data of the robot 11. The motion data recording unit 24 outputs the input state information to the learning model 31 of the model building unit 23. The learning model 31 operates in the inference phase, and its output is input to the robot control device 10. This allows the robot 11 to operate autonomously.

[0092] The operation data recording unit 24 can also record whether or not each operation has been verified when the learning model 31 is operated in the inference phase. A suitable storage unit provided by the management device 20 is used for this recording.

[0093] The following provides a detailed explanation. Regarding the operation of machines such as robot 11 based on machine learning models, due to the nature of machine learning, it is difficult to reproduce all of its behavior through repeated testing beforehand. Therefore, for example, when the learning model 31 is constructed using clustering as described above, the operation data recording unit 24 is configured to store whether or not each plot in the multidimensional space has been verified. Each plot in the clustering can potentially be adopted as a node. As mentioned above, a node is data that represents the cluster. Each data (plot) and node corresponds to an individual learning element in the learning model 31.

[0094] Figure 9 schematically shows the changes in the contents stored in the operation data recording unit 24 in relation to the multiple stages that occur after the construction of the learning model 31.

[0095] First, let's explain the operational testing phase. After the learning model 31 is constructed by the model building unit 23, operational testing is performed using the actual robot 11 with a certain number of trials.

[0096] As described above, in the inference phase of the learning model 31, state information at a certain point in time is input to the learning model 31. When the learning model 31 is constructed by clustering, the learning model 31 finds nodes that have features similar to the state information in question.

[0097] The learning model 31 has numerous nodes that may serve as the basis for outputting operational information during the inference phase. In Figure 9, the nodes are schematically represented by small ellipses. The storage unit of the management device 20 can record in a table format whether or not each node has already been verified. However, the recording may be done in a format other than a table. Hereinafter, nodes whose verification status is recorded in the table will be referred to as verified nodes, and nodes whose status is not recorded will be referred to as unverified nodes.

[0098] During the operational testing phase, when the learning model 31 outputs operational information contained in the data of the node as an inference result, it outputs information identifying the node to the operational data recording unit 24. Hereinafter, this information may be referred to as node identification information. Node identification information can, for example, be the identification number of the training data corresponding to the node, but is not limited to this.

[0099] Before starting the operational test, all nodes are in an unverified state. The user monitors the experimental operation of the robot 11 based on the output of the learning model 31. The user determines whether or not there are any problems with the operation of the robot 11. In making this determination, the user may refer to the operation data recorded by the operation data recording unit 24 in an appropriate manner.

[0100] Once it is determined that there are no problems with the operation of the robot 11, the user operates the input device of the management device 20 as appropriate to indicate that the verification is complete. As a result, the operation data recording unit 24 updates the table described above to record that verification has been performed for the nodes that produced output during the test operation. In Figure 9, verified nodes are indicated with hatching.

[0101] As the operational tests of robot 11 are repeated, the number of unverified nodes decreases, and the proportion of verified nodes gradually increases. However, it is practically impossible to make all nodes verified during the operational testing phase, and it is unavoidable that some unverified nodes remain when moving to the operational phase.

[0102] Next, we will explain the operational phase. In the operational phase of the robot 11, the learning model 31 operates in the inference phase, just as in the operational test phase described above. When state information at a certain point in time is input to the learning model 31, the learning model 31 finds nodes that have features similar to that state information.

[0103] The learning model 31 outputs the aforementioned node identification information, which identifies the requested node, to the operation data recording unit 24. This output of node identification information is performed before outputting the operation information contained in the data of that node as an inference result to the robot 11.

[0104] The operation data recording unit 24, based on the node identification information input from the learning model 31, refers to the aforementioned table and determines whether the node that is about to output operation information has been verified.

[0105] If the identified node is a verified node, the operation data recording unit 24 controls the robot 11 to operate using the output of that node in the learning model 31.

[0106] If the identified node is an unverified node, the operation data recording unit 24 searches for a verified node similar to the unverified node. This search corresponds to finding a verified node within a predetermined distance from the unverified node in the multidimensional feature space where the training data was plotted during clustering. This distance corresponds to the similarity and can be, for example, the Euclidean distance.

[0107] If a verified node with a similarity score above a predetermined level is obtained, the operation information corresponding to the output of the verified node, which is the search result, is compared with the operation information corresponding to the output of the unverified node. This comparison can also be performed, for example, using the Euclidean distance described above.

[0108] If the comparison determines that the similarity between the two outputs is above a predetermined level, it is considered that even if the robot 11 is operated based on the output of the unverified node, the operation will not differ significantly from previously verified operations. Therefore, the operation data recording unit 24 controls the robot 11 to operate using the output of the unverified node. It is preferable that the above processes, such as determining whether an output is verified or not, searching for a node, and comparing outputs, are performed within the control cycle time of the robot 11.

[0109] The operation of the robot 11 based on the output of the unverified node is monitored by the user as appropriate. If the user determines that no problems have occurred, the user operates the appropriate input device to instruct the management device 20 to complete the verification, similar to the operation test phase. In response, the operation data recording unit 24 updates the table described above to record that verification has been performed for the relevant unverified node. Thus, in this embodiment, unverified nodes can be changed to verified nodes not only in the operation test phase but also in the operation phase.

[0110] During operation, the robot 11 does not necessarily have to operate based on the output of the unverified node itself. For example, the operation data recording unit 24 may operate the robot 11 using an output that combines the output of the unverified node and a similar verified node. Output combination can be performed, for example, by calculating the average and median values ​​of the outputs. If preventing unexpected behavior is a priority, the operation data recording unit 24 can also be controlled to completely replace the output of the unverified node with the output of a similar verified node.

[0111] In the search process described above, it is possible that no verified node similar to an unverified node may be found. Furthermore, even if a similar verified node is found, the output of the unverified node and the output of the verified node may not be similar. In either of the above cases, it is preferable for the operation data recording unit 24 to forcibly change the output of the learning model 31 to an output predetermined for the operation stability of the robot 11.

[0112] For example, consider a master-slave control system that operates to eliminate the force difference between a master robot and a slave robot, where the learning model 31 is trained to perform operations on the master robot. In this configuration, if the learning model 31 attempts to output operation information based on an unverified node and cannot find a similar verified node, the operation data recording unit 24 forcibly and continuously changes the user operation force output by the learning model 31 thereafter to zero. As a result, the master robot's operation output becomes zero, causing the slave robot to move in a direction that reduces the external force to zero. Therefore, the slave robot can transition to a stable state where the external force is zero.

[0113] The operation data recording unit 24 can also be configured to output an alarm if it detects that the learning model 31 has attempted to output operation information based on an unverified node. The alarm can be displayed on a screen, for example, but notification can also be made by other means, such as voice. This allows the user to understand the situation early.

[0114] In addition to information on whether or not a node has been verified, additional information may be stored in the memory of the management device 20 in association with each node that may serve as the basis for the learning model 31 to output operational information during the inference phase. This additional information could include, for example, whether or not the node is a node newly added through additional learning, whether or not the node has an output that would cause the robot 11 to operate with a force greater than predetermined, or whether or not damage to the workpiece 81 has occurred in the past. This information can be registered, for example, by the user appropriately operating the input device of the management device 20.

[0115] The operation data recording unit 24 provides additional information to the user, for example, when outputting the above-mentioned alarm. This information is provided to the user, for example, by displaying it on a screen. This allows the user to obtain useful information about unverified nodes, making it easier for them to make appropriate decisions regarding the operation of the learning model 31.

[0116] As described above, the movements of the robot 11 based on the machine learning model are diverse, making it difficult to verify all movements. In particular, when the robot 11 performs tasks involving force contact as shown in Figure 3, it is impossible to test all possible situations in advance. Taking this into consideration, when the learning model 31 outputs a movement that has not been recorded as verified, the motion data recording unit 24 forcibly changes the output of the learning model 31 to, for example, a similar movement that has been recorded as verified. In this way, in the robot operation system 100 of this embodiment, the motion data recording unit 24 functions as a model output control unit and performs interferential control over the output of the learning model 31. This prevents unexpected movements of the robot 11.

[0117] It is possible that additional training of the learning model 31 may be required during the operational phase. In this case, the user operates the control device 12 to add new training data. As a result of clustering being performed again, new nodes are added to the learning model 31. In Figure 9, the added nodes are shown with dashed lines. The added nodes are recorded as unverified nodes in the aforementioned table. Once the additional training is complete, the system returns to the operational testing or operational phase. Once the verification of the added nodes is completed in the same manner as above, the record of those nodes in the table changes from unverified to verified.

[0118] Next, we will explain the flow of building the machine learning model shown above.

[0119] In this embodiment, as shown in Figure 10, the workflow from building the learning model 31 to putting it into operation is divided into four stages: (A) acquiring data to be collected, (B) selecting training data, (C) building the trained model, and (D) acquiring operation data. In Figure 10, the numbers [1] to [8] represent the first to eighth stages. An evaluation is performed at each stage, and if the evaluation criteria are not met, the work at that stage is repeated, as shown by the dashed arrow in Figure 10. Only if it is determined that the evaluation criteria are met can the work proceed to the next stage.

[0120] Therefore, the evaluation of each stage is carried out on the premise that the evaluation criteria were met in the upstream work. Thus, for example, if the trained model constructed in stage (C) does not meet the evaluation criteria, it can be basically assumed that there is no problem with the work in the upstream stage, and that there is a problem with the work of constructing the trained model.

[0121] When problems occur with the autonomous operation of a machine learning model, it is often extremely difficult to determine whether the cause lies in the training data or in the construction of the trained model. In this embodiment, however, since the work is carried out while performing evaluations at each stage, it is easy to narrow down the cause when a problem is found. Therefore, the machine learning model can be built and operated smoothly.

[0122] Since each stage of the work is independent, it is easy to divide the work among four people, for example, four stages, and the boundaries of each person's responsibility can be clearly defined.

[0123] Of course, for example, if the trained model constructed in stage (C) does not meet the evaluation criteria, it may be discovered that the cause was a problem with the selection of training data, which was overlooked in the evaluation in stage (B). In that case, as shown by the dashed arrow in Figure 10, the process is repeated by going back to the previous stage and redoing the work and evaluation. This rule helps prevent major rework and improves work efficiency.

[0124] As described above, in this embodiment, a trained model is constructed by a method including the following six steps. In the first step, data is collected for machine learning of the user's operation of the robot 11. In the second step, the collected data is evaluated, and if it does not meet predetermined evaluation criteria, the data is collected again. In the third step, training data is selected from the collected data that meets the evaluation criteria. In the fourth step, the training data is evaluated, and if it does not meet predetermined evaluation criteria, the training data is re-selected. In the fifth step, a trained model is constructed by machine learning using the training data that meets the evaluation criteria. In the sixth step, the trained model is evaluated, and if it does not meet predetermined evaluation criteria, the trained model is retrained.

[0125] By proceeding with the work in this step-by-step manner and conducting evaluations at each stage, it becomes easier to narrow down the cause of any problems found at each stage. Therefore, the learning model 31 can be constructed smoothly.

[0126] In this embodiment, if the training data does not meet the evaluation criteria in the fourth step and there is a problem with the collected data, the process returns to the first step. If the trained model does not meet the evaluation criteria in the sixth step and there is a problem with the training data, the process returns to the third step. However, the process may return to the second step instead of the first step, or to the fourth step instead of the third step.

[0127] If a problem is found in the work performed at a previous stage, the problem can be properly resolved by redoing the work at that stage.

[0128] In this embodiment, in the first step, when the user operates the robot 11, information including that operation is collected as data. In the second step, it is determined whether the collected data is suitable as training data based on predetermined rules, and the result of the determination is presented to the user.

[0129] This allows users to easily understand whether the collected data is suitable as training data, and also simplifies the processing.

[0130] In this embodiment, when the learning model 31 constructed in step 5 operates in the inference phase, the training data used to construct the learning model 31, which serves as the basis for the output of the learning model 31, is identified and output.

[0131] This allows users to understand, to a certain extent, the scope of the impact of editing, for example, the learning results of the learning model 31, if they wish to edit some of them. Therefore, they can accurately modify and delete parts of the learning content of the learning model 31.

[0132] In this embodiment, in the seventh step, the robot 11 is operated based on the output of the learning model 31 that satisfies the evaluation criteria, and the operation data is recorded. In the eighth step, the operation data is evaluated, and if the predetermined evaluation criteria are not met, the robot 11 is operated again and the operation data is re-recorded.

[0133] By performing evaluations at each stage leading up to the actual operation of the robot 11, it becomes easier to pinpoint the cause of any problems found in each process.

[0134] In this embodiment, if the operation data does not meet the evaluation criteria in step 8 and there is a problem with the learning model 31, the process returns to step 5. However, the process may be returned to step 6 instead of step 5.

[0135] If a problem is found in the work performed at a previous stage, the problem can be properly resolved by redoing the work at that stage.

[0136] In this embodiment, the learning model 31 includes multiple learning elements. If the learning model 31 is constructed, for example, by clustering, then the nodes, which are data representing the cluster, correspond to the learning elements. In step 7, if the operation of the robot 11 based on the output of the trained model is verified, the verified learning elements that formed the basis of the operation are recorded. During actual operation of the trained model, the output of the trained model based on unverified learning elements can be changed to a predetermined output, or to an output based on similar verified learning elements.

[0137] This prevents unexpected movements of the robot 11 due to unverified learning elements.

[0138] In this embodiment, the object to be made to operate autonomously by the learning model 31 is the robot 11.

[0139] This makes it possible to smoothly construct a learning model 31 for the autonomous operation of the robot 11.

[0140] Preferred embodiments of the present disclosure have been described above, but the above configuration can be modified as follows, for example. Modifications may be made individually or in any combination of multiple modifications.

[0141] Steps 7 and 8 in Figure 10 can be omitted, and the constructed learning model 31 can be used immediately for practical application.

[0142] The evaluation of the collected data shown in the second step may be performed using a machine learning-based method rather than a so-called rule-based method.

[0143] In steps 2, 4, 6, and 8, the evaluation may be performed by a computer or by a human.

[0144] The output device used by the management device 20 for various displays could be, for example, a liquid crystal display, but a projector, a head-mounted display, etc., can also be used. For example, if a head-mounted display is used, a display using known augmented reality (AR) may be performed.

[0145] The operation information estimated by the trained model may be, for example, the change in operating speed or position of the operating device 12 operated by the user, instead of the user's operating force. Instead of estimating the relationship between state information and operation information, the trained model may estimate the relationship between state information and control signals to the robot 11.

[0146] The constructed learning model 31 can also be applied to controlled machines other than the robot 11. [Explanation of Symbols]

[0147] 1. Robot System 10 Robot control device 11. Robots (controlled machines) 20 Management device 31. Trained Models (Pre-trained Models) 100 Robot Operation Systems

Claims

1. The first step involves the computer collecting data for machine learning to control the robot by humans, A second step involves evaluating the collected data based on a first evaluation criterion, and if the first evaluation criterion is not met, the computer recollects the data. A third step in which the computer selects training data from the collected data that meets the first evaluation criterion, A fourth step in which the computer evaluates the training data based on a second evaluation criterion, and if the second evaluation criterion is not met, the computer re-selects the training data. A fifth step in which the computer constructs a trained model by machine learning using the training data that satisfies the second evaluation criterion, Includes, The sorting in the third step is performed by a sorting model, which is a machine learning model constructed by learning the collected data. A method for constructing a pre-trained model, wherein the selection model is constructed at an earlier stage than the construction of the pre-trained model.

2. A method for constructing a trained model according to claim 1, The collected data includes state information indicating the surrounding environment data of the robot when a human operates the robot to have the robot perform a series of tasks, and time-series data of operation information reflecting the human operating force corresponding to the surrounding environment data of the robot. The sorting model can estimate and output a reference operation, which is a human operation corresponding to the state information of the data input to the sorting model, and can also output the similarity between the reference operation and the operation information of the input data. A method for constructing a trained model, wherein the selection in the third step is performed based on the similarity score output by the selection model when the collected data is input to the selection model.

3. A method for constructing a trained model according to claim 2, In the third step, the output of the similarity by the selection model is performed for the operation information within each predetermined time range included in the collected data that satisfies the first evaluation criterion. A method for constructing a trained model, wherein in the third step, the time range in which the similarity output by the selection model is greater than or equal to a predetermined threshold is adopted as the training data.

4. A method for constructing a trained model according to claim 3, The aforementioned selection model is constructed by classifying the collected data using clustering, and the method for constructing a trained model is also described.

Citation Information

Patent Citations

  • Model generation method and device, electronic equipment and storage medium

    CN111310934A

  • Operation prediction system and operation prediction method

    JP2018206286A

  • Data selection device, learning device, and program

    JP2021086558A

  • Training data selection device, robot system and training data selection method

    JP2021107970A

  • Learning system for robotic object manipulation

    KR1020210069410A