Robot control device, robot system, and robot control method
The robot control device uses a trained model to track detailed progress of tasks by learning from user operations, enhancing the monitoring and evaluation of robot performance.
Patent Information
- Application Number
- JP2021185278
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-11-12
AI Technical Summary
Existing robot control systems struggle to provide detailed progress indicators for robot operations, making it difficult to accurately assess the progress of tasks.
A robot control device equipped with a trained model that learns task data, including input data from robot states and surroundings, and outputs control data and progress levels based on user operations, allowing for detailed progress tracking through a series of tasks.
Enables users to grasp the detailed progress of robot operations, facilitating better monitoring and evaluation of task performance, and providing intuitive feedback on task completion.
Smart Images

Figure 0007759770000001 
Figure 0007759770000002 
Figure 0007759770000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to controlling robots through machine learning. [Background technology]
[0002] Conventionally, robot control devices equipped with machine learning devices that can construct models related to the work operations of robots have been known. Patent Document 1 discloses this type of robot control device.
[0003] The robot control device in Patent Document 1 includes a trained model constructed by learning task data, and the robot is controlled based on the output of the trained model.
[0004] Task data is a set of input data and output data. The input data is the state of the robot and its surroundings when a human operates the robot to perform a series of tasks. The output data is the corresponding human operation or the robot's behavior as a result of that operation.
[0005] In Patent Document 1, a base trained model is constructed by learning task data for each of a plurality of simple actions. An action label is associated with the base trained model. Based on the similarity between the trained model and each of the base trained models, a combination is obtained when the trained model is represented by a plurality of base trained models. An action label corresponding to each of the base trained models that represent the trained model is output.
[0006] In Patent Document 1, a trained model operating in the inference phase outputs output data when input data is received. At this time, the trained model can output a progress indicator indicating the degree of progress in a series of tasks to which the output data corresponds. As described above, the trained model that has learned a series of tasks is represented by multiple base trained models. At this time, the chronological relationship between the multiple base trained models can be obtained based on the progress indicator. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Patent Publication No. 2020-104215 Summary of the Invention [Problem to be solved by the invention]
[0008] In the configuration of Patent Document 1, the progress level changes in units of the base trained model. Therefore, it cannot necessarily be said that it is suitable for grasping the progress level in detail, and there is room for improvement in this respect.
[0009] The present disclosure has been made in consideration of the above circumstances, and its purpose is to provide a robot control device and the like that is capable of grasping the detailed progress of an operation. [Means for solving the problem]
[0010] The problem to be solved by the present disclosure is as described above. Next, the means for solving this problem and the effects thereof will be described.
[0011] According to a first aspect of the present disclosure, there is provided a robot control device having the following configuration. Specifically, the robot control device includes a trained model, a control data acquisition unit, a progress acquisition unit, and an information output unit. The trained model is constructed by learning task data in which input data indicates the state of the robot and its surroundings when a human operates the robot to perform a series of tasks, and output data indicates the corresponding human operation or the robot's behavior as a result of the operation. When input data regarding the state of the robot and its surroundings is input to the trained model, the control data acquisition unit obtains output data from the trained model regarding the corresponding estimated human operation or robot behavior, thereby obtaining control data for the robot to perform the tasks. When the input data is input to the trained model, the progress acquisition unit obtains a progress level indicating which progress level in the series of tasks corresponds to the output data output in response to the input data. The information output unit is capable of outputting the progress level. The trained model is capable of determining which of multiple process operations corresponding to the division of the series of tasks the input data will be classified into. In the trained model, an output transition, which is the time progression of the human operation for realizing the process operation or the robot operation resulting from the operation, is defined in association with each classification. The trained model determines an output corresponding to the input data from the output transition associated with the classification result of the input data, and sets the output data as the output. The range of change in the progress is divided into a plurality of progress ranges so as to correspond to the plurality of process operations. The order of the plurality of progress ranges corresponds to the order of the plurality of process operations in the series of work. The progress acquired by the progress acquisition unit changes depending on the position of the output corresponding to the input data in the output transition associated with the process operation, within the progress range corresponding to the process operation, which is the classification result of the input data performed by the trained model.
[0012] According to a second aspect of the present disclosure, the following robot control method is provided. Specifically, in this robot control method, a trained model is constructed by learning task data in which input data represents the state of the robot and its surroundings when a human operates the robot to perform a series of tasks, and output data represents the corresponding human operation or the robot's behavior as a result of the operation. When input data related to the state of the robot and its surroundings is input to the trained model, output data related to the estimated human operation or robot behavior is obtained from the trained model, thereby acquiring control data for the robot to perform the tasks. When the input data is input to the trained model, the output data output in response to the input data acquires a progress level indicating which progress level in the series of tasks corresponds to the input data. The trained model outputs the progress level. The trained model is capable of determining to which of multiple process actions corresponding to a division of the series of tasks the input data is classified. In the trained model, an output transition, which represents the temporal progression of the human operation to achieve the process action or the robot's behavior as a result of the operation, is defined in association with each classification. The trained model determines an output corresponding to the input data from the output transition associated with the classification result of the input data, and sets the output data as the output. The range of change in progress is divided into a plurality of progress ranges so as to correspond to a plurality of the process operations. The order of the plurality of progress ranges corresponds to the order of the plurality of process operations in the series of work. The progress changes depending on the position of the output corresponding to the input data in the output transition associated with the process operation within the progress range corresponding to the process operation which is the result of the classification of the input data performed by the trained model.
[0013] This makes it possible to obtain a progress indicator that represents the progress of a series of tasks in a form that reflects the detailed progress within a single process operation. [Effects of the Invention]
[0014] According to the present disclosure, when a robot performs a series of actions based on predictions from a trained model, the user can grasp the progress of the detailed actions. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a block diagram showing an electrical configuration of a robot system according to an embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a series of tasks to be performed by a robot and the process operations that constitute the tasks. [Figure 3] FIG. 4 is a schematic diagram showing how state values and user operating forces constituting task data are repeatedly acquired. [Figure 4] FIG. 4 is a diagram illustrating process operation data. [Figure 5] FIG. 1 is a diagram illustrating clustering performed to build a trained model. [Figure 6] A diagram explaining the inference performed by a trained model. [Figure 7] 10 is a graph showing a comparative example in which the progress rate changes in units of process operations. [Figure 8] 10 is a graph showing an example of changes in progress during autonomous driving of a robot. [Figure 9] 10 is a graph showing an example of a change in the progress level when the robot's operation stalls midway during autonomous driving. [Figure 10] 10 is a graph showing an example of a change in progress when the robot behaves as if it is lost during autonomous driving. [Figure 11] FIG. 10 is a diagram illustrating a modified example of process operation data. [Figure 12] 10 is a conceptual diagram illustrating a comparison between time-based learning and progress-based learning of multiple task data. DETAILED DESCRIPTION OF THE INVENTION
[0016] Next, the disclosed embodiments will be described with reference to the drawings. Fig. 1 is a block diagram showing the electrical configuration of a robot system 1 according to an embodiment of the present disclosure.
[0017] 1 is a system that performs work using a robot 10. The work that the robot 10 can perform is various, including, for example, assembly, processing, painting, cleaning, and the like.
[0018] As will be described in detail later, the robot 10 is controlled using a model (trained model 43) constructed by machine learning data. Therefore, the robot system 1 can basically perform tasks autonomously without requiring assistance from a user. Furthermore, the robot 10 can not only perform tasks autonomously, but also perform tasks in response to user operation. In the following description, the robot 10 performing tasks autonomously may be referred to as "autonomous operation," and the robot 10 performing tasks in response to user operation may be referred to as "manual operation."
[0019] 1, the robot system 1 includes a robot 10 and a robot control device 15. The robot 10 and the robot control device 15 are connected to each other by wire or wirelessly, and can exchange signals.
[0020] The robot 10 has an arm attached to a base. The arm has multiple joints, each of which is equipped with an actuator. The robot 10 operates the actuators in response to an externally input operation command, thereby operating the arm.
[0021] An end effector selected according to the type of work is attached to the tip of the arm, and the robot 10 can operate the end effector in response to an operation command input from the outside.
[0022] The robot 10 is equipped with sensors for detecting the motion of the robot 10 and the surrounding environment, etc. In this embodiment, a motion sensor 11, a force sensor 12, and a camera 13 are attached to the robot 10.
[0023] The motion sensors 11 are provided for each joint of the arm section of the robot 10 and detect the rotation angle or angular velocity of each joint. The force sensors 12 detect the force received by the robot 10 when the robot 10 is operating. The force sensors 12 may be configured to detect the force applied to the end effector, or may be configured to detect the force applied to each joint of the arm section. The force sensors 12 may also be configured to detect a moment instead of or in addition to a force. The camera 13 detects an image of the workpiece to be worked on (the progress of work on the workpiece).
[0024] The data detected by the motion sensor 11 is motion data that indicates the motion of the robot 10. The data detected by the force sensor 12 and the camera 13 is ambient environment data that indicates the environment around the robot 10. A set of values of motion data and ambient environment data acquired at a certain timing may be referred to as a state value in the following explanation. The state value indicates the state of the robot 10 and its surroundings.
[0025] In the following description, the motion sensor 11, force sensor 12, and camera 13 provided on the robot 10 may be collectively referred to as the "state detection sensors 11 to 13." A collection of values detected by the state detection sensors 11 to 13 at a certain timing corresponds to a state value. The state detection sensors 11 to 13 may be provided around the robot 10 instead of being attached to the robot 10.
[0026] The robot control device 15 includes a user interface unit 20, an action switching unit (control data acquisition unit) 30, an AI unit 40, and an action label storage unit 50.
[0027] Specifically, the robot control device 15 is a computer equipped with a CPU, ROM, RAM, and HDD. The computer is equipped with devices such as a mouse for user operation. It is preferable for the computer to be equipped with a GPU, as this may enable the machine learning described below to be performed in a short time. The HDD stores programs for operating the robot control device 15. The above hardware and software work together to enable the robot control device 15 to function as a user interface unit 20, an action switching unit 30, an AI unit 40, and an action label storage unit 50.
[0028] The user interface unit 20 realizes the user interface function of the robot control device 15. The user interface unit 20 includes an operation unit 21 and a display unit (information output unit) 22.
[0029] The operation unit 21 is a device used to manually operate the robot 10. The operation unit 21 can be configured to include, for example, a lever, a pedal, and the like.
[0030] Although not shown, the operation unit 21 includes a known operation force detection sensor that detects the force (operation force) applied to the operation unit 21 by the user.
[0031] When the operation unit 21 is configured to be movable in various directions, the operation force may be a value including the direction and magnitude of the force, for example, a vector. Furthermore, the operation force may be detected not only as the force (N) applied by the user, but also as acceleration, which is a value linked to the force (i.e., the value obtained by dividing the force applied by the user by the mass of the operation unit 21).
[0032] In the following description, the operating force applied by the user to the operation unit 21 may be particularly referred to as a “user operating force.” The user operating force output by the user operating the operation unit 21 is converted into an operation command by the operation switching unit 30, as will be described later.
[0033] The display unit 22 can display various information in response to instructions from the user. The display unit 22 can be, for example, a liquid crystal display. The display unit 22 is disposed near the operation unit 21. If the operation unit 21 is disposed away from the robot 10, the display unit 22 may display images of the robot 10 and its surroundings.
[0034] The robot 10, the operation unit 21, and the AI unit 40 are connected to the action switching unit 30. The action switching unit 30 receives as input the user operation force output by the operation unit 21 and an estimated operation force (described below) output by the AI unit 40.
[0035] The action switching unit 30 outputs an action command to the robot 10 to operate the robot 10. The action switching unit 30 includes a switching unit 31 and a conversion unit 32.
[0036] The switching unit 31 is configured to output one of the input user operating force and the estimated operating force to the conversion unit 32. The switching unit 31 outputs the user operating force or the estimated operating force to the conversion unit 32 based on a selection signal indicating which of the user operating force and the estimated operating force is to be converted. The selection signal is output from the user interface unit 20 to the operation switching unit 30 when the user operates the user interface unit 20 as appropriate.
[0037] This allows switching between a state in which the user operates the robot 10 (manual operation) and a state in which the robot system 1 causes the robot 10 to perform work autonomously (autonomous operation). In the manual operation mode, the robot 10 operates based on the user's operating force output by the operation unit 21. In the autonomous operation mode, the robot 10 operates based on the estimated operating force output by the AI unit 40.
[0038] The conversion unit 32 converts either the user operation force or the estimated operation force input from the switching unit 31 into an operation command for operating the robot 10, and outputs the operation command to the robot 10. The operation command can also be referred to as control data for controlling the robot 10.
[0039] The AI unit 40 includes a trained model 43, a data input unit 41, an estimated data output unit 42, and a progress calculation unit (progress acquisition unit) 46.
[0040] The trained model 43 is constructed to allow the robot 10 to perform a series of tasks through autonomous operation. The trained model 43 used in the AI unit 40 may be in any format, and for example, a clustering model may be used. The trained model 43 may be constructed in the robot control device 15 or in another computer.
[0041] The data input unit 41 functions as an interface on the input side of the AI unit 40. To the data input unit 41, sensor information output from the state detection sensors 11-13 is input.
[0042] The estimated data output unit 42 functions as an interface on the output side of the AI unit 40. The estimated data output unit 42 can output the data output by the trained model 43.
[0043] The progress calculation unit 46 calculates how much progress in a series of tasks the output of the trained model 43 corresponds to. Details of the process performed by the progress calculation unit 46 will be described later.
[0044] In this embodiment, the AI unit 40 learns the operation of the robot 10 performed by the user via the operation unit 21, and constructs a learned model 43. Specifically, the AI unit 40 receives input of state values obtained by the state detection sensors 11 to 13 and a user operation force obtained by an operation force detection sensor.
[0045] Here, a series of operations performed by the autonomous driving of the robot 10 will be described with reference to FIG.
[0046] Consider a case where the robot 10 is made to perform a series of operations to place a workpiece 81 into a recess 82, as shown in Figure 2. From the start to the end of this series of operations, it can be considered that four operation states appear: airborne, contact, inserted, and completed.
[0047] Working state 1 (air) is a state in which the robot 10 holds the workpiece 81 and positions it above the recess 82. Working state 2 (contact) is a state in which the workpiece 81 held by the robot 10 is in contact with the surface on which the recess 82 is formed. Working state 3 (insertion) is a state in which the workpiece 81 held by the robot 10 is slightly inserted into the recess 82. Working state 4 (complete) is a state in which the workpiece 81 held by the robot 10 is completely inserted into the recess 82.
[0048] The four work states correspond to one of the start state, midway state, and end state of a series of work performed by the robot 10. A series of work performed by the robot 10 is divided into multiple processes, with work states as boundaries. As the robot 10 performs the action corresponding to each process, the work state transitions in the following order: work state 1 (air), work state 2 (contact), work state 3 (insertion), and work state 4 (complete).
[0049] Hereinafter, the actions corresponding to each step in a series of tasks divided by work states will be referred to as "step actions." In the example of Figure 2, the step actions are action 1 (lowering action), action 2 (rubbing action), and action 3 (lowering action in hole). By repeating step actions, an action that naturally transitions to the next task is called. For example, when action 1 (lowering action) is performed in work state 1 (air), a transition to work state 2 (contact) occurs. The same applies to action 2 and action 3.
[0050] The four work states are registered in advance in the AI unit 40 by the user when the AI unit 40 is made to learn a series of work.
[0051] At this time, the user sets a numerical value of the degree of progress for each task state. The degree of progress is a parameter used to evaluate the degree of progress in a series of tasks that the action performed by the robot 10 based on the output of the trained model 43 corresponds to. In this embodiment, the degree of progress takes a value in the range from 0 to 1, and the closer to 1 the value is, the more progress has been made in the series of tasks. The range within which the degree of progress changes is arbitrary, and can be set to a range from 0 to 100, for example.
[0052] The progress value is predetermined for each task state so that it monotonically increases in the order in which the task state appears. The progress value may be determined by the user or automatically by the AI unit 40.
[0053] In the following description, it is assumed that the progress level is set to 0 for operation state 1 (air), 0.5 for operation state 2 (contact), 0.8 for operation state 3 (insertion), and 1 for operation state 4 (completed). This necessarily assigns a progress level range of 0 to 0.5 to operation 1 (lowering operation), a progress level range of 0.5 to 0.8 to operation 2 (rubbing operation), and a progress level range of 0.8 to 1 to operation 3 (lowering operation in hole).
[0054] Data for machine learning can be acquired by the user actually operating the operation unit 21 to make the robot 10 perform a series of tasks. Hereinafter, data obtained by making the robot 10 perform the series of tasks shown in Fig. 2 once may be referred to as task data.
[0055] The user operates the operation unit 21 to make the robot 10 perform a series of tasks, and when each task state is reached, the user instructs the AI unit 40 in real time that the task state has changed. The instruction can be given, for example, by the user uttering a specific word into a microphone. The operations between the times when the user instructs a change in the task state are treated as one process operation.
[0056] The instruction does not have to be given in real time. For example, after the operation data is obtained, the user can specify the timing at which the operation state was changed while viewing the data.
[0057] A single process operation is defined as a series of operations during which a change in work state is specified by the user. Therefore, each process operation is essentially specified by the user's manual operation (manual specification mode). However, the process operations that make up a series of operations can also be automatically recognized (automatic specification mode).
[0058] Generally, the manner in which the user applies force to the operation unit 21 changes as the process operation changes. For example, the AI unit 40 detects that a change in the process operation has occurred based on the change in the way the force is applied. This makes it possible to realize an automatic designation mode. A change in the process operation can also be detected using known machine learning (e.g., supervised learning).
[0059] 3 shows a schematic diagram of how task data for learning is acquired from various sensors when a user operates the robot 10 to perform a series of tasks. Data is repeatedly acquired at appropriate time intervals from the state detection sensors 11 to 13 and the operating force detection sensor. In this embodiment, the data acquisition cycle is set to one second, but this can be changed as appropriate.
[0060] A data set is composed of a state value and a user's operating force acquired at a certain timing. FIG. 3 shows an example in which the time interval for acquiring data from the operating force detection sensors is equal to the data acquisition cycle, while the time interval for acquiring data from the state detection sensors 11 to 13 is shorter than the data acquisition cycle. In the example of FIG. 3, one data set includes a transition of the state value over a short period from the previous data acquisition timing to the current data acquisition timing. In this way, one data set may include a transition over time of at least one of the state value and the user's operating force.
[0061] The data set is acquired both when the robot 10 is in any working state and when the robot 10 is performing any process operation. The work data includes a plurality of data sets in a form associated with the time when the data set was acquired.
[0062] Next, the process operation data obtained by dividing the work data will be described. FIG. 4 shows an example of data for constructing the learned model 43.
[0063] The data in FIG. 4 indicates that at the time when k seconds have elapsed since the start of a series of operations, the working state is the working state 2 (contact), and the state value is S 21,k However, as a result of the robot 10 operating for n seconds, the working state has transitioned to the working state 3 (insertion), and the state value has become S 31,k+n . Here, n is an integer of 1 or more. The data in FIG. 4 corresponds to a process operation that is part of the above series of operations, specifically, the operation 2 (rubbing operation) described above. Hereinafter, data corresponding to one process operation may be referred to as process operation data.
[0064] In the AI unit 40, the machine learning model learns using the process operation data shown in FIG. 4 as a unit. One piece of process operation data is a set of the following [1] to [4]. [1] Data indicating that the working state before the transition is the working state 2 (contact). [2] Data indicating that the destination working state after the transition is the working state 3 (insertion). [3] The state value S 21,m and the user operation force I 21,m of the data after m seconds from the start of a series of operations during the process of the working state transition. Here, when the time from the start of a series of operations to the start of the process operation is k seconds, m is an integer that satisfies k ≤ m < k + n. This data corresponds to a plurality of data sets consisting of the state value S 21 and the value I of the user operation force 21 arranged while being associated with the value of m respectively. [4] The state value S 31,k+n of the data after n seconds from the start of the transition as a result of the completion of the transition of the working state.
[0065] The above set of [1] to [4] can be obtained by the AI unit 40 performing processing to extract a portion from task data corresponding to a series of tasks.
[0066] Since there are variations in user operations and situations, there are many variations in operation 2 (rubbing operation) as a process operation. For example, if the current work state is work state 2 (contact) and the current state value is S 22,k When the state is changed to work state 3 (insertion), the state value is S 32,k+n The machine learning model also learns this process operation data in the same way as described above.
[0067] The series of tasks includes not only the transition from task state 2 (contact) to task state 3 (insertion), but also transitions between other task states, such as a transition from task state 1 (air) to task state 2 (contact). In other words, the series of tasks includes not only action 2 (rubbing) but also other process actions. The machine learning model also learns the process actions corresponding to these transitions in the same way as above.
[0068] The user repeatedly operates the operation unit 21 to make the robot 10 repeatedly perform the same series of tasks. This provides multiple task data, allowing the machine learning model to learn variations in each process operation.
[0069] In this embodiment, a machine learning model using clustering is adopted. Learning one piece of process operation data by the machine learning model means learning one feature vector corresponding to the set of [1] to [4] described above with reference to FIG. 4. In the machine learning model of this embodiment, a map based on a multidimensional feature space is defined. Each time one process operation is learned, the AI unit 40 plots one feature vector in the feature space of the machine learning model. The upper part of FIG. 5 conceptually shows how the feature vector corresponding to the process operation data is plotted in the feature space.
[0070] After all feature vectors to be learned are input to the AI unit 40, clustering is performed on the feature vectors plotted in the feature space.
[0071] Clustering is a technique for learning distribution laws from a large amount of data and obtaining multiple clusters that are groups of data with similar characteristics. Here, data refers to feature vectors. As a clustering method, for example, a known non-hierarchical clustering method can be used as appropriate. As a result, multiple clusters that group similar feature vectors together can be obtained. The number of clusters in clustering can be determined as appropriate. When there are two feature vectors that have different combinations of an operation state at a certain point in time and an operation state to which they will transition, it is preferable that the clustering be performed so that the two feature vectors belong to different clusters.
[0072] The AI unit 40 then calculates a feature vector (in other words, process operation data) that represents each cluster. Hereinafter, this representative feature vector may be referred to as a node. There are various ways to determine the node, but for example, it can be data corresponding to the center of gravity of each cluster. The nodes are indicated by an X mark at the bottom of Figure 5.
[0073] This completes the training phase, allowing the construction of a trained model 43. The trained model 43 constructed by the AI unit 40 is essentially a distribution of nodes in a multidimensional feature space, as schematically shown in FIG.
[0074] Next, the output from the AI unit 40 based on the trained model 43 will be described.
[0075] In the inference phase, a feature vector including the current state values obtained by the state detection sensors 11 to 13 and the current working state is generated in the AI unit 40. This single feature vector is input to the trained model 43. Hereinafter, this feature vector may be referred to as an input feature vector. The input feature vector may further include the most recent past transitions of the state values and the most recent past transitions of the user's operating force.
[0076] The trained model 43 finds one or more nodes similar to the input feature vector. Because the current operating force is the subject of inference, the input feature vector does not contain information about the current operating force. Therefore, when calculating the similarity between the input feature vector and a node, the dimension related to the current operating force is ignored. The similarity can be calculated using a known calculation formula (e.g., Euclidean distance). Hereinafter, the found node may be referred to as the output target node. Finding the output target node is synonymous with obtaining a classification result indicating which cluster the input feature vector will be classified into by clustering.
[0077] Once the output target node is determined, the trained model 43 further determines which state value at which time in the transition [3] above contained in the output target node is most similar to the current state value contained in the input feature vector. The similarity can be determined using a known method (e.g., Euclidean distance) as described above. In this case, not only the similarity of the state values but also the similarity of the time are taken into consideration.
[0078] In the example of Figure 6, the state value S at the 5th second 2X,5 is most similar to the state value of the input feature vector. When the most similar state value is found, the trained model 43 outputs the user operation force that belongs to the same data set as the state value in the transition of [3] above as the estimated operation force. In the example of FIG. 6, the state value S 2X,5 The user operation force I 2X,5 is output as the estimated operating force.
[0079] In the output target node, when the estimated operating force corresponds to the last user operating force in the transition [3], the AI unit 40 updates the current working state to the working state after the transition.
[0080] The above process is repeated, for example, every second, and the robot 10 is operated based on the estimated operating force, thereby enabling the robot 10 to operate autonomously.
[0081] Next, the processing of the above-mentioned progress degree calculation unit 46 will be described.
[0082] The progress calculation unit 46 can obtain, by calculation, a progress indicating the degree of progress in a series of tasks to which the estimated operating force output by the trained model 43 corresponds. This progress can be calculated based on the progress value corresponding to the task state before the transition indicated by the output target node, the progress value corresponding to the task state after the transition, and information on the chronological order of the estimated operating forces at the output target node.
[0083] A specific description will be given below. The output target node determined by the trained model 43 as being similar to the feature vector input to the AI unit 40 includes the above [1] to [4] as information. In this example, as shown in the lower part of Fig. 6, the output target node indicates that the work state before the transition is work state 2 (contact), the work state to be transitioned is work state 3 (insertion), and the state value S for 15 seconds for the transition between work states. 2X,k ~S 2X,k+14 , and user operation force I for 15 seconds 2X,k ~I 2X,k+14 Furthermore, for the current state value of the input feature vector, the state value S of the node at the 5th second 2X,k+5 In this case, the AI unit 40 calculates the user operation force I 2X,k+5 It is also possible to omit the conditions regarding the pre-transition working state and the post-transition working state, and select the node that is most similar to the current state value from all the nodes and output it as the estimated operating force.
[0084] As described above, a progress level of 0.5 is defined for operation state 2 (contact), and a progress level of 0.8 is defined for operation state 3 (insertion). In the above node, the state 5 seconds after the start of the process operation is, in terms of time, a state in which 5 / 15 of the progress level range from 0.5 to 0.8 has been completed, so the progress level can be calculated to be equivalent to 0.6. The progress level calculation unit 46 performs the calculation to convert time into progress level in this way. The AI unit 40 outputs the obtained progress level value of 0.6 to a display or the like. It is also possible to perform the above calculations in advance to create a conversion table, and use this table to convert time into progress level.
[0085] Let us consider a case in which the progress degree is determined only according to the order of the work states. In this case, as shown in the graph of the comparative example in FIG. 7 , even if the robot 10 starts a series of tasks from work state 1 (air), the progress degree remains at 0 until immediately before the transition to work state 2 (contact) is completed. The progress degree discontinuously changes to 0.5 at the timing when the transition to work state 2 (contact) is completed. Thereafter, the progress degree remains at 0.5 until immediately before the transition to work state 3 (insertion) is completed. The progress degree discontinuously changes to 0.8 at the timing when the transition to work state 3 (insertion) is completed. Thereafter, the progress degree remains at 0.8 until immediately before the transition to work state 4 (completed). The progress degree discontinuously becomes 1.0 at the timing when the transition to work state 4 (completed) is completed.
[0086] In this way, when the progress is changed in units of task states (in other words, in units of clusters), the changes in the progress are coarse. Therefore, although it is possible to grasp the progress at a large granularity, it is difficult to grasp the detailed progress.
[0087] In particular, process operations may include situations in which similar states continue for a relatively long period of time, such as an operation in which the robot 10 continues to move a workpiece held at its tip in the air, or an operation in which the held workpiece continues to be pressed against another member, etc. In such cases, the progress may not change as if the operation is stagnating.
[0088] In this regard, in this embodiment, the progress degree is calculated not only in a manner that reflects a change in the task state but also in a manner that reflects the order in which the estimated operational force corresponds in the time-series call-up of the user operational forces at the output target node. Therefore, as shown in Fig. 8, the progress degree increases in a relatively smooth manner as the series of tasks progresses, allowing the user to easily understand the detailed progress of the tasks.
[0089] When the robot 10 operates autonomously, the progress is displayed, for example, in real time on the display unit 22. It is preferable to display the progress in the form of a graph, as shown in FIG. 8, for example, since this allows the user to intuitively understand. If the autonomous operation by the trained model 43 in the inference phase is performed without any problems, the progress changes in small steps from 0 to 1. The unit of progress change is the progress equivalent to one second, which is the data acquisition period mentioned above. The data acquisition period can also be said to be the output period of the estimated operating force.
[0090] If the robot 10 stops in the middle of any of the process operations, the value of the progress degree will no longer increase from a value corresponding to the middle of that process operation (e.g., 0.6), as shown in Fig. 9. Also, if the autonomous operation of the robot 10 becomes hesitant or moves in a trial-and-error manner, the value of the progress degree will fluctuate repeatedly between, for example, approximately 0.55 and 0.7, as shown in Fig. 10. By monitoring the progress degree, which fluctuates minutely in this way, it is possible to easily grasp that an unexpected situation has occurred in the autonomous operation of the robot 10.
[0091] 8 to 10, in addition to the graph, the action labels stored in the action label storage unit 50 may be displayed on the display unit 22. This allows the user to intuitively understand which process action is being performed.
[0092] Generally, the operation of a machine learning model cannot be grasped from the outside, and it is unclear which of the process operations the model is intended to perform. Therefore, it is difficult to determine whether the machine learning model is performing the operations in the correct order without making any mistakes, or to evaluate which operation failed in the event of a failure. However, in this embodiment, as shown in FIGS. 8 to 10, the detailed order of operations performed in a series of tasks is visualized. Therefore, it becomes easy to correctly evaluate the performance of the constructed trained model 43, and further, it is possible to easily obtain clues for improving the model.
[0093] Next, a modified example of the process operation data will be described with reference to FIG.
[0094] In the process operation data of the above embodiment, as shown in FIG. 21 and user operation force I 21 In this modification, the data set is associated with the progress level as shown in FIG.
[0095] The progress rate represents the progress of a series of tasks as a ratio, so by using the progress rate instead of the time, it is possible to reduce the influence of variations in the length of task data.
[0096] For example, consider the case where the three task data shown in the upper part of Figure 12 are trained by a machine learning model to build a trained model 43. Because user operations and situations vary, there is variation in the time at which data corresponding to each process action appears and in the chronological length of that data. Therefore, if the time from the start of a series of tasks is used as is for the process action data, as in Figure 4, this can have a significant impact on clustering. Since temporal variations in actions accumulate in later steps, later operation steps are particularly susceptible to the impact.
[0097] In this regard, in this modified example, progress is used instead of time. By using progress as a criterion, the task data can be expanded or contracted in the time axis direction, and machine learning of the process operation data can be performed with the time series lengths aligned as shown in the lower part of Figure 12. This makes it easier for process operation data indicating similar process operations to form the same cluster in machine learning clustering. Therefore, it is expected that the trained model 43 will perform more appropriate inference.
[0098] It is also possible to set a progress level that is not affected by variations in the time taken for the work steps included in a series of tasks. Specifically, some work steps, such as pushing a workpiece against other components, can take a long time depending on the circumstances. When a series of tasks includes multiple work steps that tend to have large variations in the time required to complete, simply assigning a progress level from 0 to 1 to the entire series of tasks can make it difficult for the progress levels of the work data to match, even when the same series of tasks are performed.
[0099] To solve this situation, the following can be done. That is, the work processes expected to be included in a series of work are roughly classified in advance, and each classification is assigned a label, for example, 0, 1, 2, ..., according to the order of the work. The above-mentioned pre-classification can be performed based on the work status or by machine learning such as clustering. For work labeled 1, a progress level is assigned from 0.0 to 1.0, and for work labeled 2, a progress level is assigned from 1.0 to 2.0, .... This ensures that the same work processes will show similar progress levels even when comparing progress data with large variations in the time length of the work processes. This makes it easier for the trained model 43 to make appropriate inferences.
[0100] As described above, the robot control device 15 of this embodiment includes the trained model 43, the operation switching unit 30, the progress calculation unit 46, and the display unit 22. The trained model 43 is constructed by learning task data in which state values are input data when a human operates the robot 10 to perform a series of tasks, and corresponding user operation forces are output data. When state values are input to the trained model 43, the operation switching unit 30 obtains control data for the robot 10 to perform the task by obtaining, from the trained model 43, an estimated user operation force corresponding to the state values. When input data is input to the trained model 43, the progress calculation unit 46 obtains a progress level indicating which progress level in the series of tasks corresponds to the output data output in response to the input data. The display unit 22 can output the progress level. The trained model 43 can determine which of multiple process operations corresponding to a divided series of tasks the input data will be classified into. In the trained model 43, an output transition, which is the time transition of the user's operating force for realizing the process operation, is associated with a node representing each classification. The trained model 43 determines the user's operating force corresponding to the state value of the input data from the output transition of the node associated with the classification result of the input data, and sets this as output data. The range of change of the progress degree (from 0 to 1) is divided into multiple progress ranges so as to correspond to multiple process operations. The order of the multiple progress ranges corresponds to the order of the multiple process operations in a series of tasks. The progress degree acquired by the progress degree calculation unit 46 changes depending on the position of the user's operating force corresponding to the state value in the output transition associated with the process operation within the progress range corresponding to the process operation, which is the classification result of the input data performed by the trained model 43.
[0101] This allows the progress of a series of tasks to be displayed in a form that reflects the detailed progress of each step. Therefore, the user can refer to the progress and more appropriately understand the status of autonomous driving.
[0102] The robot control device 15 of this embodiment includes an action label storage unit 50. The action label, which includes information expressing the content of the process action, is stored in association with the process action. The display unit 22 is capable of outputting the action label.
[0103] This allows the user to more intuitively understand the autonomous driving situation.
[0104] In the example of FIG. 11, the output data of the task data includes the user's operating force associated with the progress level.
[0105] This makes it possible to prevent variations in the time length of the task data from becoming noise in machine learning, thereby enabling the trained model 43 to make more appropriate inferences.
[0106] In the robot control device 15 of this embodiment, the work data is divided into process operation data corresponding to one process operation. The trained model 43 is constructed by clustering each of the process operation data.
[0107] This makes it possible to construct a trained model 43 that can flexibly respond to various situations.
[0108] The robot system 1 of this embodiment includes a robot control device 15 and a robot 10.
[0109] This allows the user to understand the progress of the operation of the robot 10 when it is operating autonomously, in a manner that reflects the situation in more detail.
[0110] The preferred embodiment of the present disclosure has been described above, but the above configuration can be modified, for example, as follows.
[0111] In the above embodiment, the trained model 43 learns the relationship between the state value and the user's operating force. Alternatively, the trained model 43 may be configured to learn the relationship between the state value and the operation command to the robot 10.
[0112] The division of task data into process operation data may be performed automatically using multiple base trained models. Specifically, each base trained model is pre-constructed for each standard process operation. A standard process operation is an operation simpler than a series of operations, and many standard process operations are possible, including the above-mentioned lowering operation, rubbing operation, and hole-descent operation. Each base trained model learns the time-series order of state values and user operation forces when a user operates the operation unit 21 to cause the robot 10 to perform a standard process operation. When task data is obtained, the state values included in the task data are input to each base trained model, and the user operation forces output by the base trained model are compared with the user operation forces of the task data. If there is a section in the time series of inputs and outputs of the task data that is similar to the time series of inputs and outputs of the base trained model, that section can be determined to be the process operation.
[0113] The progress may be displayed numerically, for example, instead of being displayed in a graph format as in FIG.
[0114] The robot control device 15 may be provided with an audio output unit as an information output unit instead of or in addition to the display unit 22. In this case, the progress level can be output by audio. For example, the robot control device 15 may be configured to read out the progress level numerical value by voice synthesis. Similarly, the action label may be output by audio.
[0115] The action label may be any mark that can identify the content of the action. Instead of a character string such as "pull out," the action label may be an image such as an icon.
[0116] As the trained model 43, instead of a machine learning model based on clustering, a machine learning model of another type, for example, a neural network model, can be used.
[0117] As a sensor (status sensor) for acquiring the status of the robot 10 and its surroundings, a sensor other than the motion sensor 11, the force sensor 12, and the camera 13 may be used.
[0118] An operation position detection sensor may be provided in place of or in addition to the operation force detection sensor in the operation unit 21. Similar to the operation force detection sensor, the operation position detection sensor can also be said to be a sensor that detects human operation.
[0119] The robot system 1 may be one in which the operation unit 21 is a master arm used for remote control and the robot 10 is a slave arm. In this case, the AI unit 40 can construct a trained model 43 that is trained based on the operation of the master arm by the user. [Explanation of symbols]
[0120] 1. Robot System 10. Robot 15 Robot control device 22 Display unit (information output unit) 30 Operation switching unit (control data acquisition unit) 43 trained models 46 Progress calculation unit (progress acquisition unit) 50 Action label memory section 81 Work (condition) 82 Recess (condition)
Claims
1. A trained model constructed by learning task data in which input data is the state of the robot and its surroundings when a human operates a robot to perform a series of tasks, and output data is the corresponding human operation or the behavior of the robot as a result of the operation; and a control data acquisition unit that, when input data regarding the state of the robot and its surroundings is input to the trained model, obtains output data regarding the human operation or robot movement estimated in response to the input data from the trained model, thereby obtaining control data for the robot to perform the task; and a progress acquisition unit that acquires a progress level indicating a degree of progress in the series of tasks to which the output data output in response to the input data corresponds when the input data is input to the trained model; an information output unit capable of outputting the progress; Equipped with The trained model is capable of determining to which of a plurality of process actions corresponding to divisions of the series of operations the input data is classified; In the trained model, an output transition is defined, which is a time transition of the human operation for realizing the process operation or the operation of the robot due to the operation, in association with each classification; The trained model determines an output corresponding to the input data from the output transition associated with the classification result of the input data, and sets the output data as the output data; The change range of the progress degree is divided into a plurality of progress degree ranges corresponding to a plurality of the process operations; the order of the plurality of progress ranges corresponds to the order of the plurality of process actions in the series of works; A robot control device characterized in that the progress acquired by the progress acquisition unit changes depending on the position of the output corresponding to the input data in the output transition associated with the process operation within the progress range corresponding to the process operation, which is the classification result of the input data performed by the trained model.
2. The robot control device according to claim 1, an action label storage unit that stores an action label including information that expresses the content of the process action in association with the process action; The robot control device is characterized in that the information output unit is capable of outputting the action label.
3. The robot control device according to claim 1 or 2, The robot control device is characterized in that the output data of the work data includes the operation of the human or the movement of the robot resulting from the operation, in association with the progress level.
4. The robot control device according to any one of claims 1 to 3, The work data is divided into process operation data corresponding to one of the process operations, A robot control device characterized in that the learned model is constructed by clustering each of the process operation data.
5. The robot control device according to any one of claims 1 to 4, The robot; A robot system comprising:
6. A trained model is constructed by learning task data in which the state of the robot and its surroundings when a human operates a robot to perform a series of tasks is used as input data, and the corresponding human operation or the robot's behavior as a result of the operation is used as output data; When input data relating to the state of the robot and its surroundings is input to the trained model, output data relating to the human operation or robot movement estimated in response to the input data is obtained from the trained model, thereby obtaining control data for the robot to perform the task; When the input data is input to the trained model, a progress level indicating to which degree of progress in the series of tasks the output data output in response to the input data corresponds is acquired; outputting the progress; The trained model is capable of determining to which of a plurality of process actions corresponding to divisions of the series of operations the input data is classified; In the trained model, an output transition is defined, which is a time transition of the human operation for realizing the process operation or the operation of the robot due to the operation, in association with each classification; The trained model determines an output corresponding to the input data from the output transition associated with the classification result of the input data, and sets the output data as the output data; The change range of the progress degree is divided into a plurality of progress degree ranges corresponding to a plurality of the process operations, the order of the plurality of progress ranges corresponds to the order of the plurality of process actions in the series of works; A robot control method characterized in that the progress changes depending on the position of the output corresponding to the input data in the output transition associated with the process operation within the progress range corresponding to the process operation, which is the result of the classification of the input data performed by the trained model.
Citation Information
Patent Citations
Work progress estimating apparatus and method utilizing id medium and sensor
JP2011107836A
Work management device, work management method, and program
JP2018163556A
Robot control device, robot system and robot control method
JP2020104215A
Robot system and supplemental learning method
WO2019225746A1