Information processing device, information processing method, and storage medium
Patent Information
- Application Number
- JP2025503278
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-05
AI Technical Summary
Existing motion planning techniques for robots cannot optimize a combination of multiple motions effectively, leading to suboptimal performance in tasks like grasping, moving, and placing objects.
An information processing device comprising a prediction unit, calculation unit, and optimization unit that predicts the state of the operation target and robot, calculates the success score of motion plans, and adjusts the motion plan group based on the success score, using learning models to refine and optimize the motion plans.
This approach enables the optimization of motion plans by identifying and improving specific subtasks, resulting in higher success rates for robot operations by adjusting movement paths and plan contents without altering start and end states.
Abstract
Description
Information processing device, information processing method, and storage medium
[0001] The present disclosure relates to an information processing device, an information processing method, and a storage medium.
[0002] In recent years, robots have been introduced in various situations, and tasks are being automated using robots. As a result, it is necessary to generate control plans for robots, and there is a demand for generating operation plans with a high success rate. As an example, Patent Literature 1 describes a method for determining the success or failure of a picking operation by a robot.
[0003] International Publication No. 2022 / 123978
[0004] However, the technology described in the above-mentioned Patent Document 1 does not describe an action plan that combines a plurality of actions, and there is a problem in that it is not possible to optimize such an action plan.
[0005] Therefore, an object of the present disclosure is to provide an information processing device that can solve the above-mentioned problem of being unable to optimize an action plan consisting of a combination of multiple actions.
[0006] An information processing device according to one embodiment of the present disclosure includes: a prediction unit that predicts, based on a group of motion plans consisting of a plurality of motion plans for the robot with respect to a target, a state of at least one of each of the operation targets and the robot when each of the motion plans is executed; a calculation unit that calculates a degree of success of the group of motion plans based on each of the predicted states; and a modification unit that modifies the group of motion plans based on the calculated degree of success.
[0007] Moreover, an information processing method according to one embodiment of the present disclosure is configured to: predict, based on a group of motion plans consisting of a plurality of motion plans for a robot with respect to an object to be operated, a state of at least one of each of the object to be operated and the robot when each of the motion plans is executed; calculate, based on each of the predicted states, a degree of success of the group of motion plans; and modify the group of motion plans based on the calculated degree of success.
[0008] Furthermore, a program according to one embodiment of the present disclosure has a configuration for causing a computer to execute the following processes: predicting, based on a group of motion plans consisting of a plurality of motion plans for a robot with respect to a manipulation target, a state of at least one of each of the manipulation targets and the robot when each of the motion plans is executed; calculating a degree of success of the group of motion plans based on each of the predicted states; and modifying the group of motion plans based on the calculated degree of success.
[0009] With the above-described configuration, the present disclosure can optimize an action plan consisting of a combination of multiple actions.
[0010] FIG. 1 is a block diagram showing the configuration of an information processing device according to a first embodiment of the present disclosure. FIG. 2 is a diagram showing the state of processing by the information processing device disclosed in FIG. 1. FIG. 3 is a diagram showing the state of processing by the information processing device disclosed in FIG. 1. FIG. 4 is a diagram showing the state of processing by the information processing device disclosed in FIG. 1. FIG. 5 is a flowchart showing the operation of the information processing device disclosed in FIG. 1. FIG. 6 is a block diagram showing the hardware configuration of an information processing device according to a second embodiment of the present disclosure. FIG. 7 is a block diagram showing the configuration of an information processing device according to a second embodiment of the present disclosure.
[0011] First Embodiment A first embodiment of the present disclosure will be described with reference to Fig. 1 to Fig. 6. Fig. 1 is a diagram for explaining the configuration of an information processing device, and Fig. 2 to Fig. 6 are diagrams for explaining the processing operation of the information processing device.
[0012] [Configuration] The information processing device 10 in this embodiment has a function for optimizing a robot motion plan. In this embodiment, the robot is assumed to be a robot arm, as an example. The robot motion plan to be optimized is formed as a group of motion plans consisting of a combination of multiple motion plans, such as "grasping," "moving," and "placing" an object to be manipulated by the robot arm. In this embodiment, each motion plan of the robot is called a subtask, and for example, a subtask may include control commands for an action such as "grasping" that specify the movement position, movement speed, and movement trajectory of the robot's joints, etc., to achieve the action.
[0013] The information processing device 10 is composed of one or more information processing devices each including a calculation device and a storage device. As shown in FIG. 1 , the information processing device 10 includes a motion planning unit 11, a prediction unit 12, a determination unit 13, and an optimization unit 14. The functions of the motion planning unit 11, the prediction unit 12, the determination unit 13, and the optimization unit 14 can be realized by the calculation device executing a program for realizing each function stored in the storage device. The information processing device 10 also includes a goal storage unit 16, a motion plan storage unit 17, and a learning model storage unit 18. The goal storage unit 16, the motion plan storage unit 17, and the learning model storage unit 18 are each composed of a storage device. Each component will be described in detail below.
[0014] The target memory unit 16 stores target information for the robot's operation plan. For example, the target information includes initial information representing the initial state of each object to be operated by the robot and the robot, and target information representing the target state after movement relative to the initial state. For example, as shown in FIG. 3, the initial information for the object and the robot may include an initial image I (image I) of the space in which the object and the robot are located. 0However, the initial state information may be coordinate information representing the initial position of an object or a robot, or any information representing the initial state of an object or a robot. Furthermore, the target information of an object may include coordinate information of the target position of each object to be operated, and may also include information representing the shape and size of each object and information about the space in which the object is located. Alternatively, the target information may include a target movement position of the robot. However, the target information is not limited to the above-mentioned information, and may include any information that serves as a target for the robot's operation plan. Note that, although the above-mentioned initial information representing the initial state and the target information representing the target state include information representing the state of each object or the robot, they may also include information representing the state of each object and the robot, or may include information representing the state of either the object or the robot.
[0015] As will be described later, the motion plan storage unit 17 stores the motion plan information of the robot planned by the motion plan unit 11. For example, as shown in FIG. 3, the motion plan information includes subtasks a, b, c, and d, which are motion plans for each of the robots. 1 , a 2 , .... The subtasks include, for example, control commands for a robot arm, and include the start position, end position, and movement trajectory from the start position to the end position of the robot, along with an operation such as "grasping" an object.
[0016] The learning model storage unit 18 stores learning models generated by prior learning to be used to optimize the robot's motion plans, as described below. As the learning models, a first learning model Ma is stored, which predicts the state of an object to be operated and the robot from a group of motion plans for the robot. Here, a method for generating the first learning model Ma will be described with reference to FIG. 2.
[0017] First, as shown in FIG. 2, a plurality of motion plans (a) are used as learning data, which are a plurality of subtasks that actually execute the motion of the robot. 1 , a 2, ...), and an image (I) of the state of the robot and the object to be operated when the robot actually executes the operation plan group A. 0 , I 1 , ...) are prepared in advance. 0 is the initial image I before executing the motion plan. 0 and image I n is the n-th motion plan a in the motion plan group A. n The images are taken when the motion plan is executed. n A feature extractor m is provided to extract feature quantities from each of the motion plans a n The state feature Z represents the state when n In this situation, the first learning model Ma is configured to estimate the state features (Z hat mark n-1 ) and a predetermined motion plan to be executed (a n ) and a predetermined motion plan (a n ) The predicted state features of the object or robot after executing n ) for each of the above motion plans a n The state feature Z when n The image and state feature Zn described above may include information representing the state of each object and the robot, or may include information representing the state of either each object or the robot.
[0018] Specifically, as shown in FIG. 2, an initial image I 0 The feature extractor m extracts the initial state feature (Z 0 ) is extracted, and the first learning model Ma extracts the initial state features (Z 0 ) and the first motion plan (a 1 ) and input the first predicted state feature (Z hat mark 1 ) is output. Then, the first learning model Ma outputs the first predicted state feature (Z hat mark 1 ) and the second motion plan (a 2 ) to input the second predicted state feature (Z hat mark 2By repeating this process, the first learning model Ma generates a predicted state feature set (Z hat mark) shown in the following equation 1. The first learning model Ma is a combination of the predicted state feature set (Z hat mark) and each motion plan a n When each image I n The parameters are updated by learning to reduce the error between the state feature Z extracted from the state feature Z. For example, machine learning is performed using the Kullback-Leibler divergence between the predicted state feature group (Z hat mark) and the state feature Z to generate a first learning model Ma.
[0019] The learning model storage unit 18 also stores a second learning model Mb that calculates the success rate of a group of operation plans as a learning model. Here, a method for generating the second learning model Mb will be described with reference to FIG.
[0020] First, as shown in FIG. 2, each motion plan a in the motion plan group A is generated using the first learning model Ma described above as learning data. n A set of predicted state features (Z hat marks) of the object or robot corresponding to the object or robot and a success score representing the degree of success when the group of motion plans A is actually executed are prepared in advance. The second learning model Mb is trained using the success score when the group of motion plans A is actually executed as training data, in contrast to the success score, which is the output when the group of predicted state features is input. Specifically, as shown in FIG. 2 , the second learning model Mb updates its parameters by training to minimize the error between the output when the group of predicted state features (Z hat marks) shown in Equation 1 are input and the degree of success when the group of motion plans A is actually executed. For example, the second learning model Mb is generated by machine learning using the cross-entropy error between the success score and the training data. Note that in this embodiment, the second learning model Mb is configured to output a success score indicating a higher degree of success, and therefore the training label may also be configured with information representing the success score. However, the second learning model Mb may be configured to output only the presence or absence of success, and accordingly, the training data may also be information representing only the presence or absence of success.
[0021] The motion planning unit 11 generates motion planning information representing a motion plan for the robot based on the initial information and target information stored in the target storage unit 16. For example, the motion planning unit 11 generates an initial image I as initial information as shown in FIG. 0 The motion planning unit 11 generates a motion plan for the robot so that each object moves to a target position specified by the target information, based on information on the object to be operated, the initial position of the robot, and the like. At this time, as shown in FIG. 3, the motion planning unit 11 generates a motion plan for each of the subtasks (a 1 , a 2 The motion planning unit 11 generates a motion plan group A consisting of a set of subtasks (subtasks) (such as "grasp" an object, ...) as motion plan information. For example, the motion planning unit 11 generates a control command including a motion start position, an end position, a motion trajectory, etc. of the robot, as well as a subtask such as "grasp" an object. The motion planning unit 11 then stores the generated motion plan group A consisting of the plurality of subtasks in the motion plan storage unit 17.
[0022] The motion planning unit 11 is not limited to generating detailed control commands for the robot including a movement trajectory as described above as subtasks of the motion planning information, but may generate motion planning information consisting of a combination of subtasks that represent rough control contents. For example, the motion planning unit 11 may generate rough motion planning information that combines subtasks such as grasping object x and moving object x in a chronological order. The motion planning unit 11 may also acquire preset motion planning information.
[0023] The prediction unit 12 predicts the state when each of the subtasks, which are the plurality of operation plans, is executed based on the generated operation plan group, using the first learning model Ma. Specifically, as shown in FIG. 3, the prediction unit 12 first generates an initial image I representing the initial state of the object or robot to be operated. 0 The feature extractor m extracts the initial state feature (S 0 ) and extract the initial state features (S 0 ) and the first motion plan (a 1 ) and input them into the first learning model Ma to obtain the first predicted state feature (S 1 ) and then the prediction unit 12 outputs the first predicted state feature (S 1) and the second motion plan (a 2 ) is input to the first learning model Ma to obtain the second predicted state feature (S 2 By repeating this process, the prediction unit 12 outputs a predicted state feature set (S 1 , S 2 , .. . , S T ) to generate the
[0024] The determination unit 13 (calculation unit) calculates each predicted state feature set (S 1 , S 2 , .. . , S T ) into the second learning model Mb to calculate a success score representing the degree of success of the action plan group A. The determination unit 13 then determines whether the calculated success score is equal to or greater than a preset threshold, and determines whether control of the robot by the action plan group A is successful. If the success score is equal to or greater than the threshold, the determination unit 13 determines that optimization of the action plan group A is not necessary, but if the success score is less than the threshold, the determination unit 13 determines that optimization of the action plan group A is necessary, and notifies the optimization unit 14 of this.
[0025] If the success score is less than the threshold, the determination unit 13 further determines whether each subtask (a 1 , a 2 , ...) and identify the subtasks that need to be optimized. For example, as shown in FIG. 4, the determination unit 13 calculates the success score for each subtask (a 1 , a 2 , ...) 1 , S 2 , ...) is input to the second learning model Mb, and the output (P 1 , P 2 , ...) for each subtask (a 1 , a 2 Then, the determination unit 13 calculates the success score corresponding to each subtask (a 1 , a 2 , ...), each subtask (a 1 , a 2 , ...) corresponding success scores (P 1 , P 2, . . . ) is below a threshold, is at the lowest rank, or is within the last few ranks, etc., and is identified as a subtask that needs optimization, and notifies the optimization unit 14 of this.
[0026] However, even if the success score is less than the threshold, the determination unit 13 may not specify a subtask to be optimized. In this case, the determination unit 13 may simply notify the optimization unit 14 that optimization of the action plan group A is necessary.
[0027] As described above, when the optimization unit 14 (change unit) receives a notification from the determination unit 13 that optimization is necessary, the optimization unit 14 changes the operation plan group A and performs optimization. For example, when the optimization unit 14 receives a notification that a subtask to be optimized has been specified, the optimization unit 14 changes the operation plan content of the specified subtask and performs optimization. Here, as shown in FIG. 5 , in the operation plan group A, subtask a 2 In this case, the optimization unit 14 optimizes the dotted line a in FIG. 2 From solid line a 2 As shown in ', subtask a 2 The optimization unit 14 changes the motion plan content by changing the movement path of the robot without changing the start and end states of each object or robot that is the object to be operated in subtask a. 2 The optimization unit 14 does not change the coordinates of the start position and the end position of the robot within a subtask, but changes only the movement paths of each object and the robot between them. However, the optimization unit 14 is not limited to changing the movement paths of each object and the robot in a subtask as described above, and may change the content of any movement plan. Note that the optimization unit 14 may change the content of movement plans of multiple subtasks, or may change the content of movement plans of subtasks that are not specified to be optimized. Note that the start state and end state of the subtask described above may include information representing the start state and end state of each object and the robot, or may include information representing the start state and end state of either one of the objects and the robot.
[0028] Furthermore, the optimization unit 14 is not necessarily limited to optimizing by changing the content of the motion plans of the subtasks, and may change the content of the motion plan group A by other methods. For example, when a subtask to be optimized has not been identified, the optimization unit 14 may change the order of the subtasks in the motion plan group A or add a new subtask, and accordingly, may change the content of the motion plan of each subtask so that the motion plan group A can be realized by a series of subtasks. As an example, the optimization unit 14 may change the order of operations for each object to be operated, and accordingly, may add a new subtask that becomes necessary.
[0029] [Operation] Next, the operation of the information processing device 10 described above will be described mainly with reference to the flowchart of FIG.
[0030] The information processing device 10 generates motion plan information representing a motion plan for a robot based on initial information and target information (step S1). At this time, the information processing device 10 generates motion plan information representing a motion plan for a robot based on a plurality of subtasks (a 1 , a 2 , ...) as motion plan information. As an example, the information processing device 10 generates control commands including the movement start position, end position, and movement trajectory of each object or robot in the subtask. However, the information processing device 10 may also generate motion plan information consisting of a combination of subtasks that roughly represent the control content.
[0031] Next, the information processing device 10 predicts the state when each of the subtasks, which are the plurality of operation plans, is executed based on the generated operation plan group, using the first learning model Ma (step S2). At this time, the information processing device 10 first generates an initial image I representing the initial state of the object or robot to be operated, as shown in FIG. 0 The feature extractor m extracts the initial state feature (S 0 ) and extract the initial state features (S 0 ) and the first motion plan (a 1 ) and input them into the first learning model Ma to obtain the first predicted state feature (S 1 ) is output. Then, the first predicted state feature (S 1 ) and the second motion plan (a2 ) is input to the first learning model Ma to obtain the second predicted state feature (S 2 By repeating this process, a predicted state feature set (S 1 , S 2 , .. . , S T ) to generate the
[0032] Next, the information processing device 10 calculates each predicted state feature set (S 1 , S 2 , .. . , S T ) into the second learning model Mb to calculate a success score representing the degree of success of the action plan group A (step S3). Then, the information processing device 10 determines whether the calculated success score is equal to or greater than a preset threshold (step S4) and determines whether the control of the robot by the action plan group A is successful. If the success score is equal to or greater than the threshold (Yes in step S4), the information processing device 10 continues to control the action of the robot based on the action plan group A (step S7).
[0033] On the other hand, if the success score is less than the threshold (No in step S4), the information processing device 10 optimizes the operation plan group A. At this time, the information processing device 10 optimizes each subtask (a 1 , a 2 , ...) and identify the subtasks that need to be optimized (step S5). For example, the information processing device 10 calculates the success score for each subtask (a 1 , a 2 , ...) predicted state features (S 1 , S 2 , ...) is input to the second learning model Mb, and the output (P 1 , P 2 , ...) for each subtask (a 1 , a 2 , ...), and calculate the success score (P 1 , P 2 , ...) to identify the subtasks that need to be optimized.
[0034] Then, the information processing device 10 changes the operation plan content of the identified subtask and performs optimization (step S6). 2 If optimization of 2 From solid line a 2 As shown in ', subtask a 2 The optimization unit 14 changes the content of the motion plan by, for example, changing the movement path without changing the start state and end state of the robot in the motion plan group A. Note that the information processing device 10 is not necessarily limited to optimizing by changing the content of the motion plans of the subtasks, and may change the content of the motion plan group A by other methods. For example, when a subtask to be optimized has not been identified, the optimization unit 14 may change the order of the subtasks in the motion plan group A or add a new subtask, and accordingly, may change the content of the motion plan of each subtask so that the motion plan group A can be realized by a series of subtasks.
[0035] Thereafter, the information processing device 10 changes and optimizes the operation plan contents of the identified subtasks, and then controls the operation of the robot based on the operation plan group A (step S7). Note that the information processing device 10 may calculate a success score for the operation plan group A after the operation plan contents have been changed, and repeat the above-described process until the success score exceeds a threshold value.
[0036] As described above, according to this embodiment, the states of the operation target and the robot in the group of motion plans are predicted, the degree of success is calculated, and the group of motion plans is modified based on the degree of success. This makes it possible to optimize a group of motion plans consisting of multiple subtasks. In particular, in this embodiment, subtasks are identified based on the degree of success of each subtask included in the group of motion plans, and the content of the identified subtasks is modified. As a result, the group of motion plans can be optimized more appropriately.
[0037] <Embodiment 2> Next, a second embodiment of the present disclosure will be described with reference to Fig. 7 and Fig. 8. Fig. 7 and Fig. 8 are block diagrams showing the configuration of an information processing device in embodiment 2. Note that this embodiment shows an outline of the configuration of the information processing device described in the above embodiment.
[0038] First, the hardware configuration of the information processing device 100 in this embodiment will be described with reference to Fig. 7. The information processing device 100 is configured as a general information processing device, and is equipped with the following hardware configuration, for example: CPU (Central Processing Unit) 101 (arithmetic unit); ROM (Read Only Memory) 102 (storage device); RAM (Random Access Memory) 103 (storage device); programs 104 loaded into RAM 103; a storage device 105 that stores the programs 104; a drive device 106 that reads and writes data from and to a storage medium 110 external to the information processing device; a communication interface 107 that connects to a communication network 111 external to the information processing device; an input / output interface 108 that inputs and outputs data; and a bus 109 that connects the various components.
[0039] 7 shows an example of the hardware configuration of the information processing device 100, and the hardware configuration of the information processing device is not limited to the above-described case. For example, the information processing device may be configured with only a part of the above-described configuration, such as excluding the drive device 106. Furthermore, the information processing device may use a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating Point Number Processing Unit), a PPU (Physics Processing Unit), a TPU (Tensor Processing Unit), a quantum processor, a microcontroller, or a combination thereof, instead of the above-described CPU.
[0040] The information processing device 100 can be equipped with the prediction unit 121, calculation unit 122, and change unit 123 shown in FIG. 8 by having the CPU 101 acquire and execute the program group 104. The program group 104 is stored in advance in the storage device 105 or the ROM 102, for example, and is loaded into the RAM 103 and executed by the CPU 101 as needed. The program group 104 may be supplied to the CPU 101 via the communication network 111, or may be stored in advance in the storage medium 110, and the drive device 106 may read the program and supply it to the CPU 101. However, the prediction unit 121, calculation unit 122, and change unit 123 described above may be constructed using dedicated electronic circuits for realizing such means.
[0041] The prediction unit 121 predicts a state of at least one of each operation target and the robot when each operation plan is executed, based on a group of operation plans consisting of a plurality of operation plans for the robot with respect to the operation target. For example, the prediction unit 121 predicts a state when the operation plan of the prediction target is executed by inputting a state before the operation plan of the prediction target and the operation plan of the prediction target into the first learning model.
[0042] The calculation unit 122 calculates the success rate of the group of operation plans based on each predicted state. At this time, the calculation unit 122 identifies an operation plan from the group of operation plans based on the success rate.
[0043] The change unit 123 changes the group of operation plans based on the calculated degree of success. At this time, the change unit 123 changes the content of the identified operation plan, for example.
[0044] With the above-described configuration, the present disclosure predicts the state of an operation target or a robot in a group of motion plans, calculates the degree of success, and modifies the group of motion plans based on the degree of success, thereby optimizing a group of motion plans consisting of multiple motion plans.
[0045] The above-described program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can be supplied to a computer via wired communication paths such as electric wires and optical fibers, or via wireless communication paths.
[0046] Although the present disclosure has been described above with reference to the above-described embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, at least one or more of the functions of the prediction unit 121, calculation unit 122, and change unit 123 described above may be executed by an information processing device installed and connected anywhere on a network, that is, may be executed by so-called cloud computing.
[0047] <Supplementary Notes> Some or all of the above embodiments can also be described as in the following supplementary notes. Below, an outline of the configurations of an information processing device, an information processing method, and a program according to the present disclosure will be described. However, the present disclosure is not limited to the following configurations. (Supplementary Note 1) An information processing device comprising: a prediction unit that predicts, based on a group of motion plans consisting of a plurality of motion plans for the robot with respect to a manipulation target, a state of at least one of the robot and each of the motion plans when each of the motion plans is executed; a calculation unit that calculates a degree of success of the group of motion plans based on each of the predicted states; and a modification unit that modifies the group of motion plans based on the calculated degree of success. (Supplementary Note 2) The information processing device according to Supplementary Note 1, wherein the calculation unit identifies at least one of the motion plans from the group of motion plans based on the calculated degree of success; and the modification unit modifies the group of motion plans based on the identified motion plan. (Supplementary Note 3) The information processing device according to Supplementary Note 2, wherein the modification unit modifies plan content of the identified motion plan. (Supplementary Note 4) The information processing device according to Supplementary Note 3, wherein the change unit changes the content of the action plan without changing the start state and the end state in the identified action plan. (Supplementary Note 5) The information processing device according to any of Supplements 2 to 4, wherein the calculation unit calculates a degree of success of each of the action plans in the group of action plans, and identifies the action plan based on the degree of success of each of the action plans. (Supplementary Note 6) The information processing device according to Supplementary Note 5, wherein the calculation unit calculates the degree of success of the group of action plans based on a state group including each of the states corresponding to each of the action plans and a predetermined learning model, and calculates the degree of success of each of the action plans based on each of the states corresponding to each of the action plans and the learning model. (Supplementary Note 7) The information processing device according to any of Supplements 1 to 6, wherein the prediction unit predicts each of the states when each of the action plans is executed by inputting the state before executing the action plan and the action plan to be executed into a first learning model.(Supplementary Note 8) An information processing method comprising: based on a group of motion plans consisting of a plurality of motion plans for a robot with respect to a manipulation target, predicting a state of at least one of each of the manipulation target and the robot when each of the motion plans is executed; calculating a degree of success of the group of motion plans based on each of the predicted states; and modifying the group of motion plans based on the calculated degree of success. (Supplementary Note 9) An information processing method according to Supplementary Note 8, identifying at least one of the motion plans from the group of motion plans based on the calculated degree of success, and modifying the group of motion plans based on the identified motion plan. (Supplementary Note 10) An information processing method according to Supplementary Note 9, changing plan contents of the identified motion plan. (Supplementary Note 11) An information processing method according to Supplementary Note 9 or 10, calculating a degree of success of each of the motion plans in the group of motion plans, and specifying the motion plan based on the degree of success of each of the motion plans. (Supplementary Note 12) The information processing method according to Supplementary Note 11, calculating a degree of success of the group of action plans based on a state group including each of the states corresponding to each of the action plans and a predetermined learning model, and calculating the degree of success of each of the action plans based on each of the states corresponding to each of the action plans and the learning model. (Supplementary Note 13) The information processing method according to any of Supplements 8 to 12, predicting each of the states when each of the action plans is executed by inputting the state before the action plans and the action plans to be executed into a first learning model. (Supplementary Note 14) A computer-readable storage medium storing a program for causing a computer to execute processes of predicting a state of at least one of each of the operation target and the robot when each of the operation plans is executed, based on a group of action plans consisting of a plurality of action plans of the robot for an operation target, calculating a degree of success of the group of action plans based on each of the predicted states, and changing the group of action plans based on the calculated degree of success.
[0048] REFERENCE SIGNS LIST 10 Information processing device 11 Motion planning unit 12 Prediction unit 13 Determination unit 14 Optimization unit 16 Goal storage unit 17 Motion plan storage unit 18 Learning model storage unit 100 Information processing device 101 CPU 102 ROM 103 RAM 104 Program group 105 Storage device 106 Drive device 107 Communication interface 108 Input / output interface 109 Bus 110 Storage medium 111 Communication network 121 Prediction unit 122 Calculation unit 123 Change unit
Claims
1. a prediction unit that predicts a state of at least one of each of the operation targets and the robot when each of the operation plans is executed based on a group of operation plans including a plurality of operation plans for the robot with respect to the operation targets; a calculation unit that calculates a success rate of the group of operation plans based on each of the predicted states; a change unit that changes the group of operation plans based on the calculated degree of success; An information processing device comprising:
2. 2. The information processing device according to claim 1, the calculation unit identifies at least one of the operation plans from the group of operation plans based on the calculated degree of success; the change unit changes the group of operation plans based on the identified operation plan. Information processing device.
3. 3. The information processing device according to claim 2, the change unit changes the plan content of the identified operation plan. Information processing device.
4. 4. The information processing device according to claim 3, the change unit changes the content of the operation plan without changing the start state and the end state in the identified operation plan. Information processing device.
5. 3. The information processing device according to claim 2, the calculation unit calculates a degree of success of each of the operation plans in the group of operation plans, and identifies the operation plan based on the degree of success of each of the operation plans. Information processing device.
6. 6. The information processing device according to claim 5, the calculation unit calculates a success degree of the group of operation plans based on a state group including each of the states corresponding to each of the operation plans and a predetermined learning model, and calculates a success degree of each of the operation plans based on each of the states corresponding to each of the operation plans and the learning model. Information processing device.
7. 2. The information processing device according to claim 1, the prediction unit predicts each of the states when each of the operation plans is executed by inputting the state before the operation plan is executed and the operation plan to be executed into a first learning model; Information processing device.
8. predicting a state of at least one of each of the operation targets and the robot when each of the operation plans is executed based on a group of operation plans including a plurality of operation plans for the robot with respect to the operation targets; Calculating a success rate of the group of motion plans based on each of the predicted states; modifying the group of operation plans based on the calculated degree of success; Information processing methods.
9. 9. The information processing method according to claim 8, identifying at least one of the operation plans from the group of operation plans based on the calculated degree of success; modifying the group of action plans based on the identified action plan; Information processing methods.
10. predicting a state of at least one of each of the operation targets and the robot when each of the operation plans is executed based on a group of operation plans including a plurality of operation plans for the robot with respect to the operation targets; Calculating a success rate of the group of motion plans based on each of the predicted states; modifying the group of operation plans based on the calculated degree of success; A program that causes a computer to execute a process.