Information processing device, information processing method, and program

By optimizing point cloud extraction and grasp scoring for reachable areas, the method addresses the computational inefficiencies of inverse kinematics, enabling accurate and cost-effective object grasping in resource-constrained robots.

WO2025177835A1PCT designated stage Publication Date: 2025-08-28SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/003708
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-02-05
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing technologies for object grasping by robots, such as inverse kinematics, are computationally expensive and time-consuming, particularly in resource-constrained environments like entertainment robots, and often require costly tactile sensors with changing characteristics, limiting their ability to grasp a variety of objects effectively.

Method used

The method involves extracting point cloud information of a pre-calculated reachable area using an optimized extraction shape, setting a grasp score to determine the probability of grasping, and selecting a motion based on this score, thereby reducing computational costs and improving accuracy without the need for tactile sensors.

Benefits of technology

This approach reduces computational costs and enhances the accuracy of object grasping by efficiently extracting point cloud information, allowing robots to grasp objects with consideration for their reach and movement, even in environments with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025003708_28082025_PF_FP_ABST
    Figure JP2025003708_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing device, an information processing method, and a program that make it possible to reduce the calculation cost required for object gripping. From point group information acquired by a depth sensor of a mobile robot, a point group cut-out unit cuts out, to obtain a cut-out shape optimized by learning, point group information of a reachable region of a grip part of the mobile robot, the region having been calculated in advance. A point group determination unit sets a grip score indicating the possibility of gripping an object by the grip part with respect to the cut-out point group information. A motion selection unit selects, on the basis of the grip score, a motion for the grip part to reach the object, the motion being associated with the reachable region. The present disclosure can be applied to, for example, a four-leg dog type entertainment robot.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program that enable a reduction in the calculation cost required for grasping an object.

[0002] Many technologies related to object grasping by a robot have been proposed. For example, Patent Literature 1 discloses an object grasping system that moves a gripper toward a target object while repeatedly identifying the relative position of the target object with respect to the gripper based on images captured by a camera.

[0003] There is also known a technique for making a gripper reach a target object by observing point cloud information of the object, extracting its surroundings, and solving inverse kinematics.

[0004] Japanese Patent Application Laid-Open No. 2020-121352

[0005] However, solving the inverse kinematics is computationally expensive and takes a long time to process.

[0006] The present disclosure has been made in consideration of such circumstances, and aims to realize a reduction in the calculation cost required for object grasping.

[0007] The information processing device disclosed herein is an information processing device that includes a point cloud extraction unit that extracts point cloud information of a pre-calculated reachable area of ​​a gripping unit of the mobile robot from point cloud information acquired by the mobile robot's depth sensor using an extraction shape optimized by learning, a point cloud determination unit that sets a grasp score representing the probability that the object can be grasped by the gripping unit for the extracted point cloud information, and a motion selection unit that selects a motion for the gripping unit to reach the object, associated with the reachable area, based on the grasp score.

[0008] The information processing method disclosed herein includes: cutting out point cloud information of a pre-calculated reachable area of ​​a gripping unit of a mobile robot from point cloud information acquired by a depth sensor of the mobile robot using a cut-out shape optimized by learning; setting a grasping score for the cut-out point cloud information that represents the probability that the gripping unit can grasp an object; and selecting a motion for the gripping unit to reach the object, associated with the reachable area, based on the grasping score.

[0009] The program disclosed herein causes a computer to execute a process of cutting out point cloud information of a pre-calculated area that can be reached by the gripping unit of a mobile robot from point cloud information acquired by the depth sensor of the mobile robot using a cutting shape optimized through learning, setting a grasping score for the cut-out point cloud information that represents the probability that the gripping unit can grasp an object, and selecting a motion for the gripping unit to reach the object that is associated with the reachable area based on the grasping score.

[0010] In the present disclosure, point cloud information of the reachable area of ​​the gripping unit of the mobile robot, which has been calculated in advance, is extracted from point cloud information acquired by the depth sensor of the mobile robot using an extraction shape optimized by learning, a grasping score representing the probability that the object can be grasped by the gripping unit is set for the extracted point cloud information, and a motion for the gripping unit to reach the object, associated with the reachable area, is selected based on the grasping score.

[0011] FIG. 1 is a diagram illustrating an example hardware configuration of a robot according to an embodiment of the present disclosure. FIG. 2 is a block diagram illustrating an example functional configuration of a controller. FIG. 3 is a flowchart illustrating a process for determining whether a grasp is possible using object information. FIG. 4 is a diagram illustrating object detection. FIG. 5 is a diagram illustrating object search. FIG. 6 is a diagram illustrating input and output of a measurement part estimation module. FIG. 7 is a diagram illustrating learning of the measurement part estimation module. FIG. 8 is a diagram illustrating an example of an application display screen on a user terminal. FIG. 9 is a flowchart illustrating a process for determining whether a grasp is possible using point cloud information. FIG. 9 is a diagram illustrating input and output of a graspability estimation module. FIG. 10 is a diagram illustrating clipping of point cloud information at an arrival position and posture. FIG. 11 is a diagram illustrating determination of a clipping shape of point cloud information. FIG. 12 is a block diagram illustrating an example hardware configuration of a computer.

[0012] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described below in the following order.

[0013] 1. Prior art and its problems 2. Configuration of a robot to which the technology disclosed herein is applied 3. Graspability determination process using object information 4. Graspability determination process using point cloud information 5. Example of computer hardware configuration

[0014] 1. Prior Art and Issues There is a known technique for grasping an object by a robot, in which a grasping part is made to reach the target object by observing point cloud information of the object, extracting the periphery of the object, and solving inverse kinematics.

[0015] However, solving the inverse kinematics is computationally expensive and takes a long time to process.

[0016] For example, in the case of a four-legged dog-shaped entertainment robot, it is difficult to equip it with sufficient computing resources due to cost considerations, and large-scale calculations take a long time. While it is possible to perform calculations on a server using communications, this approach poses security and other issues. Furthermore, tactile sensors used to identify objects to be grasped are generally expensive, and their characteristics often change during use, making them difficult to incorporate into products. Furthermore, entertainment robots place a premium on appearance and design, making it difficult to adopt a design that prioritizes functionality alone. This inevitably limits the reach of the grasping mouth. Given these constraints, no technology has been proposed for grasping a variety of objects suitable for such entertainment robots.

[0017] In contrast, the technology disclosed herein aims to reduce the computational costs required to grasp an object and improve the accuracy of determining whether or not it can be grasped by efficiently extracting point cloud information while maintaining the necessary information using a pre-calculated reachable range of the grasping part and an optimized extraction shape.

[0018] 2. Configuration of a Robot to which the Technology According to the Present Disclosure is Applied (Example of Hardware Configuration of Robot) FIG. 1 is a diagram illustrating an example of a hardware configuration of a robot according to an embodiment of the present disclosure.

[0019] 1 is configured as, for example, a four-legged dog-type entertainment robot, but is not limited to this, and the robot 1 may be any mobile robot that is equipped with a manipulator and a gripping unit and is configured to be movable.

[0020] The robot 1 is configured to include an RGB camera 10 , a depth sensor 20 , a controller 30 , an actuator 40 , and a communication module 50 .

[0021] The RGB camera 10 is one of the sensors provided in the robot 1, and is configured as an imaging device equipped with an imaging element such as a CMOS (Complementary Metal Oxide Semiconductor) image sensor. The RGB camera 10 captures images of the environment surrounding the robot 1 and supplies the resulting RGB images (color images) to the controller 30. If the robot 1 is configured as a four-legged dog-type entertainment robot, the RGB camera 10 is provided, for example, at a location corresponding to the nose. In this case, by providing the RGB camera 10 with a fisheye lens, it is possible to obtain RGB images capturing a wide range centered on the front of the robot 1.

[0022] The depth sensor 20 is one of the sensors included in the robot 1, and is configured as, for example, a 3D Time of Flight (ToF) sensor that measures the distance to a wide range of objects by measuring the time of flight of light, or other distance measuring sensors. The depth sensor 20 measures the distance to the surrounding environment of the robot 1, and supplies the obtained distance information to the controller 30 as depth information (depth image). When the robot 1 is configured as a four-legged dog-type entertainment robot, the depth sensor 20 is provided, for example, in a location corresponding to the mouth or chest.

[0023] The controller 30 has a built-in CPU (Central Processing Unit), memory, etc., and performs various processes by the CPU executing programs stored in the memory. Specifically, the controller 30 generates control signals for controlling the actuator 40 based on the RGB image from the RGB camera 10, the depth information from the depth sensor 20, and various information supplied from outside the robot 1 via the communication module 50. Note that the various processes performed by the controller 30 are not limited to being performed by the CPU, but can also be performed by a GPU (Graphics Processing Unit) or a DSP (Digital Signal Processor).

[0024] The actuator 40 is configured with a motor or the like, and functions as a joint of the manipulator 41 and the gripping unit 42. In other words, when the actuator 40 is driven by a control signal from the controller 30, various motions of the robot 1 are realized by the manipulator 41 and the gripping unit 42. When the robot 1 is configured as a four-legged dog-type entertainment robot, the manipulator 41 is configured as legs and a head connected to the body, and the gripping unit 42 is configured as a mouth that can grip an object.

[0025] The communication module 50 is composed of electronic components including a wireless chip and peripheral circuits that communicate directly or via a network with external devices such as the server 100 and the user terminal 200. The communication module 50 supplies information received via wireless communication to the controller 30, and transmits information supplied from the controller 30 via wireless communication.

[0026] The server 100 is configured as a cloud server or the like built on a network such as the Internet. The server 100 stores various information used for motion control of the robot 1, and the information is referenced and updated by the robot 1 and the user terminal 200 as appropriate.

[0027] The user terminal 200 is configured as a smartphone, tablet terminal, or the like that is operated by the user who owns the robot 1. The user terminal 200 can acquire the status of the robot 1 and present it to the user by communicating with the robot 1 based on the control of, for example, an installed dedicated application.

[0028] (Example of Functional Configuration of Controller) FIG. 2 is a block diagram showing an example of the functional configuration of the controller 30. As shown in FIG.

[0029] Each functional block of the controller 30 shown in FIG. 2 is realized by the CPU executing a program stored in the memory.

[0030] The controller 30 is configured to include a graspability determination unit 310 , a graspability determination unit 320 , and a motion control unit 330 .

[0031] The graspability determining unit 310 executes a graspability determining process that determines whether an object can be grasped by the gripping unit 42 based on object information that indicates the characteristics of the object that can be grasped.

[0032] When the graspability determination unit 310 determines that the object can be grasped, the graspability determination unit 320 executes a graspability determination process to determine whether the object has a part that can be grasped by the grasping unit 42 using point cloud information converted from the depth information acquired by the depth sensor 20.

[0033] The motion control unit 330 generates a control signal for causing the robot 1 to execute a motion according to the determination results of the graspability determination unit 310 and the graspability determination unit 320 .

[0034] The graspability determining unit 310 is composed of an object detecting unit 311 , an object searching unit 312 , a hardness measuring unit 313 , and an object determining unit 314 .

[0035] The object detection unit 311 detects objects that can be grasped based on sensor information obtained by sensors mounted on the robot 1, i.e., at least one of the RGB image from the RGB camera 10 and the depth information from the depth sensor 20.

[0036] The object search unit 312 searches for the object detected by the object detection unit 311 from among objects registered in an object database (DB) 340. The object DB 340 may be configured within the robot 1 or may be configured within the server 100. In the object DB 340, images of registered objects are associated with object information indicating the characteristics of the objects.

[0037] If the object detected by the object detection unit 311 is not registered in the object DB 340, the hardness measurement unit 313 calculates a hardness score representing the hardness of the object as object information of the object.

[0038] The object search unit 312 associates the object information including the hardness score calculated in this manner with the object detected by the object detection unit 311, and registers the associated information in the object DB 340. In other words, the object search unit 312 also functions as an object registration unit that associates the object information with the object, and registers the associated information in the object DB 340.

[0039] The object determination unit 314 determines whether or not an object is a graspable object based on the object information of the object that has been searched for by the object search unit 312 or that has been newly registered in the object DB 340 .

[0040] The graspability determining unit 320 is configured to include a point cloud cutting unit 321 , a point cloud determining unit 322 , and a motion selecting unit 323 .

[0041] The point cloud cutout unit 321 cuts out point cloud information of the reachable area of ​​the gripping unit 42 of the robot 1, which has been calculated in advance, from point cloud information converted from the depth information acquired by the depth sensor 20, in a cutout shape optimized by learning.

[0042] The point cloud determination unit 322 determines whether the cut-out point cloud information is a graspable point cloud pattern by setting a grasp score representing the probability that the object can be grasped by the grasping unit 42 for the point cloud information cut-out by the point cloud cut-out unit 321.

[0043] The motion selection unit 323 selects a motion for the grip unit 42 to reach the object, which is associated with the reachable area, based on the grip score set by the point cloud determination unit 322. Information representing the selected motion is supplied to the motion control unit 330, which generates a control signal for causing the robot 1 to execute the selected motion.

[0044] Note that some of the functional blocks implemented in the controller 30 (for example, the graspability determination unit 310 and the graspability determination unit 320) may be implemented on a cloud such as the server 100.

[0045] The grippability determination process executed by each of the grippability determination unit 310 and the grippability determination unit 320 will be described in detail below.

[0046] 3. Graspability Determination Process Based on Object Information The graspability determination process based on object information, which is executed by the graspability determination unit 310, will be described with reference to the flowchart in Fig. 3. The process in Fig. 3 starts when, for example, a user issues a command to start grasping.

[0047] In step S11, when the robot 1 is in a shooting posture under the control of the motion control unit 330, the object detection unit 311 acquires sensor information, i.e., an RGB image from the RGB camera 10 and depth information from the depth sensor 20.

[0048] In step S12, the object detection unit 311 detects an object in the sensor information (RGB image or depth information). For example, as shown in Fig. 4, when an image IMG including an object OB is obtained, the image IMG is input to an object detection model M1 constructed by machine learning, and a rectangular image RT of the object OB is extracted.

[0049] In step S13, the object search unit 312 determines whether the object detected by the object detection unit 311 is an object registered in the object DB 340. If it is determined that the object detected by the object detection unit 311 is an object registered in the object DB 340, the process proceeds to step S14.

[0050] In step S14, the object determination unit 314 determines whether the object is a graspable object based on the object information of the object registered in the object DB 340. For example, as shown in FIG. 5 , when an object OB extracted as a rectangular image RT is registered in the object DB 340, object information INF associated with the object OB is acquired. The object information INF includes, in addition to the hardness score of the object OB, additional information indicating the past grasping success rate and whether the object is suitable for being held (grasped). The additional information may be input to the user terminal 200 operated by the user.

[0051] If it is determined that the object is a graspable object, the graspability determination process based on the object information ends, and the process proceeds to graspability determination process based on the point cloud information.

[0052] If it is determined that the object detected by the object detection unit 311 is not an object registered in the object DB 340, the process proceeds to step S15.

[0053] In step S15, the stiffness measurement unit 313 measures the stiffness of the object. Specifically, the stiffness measurement unit 313 calculates a stiffness score based on a change in the joint angle of the manipulator 41 when the manipulator 41 of the robot 1 presses a portion of the object detected by the object detection unit 311 as a measurement site. At this time, the stiffness measurement unit 313 determines the position of the measurement site and the pressing direction based on the rolling characteristics that indicate the rolling state of the object.

[0054] When calculating the hardness score, it is required that the object does not roll when the measurement site is pressed. The position of the measurement site and the pressing direction that will prevent the object from rolling can be determined using a measurement site estimation module M10, which inputs point cloud information PC including the object and the position and direction in which the object is pressed, and outputs the rolling state of the object, as shown in Figure 6. In other words, by using the measurement site estimation module M10, it is possible to infer from the point cloud information PC which position and which direction should be pressed to prevent the object from rolling.

[0055] The measurement site estimation module M10 can be obtained by using a simulation environment and machine learning. Specifically, as shown in Figure 7, the rolling behavior of objects of various shapes is calculated using physical simulation, and a data set DS11 is created that pairs the depth information (point cloud information) of the object with the position and direction to press to make it difficult to roll. The measurement site estimation module M10 can be obtained by performing supervised learning using such a data set DS11.

[0056] The hardness score is calculated from the difference between the ideal joint angle and the actual joint angle of the manipulator 41 when the object is pushed by the manipulator 41 in the position and posture of the robot 1 calculated based on the position of the measurement site and the pushing direction determined as described above. The harder the object, the less its surface will be depressed when pushed, and therefore the greater the difference between the ideal joint angle and the actual joint angle.

[0057] Returning to the flowchart of Figure 3, in step S16, the object search unit 312 associates the hardness score calculated by the hardness measurement unit 313 with the object (rectangular image) detected by the object detection unit 311 as object information and registers them in the object DB 340.

[0058] Then, in step S17, the object determination unit 314 determines whether or not the object is a graspable object based on the object information of the object newly registered in the object DB 340. If it is determined that the object is a graspable object, the process proceeds to step S18.

[0059] In step S19, the motion control unit 330 operates the robot 1 to again assume the photographing posture from the position and posture in which the object is pushed by the manipulator 41. This ends the graspability determination process based on the object information, and the process proceeds to graspability determination process based on the point cloud information.

[0060] On the other hand, if it is determined in step S14 or step S17 that the object cannot be grasped, the process proceeds to step S19. In step S19, the motion control unit 330 causes the robot 1 to end the operation for grasping the object.

[0061] According to the above process, by calculating, storing, and using the hardness score of an object, it is possible to grasp an object while taking into account the hardness of the object, without installing a tactile sensor.

[0062] While the robot 1 is measuring the hardness of an object by pressing it, the state of the robot 1 may be presented to the user by a dedicated application installed on the user terminal 200.

[0063] Specifically, while the hardness of the object is being measured, the display screen (display unit) of the user terminal 200 displays the object (rectangular image RT), the position of the measurement site on the object, and the pressing direction, as shown in the left diagram of Fig. 8. After the hardness measurement of the object is completed, the object (rectangular image RT) and the measured hardness score stored in the object DB 340 are displayed, as shown in the right diagram of Fig. 8. At this time, the user may input additional information indicating whether the object is an object that the user may add to their collection, whether to add the object to their favorites, etc. This additional information, along with the hardness score, is associated with the object as object information and stored in the object DB 340.

[0064] By adding such additional information as object information, the four-legged dog-type robot can approach an object and pick it up, or conversely, avoid an object, making the robot appear to the user as if it has emotions and is growing. For example, when the four-legged dog-type robot is playing a prank, it can be made to behave in such a way that it picks up an object that the user does not want it to pick up.

[0065] 9, the graspability determination process based on point cloud information, which is executed by the graspability determination unit 320, will be described. The process in Fig. 9 starts after it is determined that the detected object is a graspable object in the graspability determination process based on object information, which has been described with reference to the flowchart in Fig. 3.

[0066] In step S21, the point cloud cutout unit 321 converts the depth information from the depth sensor 20 into point cloud information.

[0067] In step S22, the point cloud cutout unit 321 cuts out point cloud information of the reachable area of ​​the gripper 42 from the point cloud information converted from the depth information, in a cutout shape optimized by learning. The reachable area here is defined as one or more areas (first areas) that the gripper 42 can reach by changing the posture of the robot 1 without changing the position of the robot 1.

[0068] In step S23, the point cloud determination unit 322 determines whether the cut-out point cloud information is a graspable point cloud pattern by setting a grasp score representing the probability that the object can be grasped by the grasper 42 for the point cloud information cut-out by the point cloud cut-out unit 321.

[0069] 10, whether or not a point cloud pattern is graspable can be determined using a graspability estimation module M20 that receives point cloud information PC1, PC2 extracted from the entire point cloud information PC as input and outputs the probability that the object OB can be grasped by the robot 1 (gripping unit 42). The graspability estimation module M20 can be obtained by using an actual robot to collect a data set that pairs the extracted point cloud information with whether or not the object OB can be grasped (graspability probability), and performing supervised learning using such a data set.

[0070] In this way, the point cloud determination unit 322 determines whether the extracted point cloud information is a graspable point cloud pattern by setting a grasping score for the extracted point cloud information using an inference model that has learned the point cloud information that the grasping unit 42 can grasp.

[0071] If the extracted point cloud information is determined to be a graspable point cloud pattern, the process proceeds to step S24, where the motion selection unit 323 selects a gripping motion corresponding to the point cloud pattern. That is, if a gripping score capable of gripping an object is set for point cloud information from which a reachable region (first region) is extracted, the motion selection unit 323 selects a gripping motion that causes the gripping unit 42 to grip the object. In other words, if the gripping score set for the extracted point cloud information is equal to or greater than a predetermined threshold, the motion selection unit 323 selects a gripping motion associated with the point cloud information (reachable region). If there are multiple gripping scores equal to or greater than the threshold, the gripping motion associated with the point cloud information (reachable region) with the highest gripping score is selected. The motion control unit 330 then causes the robot 1 to execute the selected gripping motion. That is, if the robot 1 is configured as a four-legged dog-type entertainment robot, the robot 1 holds the object in its mouth.

[0072] On the other hand, if it is determined in step S23 that all of the extracted point cloud information is not a graspable point cloud pattern, the process proceeds to step S26.

[0073] In step S26, the point cloud cutout unit 321 cuts out point cloud information of a wide range of reachable areas from the point cloud information obtained by converting the depth information, in a cutout shape optimized by learning. The reachable areas here are one or more areas (second areas) that can be reached by the gripper 42 by changing the position and posture of the robot 1.

[0074] In step S26, the point cloud determination unit 322 determines whether the cut-out point cloud information is a graspable point cloud pattern by setting a grasp score representing the probability that the object can be grasped by the grasper 42 for the point cloud information cut-out by the point cloud cut-out unit 321.

[0075] If the extracted point cloud information is determined to be a graspable point cloud pattern, the process proceeds to step S27, where the motion selection unit 323 selects a movement motion corresponding to the point cloud pattern. That is, if a grasp score enabling grasping of an object is set for point cloud information from which a reachable region (second region) is extracted, the motion selection unit 323 selects a movement motion for moving the robot 1. In other words, if the grasp score set for the extracted point cloud information is equal to or greater than a predetermined threshold, the motion selection unit 323 selects a movement motion associated with the point cloud information (reachable region). If there are multiple grasp scores equal to or greater than the threshold, the movement motion associated with the point cloud information (reachable region) with the highest grasp score is selected. The motion control unit 330 then causes the robot 1 to execute the selected movement motion. That is, if the robot 1 is configured as a four-legged dog-type entertainment robot, the robot 1 walks to approach the object.

[0076] In step S28, the motion control unit 330 operates the robot 1 to take a photographing posture at a position close to the object. Then, the process returns to step S21, and the subsequent processes are repeated.

[0077] On the other hand, if it is determined in step S26 that the extracted point cloud information does not all constitute a graspable point cloud pattern, the process proceeds to step S29. In step S29, the motion control unit 330 causes the robot 1 to end its operation for grasping the object.

[0078] In this way, by using two-stage prediction of whether or not a robot can grasp an object, even if the robot has a narrow reach of the grasping part, it can achieve grasping that takes into account the movement of the base (robot body), making it possible to grasp objects placed over a wide range.

[0079] In cutting out point cloud information, there are a first point that determines the position and orientation at which the point cloud information is cut out, and a second point that determines the shape at which the point cloud information is cut out.

[0080] Regarding the first point (what position and orientation to extract), point cloud information can be extracted using the arrival position and orientation of the gripper 42 for each motion calculated in advance. As a result, as shown in the left diagram of Fig. 11 , when the robot 1 grasps a nearby object OB, point cloud information PC11 and PC12 can be extracted using the overall point cloud information PC and the arrival position and orientation of the gripper 42 for each motion (grasp motion) as input. Also, as shown in the right diagram of Fig. 11 , when the robot 1 grasps a distant object OB, point cloud information PC21 and PC22 can be extracted using the overall point cloud information PC and the arrival position and orientation of the gripper 42 for each motion (movement motion and grip motion) as input.

[0081] In this way, by calculating candidates for the arrival position and orientation of the gripping part in advance, it becomes possible to extract point cloud information without the need to solve inverse kinematics, which requires high computational costs.

[0082] Regarding the second point (the cut-out shape to be used for the cut-out), the point cloud information can be cut out using a cut-out shape determined based on the output of the graspability estimation module M20 (inference model), which includes internal feature values ​​of the graspability estimation module M20. The cut-out shape is determined repeatedly during supervised learning by the graspability estimation module M20. Specifically, as shown in FIG. 12 , supervised learning is performed using a dataset DS21 in which the cut-out point cloud information and graspability (graspability probability) are paired, with the length, width, height, and other parameters representing the shape of the grip portion 42 used as cut-out parameters for learning. At this time, a label indicating graspability is obtained for each of the point cloud information PC31 and PC32 cut out from the entire point cloud information PC. Then, the cut-out parameters (cut-out shape) are optimized so that the feature value Z of the graspable point cloud information PC31 and the feature value Z′ of the ungraspable point cloud information PC32 are separated from each other. Alternatively, as a different method, the cutout parameters (cutout shape) may be optimized so that the graspability estimation module M20 outputs a high grasp score for the point cloud information PC31 and a low grasp score for the point cloud information PC32. Note that the cutout shape is determined during prior learning and is not updated during inference.

[0083] In this way, by optimizing the cutout shape, it becomes possible to efficiently cut out point cloud information while maintaining necessary information.

[0084] According to the above processing, by using the reachable range of the grasping part calculated in advance and the optimized extraction shape, it is possible to efficiently extract point cloud information while maintaining the necessary information, thereby reducing the computational cost required for grasping an object and improving the accuracy of determining whether or not it can be grasped.

[0085] 5. Example of Computer Hardware Configuration The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware, a general-purpose personal computer, or the like.

[0086] 13 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program. The controller 30 of the robot 1, the server 100, and the user terminal 200 may be configured, for example, by a computer 500 having a configuration similar to that shown in FIG.

[0087] A CPU 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .

[0088] An input / output interface 505 is also connected to the bus 504. An input unit 506 including a keyboard, a mouse, etc., and an output unit 507 including a display, a speaker, etc. are connected to the input / output interface 505. Also connected to the input / output interface 505 are a storage unit 508 including a hard disk, a nonvolatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 that drives removable media 511.

[0089] In the computer 500 configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program stored in the memory unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.

[0090] The program executed by the CPU 501 is installed in the storage unit 508 by being recorded on, for example, removable media 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.

[0091] The program executed by computer 500 may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0092] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0093] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0094] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.

[0095] For example, the embodiment of the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.

[0096] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0097] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0098] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0099] Furthermore, the technology disclosed herein may have the following configuration: (1) An information processing device comprising: a point cloud cutout unit that cuts out point cloud information of a reachable area of ​​a gripper of the mobile robot, calculated in advance, from point cloud information acquired by a depth sensor of the mobile robot, using a cutout shape optimized by learning; a point cloud determination unit that sets a grasp score representing a probability that an object can be grasped by the gripper, for the cutout point cloud information; and a motion selection unit that selects a motion associated with the reachable area, for the gripper to reach the object, based on the grasp score. (2) The information processing device described in (1), wherein the reachable area includes a first area that the gripper can reach by changing the posture of the mobile robot without changing its position. (3) The information processing device described in (2), wherein the reachable area further includes a second area that the gripper can reach by changing the position and posture of the mobile robot. (4) The information processing device according to any one of (1) to (6), wherein the motion selection unit selects a gripping motion that causes the gripping unit to grip the object when the grip score that allows the object to be gripped is set for the point cloud information from which the first region is clipped. (5) The information processing device according to (4), wherein the point cloud clipping unit clips the point cloud information of the second region when the grip score that allows the object to be gripped is not set for the point cloud information from which the first region is clipped. (6) The information processing device according to (5), wherein the motion selection unit selects a movement motion that causes the mobile robot to move when the grip score that allows the object to be gripped is set for the point cloud information from which the second region is clipped. (7) The information processing device according to any one of (1) to (6), wherein the point cloud determination unit sets the grip score for the clipped point cloud information using an inference model that has learned the point cloud information that can be gripped by the gripping unit. (8) The information processing device according to (7), wherein the clipping shape is determined based on the inference model.(9) The information processing device according to any one of (1) to (8), further comprising: an object detection unit that detects the object based on sensor information obtained by a sensor mounted on the mobile robot; and an object determination unit that determines whether the object can be grasped based on object information representing characteristics of the detected object. (10) The information processing device according to (9), further comprising: a stiffness measurement unit that calculates a stiffness score of the object as the object information. (11) The information processing device according to (10), wherein the stiffness measurement unit calculates the stiffness score based on a change in a joint angle of a manipulator possessed by the mobile robot when the manipulator presses a portion of the object as a measurement portion. (12) The information processing device according to (11), wherein the stiffness measurement unit determines the position of the measurement portion and the pushing direction based on rolling characteristics that represent a rolling manner of the object. (13) The information processing device according to any one of (10) to (12), further comprising an object registration unit that, if the detected object is not registered in a database, associates the object information including the hardness score with the object and registers it in the database. (14) The information processing device according to (13), wherein the object information further includes a success rate of grasping the object by the gripping unit, or additional information input to a user terminal operated by a user. (15) The information processing device according to (14), wherein the object information registered in the database is displayed on a display unit of the user terminal. (16) The information processing device according to any one of (9) to (15), wherein the sensor information is at least one of an RGB image and depth information. (17) An information processing method including: extracting point cloud information of a reachable area of ​​a gripping unit of a mobile robot, which has been calculated in advance, from point cloud information acquired by a depth sensor of the mobile robot using an extraction shape optimized by learning; setting a grasping score for the extracted point cloud information, which represents the probability that the gripping unit can grasp an object; and selecting a motion for the gripping unit to reach the object, which is associated with the reachable area, based on the grasping score.(18) A program for causing a computer to execute the following process: cutting out point cloud information of a reachable area of ​​a gripping unit of a mobile robot, which has been calculated in advance, from point cloud information acquired by a depth sensor of the mobile robot, using a cutout shape optimized by learning; setting a gripping score for the cutout point cloud information, which indicates the probability that the gripping unit can grasp an object; and selecting a motion for the gripping unit to reach the object, which is associated with the reachable area, based on the gripping score.

[0100] REFERENCE SIGNS LIST 1 Robot, 10 RGB camera, 20 Depth sensor, 40 Actuator, 41 Manipulator, 42 Grasping unit, 50 Communication module, 100 Server, 200 User terminal, 310 Grasping feasibility determination unit, 311 Object detection unit, 312 Object search unit, 313 Hardness measurement unit, 314 Object determination unit, 320 Grasping feasibility determination unit, 321 Point cloud extraction unit, 322 Point cloud determination unit, 323 Motion selection unit, 330 Motion control unit, 340 Object DB

Claims

1. An information processing device comprising: a point cloud extraction unit that extracts point cloud information of a pre-calculated reachable area of ​​a gripping unit of a mobile robot from point cloud information acquired by a depth sensor of the mobile robot using an extraction shape optimized by learning; a point cloud determination unit that sets a grasp score representing the probability that an object can be grasped by the gripping unit for the extracted point cloud information; and a motion selection unit that selects a motion for the gripping unit to reach the object, associated with the reachable area, based on the grasp score.

2. The information processing device according to claim 1, wherein the reachable area includes a first area that can be reached by the gripper by changing the posture of the mobile robot without changing the position of the mobile robot.

3. The information processing device according to claim 2, wherein the reachable area further includes a second area that can be reached by the gripper by changing the position and posture of the mobile robot.

4. The information processing device of claim 3, wherein the motion selection unit selects a grasping motion that causes the grasping unit to grasp the object when a grasping score that allows the object to be grasped is set for the point cloud information from which the first region is cut out.

5. The information processing device according to claim 4, wherein the point cloud cutout unit cuts out the point cloud information of the second region when the grasp score that allows the object to be grasped is not set for the point cloud information from which the first region is cut out.

6. The information processing device according to claim 5, wherein the motion selection unit selects a movement motion for moving the mobile robot when the grasping score capable of grasping the object is set for the point cloud information from which the second region is cut out.

7. The information processing device according to claim 1, wherein the point cloud determination unit sets the grasping score for the extracted point cloud information using an inference model that has learned the point cloud information that can be grasped by the grasping unit.

8. The information processing device according to claim 7, wherein the cutout shape is determined based on the inference model.

9. The information processing device according to claim 1, further comprising: an object detection unit that detects the object based on sensor information obtained by a sensor mounted on the mobile robot; and an object determination unit that determines whether the object can be grasped based on object information that represents the characteristics of the detected object.

10. The information processing device according to claim 9, further comprising a hardness measurement unit that calculates a hardness score of the object as the object information.

11. The information processing device according to claim 10, wherein the stiffness measurement unit calculates the stiffness score based on a change in the joint angle of the manipulator of the mobile robot when the manipulator presses a part of the object as a measurement site.

12. The information processing device according to claim 11, wherein the hardness measurement unit determines the position of the measurement site and the pressing direction based on rolling characteristics that indicate the rolling condition of the object.

13. The information processing device according to claim 10, further comprising an object registration unit that, if the detected object is not registered in the database, associates the object information including the hardness score with the object and registers the object in the database.

14. The information processing device according to claim 13, wherein the object information further includes a success rate of grasping the object by the grasping unit, or additional information input to a user terminal operated by a user.

15. The information processing device according to claim 14, wherein the object information registered in the database is displayed on a display unit of the user terminal.

16. The information processing device according to claim 9, wherein the sensor information is at least one of an RGB image and depth information.

17. An information processing method comprising: extracting point cloud information of a reachable area of ​​a gripping unit of a mobile robot, which has been calculated in advance, from point cloud information acquired by a depth sensor of the mobile robot using an extraction shape optimized by learning; setting a grasping score for the extracted point cloud information, which represents the probability that the gripping unit can grasp an object; and selecting a motion for the gripping unit to reach the object, which is associated with the reachable area, based on the grasping score.

18. A program for causing a computer to execute the following process: from point cloud information acquired by a depth sensor of a mobile robot, cut out point cloud information of a reachable area of ​​a gripping unit of the mobile robot, which has been calculated in advance, using a cut-out shape optimized by learning; for the cut-out point cloud information, set a grasp score indicating the probability that the gripping unit can grasp an object; and based on the grasp score, select a motion for the gripping unit to reach the object, which is associated with the reachable area.

Citation Information

Patent Citations

  • Learning device, learning method, learning model, detection device and gripping system

    JP2018205929A

  • Remote-controlled device, remote control system, remote control support method, program and non-temporary computer readable medium

    JP2021070140A

  • Information processing device and program

    WO2023203896A1