Handling device and program

The handling device uses a force and imaging sensor combination with machine learning to rapidly and accurately identify the position of irregularly shaped objects, addressing the challenges of incomplete data from single-sensor systems.

WO2026009706A1PCT designated stage Publication Date: 2026-01-08SUMITOMO HEAVY IND LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-18
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing handling devices struggle with identifying the position and condition of irregularly shaped objects due to reliance on incomplete or inaccurate information from force sensors or image sensors, leading to prolonged identification times, especially in challenging environments or with objects like translucent or glossy materials.

Method used

A handling device that combines a force sensor and an imaging sensor to identify the position of an object by integrating contact information with image data, utilizing machine learning to enhance estimation accuracy through a force sense state estimator and image state estimator, enabling rapid handling of irregularly shaped objects.

Benefits of technology

The combined use of force and image data allows for quicker and more accurate identification of object positions, even in environments where image data is incomplete or difficult to obtain, facilitating efficient handling of irregularly shaped objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025021914_08012026_PF_FP_ABST
    Figure JP2025021914_08012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a handling device and a program that make it possible to swiftly handle an object, even when the object has an appearance which is difficult to detect satisfactorily and also has an indefinite shape. The handling device identifies the position of an object on the basis of an image of the object and information pertaining to contact with the object.
Need to check novelty before this filing date? Find Prior Art

Description

Handling device and program

[0001] The present invention relates to a handling device and a program.

[0002] Patent Document 1 describes a system that detects the state of an object using a sensor incorporated in a robot hand.

[0003] JP 2010-280054 A

[0004] When handling irregularly shaped objects with a handling device, contact information from a force sensor alone is needed to estimate the object's position and condition, as the shape is estimated completely by feel with no prior information. Using an image sensor, however, can also result in the same long time required for estimation, even if incomplete information is gathered haphazardly, such as when the object is exposed to excessive light and dark or has a transparent exterior. Handling of the object cannot proceed until the above identification is complete.

[0005] An object of the present invention is to provide a handling device and a program that enable quick handling of objects.

[0006] A handling device according to the present invention is a handling device that identifies a position of an object, and identifies the position of the object based on an image of the object and contact information on the object.

[0007] The program of the present invention enables a computer that controls a handling device that identifies the position of an object to perform the following functions: acquire an image of the object; acquire contact information on the object; and identify the position of the object based on the image and the contact information.

[0008] According to the present invention, the position of an object is identified using not only information from a force sensor but also image information. Image information allows the object's position to be identified by relying on contours and other information obtained from the image, compared to identifying the object blindly with no prior information. Obtaining image information makes it easier to pinpoint the object compared to using a force sensor alone, and if force information is obtained from the force sensor, the object's position can be identified more quickly. The above provides the effect of providing a handling device and program that enable the rapid handling of objects.

[0009] It is a diagram showing a handling device of an embodiment of the present invention. It is a block diagram showing a functional module realized by a calculation device. It is a flowchart showing the procedure of a learning process. It is a block diagram explaining the flow of information during pre-learning. It is a flowchart showing the procedure of a handling process.

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0011] FIG. 1 is a diagram showing a handling device 1 according to an embodiment of the present invention. The handling device 1 according to this embodiment is a device that enables quick handling of the processing object 51, even if the processing object 51 is an object that has an irregular shape and is difficult to detect due to its good appearance. The handling device 1 according to this embodiment is a device that performs different operations and processes during "pre-learning" and "operation." Pre-learning refers to the time when machine learning, which will be described later, is performed, and operation refers to the time when the processing object 51 is actually handled. The processing object 51 corresponds to an example of an object according to the present invention.

[0012] Here, the term "irregular shape" for the processing object 51 means that the shape does not have a fixed shape due to changes in shape. The above-mentioned shape changes may include changes due to bending as well as changes due to expansion and contraction. Alternatively, objects packaged in bubble wrap or similar materials are also considered irregular objects. Furthermore, the above-mentioned "difficult-to-detect appearance" cases include cases where detection is difficult due to the characteristics of the processing object 51's appearance and cases where detection is difficult due to the environment. Examples of "appearance characteristics that are difficult to detect well" include objects with transparent, translucent, black, or glossy appearances when detecting the appearance using a range camera, and transparent or translucent objects when detecting the appearance using an imaging camera. A range camera is a photography sensor that scans space with a range sensor such as LiDAR (Light Detection and Ranging) to obtain three-dimensional coordinate data of each point on the target surface. A range camera may also be called a spatial scanner. Transparent, translucent, or black appearances significantly reduce the amount of laser light reflected, making accurate detection by a range camera difficult. If the object has a glossy appearance, it is difficult to obtain uniform reflection of the laser light, making it difficult to perform good detection using a distance camera. The above-mentioned "environment where good detection is difficult" corresponds to a situation where the object 51 is located in a place where light does not easily reach, a situation where the object 51 is located in darkness, etc.

[0013] The following describes, as a specific example, an example in which the handling device 1 of this embodiment handles a translucent, irregularly shaped bag as the processing object 51 during operation. More specifically, the handling device 1 estimates the center point (specifically, the geometric center of gravity) of the opening of the bag (i.e., the processing object 51) as the target position, and based on the estimation result, moves the tip of the articulated robot 12 into the processing object 51 through the opening. Hereinafter, the center point of the opening, which is the target position, will be referred to as the "opening position." The opening position is determined by the overall position, orientation, and deflection state of the processing object 51, and estimating the opening position corresponds to estimating the position and state of the processing object 51. The target position may be set to a position suitable for handling irregularly shaped objects.

[0014] 1, the handling device 1 of this embodiment includes a force sensor 11 that acquires contact information, an articulated robot 12 equipped with the force sensor 11, and an imaging sensor 13 that acquires imaging data of the processing object 51 (specifically, imaging data of an area including the processing object 51). The handling device 1 is controlled by a computer 20. The handling device 1 may also include the computer 20. The articulated robot 12 corresponds to an example of a movable part according to the present invention.

[0015] The force sensor 11 can detect contact with an object. The force sensor 11 may further have a function of detecting the direction of the reaction force against the contact (i.e., the direction of the contact surface). This detected information corresponds to contact information. From the contact information of the force sensor 11 and information on the operating position of the articulated robot 12, it is possible to calculate information such as the position and direction of the contact surface and the position where an object is not present.

[0016] A force sensor 11 is attached to the tip of the articulated robot 12. By rotating the arm around the joint, the articulated robot 12 can change the position of the tip to a plurality of orientations and positions, and can move the tip in a plurality of directions.

[0017] The imaging sensor 13 is the aforementioned distance camera and can acquire point cloud data of the object surface as imaging data. The point cloud data is a collection of three-dimensional coordinate values ​​of each point. The imaging sensor 13 is attached to the arm of the articulated robot 12. While the imaging sensor 13 may be located separately from the articulated robot 12, being attached to the articulated robot 12 allows the object 51 to be imaged from multiple positions and orientations as the articulated robot 12 moves. Furthermore, the imaging sensor 13 may be attached to a location whose relative position with respect to the force sensor 11 does not change significantly. With this configuration, imaging data can be acquired without significant change in the position and orientation relative to the tip of the articulated robot 12 when the tip of the articulated robot 12 is in contact with the object 51. Furthermore, imaging data can be obtained in which the tip of the articulated robot 12 is captured in the same way every time. As will be described in detail later, the image state estimator 24 performs machine learning using the imaging data, and the image state estimator 24, after undergoing machine learning, performs estimation calculations using the imaging data. Since the tip of the articulated robot 12 is always captured in the same location in the photographed data, this reduces the possibility that this part will become noise in the estimation calculation, contributing to improving the estimation accuracy of the image state estimator 24. In addition to the above-mentioned contact information with the processing object 51, by using an image that captures the processing object 51 and a part of the articulated robot 12 (for example, the tip of the articulated robot 12), it is possible to "guess" the relative direction of the processing object 51 with respect to the articulated robot 12 even in a situation where the processing object 51 is not captured accurately. Therefore, the position of the processing object 51 can be detected more quickly compared to position detection using an image that does not capture a part of the articulated robot 12.

[0018] The photographing sensor 13 is not limited to a distance camera, but may be any sensor that can obtain photographing data representing an image captured within a specified angle of view, such as an imaging camera that acquires an image of the subject area as image data.

[0019] The computer 20 includes an interface 201 for inputting and outputting commands or data, a storage device 202 for storing a control program, a calculation device 206 for performing calculation processing, and an input device 208 such as a keyboard, a mouse, or a storage medium reader. The calculation device 206 inputs and outputs commands or data via the interface 201 in accordance with the control program stored in the storage device 202. Contact information from the force sensor 11 and image data from the image sensor 13 are input to the computer 20. The computer 20 can cause the articulated robot 12 to perform a specified movement by outputting a command to the articulated robot 12. The computer 20 causes the articulated robot 12 to perform a movement while retaining information about the position and orientation of the tip of the articulated robot 12. An operator can input information for the learning process described below via the input device 208 and the interface 201. The input device 208 and the interface 201 correspond to an example of an input unit according to the present invention.

[0020] 2 is a block diagram showing functional modules implemented by the arithmetic unit 206. The arithmetic unit 206 implements a plurality of functional modules by executing a control program. The functional modules include a first data storage unit 21, a force sense state estimator 22, a second data storage unit 23, an image state estimator 24, and a controller 25.

[0021] The first data storage unit 21 stores contact information detected by the force sensor 11 in association with position information indicating the position where the contact information was obtained. Hereinafter, a set of data including contact information and position information will be referred to as a "force data sample." The position information corresponds to information indicating position and orientation.

[0022] The force sense state estimator 22 estimates the target position of the processing object 51 based on the data in the first data storage unit 21. The target position is the center point of the opening of the processing object 51 (specifically, a translucent bag). The center point of the opening, which is the target position, is called the "opening position."

[0023] A specific example of the estimation method will be described. The force sense state estimator 22 first creates observation distribution data from the data in the first data storage unit 21. The observation distribution refers to a collection of the force sense data samples expressed in a predetermined coordinate system. This coordinate system is a coordinate system in which the horizontal axis represents the position in space in which the tip of the articulated robot 12 can move, and the vertical axis represents the sensor value of the force sensor 11. The horizontal axis is not limited to one axis representing one-dimensional position, but may include three axes representing three dimensions. The vertical axis may include two axes representing the value and direction of the reaction force.

[0024] Next, the force sense state estimator 22 calculates a probability distribution of the opening position based on the structural data of the processing object 51 given in advance and the above-mentioned observation distribution. The structural data includes a reference shape of the processing object 51 and data that can be used to determine various shapes that can be deformed from the reference shape. The above-mentioned "probability distribution of the opening position" is expressed as a probability value for each point in a predetermined coordinate system. The above-mentioned predetermined coordinate system corresponds to a coordinate system in which the horizontal axis represents the position in space in which the tip of the articulated robot 12 can move, and the vertical axis represents probability. The above-mentioned horizontal axis is not limited to one axis representing one-dimensional position, but may include three axes representing three-dimensional position. From the probability distribution of the opening position, a position with a high probability of being the opening position can be determined. If there are multiple positions with high probability, it indicates that the opening position is likely to be one of the multiple positions. If there is only one position with high probability, it indicates that the position is likely to be the opening position.

[0025] The force sense state estimator 22 may use a technique called a particle filter to calculate the probability distribution of the opening position from the observation distribution. In the particle filter technique, the processing object 51 is represented as a set of multiple particles located in space. Based on pre-given structural data of the processing object 51, the force sense state estimator 22 can determine the processing object 51 in one position and one state as one set of particle sets. Furthermore, in the particle filter technique, as initial values, virtual processing objects 51 in various positions and states (i.e., multiple processing objects 51 with different shapes due to deformation and different positions and orientations) are represented as a composite of multiple sets of particle sets. Then, once the observation distribution is obtained, particle sets that do not match the observation distribution are removed using the observation distribution. This technique narrows down the particle sets that match the observation distribution, i.e., the positions and states of the processing object 51 that match the observation distribution, and makes it possible to calculate the probability of the location of each part of the processing object 51 using the remaining particles.

[0026] The second data storage unit 23 stores the imaging data acquired by the imaging sensor 13 as needed. Each piece of imaging data is called a "photography data sample." In this embodiment, the imaging data is point cloud data of three-dimensional coordinates. Therefore, by associating the three-dimensional coordinates with absolute coordinates based on the position information of the articulated robot 12, the "photography data sample" becomes data representing the object surface in an absolute three-dimensional coordinate system. Note that, if the imaging data is image data acquired by an imaging sensor, a set of data combining the imaging data with position information (specifically, information indicating the position and orientation) at which the imaging data was acquired may also be called a "photography data sample."

[0027] The image state estimator 24 is a machine learning model that, during operation, generates a probability distribution of the opening position (i.e., the target position) of the processing object 51 from one or more imaging data samples. A neural network can be applied as the above-mentioned machine learning model. During pre-learning, the image state estimator 24 is trained by machine learning, using imaging data samples and the probability distribution of the opening position obtained when the imaging data samples were acquired as learning data. Through this machine learning, during operation, when imaging data samples are input, the image state estimator 24 can output a probability distribution of the opening position that matches the imaging data sample.

[0028] Note that the image state estimator 24 is not limited to a configuration in which one photographic data sample is input and a probability distribution of the aperture position is output during operation. A configuration in which multiple photographic data samples are input and a probability distribution of the aperture position appropriate for the multiple photographic data samples is also applicable. In this case, during pre-learning, multiple photographic data samples obtained by photographing a training object 56 (described later) and the probability distribution of the aperture position obtained for the training object 56 are provided as training data and machine learning is performed. During pre-learning, the training object 56 is used as a contact and photograph target for acquiring contact information and photographic data, instead of the processing object 51. The training object 56 may be the same as the processing object 51, or a similar object with similar characteristics may be used.

[0029] Furthermore, the image state estimator 24 may be configured, during operation, to output the probability distribution of the aperture position (i.e., the target position) as image data representing the probability distribution (e.g., grayscale image data). Image data representing the probability distribution corresponds to information on the probability distribution in three-dimensional space projected onto a two-dimensional plane. Furthermore, the image state estimator 24 may be configured, during operation, to convert and input photographic data samples, which are point cloud data including the three-dimensional coordinate values ​​of each point, into photographic image data projected onto a two-dimensional plane (e.g., grayscale image data), and output the data as the above-mentioned "image data representing the probability distribution." These configurations can be realized by training the image state estimator 24 using training data corresponding to the configuration during pre-training.

[0030] During operation, the controller 25 receives the probability distribution data of the opening position output from the force sense state estimator 22 and the probability distribution data of the opening position output from the image state estimator 24, and moves the articulated robot 12 based on this data. More specifically, the controller 25 determines whether the estimated opening position meets the desired condition (e.g., the condition that the opening position can be identified at approximately one location) based on the probability distribution data of the opening position. If this condition is not met, the controller 25 calculates a movement path that will obtain contact information with the processing object 51, moves the articulated robot 12 so that the tip of the robot moves along this movement path, and further obtains contact information to increase the number of force sense data samples. The controller 25 calculates the movement path so that the robot randomly contacts the processing object 51 or its opening.

[0031] Instead of calculating the movement path as described above, the controller 25 may calculate the movement path using Bayesian optimization. Bayesian optimization is a calculation method that attempts to contact a point with a large variance in the probability distribution of the current opening position, thereby efficiently narrowing the probability distribution of the opening position to one point.

[0032] During pre-learning, the controller 25 further has the function of creating multiple movement paths that randomly attempt to contact the learning object 56 by inputting known position data representing the position of the learning object 56 from outside.

[0033] <Processing During Pre-Learning> Fig. 3 is a flowchart showing the procedure of the learning process. Fig. 4 is a block diagram showing the flow of information during pre-learning. Next, the learning process performed during pre-learning using the handling device 1 will be described. The learning process is a process for training the image state estimator 24 using a learning object 56. The learning object 56 is an object having characteristics that are the same as or similar to those of the processing object 51.

[0034] In the learning process, an operator arranges a learning object 56 in an arbitrary shape and inputs information about the position of the learning object 56 as "known position data" into the computer 20 via the interface 201 (step S1). The known position data may be data indicating the overall position of the learning object 56 or data indicating the opening position. The input known position data is sent to the controller 25, which calculates multiple movement paths that randomly contact the learning object 56 based on the information (step S2).

[0035] Next, the controller 25 executes a loop process of steps S3 to S5. In this loop process, the controller 25 repeats the process of step S4 while changing the movement path by the number of movement paths calculated in step S2. In step S4, the controller 25 drives the articulated robot 12 so that the tip moves along the calculated movement path, and acquires contact information and photographic data (step S4). Through this loop process, a plurality of force sense data samples are stored in the first data storage unit 21, and a plurality of photographic data samples are stored in the second data storage unit 23.

[0036] When the loop processing of steps S3 to S5 is completed, the haptic state estimator 22 performs an estimation calculation of the mouth opening position using the haptic data samples in the first data storage unit 21 (step S6). As a result, data of the probability distribution of the mouth opening position, which is the estimation result, is output from the haptic state estimator 22. The estimation calculation method and the probability distribution of the mouth opening position are as described above.

[0037] Next, the image state estimator 24 receives the multiple imaging data samples acquired in the loop processing of steps S3 to S5 and the probability distribution of the aperture positions output in step S6, and performs machine learning based on these (step S7). Through this learning processing, the image state estimator 24 can perform operations during operation, that is, when imaging data is input, output data of the probability distribution of the aperture positions that matches the imaging data.

[0038] The worker repeats the learning process of steps S1 to S7 multiple times by changing the position and state of the learning object 56. By repeating the learning process multiple times in this way, the estimation accuracy of the image state estimator 24 improves to a certain extent.

[0039] There are several variations in how the known position data input in step S1 of the learning process is used. As a first variation, for example, in step S1, the known position data may be converted into a haptic data sample format and the haptic data sample may be stored in the first data storage unit 21. Because the known position data does not include information about the shape of the amorphous learning object 56, the haptic data sample calculated based on the known position data is an incomplete haptic data sample rather than a complete haptic data sample that can determine the overall position of the learning object 56. The haptic state estimator 22 may then estimate the opening position based on the haptic data sample, and the controller 25 may calculate a movement path based on the estimation result and drive the articulated robot 12. By repeating this movement path calculation and driving the articulated robot 12, the haptic state estimator 22 obtains probability distribution data of the opening position with high estimation accuracy. Furthermore, multiple pieces of image data may be acquired during the repetition. The image state estimator 24 may then perform machine learning using this as learning data.

[0040] As a second variation, the following data may be provided to the image state estimator 24 as answer data for machine learning, instead of the data of the probability distribution of the opening position output by the force sense state estimator 22. That is, the known position data input in step S1, or data obtained by converting this data into the format of the probability distribution of the opening position or the format of image data, may be provided to the image state estimator 24 as answer data. Then, the image state estimator 24 may perform machine learning based on the photographed data and the above answer data.

[0041] In the learning process of the above embodiment and the first variation, an example has been shown in which the haptic state estimator 22 estimates and calculates the probability distribution of the opening position based on all the haptic data samples obtained in the loop process (S3 to S5), and provides this data to the image state estimator 24 as answer data. However, the haptic state estimator 22 may estimate and calculate the probability distribution of the opening position each time a haptic data sample is added, and provide the data of the probability distribution of the multiple opening positions calculated in this way as answer data to the image state estimator 24. Alternatively, only the data of the probability distribution of the multiple opening positions calculated as above, whose estimation accuracy exceeds a preset threshold, may be provided to the image state estimator 24 as answer data.

[0042] A learning processing program 203 that realizes the above learning processing is stored in the storage device 202 .

[0043] <Processing During Operation> Fig. 5 is a flowchart showing the procedure of the handling process. Next, the handling process executed during operation will be described. The handling process is a process in which, during operation, the handling device 1 estimates the opening position of the processing object 51 and then performs predetermined handling on the processing object 51. The predetermined handling is, for example, carrying in or out an object through the opening of the processing object 51. During operation, the processing object 51 is placed in a roughly determined location (for example, the nth shelf), and the computer 20 has information on the rough location where the processing object 51 is placed.

[0044] When the handling process begins, the controller 25 drives the articulated robot 12 so that the photographing sensor 13 faces the approximate location where the object to be processed 51 will be placed, and acquires the photographing data (specifically, one or more photographing data samples) output from the photographing sensor 13 (step J1).

[0045] When the photographic data is acquired, the data is stored as a photographic data sample in the second data storage unit 23 and sent to the image state estimator 24. The image state estimator 24 then performs an estimation calculation of a machine learning model that uses the photographic data sample as input and outputs data on the probability distribution of the opening position, thereby estimating the opening position of the processing object 51 (step J2).

[0046] By the estimation calculation in step J2, the image state estimator 24 outputs data on the probability distribution of the opening positions, which is the estimation result, to the controller 25. The probability distribution of the opening positions is a calculation result of a machine learning model, and therefore may contain errors.

[0047] The controller 25 calculates a movement path for the tip of the articulated robot 12 that will allow the opening position to be narrowed down further based on the probability distribution data of the opening position in step J2 (step J3), and then drives the articulated robot 12 according to the calculated movement path (step J4).

[0048] Then, the force sensor 11 comes into contact with the processing object 51, and contact information is acquired (step J5). The contact information is combined with the position information and stored in the first data storage unit 21 as a force data sample.

[0049] The driving in step J4 is not limited to driving that contacts one point on the processing object 51, but may be driving that contacts multiple points or across multiple locations. The force sense data samples obtained in step J5 may be force sense data samples from multiple points or across multiple locations that are obtained by contacting multiple points or across multiple locations.

[0050] Next, the force sense state estimator 22 estimates and calculates a probability distribution of the opening position based on the force sense data samples stored in the first data storage unit 21 (step J6).The controller 25 then determines whether the probability distribution of the opening position calculated in step J6 satisfies a predetermined condition (step J7).The predetermined condition may be a condition that allows the opening position to be identified to approximately one location, such as the opening position being narrowed down to one location within a threshold volume in three-dimensional space, or any other condition that makes it possible to handle the subsequent step J8.

[0051] If the determination result in step J7 is NO, the process returns to step J1, and the process is repeated from step J1 onwards. This repeated process increases the number of accumulated haptic data samples, thereby improving the estimation accuracy of the haptic state estimator 22 in step J6 and further narrowing down the opening position.

[0052] If the determination result in step J7 is YES, the controller 25 drives the articulated robot 12 along a movement path based on the estimated opening position, and performs a predetermined handling operation, such as carrying an object into the bag-shaped processing object 51 or carrying an object out of the bag-shaped processing object 51 (step J8). Then, one handling operation is completed.

[0053] 5 shows an example in which the process is repeated from step J1 if the determination result in step J7 is NO. However, the acquisition of photographic data in steps J1 and J2 and the estimation calculation by the image state estimator 24 may be performed the first time and omitted in the repeated processes. In other words, if the determination result in step J7 is NO, the process may return to step J3. In this case, in step J3, the controller 25 may calculate a movement path based on the probability distribution of the opening position calculated in the previous step J6 in the repeated processes and the force data samples accumulated in the first data storage unit 21.

[0054] A handling processing program 204 that realizes the above handling processing is stored in the storage device 202 .

[0055] As described above, the handling device 1 according to the present embodiment is a handling device 1 that identifies the position of the processing object 51, and identifies the position of the processing object 51 based on photographic data (equivalent to an image) of the processing object 51 and contact information with the processing object 51. By identifying the position using a combination of the photographic data and the contact information, it is possible to quickly identify the position.

[0056] Furthermore, according to the handling device 1 of this embodiment, the position of the processing object 51 is identified using an image that shows the processing object 51 and a part of the articulated robot 12 (for example, the tip of the articulated robot 12). By using such an image, even in a situation where the processing object 51 is not accurately shown, it is possible to "guess" the relative direction of the processing object 51 with respect to the articulated robot 12. Therefore, the position of the processing object 51 can be detected more quickly compared to position detection using an image that does not show a part of the articulated robot 12.

[0057] Furthermore, the handling device 1 according to this embodiment includes a force sensor 11 that acquires contact information, an articulated robot 12 equipped with the force sensor 11, an imaging sensor 13 that acquires imaging data, and an interface 201 that can input data. During pre-learning, the handling device 1 performs a learning process to estimate the position of the training object 56 based on position information of the irregular training object 56 input via the interface 201, contact information of the training object 56 acquired by the force sensor 11, and imaging data acquired by the imaging sensor 13. As a more specific example, the image state estimator 24, which is a machine learning model, performs machine learning based on the force data samples in the first data storage unit 21 and the imaging data samples in the second data storage unit 23 so that, when the imaging data samples are input, data of a probability distribution of opening positions that matches the input imaging data samples is output.

[0058] According to this configuration, even in a situation where good photographing data cannot be obtained and the training object 56 has an irregular configuration, the above-mentioned learning process can be performed efficiently using learning data based on the input position information of the training object 56 and the contact information acquired by the force sensor 11. Furthermore, by acquiring photographing data through this learning process, even in a situation where good photographing data cannot be obtained and the processing object 51 has an irregular shape, it becomes possible to perform a relatively accurate estimation calculation of the position and state of the processing object 51 based on the photographing data.

[0059] Furthermore, during operation, the handling system 1 according to this embodiment handles the processing object 51 based on the image data of the processing object 51 acquired by the image sensor 13 and the contact information of the processing object 51 acquired by the force sensor 11. As a more specific example, the machine-learned image state estimator 24 estimates the opening position of the processing object 51 based on the image data samples and drives the articulated robot 12 along a movement path determined based on the estimation result, thereby efficiently acquiring contact information from the force sensor 11. Then, the force state estimator 22 can quickly identify the opening position of the processing object 51 based on the force data samples containing the contact information. Therefore, even if the processing object 51 is in a situation where good image data cannot be obtained and has an irregular shape, prompt handling can be achieved.

[0060] Furthermore, in the handling system 1 according to this embodiment, a semi-transparent, amorphous object is used as the processing object 51. It is difficult to obtain good image data for the processing object 51 with such characteristics using the imaging sensor 13, and it takes a long time to identify the state and position of the processing object 51 using only the contact information from the force sensor 11. The handling system 1 according to this embodiment is particularly useful because it can quickly identify the opening position even for such a processing object 51.

[0061] Furthermore, in the handling system 1 according to this embodiment, a semi-transparent, shapeless object is used as the learning object 56, similar to the processing object 51. By performing a learning process using such a learning object 56, learning suitable for handling the semi-transparent, shapeless processing object 51 is performed, and an operation suitable for handling the semi-transparent, shapeless processing object 51 is realized.

[0062] The above describes embodiments of the present invention. However, the present invention is not limited to the above embodiments. For example, in the above embodiments, bag-shaped objects are used as processing objects, but the present invention is not limited to bag-shaped objects. For example, the present invention can also be applied to estimating the gripping position of a handling device (specifically, estimating the position that is easy to grasp) when transporting a package containing substrates or the like packed in bubble wrap. Furthermore, in the above embodiments, the estimated position and state of the processing object are the opening position (e.g., the center position of the opening), but the estimated position and state are not limited to the opening position. The positions and states of various locations necessary for handling the processing object may be applied. Furthermore, in the above embodiments, the articulated robot 12 is used as the moving part that moves the force sensor 11. However, any type of moving part may be used as long as the force sensor 11 can be brought into contact with the processing object. Furthermore, in the above embodiments, the force sensor 11 is provided at the tip of the articulated robot 12, but the force sensor 11 may be provided at a location other than the tip. In addition, the details shown in the embodiments can be modified as appropriate without departing from the spirit of the invention.

[0063] The disclosures of the specification, drawings and abstract contained in Japanese Patent Application No. 2024-107225, filed on July 3, 2024, are incorporated herein by reference in their entirety.

[0064] The present invention can be used in a handling device and a program.

[0065] REFERENCE SIGNS LIST 1 Handling device 11 Force sensor 12 Articulated robot (moving part) 13 Photography sensor 51 Processing object 56 Learning object 20 Computer 21 First data storage unit 22 Force sense state estimator 23 Second data storage unit 24 Image state estimator 25 Controller 201 Interface (input unit) 202 Storage device 203 Learning processing program 204 Handling processing program 206 Arithmetic unit 208 Input device

Claims

1. A handling device that identifies the position of an object, the handling device identifying the position of the object based on an image of the object and contact information for the object.

2. The handling device according to claim 1, further comprising a movable part that can come into contact with the object, and the image of the object is an image of the object and the movable part.

3. The handling device according to claim 1, wherein the object is an amorphous object having a transparent, translucent, black, or glossy appearance, and the position of the object is a position suitable for handling the amorphous object.

4. A program for enabling a computer that controls a handling device that identifies the position of an object to perform the following functions: acquiring an image of the object; acquiring contact information on the object; and identifying the position of the object based on the image and the contact information.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, robot control apparatus, and robot system

    JP2017136677A

  • Holding system, learning device, holding method and model manufacturing method

    JP2019093461A

  • Information processor, information processing method, and program

    JP2019188516A

  • Information processing device, control method, robot system, computer program, and storage medium

    JP2019188580A

  • Information processing device and picking device

    JP2022086536A