Object recognition apparatus and object recognition method
By inputting scene information through an object recognition device, the robot sorts and detects objects based on the number of objects of each size, solving the problem of grasping and obstacle avoidance when recognizing objects of unknown posture and size, and achieving high-precision object recognition and accurate object sorting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, robots struggle to grasp and avoid obstacles with high precision when identifying objects of unknown pose and size, and creating object models is time-consuming, especially when dealing with a wide variety of objects.
The system inputs scene information through an object recognition device, sorts and detects objects based on the number of objects of each size, excludes sizes that are not allowed to be stored in the container, uses sensors to obtain the size and posture information of the objects, and generates histograms for sorting and recognition.
Even without an object model, it can accurately identify objects, ensuring that the robot can accurately grasp objects of permissible size, avoiding mishandling, and improving recognition efficiency and accuracy.
Smart Images

Figure CN117203669B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the identification of objects existing in a container. Background Technology
[0002] Previously, in logistics and other fields, warehouses have robots that retrieve goods from storage bins based on orders. These bins often contain similar items stored in various ways, requiring the robot to estimate the item's posture and size. Without knowing the item's posture, the robot cannot, for example, make a movement to approach the desired item without collision. Furthermore, without knowing the item's size, the robot cannot, for example, retrieve the item while avoiding obstacles such as other items after grasping it.
[0003] In recent years, a method has been known to use product models to estimate the posture and size of products. In this method, since the product size is known because a product model is held, there is no need to estimate the product size. However, since product models must be created for each type of product, the creation of product models becomes time-consuming as the number of product types increases.
[0004] In this regard, a position and posture recognition device is disclosed that can identify the position and posture of randomly stacked workpieces with less computation and can be used with SCARA robots that require less control information (see Patent Document 1). According to the position and posture recognition device, the workpiece to be grasped is defined as approximately cylindrical, and point group data is generated from a three-dimensional image of the workpiece in a randomly stacked state. The size and position and posture of the workpiece are calculated based on the point group data, so there is no need to prepare a model in advance.
[0005] Existing technical documents
[0006] Patent documents
[0007] Patent Document 1: Japanese Patent Application Publication No. 2020-34526 Summary of the Invention
[0008] The problem that the invention aims to solve
[0009] In the technology described in Patent Document 1, the object's posture and size are estimated by fitting the prototype shape to the input data based on information about the prototype shape of a roughly cylindrical workpiece. However, this method has the following problems: there are many cases where the prototype shape can be fitted with high accuracy even when part of the object is hidden, and there is no reference object model, so it is impossible to determine whether the estimated size of the object is correct.
[0010] The present invention was made with the above considerations in mind, and aims to provide an object recognition device that can identify objects with high accuracy even when there is no object model as a reference.
[0011] Solution for solving the problem
[0012] To address this issue, the present invention includes: an input unit that inputs scene information representing the state of a specified scene in a container storing objects; a processing unit that, based on the scene information input by the input unit, sorts objects existing in the container according to each size, and, based on the number of objects of each size sorted, performs processing to detect objects of sizes that are not allowed to be stored in the container; and an output unit that outputs the result of the processing performed by the processing unit.
[0013] In the above structure, based on the number of objects of each size, objects of sizes that are not allowed to be stored in the container are detected. Therefore, even if there is no object model as a reference, objects of sizes that are not allowed to be stored in the container can be excluded from the container.
[0014] Invention Effects
[0015] According to the present invention, objects can be identified with high accuracy even when there is no object model to serve as a reference. Attached Figure Description
[0016] Figure 1 This is a diagram illustrating an example of the processing performed by the object recognition system of the first embodiment.
[0017] Figure 2 This is a diagram illustrating an example of the structure involved in the object recognition system of the first embodiment.
[0018] Figure 3 This is a diagram illustrating an example of the hardware structure involved in the object recognition device of the first embodiment.
[0019] Figure 4 This is a diagram illustrating an example of the object recognition processing of the first embodiment.
[0020] Figure 5 This is a diagram showing the result of the first embodiment.
[0021] Figure 6 This is a diagram illustrating an example of the structure involved in the object recognition device according to the second embodiment.
[0022] Figure 7 This is a diagram illustrating an example of the object recognition processing of the second embodiment.
[0023] Figure 8 This is a diagram showing the result of the second embodiment.
[0024] Figure 9 This is an example of the structure involved in the object recognition device of the third embodiment.
[0025] Figure 10 This is a diagram illustrating an example of the object recognition processing of the third embodiment.
[0026] Figure 11 This diagram illustrates an example of the processing performed by the object recognition device according to the fourth embodiment.
[0027] Figure 12 This diagram illustrates an example of the processing performed by the object recognition device according to the fifth embodiment. Detailed Implementation
[0028] (I) First Embodiment
[0029] The following describes one embodiment of the present invention in detail. However, the present invention is not limited to this embodiment.
[0030] The object recognition device of this embodiment receives scene information related to the scene in which objects exist within a container. The scene information includes an estimated result of the object size. The object recognition device determines the size of the objects present in the scene based on the quantity of objects of each size.
[0031] Based on the above structure, it is possible to exclude the estimated dimensions of objects that are not allowed in the scene from the estimated dimensions contained in the scene information, thereby enabling the selection of the estimated dimensions of the correct objects.
[0032] The designations "first," "second," "third," etc., used in this specification are for identifying structural elements and do not necessarily limit the number or order. Furthermore, the reference numerals used to identify structural elements are used in a coherent context; a reference numeral used in one coherent context is not necessarily limited to representing the same structure in other coherent contexts. Additionally, this does not preclude a structural element identified by a particular reference numeral from also performing the functions of structural elements identified by other reference numerals.
[0033] Next, embodiments of the present invention will be described based on the accompanying drawings. The following description and drawings are illustrative of the invention, and appropriate omissions and simplifications have been made for clarity. The present invention can also be implemented in various other ways. Unless otherwise specified, each structural element may be one or more.
[0034] Furthermore, in the following description, the same reference numerals are used for the same elements in the accompanying drawings, and descriptions are omitted where appropriate.
[0035] exist Figure 1In the text, 100 generally represents the object recognition system of the first embodiment. Figure 1 This is a diagram illustrating an example of the processing performed by the object recognition system 100.
[0036] The object recognition system 100 performs input processing 110, sorting processing 120, and output processing 130.
[0037] In its input processing 110, the object recognition system 100 inputs data related to the scene of the object being recognized (the object scene) as scene information. The scene information includes an estimation result 111. The estimation result 111 includes information indicating the estimated dimensions of objects contained within the object scene. Furthermore, the estimation of the dimensions of objects contained within the object scene is illustrated using the case described later in the robot 210, but is not limited to this. For example, the estimation of the dimensions of objects contained within the object scene can be performed in the object recognition system 100, or in a system different from the object recognition system 100, such as a computer.
[0038] In the sorting process 120, the object recognition system 100 uses an estimation result 111 for the object scene (the scene as a whole) to sort objects that appear appropriate in the object scene (objects of allowed sizes in the object scene) and generates a sorting result 121. The sorting result 121 includes information 122 indicating objects of allowed sizes in the object scene. For example, the object recognition system 100 sorts objects of allowed sizes in the object scene by voting on the sizes of multiple objects.
[0039] In the output processing 130, the object recognition system 100 generates and outputs a recognition result 131 based on the sorting result 121. The recognition result 131 includes information such as information representing objects of permitted size in the object scene, information representing objects of prohibited size in the object scene, information representing the pose of objects of prohibited size in the object scene, and information representing the evaluation of the inference result 111 (e.g., whether the objects contained in the object scene are objects of permitted size in the object scene).
[0040] Figure 2 This is a diagram illustrating an example of the structure involved in the object recognition system 100.
[0041] The object recognition system 100 is configured to include a robot 210 and an object recognition device 220. The robot 210 and the object recognition device 220 are connected in a communicative manner. The robot 210 and the object recognition device 220 may or may not be connected via a network.
[0042] Robot 210 estimates the size and orientation of the object in storage box 211 based on the order (order information), includes the estimation result 111 in the scene information of the object scene, and sends it to object recognition device 220. When object recognition device 220 receives the scene information, it uses the received scene information to generate a sorting result 121, and generates a recognition result 131 based on the sorting result 121, and sends the generated recognition result 131 to robot 210. Based on the recognition result 131, robot 210 retrieves an object of an allowed size in the object scene from storage box 211 and moves the object to the designated location 212.
[0043] In this embodiment, objects of a single size (first size) are generally stored in the storage box 211. However, when storing objects in the storage box 211, it is possible that objects of the wrong size (a second size different from the first size) may be mixed in due to human error or other reasons. Even if such a situation occurs, the object recognition device 220 will sort out objects of sizes that are not allowed in the object scene, so the robot 210 can retrieve (pick) objects of the size that conform to the order (objects of the allowed size in the object scene) from the storage box 211.
[0044] More specifically, the object recognition device 220 includes an input unit 221, a processing unit 222, and an output unit 223, and identifies objects stored in the storage box 211.
[0045] The input unit 221 receives scene information sent from the robot 210 and inputs the received scene information. The processing unit 222 performs object sorting and other processing based on the scene information input by the input unit 221. The output unit 223 outputs the results of the objects sorted by the processing unit 222.
[0046] Figure 3 This is a diagram illustrating an example of the hardware structure involved in the object recognition device 220.
[0047] The object recognition device 220 includes a processor 301, a main storage device 302, an auxiliary storage device 303, an input device 304, an output device 305, and a communication device 306.
[0048] Processor 301 is a device for performing computational processing. Processor 301 may be, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), or an AI (Artificial Intelligence) chip.
[0049] Main storage device 302 is a device for storing programs, data, etc. Main storage device 302 may be, for example, ROM (Read Only Memory) or RAM (Random Access Memory). ROM may be SRAM (Static Random Access Memory), NVRAM (Non-Volatile RAM), Mask ROM (Mask Read Only Memory), PROM (Programmable ROM), etc. RAM may be DRAM (Dynamic Random Access Memory), etc.
[0050] The auxiliary storage device 303 can be a hard disk drive, flash memory, SSD (Solid State Drive), optical storage device, etc. Optical storage devices can be CD (Compact Disc), DVD (Digital Versatile Disc), etc. Programs and data stored in the auxiliary storage device 303 can be read into the main storage device 302 at any time.
[0051] Input device 304 is a user interface for receiving information from a user. Input device 304 may be, for example, a keyboard, mouse, card reader, touch panel, etc.
[0052] Output device 305 is a user interface that outputs various information (display output, voice output, typing output, etc.). Output device 305 may be, for example, a display device that visualizes various information, a voice output device (speaker), a typing device, etc. The display device may be an LCD (Liquid Crystal Display), a graphics card, etc.
[0053] Communication device 306 is a communication interface for communicating with other devices via a communication medium. Communication device 306 can be, for example, a NIC (Network Interface Card), a wireless communication module, a USB (Universal Serial Interface) module, a serial communication module, etc. The communication medium can be based on various communication standards such as USB (Universal Serial Bus), RS-232C, LAN (Local Area Network), WAN (Wide Area Network), the Internet, leased lines, etc. The structure of the communication medium is not necessarily limited. Furthermore, communication device 306 can also function as an input device for receiving information from other communicatively connected devices. Additionally, communication device 306 can also function as an output device for sending information to other communicatively connected devices.
[0054] The functions of the object recognition device 220 (input unit 221, processing unit 222, output unit 223, etc.) can be implemented, for example, by the processor 301 reading the program stored in the auxiliary storage device 303 into the main storage device 302 and executing it (software), or by dedicated hardware such as circuitry, or by a combination of software and hardware. Furthermore, one function of the object recognition device 220 can be divided into multiple functions, or multiple functions can be combined into one function. Additionally, a portion of the functions of the object recognition device 220 can be set as different functions, or included in other functions. Furthermore, a portion of the functions of the object recognition device 220 can be implemented by another computer capable of communicating with the object recognition device 220. For example, some or all of the functions of the object recognition device 220 can be implemented in the robot 210, or in cloud computing.
[0055] Figure 4 This is a diagram illustrating an example of object recognition processing (input processing 110, sorting processing 120, and output processing 130) performed by the object recognition device 220.
[0056] In step S401, the object recognition device 220 inputs scene information. Furthermore, the scene information includes an estimation result 111.
[0057] In step S402, the object recognition device 220 creates a histogram based on scene information. In the histogram, the horizontal axis represents the grade, and the vertical axis represents the frequency. The frequency of each grade is shown as bars (e.g., rectangular columns). A label value representing that grade is set. Regarding the grade, one of the following can be used: either the grade is shown as a unique value, or the grade is shown as a range (a specified interval). For example, in the case of a specified grade with a vertical range of "10.00cm to 10.99cm" and a horizontal range of "8.00cm to 8.99cm", "10cm × 8cm" is set as the label value.
[0058] Here, for example, if the scene information contains data of a specific object that is presumed to have a vertical length of "10.01cm" and a horizontal length of "8.16cm", since the specific object belongs to the aforementioned specified level, the object recognition device 220 increments the frequency of the specified level by "1".
[0059] Furthermore, the object recognition device 220 may also have an interface that allows the user to set a specified group interval. Through this interface, for example, the user can set hyperparameters (allowable error).
[0060] In step S403, the object recognition device 220 sorts objects based on the histogram created in step S402. For example, the object recognition device 220 determines the size of the object in the scene based on the principle that the majority of the levels is better than the minority, and presumes the label value of the level with the highest frequency as the size of the object allowed in the scene. For example, if the level with the highest frequency is the aforementioned predetermined level, the object recognition device 220 presumes the size of the object as "10cm × 8cm", which is assigned as the label value of the predetermined level.
[0061] In step S404, the object recognition device 220 generates and outputs the recognition result 131 based on the sorting result (sorting result 121) obtained in step S403. Furthermore, the object recognition device 220 can send the recognition result 131 to the robot 210, display the recognition result 131 on the output device 305, or send the recognition result 131 to other computers.
[0062] Figure 5 This is a diagram showing an example (result image) of the result of processing performed in object recognition.
[0063] Result image 510 shows an example of presumed result 111. As shown in result image 510, presumed result 111 includes object IDs and values representing the dimensions of each object contained in the object scene. Furthermore, result image 510 shows an example where the object is box-shaped, and an example where the parameters representing the object's dimensions are height and width. However, the object's dimensions can be combinations other than height and width, for example, including a height parameter.
[0064] Result image 520 shows an example of sorting result 121. As shown in result image 520, sorting result 121 includes information such as voting results based on strict numerical consistency, voting results based on histograms of specified group intervals (e.g., 1 cm group intervals), etc.
[0065] Result image 530 shows an example of recognition result 131. As shown in result image 530, recognition result 131 contains information about the size of the object selected as an object of an allowed size in the object scene (e.g., the object's ID).
[0066] According to this embodiment, it is possible to exclude the estimated results of objects whose sizes are not allowed in the object scene from the estimated results contained in the scene information, thereby enabling the correct selection of the estimated results of objects whose sizes are allowed in the object scene.
[0067] (II) Second Embodiment
[0068] In this embodiment, the main difference from the first embodiment lies in how the object recognition device 600 generates the estimation result 111. In this embodiment, the same reference numerals are used for structures identical to those in the first embodiment, and their descriptions are omitted.
[0069] Figure 6 This is a diagram illustrating an example of the structure involved in the object recognition device 600 of this embodiment.
[0070] The object recognition device 600 is connected to one or more sensors 601 that acquire the state (physical phenomena, chemical phenomena, etc.) of the storage box 211. The object recognition device 600 includes an input unit 611, a processing unit 222, and an output unit 223. The processing unit 222 is configured to include a size estimation unit 612 and a posture estimation unit 613.
[0071] The input unit 611 inputs sensor information acquired by the sensor 601. The sensor information is an example of scene information showing the object scene, such as a color image, point cloud, depth image, grayscale image, or a combination thereof. The size estimation unit 612 estimates the size of objects within the object scene. The pose estimation unit 613 estimates the pose of objects within the object scene.
[0072] The following describes the case where the object has a prototype shape and is stored in storage box 211 in an unordered manner. The prototype shape is a combination of parameters used to represent the shape of the object, such as length, width, and depth. In other words, the object's shape being a prototype shape can be described as the object's dimensions being a combination of values representing the prototype shape.
[0073] Here, the object recognition device 600 records prototype shape information 614, which includes information showing parameters representing the prototype shape of the object. When the prototype shape is box-shaped, the parameters representing the prototype shape are vertical, horizontal, and height, or vertical and horizontal; when the prototype shape is cylindrical, the parameters representing the prototype shape are radius and height, radius, or height; and when the prototype shape is spherical, the parameter representing the prototype shape is radius. For example, when the prototype shape is box-shaped, the object recognition device 600 can estimate box-shaped objects of various sizes. Furthermore, the prototype shape is not limited to the examples described above; it can also be a triangular prism, a combination of the above prototype shapes, or other prototype shapes.
[0074] Furthermore, the structure of the object recognition device 600 is not limited to the structure described above. For example, the object recognition device 600 may not be directly connected to the sensor 601 (direct connection), but may be connected to a device connected to the sensor 601 (indirect connection). Additionally, the object recognition device 600 may not include the posture estimation unit 613.
[0075] Figure 7 This diagram illustrates an example of object recognition processing performed by the object recognition device 600.
[0076] In step S701, the object recognition device 600 inputs scene information from the sensor 601. The scene information can be sent from the sensor 601 at a first moment, or it can be requested and obtained by the object recognition device 600 at a second moment. The first moment can be a normal time, a time when an order is accepted, a regular time, a pre-specified time, or a time indicated by the user. Furthermore, the second moment can be the same as the first moment or a different time.
[0077] In step S702, the object recognition device 600 performs region extraction based on the scene information input in step S701, and segments out the object point groups that are presumed to constitute objects. For example, the object recognition device 600 segments the region of each object using InstantSegmentation and identifies the type of object. For example, the object recognition device 600 classifies which instance a pixel in the image belongs to among multiple instances (segmenting out object point groups with (x, y, z) coordinates distributed in three-dimensional space). In addition, the object recognition device 600 can also use other techniques such as Watershed for segmentation. Furthermore, for region extraction, it is not limited to InstantSegmentation; object detection techniques such as BING (Binarized Normed Gradients) can also be used.
[0078] In step S703, the object recognition device 600 estimates the surfaces that constitute the object. For example, the object recognition device 600 reads the prototype shape information 614 and uses the prototype shape of the object to calculate at least one surface most suitable for the object point group cut out in step S702 using methods such as RANSAC (Random Sample Consensus). It then confirms whether the calculated surfaces satisfy the rules that should be satisfied for each prototype shape of the object, and cuts out the point group that constitutes the object. The rules that should be satisfied for each prototype shape of the object refer to, for example, rules that can be satisfied if the calculated surfaces are orthogonal to each other, when the prototype shape of the object is a box shape.
[0079] In step S704, the object recognition device 600 performs a size estimation. The object recognition device 600 estimates the size of the object from the point group constituting the object, which was estimated in step S703. For example, if the object's prototype shape is box-shaped, and the point group constituting the object is a single face, the minor and major axes of the point group become the object's size. Furthermore, if the point group constituting the object has two or three faces, the lengths of the axes of the three-dimensional rectangle that has the smallest volume among the three-dimensional rectangles covering the point group become the object's size.
[0080] In step S705, the object recognition device 600 performs pose estimation. For example, the object recognition device 600 generates a group of points (a model of the prototype shape) representing the size of the object estimated in step S703, and uses an algorithm (e.g., ICP (Iterative Closest Point)) that minimizes the distance between the group of points constituting the object and the group of points in the model of the prototype shape, as cut out in step S702, to estimate the pose of the object (e.g., estimate the rotation and translation matrices).
[0081] Steps S402 to S404 have been described in the first embodiment, so their description is omitted.
[0082] Figure 8 This is a diagram showing an example (result image) of the result of processing performed in object recognition.
[0083] Result image 810 shows an example of the result (scene information) after inputting the sensor information acquired by sensor 601 in step S701. Here, the sensor information included in the scene information is shown using a color image and a point group.
[0084] Result image 820 shows an example of the result (estimation result) of estimating the size of the object in step S704. As shown in result image 820, the estimation result includes the ID of each object used to identify it and a value representing the size of the object, for each object contained in the object scene. Furthermore, result image 820 shows an example where the object is box-shaped, and an example where the parameters representing the size of the object are height and width. However, the size of the object can be a combination other than height and width, for example, including a height parameter. Additionally, the estimation result may also include the result of estimating the pose of the object in step S705.
[0085] According to this embodiment, the size of an object can be estimated based on sensor information of the object scene, and objects of sizes that are not allowed in the object scene can be detected.
[0086] (III) Third Embodiment
[0087] This embodiment differs from the second embodiment primarily in that it simultaneously estimates both the size and orientation of the object. In this embodiment, the same reference numerals are used for structures identical to those in the second embodiment, and their descriptions are omitted.
[0088] Figure 9 An example of the structure involved in the object recognition device 900 of this embodiment is shown.
[0089] The object recognition device 900 includes an input unit 611, a processing unit 222, and an output unit 223. The processing unit 222 is configured to include a size and posture estimation unit 911.
[0090] The size and pose estimation unit 911 estimates the size and pose of an object based on the scene information input in the input unit 611.
[0091] Figure 10 This diagram illustrates an example of object recognition processing performed by the object recognition device 900.
[0092] In step S1001, the object recognition device 900 performs a fitting operation. More specifically, the object recognition device 900 calculates the error between the group of points constituting the object, which was cut out in step S702, and the arbitrarily created prototype shape model (e.g., a group of box points). For example, the object recognition device 900 uses an algorithm that minimizes the distance between the group of points (e.g., ICP) to estimate the pose of the object. At this time, the object recognition device 900 decomposes the distance between a point of the prototype shape model and the nearest point of the object's group of points into the x-axis, y-axis, and z-axis directions for calculation, and sets the average of the distances calculated for all points on each axis as the error for each axis.
[0093] In step S1002, the object recognition device 900 determines whether the termination condition is met. If the termination condition is met, the object recognition device 900 transfers the processing to step S402; if the termination condition is not met, the processing transfers to step S1003. The termination condition may be, for example, that the error of each axis calculated in step S1002 is below a threshold, or that the processing in step S1001 has been performed a predetermined number of times.
[0094] In step S1003, the object recognition device 900 adjusts the parameters, causing the processing to transfer to step S1001. For example, the object recognition device 900 changes the value of the parameters of the prototype shape model in the x-axis, y-axis, and z-axis directions where the error is greatest (e.g., halving or doubling the value).
[0095] In this way, the object recognition device 900 estimates the pose and size of the object by repeatedly fitting and adjusting parameters.
[0096] According to this embodiment, it is possible to simultaneously estimate the size and pose of an object, and detect objects of sizes that are not allowed in the object scene.
[0097] (IV) Fourth Embodiment
[0098] This embodiment differs from the second embodiment primarily in that it alerts the user to foreign objects when it detects objects of an unacceptable size (denoted as "foreign objects") during the estimation and sorting of object dimensions. In this embodiment, the same reference numerals are used for structures identical to those in the second embodiment, and their descriptions are omitted.
[0099] Figure 11 This diagram illustrates an example of the processing performed by the object recognition device 1100 of this embodiment.
[0100] The object recognition device 1100 performs input processing 1110, estimation processing 1120, sorting processing 1130, foreign object detection processing 1140, and output processing 1150.
[0101] In the input processing 1110, the object recognition device 1100 inputs scene information. The scene information includes sensor information 1111. The sensor information 1111 includes color images, point clusters, etc., acquired by the sensor 601.
[0102] In the estimation process 1120, the object recognition device 1100 estimates the size of the objects contained in the object scene and generates an estimation result 1121. At this time, the object recognition device 1100 can also estimate the pose of the objects.
[0103] In the sorting process 1130, the object recognition device 1100 uses an estimation result 1121 for the object scene (scene as a whole) to sort objects of permitted sizes within the object scene and generates a sorting result 1131. The sorting result 1131 includes information 1132 indicating objects of permitted sizes within the object scene and information 1133 indicating foreign objects. For example, the object recognition device 1100 sorts permitted objects and foreign objects within the object scene by voting on the sizes of multiple objects.
[0104] In the foreign object detection processing 1140, when a foreign object is detected in the sorting processing 1130, the object recognition device 1100 determines the foreign object in the object contained in the sensor information 1111 and generates foreign object information 1141 as information representing the determined foreign object. Furthermore, the foreign object information 1141 can be an ID or similar information that can identify objects of sizes not permitted in the object scene.
[0105] In the output processing 1150, the object recognition device 1100 outputs a recognition result 1151 containing foreign object information 1141. The recognition result 1151 is displayed on the output device 305. Alternatively, the recognition result 1151 can also be displayed on a user's tablet terminal, a monitor installed in a warehouse, or the like. Based on this display, the user can identify and remove foreign objects from the safe deposit box 211.
[0106] In addition, the object recognition device 1100 may also, on the basis of or instead of presenting the recognition result 1151 to the user, present the recognition result 1151 to the robot 210, which will then grasp and remove the foreign object, thereby maintaining the consistency of the objects in the storage box 211.
[0107] In this embodiment, it is determined whether there is an object of a size that should not be present in the storage box, and the determination result is output. According to this embodiment, for example, when storing an object in the storage box, even if an object mistakenly enters the storage box due to human error, the object can be removed from the storage box.
[0108] (V) Fifth Embodiment
[0109] This embodiment differs from the second embodiment mainly in that it can store a variety of objects in the storage box. In this embodiment, the same reference numerals are used for structures that are the same as in the second embodiment, and their descriptions are omitted.
[0110] Figure 12 This diagram illustrates an example of the processing performed by the object recognition device 1200 of this embodiment.
[0111] The object recognition device 1200 performs input processing 1210, estimation processing 1220, sorting processing 1230 and output processing 1240.
[0112] In the input processing 1210, the object recognition device 1200 inputs scene information. The scene information includes sensor information 1211 and presence information 1212. The sensor information 1211 includes color images, point clusters, etc., acquired by the sensor 601. The presence information 1212 includes information indicating the number of objects of the same type present in the object scene (in this example, "2").
[0113] In the estimation process 1220, the object recognition device 1200 estimates the size of the objects contained in the object scene and generates an estimation result 1221. At this time, the object recognition device 1200 can also estimate the pose of the objects.
[0114] In the sorting process 1230, the object recognition device 1200 uses the estimation result 1221 for the object scene (the scene as a whole) and the existence information 1212 to sort objects of an allowed size in the object scene and generate a sorting result 1231.
[0115] For example, if information 1212 is "2", the object recognition device 1200 selects the top two sizes of objects as objects of the allowed sizes in the object scene (the first size object and the second size object) through a vote of multiple object sizes, and classifies the remaining sizes as foreign objects. Furthermore, regarding the object recognition device 1200, even if the top two sizes of objects are present, if the number of objects of a certain size is less than a threshold (e.g., three), objects of that size can be classified as foreign objects.
[0116] In this example, sorting result 1231 includes information 1232 and 1233 indicating objects permitted in the object scene. Incidentally, sorting result 1231 may also include information indicating foreign objects.
[0117] In the output processing 1240, the object recognition device 1200 generates and outputs a recognition result 1241 based on the sorting result 1231. The recognition result 1241 is displayed on the output device 305. Alternatively, the recognition result 1241 can also be displayed on a user's tablet terminal, a monitor installed in a warehouse, or the like. Based on this display, the user can identify and remove foreign objects from the storage box 211.
[0118] In this embodiment, by adding the number of dimensions of the objects in the storage box as input, multiple dimensions are allowed, and the estimated result is selected. Therefore, even if there are more than two kinds of objects in the storage box, it is possible to detect objects of sizes that are not allowed in the storage box.
[0119] (VI) Notes
[0120] The above-described embodiments include, for example, the following.
[0121] The above embodiments describe the application of the present invention to an object recognition device, but the present invention is not limited thereto and can be widely applied to various other systems, devices, robots, methods, and programs.
[0122] Furthermore, in the above embodiments, part or all of the program can be installed from a program source onto a device such as a computer that implements the object recognition device. The program source can be, for example, a program distribution server connected via a network or a computer-readable recording medium (e.g., a non-temporary recording medium). Additionally, in the above description, two or more programs can be implemented as a single program, or a single program can be implemented as two or more programs.
[0123] Furthermore, in the above embodiments, the illustrated and explained screens are just examples. As long as the information received is the same, any design is acceptable.
[0124] Furthermore, in the above embodiments, the illustrated and explained screens are just examples. As long as the information displayed is the same, any design is acceptable.
[0125] Furthermore, in the above embodiments, the use of the average value as a statistical value is described, but the statistical value is not limited to the average value, and may also be other statistical values such as the maximum value, minimum value, the difference between the maximum and minimum values, the most frequent value, the median value, and the standard deviation.
[0126] Furthermore, in the above embodiments, the output of information is not limited to display on a monitor. The output of information may be voice output based on a speaker, output to a document, printing on paper media or the like based on a printing device, projection onto a screen or the like based on a projector, or other forms.
[0127] In addition, as described above, the programs, tables, files, and other information that implement each function can be placed in storage devices such as memory, hard disk, SSD (Solid State Drive), or recording media such as IC card, SD card, and DVD.
[0128] The above-described embodiments have, for example, the following characteristic structures. (1)
[0130] The object recognition device (e.g., object recognition device 220, object recognition device 600, object recognition device 900, object recognition device 1100, object recognition device 1200) comprises: an input unit (e.g., input unit 221, input unit 611, circuit, object recognition device, computer) that inputs scene information (object estimation result 111, sensor information 1111, sensor information 1211, color image, dot group, etc.) representing the state of a specified scene in a container (e.g., storage box 211) containing objects; a processing unit (e.g., processing unit 222, circuit, object recognition device, computer) that, based on the scene information input by the input unit, sorts objects existing in the container according to each size, and, based on the number of objects of each size sorted, performs processing to detect objects of sizes that are not allowed to be stored in the container; and an output unit (e.g., output unit 223, circuit, object recognition device, computer) that outputs the result of the processing performed by the processing unit.
[0131] Furthermore, in the above structure, the dimensional information representing the size of objects contained in the specified scene can be either dimensional information estimated by other computers and included in the scene information, or it can be estimated from sensor information contained in the scene information.
[0132] In the above structure, objects of sizes that are not allowed to be stored in the container are detected based on the number of objects of each size. Therefore, even if there is no model of a reference object, objects of sizes that are not allowed to be stored in the container can be excluded from the container. According to the above structure, objects of sizes that are not allowed to be stored in the container can be detected without creating an object model for each object and setting the object's size, color, texture, etc. (2)
[0134] The scene information input by the input unit includes at least one of a color image, a point cluster, a depth image, and a grayscale image. Based on the scene information input by the input unit, the processing unit estimates the size of the objects present in the container, and sorts the objects present in the container according to each size based on the estimated object size (see reference). Figure 7 , Figure 10 wait).
[0135] In the above structure, the size of the object is estimated from at least one of the information in the color image, point group, depth image and grayscale image, and the objects present in the container are sorted according to each size. Therefore, for example, by inputting this information into the object recognition device, the objects present in the container can be sorted according to each size. (3)
[0137] The dimensions of the objects existing in the container are combinations of values of parameters representing the prototype shape of the objects. Based on the scene information input by the input unit, the processing unit sorts the objects existing in the container according to each combination of values of the parameters representing the prototype shape (see reference). Figure 5 , Figure 8 wait).
[0138] In the above structure, objects are sorted using the values of parameters that represent the shape of the prototype. Therefore, compared to sorting objects using the area where the object exists (the area of space occupied by the object), it is easier to compare the size of the objects with each other, thus making it easier to sort the objects. (4)
[0140] When the prototype shape is a box, the parameters representing the prototype shape are vertical, horizontal and height, or vertical and horizontal. When the prototype shape is a cylinder, the parameters representing the prototype shape are radius and height, or radius or height. (5)
[0142] The aforementioned processing unit creates a histogram showing the frequency distribution of objects present in the aforementioned container, and detects objects of sizes that are not permitted to be stored in the aforementioned container based on the number of elements in each bar of the created histogram. This range is a range set for each parameter representing the aforementioned prototype shape, and is a defined range of values for parameters that are determined to be the same (see [reference]). Figure 4 , Figure 7 , Figure 10 wait).
[0143] In the above structure, objects that are determined to have a combination of parameter values that fall within a specified range are considered to be objects of the same size. Therefore, for example, by setting an interface for setting the specified range, the user can set the hyperparameter (allowable error). (6)
[0145] The scene information input by the aforementioned input unit includes sensor information (color image, point cluster, depth image, grayscale image, and combinations thereof) obtained by sensors acquiring the state of the container. Based on the sensor information contained in the scene information input by the aforementioned input unit, the aforementioned processing unit estimates the size of the objects existing in the container, and sorts the objects existing in the container according to each size based on the estimated object size (see reference). Figure 7 , Figure 10 wait).
[0146] In the above structure, the size of the object is estimated from the sensor information and the objects present in the container are sorted according to each size. Therefore, for example, by directly or indirectly connecting the sensor to the object recognition device, it is possible to sort the objects present in the container according to each size. (7)
[0148] The size of the object existing in the container is a combination of values of parameters representing the prototype shape of the object. The processing unit, based on sensor information contained in the scene information input by the input unit, estimates the values of the parameters representing the prototype shape as the size of the object existing in the container (see reference). Figure 7 , Figure 10 wait).
[0149] In the above structure, the size of the object is estimated by using the value of the parameter representing the shape of the prototype. Therefore, compared with the case where the size of the object is estimated by using the area where the object exists, it is easier to compare the size of the objects with each other, thereby making it easier to sort the objects. (8)
[0151] The aforementioned processing unit uses the estimated dimensions of the object to estimate the object's posture, and the aforementioned output unit outputs the object's dimensions and posture estimated by the processing unit (see reference). Figure 7 wait).
[0152] In the above structure, the size and orientation of the object are output. Thus, for example, a robot can use its hand to approach an object of a size that is not allowed to be stored in a container without collision, and after grasping the object, remove the object from the container while avoiding obstacles such as other objects. (9)
[0154] The aforementioned processing unit estimates the object's posture (refer to) by aligning the estimated object's dimensions with the model of the aforementioned prototype shape. Figure 7 wait).
[0155] In the above structure, the object's posture is estimated by aligning the estimated object's size with the model of the prototype shape. Therefore, even without a model of the object as a reference, the object's posture can be estimated by creating a model of the prototype shape. (10)
[0157] The size of the object existing in the container is a combination of parameters representing the prototype shape of the object. The scene information input by the input unit includes sensor information obtained by the sensors from the state of the container. The processing unit repeatedly performs fitting between the object point group obtained by cutting out the object from the sensor information contained in the scene information input by the input unit and the prototype shape model, and adjusts the parameter values of the prototype shape model to estimate the size of the object and the pose of the object (refer to...). Figure 10 wait).
[0158] Based on the above structure, even without a reference object model, the size and pose of an object can be estimated by creating a model with a prototype shape. (11)
[0160] The output unit outputs information indicating the size of objects detected by the processing unit as not permitted to be stored in the container (see reference). Figure 11 wait).
[0161] In the above structure, the output indicates information about objects of a size that are not allowed to be stored in the container. Therefore, users, robots, and others can exclude objects of a size that are not allowed to be stored in the container from the container. (12)
[0163] The input unit inputs presence information indicating the number of types of objects present in the container (e.g., presence information 1212), and the processing unit detects objects of sizes that are not permitted to be stored in the container based on the presence information (see reference). Figure 12 wait).
[0164] In the above structure, objects of sizes that are not allowed to be stored in the container are detected based on the presence information. Therefore, even when objects of various sizes are stored in the container, objects of sizes that are not allowed to be stored in the container can be detected.
[0165] Furthermore, regarding the above structure, appropriate changes, alterations, combinations, or omissions may be made without departing from the spirit of this invention.
[0166] The items contained in a list in the form of “at least one of A, B, and C” are intended to be understood as being able to represent (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Similarly, the items listed in the form of “at least one of A, B, or C” can represent (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0167] Explanation of reference numerals in the attached figures
[0168] 220……Object recognition device, 221……Input unit, 222……Processing unit, 223……Output unit.
Claims
1. An object recognition device comprising: an input section that inputs scene information indicating a state of a predetermined scene in a container that stores objects; a processing section that sorts objects present in the container by each size based on the scene information input by the input section, and performs processing for detecting objects of a size that is not allowed to be stored in the container based on the number of objects of each size that is sorted; and an output section that outputs a result of the processing performed by the processing section, wherein the size of the objects present in the container is a combination of values of parameters that express a prototype shape of the objects, and the processing section sorts the objects present in the container by each combination of values of the parameters that express the prototype shape based on the scene information input by the input section. 2.The object recognition device according to claim 1, wherein at least one of a color image, a point cloud, a depth image, and a gray scale image is included in the scene information input by the input section, and the processing section estimates the size of the objects present in the container based on the scene information input by the input section, and sorts the objects present in the container by each size based on the estimated size of the objects. 3.The object recognition device according to claim 1, wherein in a case where the prototype shape is a box shape, the parameters that express the prototype shape are length, width, and height, or are length and width, and in a case where the prototype shape is a cylindrical shape, the parameters that express the prototype shape are radius and height, or are radius or height. 4.The object recognition device according to claim 1, wherein the processing section creates a histogram that shows a frequency distribution of the objects present in the container in compliance with a range that is set for each parameter that expresses the prototype shape and is a prescribed range that determines to be the same value of the parameter, and detects the objects of the size that is not allowed to be stored in the container based on the number of elements of each bar of the created histogram. 5.The object recognition device according to claim 1, wherein the output section outputs information indicating the objects of the size that is not allowed to be stored in the container detected by the processing section. 6.An object recognition device comprising: an input section that inputs scene information indicating a state of a predetermined scene in a container that stores objects; a processing section that sorts objects present in the container by each size based on the scene information input by the input section, and performs processing for detecting objects of a size that is not allowed to be stored in the container based on the number of objects of each size that is sorted; and an output section that outputs a result of the processing performed by the processing section, wherein the scene information input by the input section includes sensor information obtained by a sensor acquiring the state of the container, the processing section estimates the size of the objects present in the container based on the sensor information included in the scene information input by the input section, and sorts the objects present in the container by each size based on the estimated size of the objects, and the size of the objects present in the container is a combination of values of parameters that express a prototype shape of the objects. The processing section estimates values of parameters representing the prototype shape as sizes of the objects present in the container, based on sensor information included in the scene information input by the input section.
7. The object recognition apparatus according to claim 6, wherein The processing section estimates the posture of the object using the estimated size of the object, The output section outputs the size of the object and the posture of the object estimated by the processing section.
8. The object recognition apparatus according to claim 7, wherein The processing section estimates the posture of the object by alignment of the estimated size of the object with the model of the prototype shape.
9. An object recognition apparatus comprising: an input section that inputs scene information representing a state of a prescribed scene in a container that stores objects; a processing section that sorts objects present in the container by size based on the scene information input by the input section, and performs processing for detecting objects of a size that is not allowed to be stored in the container based on the number of objects of each size that are sorted; an output section that outputs a result based on the processing performed by the processing section, sizes of objects present in the container are combinations of values of parameters representing a prototype shape of the objects, the scene information input by the input section includes sensor information obtained by a sensor acquiring a state of the container, the processing section repeatedly performs fitting between a point cloud of objects obtained by cutting out objects from sensor information included in the scene information input by the input section and a model of the prototype shape, and adjustment of values of parameters of the model of the prototype shape, to estimate sizes of the objects and postures of the objects.
10. An object recognition apparatus comprising: an input section that inputs scene information representing a state of a prescribed scene in a container that stores objects; a processing section that sorts objects present in the container by size based on the scene information input by the input section, and performs processing for detecting objects of a size that is not allowed to be stored in the container based on the number of objects of each size that are sorted; an output section that outputs a result based on the processing performed by the processing section, the input section inputs presence information representing a number of kinds of objects present in the container, the processing section detects objects of a size that is not allowed to be stored in the container based on the presence information.
11. An object recognition method comprising: an input section inputs scene information representing a state of a prescribed scene in a container that stores objects; a processing section sorts objects present in the container by size based on the scene information input by the input section, and performs processing for detecting objects of a size that is not allowed to be stored in the container based on the number of objects of each size that are sorted; and an output section outputs a result based on the processing performed by the processing section, sizes of objects present in the container are combinations of values of parameters representing a prototype shape of the objects, the processing section sorts objects present in the container by each combination of values of parameters representing the prototype shape based on the scene information input by the input section.
12. An object recognition method comprising: The input section inputs scene information indicating a state of a predetermined scene in a container storing objects; The processing section sorts the objects present in the container by each size based on the scene information input by the input section, and performs processing of detecting objects of a size not allowed to be stored in the container based on the number of objects of each size sorted; and The output section outputs a result of the processing performed by the processing section, The scene information input by the input section includes sensor information obtained by a sensor acquiring a state of the container, The processing section estimates the size of the objects present in the container based on the sensor information included in the scene information input by the input section, and sorts the objects present in the container by each size based on the estimated size of the objects, The size of the objects present in the container is a combination of values of parameters representing a prototype shape of the objects, The processing section estimates the values of the parameters representing the prototype shape as the size of the objects present in the container based on the sensor information included in the scene information input by the input section.
13. An object recognition method comprising: The input section inputs scene information indicating a state of a predetermined scene in a container storing objects; The processing section sorts the objects present in the container by each size based on the scene information input by the input section, and performs processing of detecting objects of a size not allowed to be stored in the container based on the number of objects of each size sorted; and The output section outputs a result of the processing performed by the processing section, The size of the objects present in the container is a combination of values of parameters representing a prototype shape of the objects, The scene information input by the input section includes sensor information obtained by a sensor acquiring a state of the container, The processing section repeatedly performs fitting between a point cloud of objects obtained by cutting out the objects from the sensor information included in the scene information input by the input section and a model of the prototype shape, and adjustment of the values of the parameters of the model of the prototype shape, to estimate the size of the objects and a posture of the objects.
14. An object recognition method comprising: The input section inputs scene information indicating a state of a predetermined scene in a container storing objects; The processing section sorts the objects present in the container by each size based on the scene information input by the input section, and performs processing of detecting objects of a size not allowed to be stored in the container based on the number of objects of each size sorted; and The output section outputs a result of the processing performed by the processing section, The input section inputs existence information indicating the number of kinds of objects present in the container, The processing section detects objects of a size not allowed to be stored in the container based on the existence information.
Citation Information
Patent Citations
Work position and orientation recognition device and picking system
JP2020034526A
Component counting device, component counting method, and program
JP2019057211A