Image processing apparatus

By automating the acquisition and processing of multiple learning images using an image processing device, the problem of cumbersome manual operation in existing technologies is solved, the object recognition rate is improved, and the automated construction of the learning model is realized.

CN120957845APending Publication Date: 2025-11-14FANUC LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380096604.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, building learning models to improve object recognition rates based on visual sensors requires a lot of manual work, including changing the configuration of objects and taking multiple images, and providing the correct answer. There is a lack of automated image acquisition methods.

Method used

An image processing device is designed, comprising an image acquisition unit, a robot position acquisition unit, a storage unit, and a feature position determination unit. It automatically acquires and processes multiple learning images and determines the feature positions of objects through the coordinated operation of robot movement and vision sensors.

Benefits of technology

It significantly reduces the workload of operators in image acquisition for learning, improves object recognition rate, and realizes automated construction of learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120957845A_ABST
    Figure CN120957845A_ABST
Patent Text Reader

Abstract

An image processing device is provided with: an image acquisition unit that acquires an image obtained by capturing an image of an object by a vision sensor; a robot position acquisition unit that acquires the position of a robot that can move while gripping the object; a storage unit that stores a position of a feature portion of the object on a first image in which the visual sensor captures an image of the object in a state in which the object is disposed at the first position by the robot; and a feature position specifying unit that specifies the position of the feature part of the object in the second image on the basis of the position of the robot when the robot arranges the object in the first position and the second position, the position of the vision sensor when the first image and the second image are captured, and the position of the feature part in the first image, the second image is an image obtained by capturing an image of the object by the vision sensor in a state in which the object is disposed at the second position by the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an image processing apparatus. Background Technology

[0002] Robotic systems are known to perform the following tasks: detecting the position of an object using a vision sensor and retrieving the object based on the detection result. In such robotic systems, as a method to improve the object recognition rate based on vision sensors, the following method is sometimes used: constructing a learning model using learning images obtained by photographing the object (e.g., see Patent Documents 1-3).

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2020-82315

[0006] Patent Document 2: Japanese Patent Application Publication No. 2019-56966

[0007] Patent Document 3: Japanese Patent Application Publication No. 2016-203293 Summary of the Invention

[0008] The problem that the invention aims to solve

[0009] To build a learning model for improving object recognition accuracy based on vision sensors, multiple training images are needed. Typically, this involves repeatedly performing tasks such as changing the object's configuration to capture multiple images and providing the correct answer regarding the object's location relative to each image. A technique that automates the acquisition of training images is desired.

[0010] Methods for solving problems

[0011] One aspect of this disclosure is an image processing apparatus comprising: an image acquisition unit that acquires an image of an object captured by a vision sensor; a robot position acquisition unit that acquires the position of a robot capable of moving the object; a storage unit that stores the position of a feature portion of the object on a first image, wherein the first image is an image of the object captured by the vision sensor when the robot has positioned the object in a first position; and a feature position determination unit that determines the position of a feature portion of the object on a second image based on the robot's position when the robot has positioned the object in the first position and a second position, the position of the vision sensor when the first image and the second image are captured, and the position of the feature portion on the first image, wherein the second image is an image of the object captured by the vision sensor when the robot has positioned the object in the second position.

[0012] These objects, features, and advantages of the invention will become more apparent from the detailed description of typical embodiments of the invention shown in the accompanying drawings. Attached Figure Description

[0013] Figure 1 This is a diagram illustrating the structure of a robot system that includes an image processing apparatus according to one embodiment.

[0014] Figure 2 It is a functional block diagram related to image processing devices and robot control devices.

[0015] Figure 3 This is a flowchart illustrating the processing of the preparation stage for collecting and learning images in the first embodiment.

[0016] Figure 4 This is a flowchart illustrating the image acquisition process for the first embodiment.

[0017] Figure 5 It is a diagram used to illustrate the calculation of feature locations on an image.

[0018] Figure 6 This is a flowchart illustrating the box removal process during the actual operation of the robot system.

[0019] Figure 7 This is a flowchart illustrating the processing of the preparation stage for collecting and learning images in the second embodiment.

[0020] Figure 8 This is a flowchart illustrating the image acquisition process for the second embodiment. Detailed Implementation

[0021] Next, embodiments of the present disclosure will be described with reference to the accompanying drawings. In the referenced drawings, the same structural or functional parts are labeled with the same reference numerals. For ease of understanding, the scale of these drawings has been appropriately changed. Furthermore, the embodiments shown in the drawings are examples for carrying out the invention, and the invention is not limited to the illustrated embodiments.

[0022] Furthermore, in this specification, when the robot, vision sensor, end effector mounted on the robot, and coordinate system are described only as "position," the concept of their pose is also included unless otherwise specified.

[0023] First Implementation Method

[0024] Figure 1 This diagram illustrates the structure of a robot system 100 including an image processing apparatus 20 according to one embodiment. The image processing apparatus 20 controls a vision sensor 70 and performs image processing functions on images captured by the vision sensor 70. Figure 1 As shown, the robot system 100 includes: a robot 10, a robot control device 50 for controlling the robot 10, a teaching device 40 connected to the robot control device 50, a vision sensor 70, and an image processing device 20. The robot system 100 can, for example, detect objects positioned in the work area using the vision sensor 70, and manipulate the objects using a robotic arm (not shown) mounted on the robot 10.

[0025] In this embodiment, robot 10 is a vertical joint robot, but other types of robots, such as parallel linkage robots or dual-arm robots, can also be used depending on the task objective. Robot 10 can perform the desired task via an end effector mounted on its wrist.

[0026] In this embodiment, the object identified by the vision sensor 70 is assumed to be a three-dimensional object like the box 90. Therefore, a three-dimensional camera capable of acquiring distance images is used as the vision sensor 70. The three-dimensional camera can be a TOF (Time of Flight) camera that captures distance images using the time-of-flight method, or a stereo camera comprising two cameras, etc. In this embodiment, as... Figure 1 As shown, it is assumed that the vision sensor 70 is fixedly mounted on the platform 3 within the workspace. The vision sensor 70 has completed calibration, and its position within the workspace is known in the robot control device 50. The image processing device 20, using the calibration data, can determine the position on the image captured by the vision sensor 70 as its position in a coordinate system (robot coordinate system, etc.) set for the workspace. The calibration data (position information of the vision sensor 70, etc.) is stored in the storage device 22 of the image processing device 20 (see reference). Figure 2)middle.

[0027] The robot system 100 can measure the position of the box 90, which is the object, loaded on the pallet 2, using the image processing device 20 (vision sensor 70), and then perform the operation of retrieving the box 90 using the robot 10. When measuring the position of the box using the vision sensor 70, if there are clear features such as printed words or patterns on the surface of the box, the position of the box can be determined by detecting them. On the other hand, if the box does not have such features, it is necessary to measure the shape of the box in order to use its shape as a feature. However, the shape of the box is sometimes difficult to measure for the following reasons.

[0028] The shape of the box can sometimes be difficult to identify due to the background.

[0029] • When boxes are in close contact with each other, it can sometimes be difficult to discern their shape.

[0030] • Sometimes, the tape on the surface of the box is mistaken for its shape.

[0031] To address the aforementioned problems and improve the object recognition rate, the image processing apparatus 20 of this embodiment has the function of learning the relationship between multiple images of an object and the feature positions on these images. The image processing apparatus 20 of this embodiment automates the acquisition of multiple images of an object and the determination of the feature positions in each of these images.

[0032] The image processing device 20 can be configured as a dedicated image processing device, or it can be configured as a general-purpose information processing device such as a PC (personal computer). The image processing device 20 can have the hardware structure of a general-purpose computer, which includes a processor 21, memory (ROM, RAM, non-volatile memory, etc.), storage device 22, input / output interface, network interface, etc. (see reference). Figure 2 The image processing device 20 may also include a display unit and an operation unit as hardware components.

[0033] exist Figure 1 The example shown is that the image processing device 20 is configured as a device different from the robot control device 50, but the function of the image processing device 20 can also be integrated into the robot control device 50.

[0034] The robot control device 50 controls the actions of the robot 10 according to the action program or instructions from the teach pendant 40. The robot control device 50 may have the hardware structure of a general computer, which includes a processor, memory (ROM, RAM, non-volatile memory, etc.), storage device, operation unit, input / output interface, network interface, etc.

[0035] The teaching pendant 40 is connected to the robot control device 50 and also connected to the image processing device 20 via the robot control device 50. The teaching pendant 40 is used as an operating terminal for teaching the robot 10's motion programs (programming), teaching programs related to image capture and processing by the vision sensor 70 (hereinafter also referred to as vision programs), displaying images captured by the vision sensor 70, and setting various other parameters. The teaching pendant 40 can be a teaching operation panel or a tablet terminal, etc. The teaching pendant 40 can have a hardware structure similar to a general computer, including a processor, memory (ROM, RAM, non-volatile memory, etc.), storage device, operation unit, display unit 41, input / output interface, network interface, etc. (see reference). Figure 2 The display unit 41 is, for example, composed of a liquid crystal display.

[0036] Figure 2 This is a functional block diagram related to the image processing device 20 and the robot control device 50. For example... Figure 2 As shown, the robot control device 50 includes a motion control unit 151. The motion control unit 151 controls the motion of the robot 10 according to a motion program or instructions from the teach pendant 40. The motion program may include a vision program. The robot control device 50 also includes a servo control unit (not shown) that executes servo control of the servo motors for each axis according to the instructions generated by the motion control unit 151 for each axis.

[0037] like Figure 2 As shown, the image processing apparatus 20 includes: a vision sensor control unit 121, an image acquisition unit 122, a camera position acquisition unit 123, a robot position acquisition unit 124, a feature position registration unit 125, a feature position determination unit 126, a detection unit 127, a learning unit 128, a learning image acquisition unit 129, and a setting unit 130. Figure 2 As shown, these functional blocks can be implemented by the processor 21 executing software.

[0038] The vision sensor control unit 121 controls the operation of the vision sensor 70. For example, the vision sensor control unit 121 can receive operation commands from the robot control device 50 (motion control unit 151) and control the vision sensor 70. The image acquisition unit 122 acquires images captured by the vision sensor 70.

[0039] The camera position acquisition unit 123 provides the function of acquiring the position of the vision sensor 70. In this embodiment, the vision sensor 70 is fixedly disposed in the work space, and the camera position acquisition unit 123 provides, for example, the position / posture of the vision sensor 70 in the work space that is pre-stored in the storage device 22.

[0040] The robot position acquisition unit 124 provides the function of acquiring the position of the robot 10. The robot position acquisition unit 124 can acquire the position / pose of the robot 10 from the motion control unit 151.

[0041] The feature position registration unit 125 accepts input during the preparation stage for collecting learning images to register the position of a feature portion of the box 90 on an image (hereinafter also referred to as a reference image) taken of the box 90 placed at a reference position. In this embodiment, the feature portion of the box 90 is its shape. The feature position registration unit 125 may have the function of providing a user interface that displays the reference image on the display unit 41 of the teaching pendant 40 and accepts operations to specify the position of the feature portion on the reference image via the operation unit of the teaching pendant 40.

[0042] The feature position determination unit 126 provides the following function: based on the position of the feature portion registered relative to the reference image and the position of the robot 10 when the robot 10 positions the object in the second position, it determines the position of the feature portion of the object in an image (second image), wherein the image is obtained by the vision sensor 70 capturing the object in the state where the robot 10 has positioned the object in the new position (second position). Details regarding the function of the feature position determination unit 126 will be described later.

[0043] The learning image acquisition unit 129 is responsible for acquiring multiple learning images while changing various conditions related to shooting. Each learning image is saved as a completed annotation image with the feature positions calculated by the feature position determination unit 126.

[0044] The learning unit 128 uses multiple learning images collected as described above to learn and construct a learning model for estimating the position of the feature unit on any image obtained from the photographed object.

[0045] The detection unit 127 has the function of detecting objects from images captured by the vision sensor 70. For example, the detection unit 127 has the function of determining the position of an object from any image obtained by the vision sensor 70 using a learning model constructed by the learning unit 128.

[0046] The setting unit 130 is responsible for accepting various settings related to image processing. The setting unit 130 can be configured to accept user input for setting conditions related to shooting for obtaining learning images. Shooting-related conditions may include one or more of the following: the configuration mode of the object (box 90) and exposure conditions.

[0047] Reference Figure 3 and Figure 4The processing used to collect learning images is described. The processing for collecting learning images includes: processing in the preparation phase (…). Figure 3 ), and on the reference image obtained by photographing the object positioned at the reference location, register the characteristic positions of the object; learn to use image processing ( Figure 4 This tool is used to save multiple learning images as completed annotated images while changing various conditions related to shooting.

[0048] Figure 3 This is a flowchart illustrating the preparation phase of the process. The preparation phase involves processing the acquired learning images later. Figure 4 The processing of information set in the document serves as a reference for determining the location of features. Therefore, in this specification, during the preparation stage of the process, the position of the housing 90 is also referred to as the reference position, and the image obtained by photographing the housing 90 positioned at the reference position is also referred to as the reference image.

[0049] First, the robot 10, through user operation via the teaching pendant 40 or according to pre-taught instructions, positions a box 90 within the field of view of the vision sensor 70 (step S11). The position (reference position) of the positioned box 90 is then obtained as the position / pose of the robot 10. The robot position acquisition unit 124 obtains the position (reference position) of the box 90 as the position / pose of the robot 10.

[0050] Next, the image processing device 20 captures an image (reference image) of the box 90 by taking a picture of the box 90 using the vision sensor 70 (step S12). Here, the operator can input a shooting command to the vision sensor 70 through the teaching device 40 to perform the shooting.

[0051] Next, the feature position registration unit 125 accepts user operations that specify the feature position (shape position) of the box 90 on the reference image (step S13). Here, the feature on the image is the shape of the box, which is the object. For example, the feature position registration unit 125 can provide a graphical user interface that displays the reference image on the display unit 41 of the teaching device 40 and accepts operations that specify the shape position of the box 90 on the reference image. In this case, for example, the operator specifies the shape position of the box by using a pointing device to draw a line along the outline of the box 90 projected in the image. The feature position registration unit 125 associates the reference image with the feature position specified on the reference image, for example, and stores it in the storage device 22.

[0052] Next, during the execution phase of acquiring the learning image, the operator registers the configuration mode, which serves as the configuration condition for configuring the robot 10 configuration box 90, with the robot control device 50 (step S14). The configuration mode of the box corresponds to the conditions related to capturing the learning image. The registration of the configuration mode with the robot control device 50 can be managed by the setting unit 130 of the image processing device 20. In this case, the setting unit 130 of the image processing device 20 can display a user interface for setting the configuration mode on the display unit 41 of the teaching device 40, and accept the setting of the configuration mode via the operation unit of the teaching device 40. Alternatively, the registration of the configuration mode can also be managed by the robot control device 50.

[0053] The configuration mode can be various information used to change the configuration state of one or more boxes 90. For example, the configuration mode can be information indicating one or more positions for configuring the boxes 90. Alternatively, the configuration mode can also be information specifying the height or number of layers of the stacking boxes. In this case, the robot control device 50 can calculate the position and posture of the robot 10 for configuring the boxes 90 at a specified height or number of layers based on model data, such as model data representing the size and shape of the boxes. The configuration mode can be stored, for example, in the storage unit of the robot control device 50.

[0054] Next, the operator registers the exposure conditions used when taking multiple images during the execution phase of obtaining learning images to the image processing device 20 (step S15). Exposure conditions may include multiple set values ​​related to exposure time and illumination intensity. Exposure conditions are stored, for example, in the storage device 22 of the image processing device 20. Registration of exposure conditions can be managed by the setting unit 130 of the image processing device 20. In this case, the setting unit 130 of the image processing device 20 can prompt the user interface for setting exposure conditions on the display unit 41 of the teaching device 40, and accept the setting of exposure conditions via the operation unit of the teaching device 40. Alternatively, registration of exposure conditions can also be managed by the robot control device 50.

[0055] Next, the operator instructs robot 10 to begin learning (step S16). Following this instruction, the learning process begins. Figure 4 The image processing for obtaining the learning image is shown. The preparation phase ends here.

[0056] Next, refer to Figure 4 The processing for acquiring learning images (learning image acquisition processing) will be described. The learning image acquisition unit 129 (processor 21) cooperates with the robot control device 50 and the teaching device 40 to manage the execution of the learning image acquisition processing.

[0057] First, robot 10 configures box 90 at a specified position within the configuration mode (step S21). Furthermore, the number of boxes 90 configured here (i.e., the number of boxes captured in a single image) can be one or more. Robot position acquisition unit 124 stores the position / pose of robot 10 when configuring the boxes according to the configuration mode in, for example, storage device 22 (step S22). The robot position / pose stored here is used as information indicating the position of the box 90 when acquiring the learning image.

[0058] Next, the image processing device 20 (vision sensor control unit 121) sets the exposure conditions according to the preset exposure conditions (step S23). Then, the image processing device 20 (vision sensor control unit 121) takes a picture of the configured box 90 through the vision sensor 70 (step S24). The image acquisition unit 122 acquires the captured image.

[0059] Next, the feature position determination unit 126 determines the position (box outline position) on the image (hereinafter also referred to as the second image) obtained by the feature unit in step S24 based on the feature position of the box 90 registered in step S13 of the preparation phase and the position / posture of the robot 10 recorded in step S22 (step S25). Specifically, the feature position determination unit 126 determines the feature position of the box 90 on the second image based on the position / posture of the robot 10 when the robot 10 positions the box 90 in the reference position, the position / posture of the robot 10 when the robot 10 positions the box 90 according to the configuration mode, the position / posture of the vision sensor 70 when the reference image and the second image are captured respectively, and the registered feature positions on the reference image. Furthermore, in this embodiment, the vision sensor 70 is fixedly configured; therefore, the position / posture of the vision sensor 70 is the same when the reference image and the second image are captured respectively.

[0060] The image processing apparatus 20 (learning image acquisition unit 129) registers the learning image as an annotated learning image by appending the feature position information obtained in step S25 to the learning image obtained in step S24 (step S26). The annotated learning image is stored, for example, in the storage device 22.

[0061] Next, the image processing device 20 (learning image acquisition unit 129) determines whether the shooting has ended for all specified exposure conditions (step S27). If the shooting has not ended for all exposure conditions (S27: No), the image processing device 20 (learning image acquisition unit 129) performs the processing from step S23 again for the next exposure condition.

[0062] If all exposure conditions have been captured (S27: Yes), the process proceeds to step S28. In step S28, the image processing device 20 (learning image acquisition unit 129) determines whether the action of configuring the box 90 and capturing images has been performed for all specified configuration modes. If capturing images in all configuration modes has not been completed (S28: No), the process is executed again from the action of configuring the box 90 to the next configuration position by the robot 10 according to the configuration mode (step S21).

[0063] If shooting is completed for all configuration modes (S28: Yes), the collection of learning images is complete, and processing proceeds to step S29. Through the above, multiple learning images are obtained according to all configuration modes and exposure conditions. Specifically, the number of learning images is obtained by multiplying the number of configuration positions included in the configuration mode by the number of exposure conditions.

[0064] In step S29, the image processing apparatus 20 (learning unit 128) performs learning using the annotated learning images collected as described above. For example, the learning unit 128 performs learning through supervised learning, a type of machine learning. Deep learning methods can also be incorporated into the learning process. The learning unit 128 takes the learning images as input and uses a neural network (NN) or a convolutional neural network (CNN) as an inferrer to learn training data that correctly identifies the feature locations for the image. Thus, a learning model for inferring feature locations on any image reflected by the box 90, which is the object, can be constructed.

[0065] Next, the image processing device 20 (learning unit 128) registers the learning model constructed through learning (step S30). The learning model is stored, for example, in the storage device 22. After the above processing, the image acquisition processing for learning ends.

[0066] Here, refer to Figure 5 The calculation of feature positions on the image performed by the feature position determination unit 126 (i.e., the determination of the shape position in step S25) will be explained. Figure 5 The diagram illustrates the states where robot 10, via manipulator 31, positions box 90 in a reference position and in a position according to a specified configuration pattern. Furthermore, a magnetic manipulator 31 is used here. For clarity, symbol 90a represents a box placed in the reference position, and symbol 90b represents a box positioned according to the configuration pattern.

[0067] Image IG1 represents in Figure 5Examples of images captured by the vision sensor 70 under such a configuration of box 90. Furthermore, boxes 90a and 90b are actually captured as different images, but are shown here in a single image IG1 for illustrative purposes. In image IG1, the dashed line indicated by symbol 91 represents the external shape position registered by the operator relative to box 90a. In image IG1, the dashed line indicated by symbol 92 represents the feature position of box 90b determined by the feature position determination unit 126.

[0068] exist Figure 5 middle,

[0069] • C represents the position of the camera coordinate system CS2 as observed from the reference coordinate system (robot coordinate system) CS1 of robot 10.

[0070] R1 represents the reference position of robot 10 as observed from the reference coordinate system CS1 when robot 10 places box 90 in the reference position.

[0071] ·V1 n This indicates the external shape and position of the reference object (box 90a positioned at the reference position) as observed from the camera coordinate system CS2.

[0072] ·P n This indicates the external position of the box 90 (90a) as observed from the reference position (reference coordinate system CS1) of the robot 10.

[0073] When the line length (number of pixels) of the outline is 100, n is 1 to 100.

[0074] The feature position of the box 90 on the image is registered by the feature position registration unit 125 (registration of the shape position in step S13), that is, the shape position V1 on the camera coordinate system CS1. n Therefore, the external position V1 n It is known. Furthermore, the position C of the camera coordinate system CS2, observed from the reference coordinate system (robot coordinate system) CS1 of robot 10, is known based on calibration data. Additionally, the reference position R1 of robot 10 can be obtained from robot 10 (robot control device 50).

[0075] The feature location determination unit 126 can determine the feature location using the following mathematical formula (1), based on C, R1, and V1. n Find the external shape position P on the reference coordinate system CS1. n In addition, C, R1, V1 n P n It is represented by a homogeneous transformation matrix.

[0076] R1 -1 *C*V1 n=P n …(1)

[0077] In addition, Figure 5 middle,

[0078] R2 represents the position of robot 10 as observed from the reference coordinate system CS1 when robot 10 is configured in the configuration box 90 (90b) during the execution phase of acquiring the learning image.

[0079] ·V2 n This represents the external position of box 90 (90b) as observed from the camera coordinate system CS2 on the learning image.

[0080] V2 n The ultimate goal is to determine the shape and position of the learning image on box 90 (90b).

[0081] The feature location determination unit 126 can determine the feature location using the following mathematical formula (2) based on R2 and P. n Find V2 n In addition, R2 and V2 n It is represented by a homogeneous transformation matrix.

[0082] C -1 *R2*P n =V2 n …(2)

[0083] Furthermore, when the robotic arm 31 holds boxes 90a and 90b, it holds the same position (e.g., the center position of the upper surface of boxes 90a and 90b).

[0084] The determined external shape position V2 n Attach to the learning image and save it as a completed annotated image.

[0085] The image processing device 20 obtains the position of the box 90 as the position of the robot 10, and therefore can automatically obtain the shape position V2 in the image for computational learning. n The required information is as follows: The operator only needs to register the feature positions (shape information) and the conditions related to the shooting with the robot control device 50 or the image processing device 20, and the acquisition of multiple learning images and the determination of feature positions on these learning images will be automatically performed. Therefore, according to this embodiment, the operator's burden in acquiring learning images is greatly reduced.

[0086] Furthermore, these multiple learning images can be used to construct learning models for improving object recognition rates.

[0087] Figure 6This is a flowchart illustrating the box retrieval process during the actual operation of the robot system 100. The box retrieval process is primarily executed under the control of the processor of the robot control unit 50. First, the robot control unit 50 uses the image processing unit 20 to have the vision sensor 70 capture an image of the box 90 (step S31). The image processing unit 20 uses a learning model to detect the position (shape) and number of the boxes 90 from the captured image (step S32). The image processing unit 20 provides the detected position and number of the boxes 90 to the robot control unit 50.

[0088] The robot control device 50 determines whether more than one box has been detected (step S33). If more than one box has been detected (S33: Yes), the robot control device 50 performs the action of retrieving the detected box 90 using the robot 10 (step S34). For all detected boxes, the robot 10 is used to retrieve the box 90 (S35: No). If all detected boxes have been retrieved (S35: Yes), a new box 90 is supplied on the tray 2, and the processing from step S31 continues for the newly supplied box 90. If no box is detected in step S33 (S33: No), the box retrieval process ends.

[0089] Second Implementation Method

[0090] The robot system having the image processing apparatus of the second embodiment will now be described. The structure of the robot system of the second embodiment is the same as that of the robot system 100 of the first embodiment, except that the vision sensor 70 is disposed in the movable part (the forearm end, etc.) of the robot 10. Furthermore, the structure of the functional blocks of each device in the second embodiment is similar to... Figure 2 The functional block diagrams shown are common. Therefore, refer to... Figure 2 The functional block diagram of the second embodiment uses the same symbols as those used in the first embodiment to describe the functions of each device.

[0091] In the second embodiment, the vision sensor 70 is also calibrated, therefore, the position / pose of the vision sensor 70 relative to the coordinate system (e.g., the flange coordinate system) set for the robot 10 is known. Furthermore, the position / pose of the flange coordinate system relative to the reference coordinate system (robot coordinate system) set for the robot 10 is known. Therefore, by obtaining the position / pose of the robot 10, the position / pose of the vision sensor 70 relative to the reference coordinate system (robot coordinate system) set for the robot 10 can be determined. In this embodiment, the camera position acquisition unit 123 acquires the position / pose of the vision sensor 70 based on the position / pose of the robot 10.

[0092] Reference Figure 7 and Figure 8 The processing for collecting learning images in the second embodiment will be explained. Figure 7 This is a preparatory stage process for registering the feature positions of an object on a reference image obtained by photographing the object positioned at a reference location, similar to the process in the first embodiment. Figure 3 The corresponding processing. Figure 8 This involves processing multiple learning images into a completed annotated learning image while simultaneously changing various conditions related to the shooting process, which is equivalent to the first embodiment. Figure 4 The processing. In Figure 7 as well as Figure 8 In this document, the same step numbers are used for processes that are the same as those in the first embodiment, and their descriptions are omitted or simplified.

[0093] like Figure 7 As shown, in the preparation phase, firstly, robot 10, through user operation via teaching device 40 or according to pre-taught instructions, positions a box 90 in the work area (step S11'). A vision sensor 70 is mounted on robot 10; the field of view of vision sensor 70 is not fixed, therefore, robot 10 can position the box 90 anywhere within the work area. Then, the box 90 is photographed using the vision sensor 70 mounted on robot 10 (step S12). Alternatively, the operator can operate teaching device 40 to position robot 10 in a position capable of photographing the box 90, and input a photographing command to vision sensor 70, thereby performing the photographing. The processing content of steps S13 to S15 is similar to... Figure 3 The processing is the same as in the first embodiment shown, therefore, the description is omitted.

[0094] In this embodiment, the vision sensor 70 is disposed at the forearm of the robot 10 (e.g., a robotic arm device). By controlling the robot 10, the position / pose of the vision sensor 70 can be positioned at any position / pose. Therefore, in this embodiment, the position / pose conditions of the vision sensor 70 can be set as conditions related to capturing learning images. This allows for changing the position / pose of the vision sensor 70 to obtain more learning images. In step S15a, the operator registers the position / pose conditions of the vision sensor 70 with the image processing device 20. In this case, the setting unit 130 of the image processing device 20 can display a user interface for setting the position / pose conditions of the vision sensor 70 on the display unit 41 of the teaching device 40, and the setting of the position / pose conditions of the vision sensor 70 can be handled via the operation unit of the teaching device 40. Alternatively, the registration of the position / pose conditions of the vision sensor 70 can also be managed by the robot control device 50.

[0095] Then, the operator instructs robot 10 to begin learning (step S16). Following this instruction, the learning process begins. Figure 8 The image processing for obtaining the learning image is shown. The preparation phase ends here.

[0096] Figure 8 The image acquisition processing for learning in the second embodiment shown is configured to... Figure 4 The learning image acquisition process of the first embodiment shown is supplemented with a loop process that repeatedly takes pictures according to the position / pose conditions of the vision sensor 70. Therefore, the loop process of repeatedly taking pictures according to the position / pose conditions of the vision sensor 70 will be described here. After the exposure conditions are set in step S23, the learning image acquisition unit 129 instructs the robot control device 50 to position the vision sensor 70 according to the position / pose conditions of the vision sensor 70 set in step S15a (step S23a).

[0097] Then, after photographing the box and determining its shape and position (steps S24-S25), the learning image acquisition unit 129 determines whether the photograph was taken with the position and pose of all the vision sensors 70 registered in step S15a (step S25a). If the photograph was not taken with the position and pose of all the registered vision sensors 70 (S25a: No), the process returns to step S23a, and the vision sensors 70 are positioned at the next registered position / pose. If the photograph was taken with the position / pose of all the registered vision sensors 70 (S25a: Yes), the process proceeds to step S26. In step S26, the images obtained for all camera positions / poses and the shape and position determined for them are saved as a completed annotation image. The processing after step S26 is the same as in the first embodiment (…). Figure 4 Since they are the same, the explanation is omitted.

[0098] Furthermore, while describing an example of repeatedly acquiring images while changing the position / pose of the vision sensor 70 according to the conditions of its position / pose pre-registered by the operator in step S15a, the image acquisition unit 129 can be configured to automatically set the conditions of the position / pose of the vision sensor 70. For example, the image acquisition unit 129 can set multiple shooting positions within the workspace and determine the pose of the vision sensor 70 based on the position / pose of the robot 10 when configuring the box 90 (i.e., the configuration information of the box 90), so that the box 90 enters the field of view of the vision sensor 70 at each shooting position. In this case, the operator may not need to register the conditions of the position / pose of the vision sensor 70 (step S15a).

[0099] Through the above processing, multiple learning images can be collected based on all the conditions of the configuration mode, exposure conditions, and the position / pose conditions of the vision sensor. Specifically, when the number of configuration positions included in the configuration mode is set to K1, the number of exposure conditions is set to K2, and the number of position / pose conditions of the vision sensor is set to K3, a number of learning images represented by K1×K2×K3 are obtained.

[0100] Furthermore, in this embodiment, during the preparation stage ( Figure 7 The position of the vision sensor 70 in the case of acquiring the reference image and the execution phase of acquiring the learning image ( Figure 8 The position of the vision sensor 70 is different. In this case, the feature position determination unit 126 can calculate the shape position P on the reference coordinate system CS1 by the following mathematical formula (3). n The position of the camera coordinate system CS2, observed from the reference coordinate system (robot coordinate system) CS1 of robot 10 when capturing the reference image, is set as C1, and other symbols R1, V1, etc. n Same as above. Position C1 can be obtained based on the position / pose of robot 10 configured in reference position box 90. Camera position acquisition unit 123 can calculate and provide position C1.

[0101] R1 -1 *C1*V1 n =P n …(3)

[0102] Furthermore, the feature location determination unit 126 can calculate the external position V2 of the box 90 in the learning image using the following mathematical formula (4). n The position of the camera coordinate system CS2, observed from the reference coordinate system (robot coordinate system) CS1 of robot 10 when capturing images for learning purposes, is set as C2. Other symbols include R2 and P. n V2 n Similar to the above, position C2 can be obtained based on the position / pose of the robot 10 configured according to the configuration mode. The camera position acquisition unit 123 can calculate and provide position C2.

[0103] C2 -1 *R2*P n =V2 n …(4)

[0104] Therefore, in this embodiment, as in the first embodiment, the acquisition of multiple learning images and the determination of feature positions on these learning images are performed automatically. Thus, according to this embodiment, the operator's burden in acquiring learning images is significantly reduced. In particular, in the second embodiment, in addition to the configuration mode conditions and exposure conditions, the position / pose conditions of the visual sensor are also used as shooting-related conditions. Therefore, more learning images can be acquired.

[0105] Furthermore, these multiple learning images can be used to construct learning models for improving object recognition rates.

[0106] In the embodiments described above, an example of the robot 10 configuring the boxes according to a configuration pattern specified by the operator is described as a condition related to the box configuration in the learning image acquisition process. However, there can be various examples of box configuration conditions. For example, there may be an example where the robot 10 randomly selects a location within the work area to configure the boxes. In this case, the learning image acquisition unit 129 of the robot control device 50 or the image processing device 20 is responsible for processing the random selection of box locations within the work area. By applying the above-described learning image acquisition process while changing the box configuration in this way, multiple learning images can be acquired.

[0107] Alternatively, the operator can pre-configure multiple boxes, and the robot 10 can acquire and capture images of its position while partially removing the boxes. By sequentially removing the boxes from the configured set, the robot 10 can obtain multiple learning images showing the changes in the box configuration. In this case, the initial box configuration position information is registered in the robot control device 50 or the image processing device 20. Based on the position / posture information of the robot 10 when removing the boxes, the robot control device 50 or the image processing device 20 identifies the position of the removed boxes, thereby enabling the continuous recognition of the box configuration. By applying the aforementioned learning image acquisition processing while the box configuration changes, multiple learning images can be acquired.

[0108] The above embodiments are structural examples of collecting learning images that correspond to the images of boxes as objects and their external shapes as feature locations. The structures of the above embodiments can be applied to the automatic collection of various types of learning images that correspond to the images of objects and their feature information as learning data.

[0109] Figure 2The functional configuration shown in the functional block diagram is an example, and various variations of the functional configuration are possible. For example, there may be a structure in which the functions of the image processing device 20 are integrated into the robot control device 50. In this case, the processor of the robot control device 50 performs the functions of the vision sensor control unit 121, the image acquisition unit 122, the camera position acquisition unit 123, the robot position acquisition unit 124, the feature position registration unit 125, the feature position determination unit 126, the detection unit 127, the learning unit 128, the learning image acquisition unit 129, and the setting unit 130. Alternatively, there may be a structure in which the functions of the image processing device 20 are configured in the teaching pendant 40. Alternatively, there may be a structure in which the functions of the image processing device 20 are distributed between the robot control device 50 and the teaching pendant 40.

[0110] It is also possible to have Figure 2 A portion of the functions configured within the image processing device 20 are configured within other devices (robot control device 50 or teaching device 40) in the system structure. For example, the functions of the setting unit 130 can be configured within the robot control device 50 or the teaching device 40.

[0111] In the embodiments described above, examples of vision sensors being three-dimensional cameras have been given. However, in applications where object detection and manipulation are possible by understanding the characteristic positions of objects as two-dimensional positional information, a two-dimensional camera can also be used as the vision sensor.

[0112] Figure 2 The functional blocks of the image processing device shown can be implemented by one or more processors of the image processing device executing various software stored in the storage device, or they can be implemented by a structure based on hardware such as ASIC (Application Specific Integrated Circuit).

[0113] The program that performs various processing steps, such as image acquisition processing, in the above embodiments can be recorded on various computer-readable recording media (e.g., semiconductor memory such as ROM, EEPROM, flash memory, magnetic recording media, CD-ROM, DVD-ROM, etc.).

[0114] As explained above, according to each implementation method, multiple learning images can be automatically acquired.

[0115] This disclosure has been described in detail, but it is not limited to the various embodiments described above. Various additions, substitutions, modifications, and partial deletions can be made to these embodiments without departing from the spirit of this disclosure, or without departing from the spirit of this disclosure derived from the claims and their equivalents. Furthermore, these embodiments can also be implemented in combination. For example, in the above embodiments, the order of actions or the order of processes has been shown as an example, but it is not a limitation. Similarly, the use of numerical values ​​or mathematical formulas in the description of the above embodiments is also relevant.

[0116] The following notes further describe the above-described embodiments and variations.

[0117] (Postscript 1)

[0118] An image processing apparatus (20) includes: an image acquisition unit (122) that acquires an image of an object captured by a vision sensor (70); a robot position acquisition unit (124) that acquires the position of a robot (10) capable of moving while holding the object; a storage unit (22) that stores the position of a feature portion of the object on a first image, wherein the first image is an image of the object captured by the vision sensor (70) when the robot (10) has the object positioned in a first position; and a feature position determination unit (126) that determines the position of a feature portion of the object on a second image based on the position of the robot (10) when the object is positioned in the first position and the second position respectively, the position of the vision sensor (70) when the first image and the second image are captured respectively, and the position of the feature portion on the first image, wherein the second image is an image of the object captured by the vision sensor (70) when the object is positioned in the second position by the robot (10).

[0119] (Postscript 2)

[0120] According to the image processing apparatus (20) described in Appendix 1, the vision sensor (70) is fixedly disposed in the work space and calibrated, and the feature position determination unit (126) obtains the position of the vision sensor (70) when the first image and the second image are captured respectively as the position of the vision sensor fixedly disposed in the work space.

[0121] (Note 3)

[0122] According to the image processing apparatus (20) described in Appendix 1, the vision sensor (70) is mounted on the movable part of the robot and is calibrated, and the feature position determination unit (126) obtains the position of the vision sensor (70) when the first image and the second image are captured respectively, based on the position of the robot (10) when the first image and the second image are captured respectively.

[0123] (Postscript 4)

[0124] According to Appendix 1, the image processing apparatus (20) further includes a learning image acquisition unit (129) which, based on conditions related to the shooting, causes the visual sensor (70) to capture a plurality of second images and saves the plurality of second images as a plurality of learning images, corresponding to the positions of the feature portions determined for the plurality of second images respectively.

[0125] (Note 5)

[0126] According to the image processing apparatus (20) described in Appendix 4, the conditions related to the shooting include configuration conditions related to the configuration of the object, and the learning image acquisition unit (129) acquires a plurality of second images by having the vision sensor (70) take pictures in various states in which the robot (10) changes the configuration position of the object according to the configuration conditions.

[0127] (Note 6)

[0128] According to the image processing apparatus (20) described in Appendix 4 or 5, wherein the conditions related to the shooting include exposure conditions, the learning image acquisition unit (129) acquires a plurality of second images by repeatedly shooting with the visual sensor (70) while changing the exposure conditions, wherein the exposure conditions are exposure conditions for the object disposed at the second position.

[0129] (Note 7)

[0130] According to any one of Appendices 4 to 6, the image processing apparatus (20) wherein the conditions related to the shooting include: the position or posture conditions of the visual sensor (70), and the learning image acquisition unit (129) acquires a plurality of second images by repeatedly shooting the visual sensor (70) while changing the position or posture of the visual sensor (70) relative to the object disposed in the second position.

[0131] (Postscript 8)

[0132] The image processing apparatus (20) according to any one of Appendices 4 to 7, wherein the image processing apparatus (20) further comprises: a setting unit (130) that accepts input for setting conditions related to the shooting.

[0133] (Note 9)

[0134] The image processing apparatus (20) according to any one of Appendices 4 to 8, wherein the image processing apparatus (20) further comprises: a learning unit (128) which constructs a learning model for estimating the position of the feature part on any image of the object by performing learning based on the plurality of learning images.

[0135] (Postscript 10)

[0136] The image processing apparatus (20) according to any one of Appendices 1 to 9 further comprises: a feature position registration unit (125) that accepts user operations for specifying the position of a feature of the object on the first image.

[0137] Symbol Explanation

[0138] 10 robots

[0139] 20 Image processing devices

[0140] 21 processors

[0141] 22 Storage devices

[0142] 31. Robotic Arm

[0143] 40 Teaching devices

[0144] 41 Display Section

[0145] 50 Robot Control Device

[0146] 70 Vision Sensors

[0147] 100 Robot Systems

[0148] 121 Vision Sensor Control Unit

[0149] 122 Image Acquisition Unit

[0150] 123 Camera Position Acquisition Unit

[0151] 124 Robot Position Acquisition Unit

[0152] 125 Feature Location Registration Department

[0153] 126 Feature Location Determination Unit

[0154] 127 Inspection Department

[0155] 128 Study Department

[0156] 129 Learning Image Acquisition Department

[0157] 130. Setting Department.

Claims

1. An image processing apparatus, characterized in that, have: The image acquisition unit acquires images of objects captured by the vision sensor; The robot position acquisition unit acquires the position of the robot that is able to move while holding the object; A storage unit stores the position of the feature portion of the object on a first image, wherein the first image is an image of the object captured by the vision sensor while the robot has positioned the object in a first position; and The feature position determination unit determines the position of the feature of the object in the second image based on the position of the robot when the robot places the object in the first position and the second position respectively, the position of the vision sensor when the first image and the second image are captured respectively, and the position of the feature on the first image. The second image is an image of the object captured by the vision sensor when the robot places the object in the second position.

2. The image processing apparatus according to claim 1, characterized in that, The vision sensor is fixedly mounted in the workspace and calibrated. The feature position determination unit obtains the position of the vision sensor when the first image and the second image are captured respectively, and uses it as the position of the vision sensor fixedly configured in the work space.

3. The image processing apparatus according to claim 1, characterized in that, The visual sensor is mounted on the movable part of the robot and is calibrated. The feature position determination unit obtains the position of the vision sensor when the first image and the second image are captured, respectively, based on the position of the robot when the first image and the second image are captured.

4. The image processing apparatus according to claim 1, characterized in that, The image processing apparatus further includes a learning image acquisition unit, which, based on conditions related to the shooting, causes the visual sensor to capture a plurality of second images, and saves the plurality of second images as a plurality of learning images by corresponding the positions of the feature portions determined for the plurality of second images.

5. The image processing apparatus according to claim 4, characterized in that, The conditions related to the shooting include: configuration conditions related to the configuration of the object. The learning image acquisition unit acquires multiple second images by having the vision sensor capture images in various states where the robot changes the configuration position of the object according to the configuration conditions.

6. The image processing apparatus according to claim 4 or 5, characterized in that, The conditions associated with the shooting include exposure conditions. The learning image acquisition unit acquires multiple second images by repeatedly taking pictures with the visual sensor while changing the exposure conditions, wherein the exposure conditions are for the object disposed at the second position.

7. The image processing apparatus according to any one of claims 4 to 6, characterized in that, The conditions related to the shooting include: the position or orientation of the visual sensor. The learning image acquisition unit acquires multiple second images by repeatedly taking pictures while changing the position or posture of the visual sensor relative to the object disposed at the second position.

8. The image processing apparatus according to any one of claims 4 to 7, characterized in that, The image processing apparatus further includes a setting unit that accepts input for setting conditions related to the shooting.

9. The image processing apparatus according to any one of claims 4 to 8, characterized in that, The image processing apparatus further includes a learning unit that constructs a learning model for estimating the position of the feature portion on any image of the object by performing learning based on the plurality of learning images.

10. The image processing apparatus according to any one of claims 1 to 9, characterized in that, The image processing apparatus further includes a feature location registration unit, which accepts user operations for specifying the location of a feature of the object on the first image.

Citation Information

Patent Citations

  • Picking device and picking method

    JP2016203293A

  • Information processing device, image recognition method and image recognition program

    JP2019056966A

  • Image generating device, robot training system, image generating method, and image generating program

    JP2020082315A