Annotating device

By employing robot control and coordinate system transformation techniques through the annotation device, the problem of inaccurate object detection caused by changes in image brightness and camera position was solved. This enabled efficient and accurate annotation and learning model generation, improving the robustness and accuracy of object detection.

CN112743537BActive Publication Date: 2025-12-09FANUC LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202011172863.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-30
Filing Date
2020-10-28
Publication Date
2025-12-09
Estimated Expiration
2040-10-28

AI Technical Summary

Technical Problem

In existing technologies, object detection struggles to maintain high robustness when image brightness and camera position relationships change. Furthermore, annotation accuracy in supervised learning is easily affected by user judgment biases, resulting in inaccurate and inefficient annotation processing.

Method used

An annotation device is used, and the robot controls the movement of the camera device and the object. By combining the transformation of the image coordinate system, sensor coordinate system and robot coordinate system, the position information of the object under multi-angle and multi-light conditions is automatically acquired. A learner is used for supervised learning to generate a high-precision learning model.

Benefits of technology

It enables efficient and accurate annotation processing in a large number of images, improves the robustness and accuracy of object detection, and reduces the impact of user judgment bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112743537B_ABST
    Figure CN112743537B_ABST
Patent Text Reader

Abstract

The present application provides an annotation device that can easily and accurately perform annotation processing even for a large number of images. The annotation device (1) includes an imaging device (50), a robot (3), a control unit (4), a designation unit (53, 55), a coordinate processing unit (56), and a storage unit (54). The control unit (4) controls the robot (3) to acquire learning images of a plurality of target objects (W) having different positional relationships with the imaging device (50). In addition, the storage unit (54) stores the position of the target object (W) in the robot coordinate system as the position of the target object (W) in the image coordinate system at the time of imaging or the position of the target object (W) in the sensor coordinate system, together with the learning images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an annotation apparatus. BACKGROUND

[0002] In the past, in a technique of detecting an object from an image, a method of learning to improve detection accuracy is known. As a document disclosing a technique related to such learning, there are Patent Documents 1 to 3, Non-Patent Document 1.

[0003] Patent Document 1 relates to a teacher data generation apparatus that generates teacher data used when performing object detection on a specific recognition target. In Patent Document 1, an identification model is described that uses reference data including a specific recognition target to make the specific recognition target by learning by an object identification method. The identification model is used to detect the specific recognition target by inferring from animation data including the specific recognition target by an object identification method, and to generate teacher data of the specific recognition target.

[0004] Patent Document 2 relates to an information processing apparatus that includes an imaging device that can capture a first distance image of an object at a plurality of angles, and a generation section that generates a three-dimensional model of the object based on the first distance image, and generates an extraction image that represents a specific part of the object corresponding to the plurality of angles based on the three-dimensional model. In Patent Document 2, it is described that the position at which a robot hand grips an object is set as a specific part of the object, and an image recognition section provides position information of the estimated specific part of the object to the robot hand as control information.

[0005] In Patent Document 3, an information processing apparatus is described that generates learning data for holding an object by using a holding position posture of a manipulator when holding the object, information on whether or not holding of the object was successful at the holding position posture, and an imaging position posture of the manipulator when imaging an image.

[0006] In Non-Patent Document 1, a technique is described in which, in a Convolution Neural Network (CNN) that is one of methods of deep learning, a three-dimensional position of an object is estimated from an image.

[0007] Prior Art Documents

[0008] Patent Documents

[0009] Patent Document 1: Japanese Patent Application Publication No. 2018-200531

[0010] Patent Document 2: Japanese Patent Application Publication No. 2019-056966

[0011] Patent Literature 3: Japanese Patent Application Laid-Open No. 2018-161692

[0012] Non-Patent Documents

[0013] Non-Patent Literature 1: Real-Time Seamless Single Shot 6D Object Pose Prediction SUMMARY

[0014] Problems to be Solved by the Invention

[0015] In the related art, a contour or the like of an object is focused on as one of features, and the object is detected from an image. There are cases where the focused feature cannot be distinguished depending on brightness at the time of imaging of the image, and the object cannot be detected from the image. Further, there are cases where the object looks largely changed due to a change in positional relationship between a camera and the object at the time of imaging of the image including the object, and the object cannot be detected from the image.

[0016] In a case where machine learning such as deep learning is used, it is possible to learn features for each object, and detection of the object is performed with higher robustness. As one of techniques of the deep learning, there is supervised learning, but in the supervised learning, it is necessary to annotate positions and poses of the object in the image for a large number of images, and this is one of obstacles in performing the deep learning. Further, in a case where the user performs the annotation processing, there is a possibility that accuracy of the annotation decreases due to a deviation in judgment criteria of the user.

[0017] It is desirable to provide an annotation apparatus capable of easily and accurately performing annotation processing even for a large number of images.

[0018] Means for Solving the Problems

[0019] One embodiment of the annotation apparatus of the present disclosure is provided with: an imaging device that captures an object to acquire an image; a robot that moves the imaging device or the object to cause the object to enter an imaging range of the imaging device; a control section that controls movement of the robot; a designation section that designates a position of the object in an image coordinate system of an image captured by the imaging device, a position of the object in a sensor coordinate system based on a position of the imaging device, or a position of the object in a robot coordinate system based on the robot; a coordinate processing section that can transform the position of the object in the image coordinate system or the position of the object in the sensor coordinate system into the position of the object in the robot coordinate system, and transform the position of the object in the robot coordinate system into the position of the object in the image coordinate system or the position of the object in the sensor coordinate system; and a storage section that stores the position of the object in the robot coordinate system acquired based on designation by the designation section, wherein the control section controls the robot to acquire a plurality of learning images of the object in which a positional relationship between the imaging device and the object is different, and the storage section transforms the position of the object in the robot coordinate system into the position of the object in the image coordinate system or the position of the object in the sensor coordinate system at the time of imaging and stores the position of the object in association with the learning image.

[0020] Effects of the Invention

[0021] According to one embodiment of the present disclosure, an annotation apparatus capable of easily and accurately performing annotation processing even for a large number of images can be provided. BRIEF DESCRIPTION OF DRAWINGS

[0022] FIG. 1 is a schematic view showing a structure of an industrial machine as an annotation apparatus of one embodiment of the present disclosure.

[0023] FIG. 2 is a view schematically showing a case where a model pattern designation region is designated in an image acquired by an industrial machine according to one embodiment of the present disclosure.

[0024] FIG. 3 is a block diagram of a function of a learning apparatus included in an industrial machine according to one embodiment of the present disclosure.

[0025] FIG. 4 is a flowchart showing a flow of annotation processing performed by an industrial machine according to one embodiment of the present disclosure.

[0026] FIG. 5 is a view schematically showing an example of an industrial machine including a plurality of image processing apparatuses according to one embodiment of the present disclosure.

[0027] Explanation of Reference Numerals

[0028] 1: industrial machine (annotation device); 3: robot; 4: mechanical control device (control section); 50: vision sensor (imaging device); 53: input section (designation section); 54: storage section; 55: image processing section (designation section); 56: coordinate processing section; W: workpiece (object). DETAILED DESCRIPTION

[0029] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. FIG. 1 is a schematic view showing the structure of an industrial machine 1 as an annotation device of one embodiment of the present disclosure.

[0030] The industrial machine 1 of the present embodiment is provided with a robot 3 that performs a prescribed process on a workpiece W placed on a table T, a mechanical control device 4 that controls the robot 3, an image processing system 5 that acquires an image containing the workpiece W and determines the position and orientation of the workpiece W, and a learning device 7 that performs a learning process.

[0031] The robot 3 is a vertical multi-joint robot having a plurality of movable members 31, 32, 33, 34 connected to each other in a rotatable manner, and a processing head 2 connected to the top ends of the plurality of movable members 31, 32, 33, 34. The processing head 2 is positioned by the plurality of movable members 31, 32, 33, 34. The kind of the robot 3 is not particularly limited. The robot 3 can be provided as an orthogonal coordinate robot, a scalar robot, a parallel robot, or the like, in addition to being provided as a vertical multi-joint robot.

[0032] As an example, the processing head 2 has an appropriate structure corresponding to the process to be performed on the workpiece W, such as a hand to hold the workpiece W to move it, a processing head capable of performing welding, laser processing, cutting processing, or the like on the workpiece W, and the like. In the illustrated industrial machine 1, the processing head 2 is a hand that is a holding portion that holds the workpiece. The workpiece W can be held by the processing head 2 to be moved to a prescribed position, or the posture of the workpiece W can be changed.

[0033] The mechanical control device 4 is, for example, a control section that determines the actions of the robot 3 and the image processing system 5 in accordance with a job program such as a processing program provided in advance. The mechanical control device 4 is constituted, for example, by appropriately programming a programmable controller, a numerical control device, or the like. The mechanical control device 4 is provided with a CPU (not shown) for comprehensively controlling the entire system, and is connected to the robot 3 and the image processing device 51 via an external device interface (not shown).

[0034] The mechanical control device 4 of the present embodiment has a program control section 41, a positioning control section 42 that controls the robot 3, and a head control section 43 that controls the processing head 2. The program control section 41, the positioning control section 42, and the head control section 43 in the mechanical control device 4 are distinguished according to their functions, and can not be clearly distinguished in physical structure and program structure.

[0035] The program control section 41 provides the robot 3 and the image processing system 5 with action instructions according to a work program such as a machining program. Specifically, the program control section 41 provides the robot 3 with an instruction for causing the processing head 2 to hold or release the workpiece W and determining a position to which the processing head 2 is moved. Also, the program control section 41 provides the image processing system 5 with an instruction for confirming the position of the workpiece W.

[0036] In addition, the program control section 41 is configured to input, as mechanical control information, to the image processing system 5 a parameter that is obtained from the positioning control section and that can determine the position and orientation of the processing head 2, such as a drive amount that indicates the relative relationship of the movable members 31, 32, 33, 34.

[0037] The mechanical control information also contains the coordinate position of the tip of the robot 3 in the robot coordinate system. As needed, information that indicates the state of the processing head 2 controlled by the head control section 43, information that indicates whether the processing head 2 has appropriately performed processing on the workpiece W, and the like can also be input as part of the mechanical control information to the image processing system 5.

[0038] The positioning control section 42 generates a drive signal that causes the movable members 31, 32, 33, 34 of the robot 3 to relatively rotate, according to the instruction from the program control section 41. In addition, the positioning control section 42 outputs a parameter that is mechanical control information. As a specific example, the parameter output by the positioning control section can be set to rotation position information of a plurality of drive motors that drive the movable members 31, 32, 33, 34, vector information that indicates the coordinate position and orientation of a reference point of the processing head 2, and the like.

[0039] The head control section 43 controls the action of the processing head 2 to perform processing on the workpiece W. In addition, it can also be configured to input a signal that indicates the state of the processing head 2 to the program control section 41.

[0040] The image processing system 5 has a vision sensor 50 that photographs an image of the workpiece W, and an image processing device 51 that performs control of the vision sensor 50 and processing of image data photographed by the vision sensor 50.

[0041] The vision sensor 50 can be configured by an imaging device having an optical system that images light from an object and a two-dimensional imaging element that converts the imaged image into an electric signal in accordance with a two-dimensional position.

[0042] The vision sensor 50 is mounted to the robot 3. The vision sensor 50 of the present embodiment is supported by the processing head 2 or the movable member 34 at the tip connected to the processing head 2.

[0043] The robot 3 can be driven to a position at which the workpiece W as an object enters the field of view of the vision sensor 50, an image is captured by the vision sensor 50, and the workpiece W detected from the captured image by image processing described later is subjected to a working operation by the processing head 2.

[0044] The image processing device 51 performs various kinds of processing on the image detected by the vision sensor 50. The image processing device 51 of the present embodiment is provided with a display section 52, an input section 53, a storage section 54, an image processing section 55, and a coordinate processing section 56.

[0045] The display section 52 can be configured to have a display panel or the like that displays information to an operator. Alternatively, the display section 52 can be a touch panel or the like that is formed integrally with the input section 53.

[0046] The input section 53 can have an input device that an operator can operate, such as a keyboard, a switch, or the like. Alternatively, the input section 53 can receive an input from another control device or a computer via a communication line or the like.

[0047] The storage section 54 is not particularly limited and can be configured by a volatile memory such as a DRAM, an SRAM, or the like. The storage section 54 stores various kinds of information related to the control of the vision sensor 50 and the image processing. For example, the storage section 54 stores image data captured and acquired by the vision sensor 50, a processing result of the image data, and mechanical control information at the time of imaging, or the like as imaging information.

[0048] Alternatively, the storage section 54 stores a model pattern obtained by modeling an image of the workpiece W, such as a model pattern that represents a feature of the image of the workpiece W.

[0049] Further, the storage section 54 stores calibration data of the vision sensor 50, for example, calibration data for transforming a two-dimensional position on an image coordinate system into a value on a three-dimensional coordinate or calibration data for performing an inverse transformation thereof. For example, the vision sensor 50 is calibrated based on the calibration data stored in the storage section 54, whereby when a three-dimensional point in a robot coordinate system (hereinafter, referred to as a gaze point) is provided, a position of an image of the three-dimensional point on an image of the vision sensor 50, that is, a two-dimensional point in a sensor coordinate system can be calculated. Further, when an image of a certain gaze point, that is, a two-dimensional point in a sensor coordinate system is provided, a line of sight in a robot coordinate system (a three-dimensional straight line passing through the gaze point and a focal point of the camera) can be calculated. As for the form of the calibration data and a method of deriving the calibration data, various ways are proposed, and any way can be used. Further, the image coordinate system is a coordinate system defined on an image (two-dimensional), and the sensor coordinate system is a coordinate system observed from the vision sensor 50 (three-dimensional). The robot coordinate system is a coordinate system observed from the robot 3 (three-dimensional).

[0050] The image processing section 55 analyzes image data captured by the vision sensor 50 by a known image processing technique to determine the position and orientation of the workpiece W. The image processing section 55 can be realized by causing a CPU or the like to execute an appropriate program.

[0051] Referring to FIG. 2 to explain a process of detecting the workpiece W from an image acquired by the vision sensor 50. FIG. 2 is a view schematically showing a case where the model pattern designation region 60 is designated in an image acquired by the industrial machine 1 according to one embodiment of the present disclosure.

[0052] The model pattern designation region 60 is set by an operator who confirms the image of the display section 52 and operates the input section 53. Further, it can be configured not to be set by the operator but to be automatically designated by the image processing section 55 through a prescribed image processing. For example, a portion where a luminance gradient is large in the image can be calculated as an outline of the image of the workpiece W, and the model pattern designation region 60 can be set to include the outline of the image of the workpiece W inside.

[0053] The coordinate processing section 56 transforms the detected position of the workpiece W in the image coordinate system (two-dimensional) or the detected position of the workpiece W in the sensor coordinate system (three-dimensional) to acquire a three-dimensional line of sight. Further, the coordinate processing section 56 performs a three-dimensional transformation process for transforming the detected position of the workpiece W in the image coordinate system (two-dimensional) or the detected position of the workpiece W in the sensor coordinate system (three-dimensional) into the detected position of the workpiece W in the robot coordinate system (three-dimensional) on the basis of the model pattern specifying region and the calibration data stored in the storage section 54 and the position of the vision sensor 50 of the robot 3 at the time of imaging. In this three-dimensional transformation process, when a two-dimensional imaging camera is used in the vision sensor 50, information for determining a position that is not determined in the line of sight direction is required. In the present embodiment, the three-dimensional transformation process is performed on the basis of a plane that is a plane in which four points that specify the position of the workpiece W on the image exist being set.

[0054] Next, the structure of the learning device 7 will be described with reference to FIG. 3 FIG. 3 is a block diagram of the function of the learning device 7 provided in the industrial machine 1 according to one embodiment of the present disclosure.

[0055] As shown in FIG. 3 , a state observation section 71 to which input data is input from the image processing device 51, a label acquisition section 72 to which a label corresponding to an image is input from the machine control device 4, and a learner 70 that performs supervised learning on the basis of the input data acquired by the state observation section 71 and the label input from the machine control device 4 to generate a learning model are provided.

[0056] The input data input from the image processing device 51 to the state observation section 71 includes an image of the workpiece W as an object, a processing result of the image, or both.

[0057] The label input from the machine control device 4 to the label acquisition section 72 is position information indicating the position of the workpiece W as an object at the time of imaging an image. The position information is imaging-time information including the detected position, posture, and size of the workpiece W corresponding to a certain image. For example, in a case where the position of the workpiece W is specified by a region closed by four corners, information indicating the positions of the four points is stored as information indicating the position of the workpiece W.

[0058] In the present embodiment, the label input to the label acquisition section 72 is acquired by transforming the position of the workpiece W in the robot coordinate system into a position in the image coordinate system or a position in the sensor coordinate system. Further, a process of acquiring position information indicating the position of the workpiece W in the image coordinate system or the position in the sensor coordinate system that becomes the label will be described later.

[0059] ​The input data acquired by the state observation section 71 and the label acquired by the label acquisition section 72 are input to the learner 70 in association with each other. In the learner 70, a learning model is generated based on a plurality of sets of input data and labels.

[0060] Next, the flow of the annotation processing will be described. FIG. 4 is a flowchart showing the flow of the annotation processing performed by the industrial machine 1 according to one embodiment of the present disclosure. Furthermore, the embodiment in which this flow is shown and described is an example.

[0061] When the annotation processing is started, the processing for learning the position of the workpiece W is executed in step S100. For example, the machine control device 4 drives the robot 3 so that the workpiece W is located within the imaging range of the vision sensor 50. Also, in a state in which the workpiece W has entered the imaging range of the vision sensor 50, an image containing the workpiece W is acquired by imaging the workpiece W by the vision sensor 50. For example, an image as shown in FIG. 2 is acquired.

[0062] Next, the processing of specifying the position of the workpiece W is performed. The specification of the position of the workpiece W can be performed, for example, by a method of specifying one point of the position of the workpiece W in the image, a method of surrounding the periphery of the workpiece W by a region closed by four corners, or the like. The specification of the position of the workpiece W can be performed by the image processing section 55 using an image processing algorithm as the specifying section, or the position of the workpiece W can be specified by the user using the input section 53 as the specifying section. Thereby, the position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system is determined.

[0063] As described above, since the calibration processing that enables the transformation from the image coordinate system to the robot coordinate system is performed in advance, the specified position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system is transformed into the robot coordinate system. Thereby, the position of the workpiece W in the robot coordinate system is acquired. That is, the position of the workpiece W has been learned. Furthermore, as the processing of learning the position of the workpiece W, a method of directly specifying the set position of the workpiece W in the robot coordinate system can also be used.

[0064] Through the above processing, the position of the workpiece W in the robot coordinate system can be learned. Then, the processing proceeds from step S100 to step S101 for acquiring a data set used in learning.

[0065] In step S101, the position relationship between the vision sensor 50 and the workpiece W is changed by the robot 3 with the condition that the workpiece W is contained in the imaging range. Since the position of the workpiece W in the robot coordinate system has been acquired in step S100, the position of the robot 3 in which the workpiece W is contained in the image can also be calculated in the control of the movement of the robot 3.

[0066] In step S101, the movement of the workpiece W can be performed using the processing head 2 in the process of changing the positional relationship between the vision sensor 50 and the workpiece W by the robot 3. The movement of the workpiece W includes, for example, moving the workpiece W front and back, left and right, changing the posture of the workpiece W, and inverting the workpiece W. Further, in the case where the workpiece W is moved, the process of updating the position of the workpiece W according to the place of movement is also performed. That is, the position of the workpiece W in the robot coordinate system is updated in conjunction with the movement of the workpiece W.

[0067] In step S102, the workpiece W is imaged in the state where the positional relationship between the vision sensor 50 and the workpiece W is changed by the robot 3 in step S101, and an image including the workpiece W is acquired.

[0068] In step S103, the position of the workpiece W in the image coordinate system corresponding to the image acquired in step S102 is acquired. The position of the workpiece W in the image coordinate system acquired in this step S103 is acquired using the position of the workpiece W in the robot coordinate system already known in step S100.

[0069] That is, the coordinate processing section 56 performs the process of transforming the position of the workpiece W in the robot coordinate system into position information indicating the position of the workpiece W in the image coordinate system or position information indicating the position of the workpiece W in the sensor coordinate system by taking into account the position of the vision sensor 50 held by the robot 3 at the time of imaging. The position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system is stored in the storage section 54 together with the image at the time of imaging. That is, the position of the workpiece W is annotated in the image.

[0070] In step S104, it is determined whether the processes of steps S101 to S103 are performed for a prescribed number of times. In the case where the processes are not performed for the prescribed number of times, the processes of steps S101 to S103 are repeated. That is, the processes of steps S101 to S103 are repeated until a sufficient number of sets of images and position information of the workpiece W for learning are acquired. In step S104, in the case where the prescribed number of times or more is reached, it is determined that a sufficient number of sets of images and position information of the workpiece W are acquired, and the process of step S105 is advanced to.

[0071] In step S105, a group of a prescribed number or more of images and position information of the workpiece W is input to the learner 70 as a learning data set. The learner 70 performs supervised learning based on the input data and the label of the data set, thereby constructing a learning model. In the learning of the learner 70, a method such as YOLO (You Only Look Once), SSD (Single Shot multibox Detector), or the like can be used. Alternatively, as described in Non-Patent Literature, it is also possible to use as the label and the output of the inference a point constituting a bounding box of the object. In this way, a publicly known method can be used in the learning of the learner 70.

[0072] For example, the learner 70 performs supervised learning using a neural network. In this case, the learner 70 provides a group of the input data and the label (teacher data) to a neural network configured by combining perceptrons, and performs so-called forward propagation of changing the weight of each perceptron included in the neural network so that the output of the neural network is the same as the label. The forward propagation is performed so that the detection result (for example, position, posture, size) of the workpiece W output from the neural network is the same as the object detection result (for example, position, posture, size) of the label.

[0073] Further, the learner 70 adjusts the weight value by a method of back propagation (also referred to as error back propagation method) after performing the forward propagation like this, so that the error of the output of each perceptron is reduced. In more detail, the learner 70 calculates the error of the output of the neural network from the label, and corrects the weight value so that the calculated error is reduced. In this way, the learner 70 learns the characteristics of the teacher data, and inductively obtains a learning model for estimating the result from the input.

[0074] In step S106, it is determined by the image processing section 55 whether sufficient detection accuracy has been obtained using the generated learning model. That is, it is determined whether the image of the workpiece W can be accurately detected from the image newly captured by the vision sensor 50. The user can determine whether the performance requirement is satisfied based on a prescribed determination method set in advance, or can determine in a manner in which an image of a positive solution or the like is determined based on a predetermined determination algorithm. The performance requirement is, for example, various conditions related to image detection, such as the accuracy of whether the position of the image is correctly detected, the error frequency being prescribed or less, and the like. When it is determined in step S106 that the performance requirement is satisfied, the processing proceeds to step S107. When it is determined in step S106 that the performance requirement is not satisfied, the processing returns to the processing of step S101 to add learning data. At this time, the following processing is performed: the value of the prescribed number of times determined in step S104 is incremented.

[0075] In step S107, a process of updating the previous learning model generated before the data set is input in step S105 to a learning model generated based on the newly input data set is performed. That is, if new teacher data is acquired after the learning model is constructed, the once-constructed learning model is updated.

[0076] The learning device 7 is able to accurately detect the image of the workpiece W from the image containing the workpiece W captured by the vision sensor 50 by using the learning model updated based on the latest data set.

[0077] Further, the machine control device 4, the image processing device 51, and the learning device 7 are constituted by, for example, an arithmetic processor such as a DSP (Digital Signal Processor), an FPGA (Field-Programmable Gate Array), or the like. The various functions of the machine control device 4, the image processing device 51, and the learning device 7 are realized, for example, by executing a prescribed software (program, application) stored in a storage. The various functions of the machine control device 4, the image processing device 51, and the learning device 7 can be realized by cooperation of hardware and software, or can be realized by hardware (electronic circuit) alone.

[0078] It is also possible that the learning model generated by the learner 70 is shared with other learning devices. If a plurality of learning devices are caused to share the learning model, the supervised learning can be performed by the respective learning devices in a distributed manner, and thus the efficiency of the supervised learning can be improved. An example of sharing the learning model will be described with reference to FIG. 5 FIG. 5 is a diagram schematically showing an example of the industrial machine 1 provided with a plurality of image processing devices 51 according to an embodiment of the present disclosure.

[0079] In FIG. 5 , the m image processing devices 51 are connected to the unit controller 101 via the network bus 102. One or more vision sensors 50 are connected to each of the image processing devices 51. The n vision sensors 50 are provided in total for the industrial machine 1 as a whole.

[0080] The learning device 7 is connected to the network bus 102. In the learning device 7, a learning model is constructed by machine learning of a set of learning data transmitted from the plurality of image processing devices 51 as a data set. The learning model can be used for detection of the workpiece W in each of the image processing devices 51.

[0081] ​As explained above, the industrial machine 1 of one embodiment of the present disclosure includes: a vision sensor (imaging device) 50 that captures an image of a workpiece (object) W; a robot 3 that moves the vision sensor 50 or the workpiece W to bring the workpiece W into the imaging range of the vision sensor 50; a machine control device (control section) 4 that controls movement of the robot 3; an image processing section (designation section) 55 or an input section (designation section) that designates the position of the workpiece W in an image coordinate system of an image captured by the vision sensor 50, the position of the workpiece W in a sensor coordinate system based on the position of the vision sensor 50, or the position of the workpiece 3 in a robot coordinate system based on the robot 3; a coordinate processing section 56 that can convert the position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system into the position of the workpiece W in the robot coordinate system, and convert the position of the workpiece W in the robot coordinate system into the position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system; and a storage section 54 that stores the known position of the workpiece W in the robot coordinate system acquired based on designation by the image processing section (designation section) 55 or the input section (designation section) 53. The machine control device (control section) 4 controls the robot 3 to acquire learning images of a plurality of workpieces W having different positional relationships with the vision sensor 50. In addition, the storage section 54 converts the position of the workpiece W in the robot coordinate system into the position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system at the time of imaging and stores the same together with the learning images.

[0082] Thus, the relative relationship between the vision sensor 50 as an imaging device and the workpiece W as an object can be grasped using the position of the robot 3. By using the known position of the workpiece W in the robot coordinate system, the position of the workpiece W in the image coordinate system or the position of the workpiece W in the sensor coordinate system corresponding to a certain image can be automatically acquired. Thus, a large amount of supervised data can be efficiently and accurately collected.

[0083] In addition, the industrial machine 1 of one embodiment of the present disclosure further includes a processing head (gripping section) 2 that holds the workpiece W, and the machine control device 4 changes the position, posture, or both of the workpiece W by gripping the workpiece W with the processing head 2, thereby changing the positional relationship between the vision sensor 50 and the workpiece W to acquire the learning images.

[0084] Thus, for the workpiece W as an object, learning data of a wide range of positional relationships of the workpiece W can be automatically acquired easily and accurately.

[0085] In addition, the industrial machine 1 of one embodiment of the present disclosure generates a learning model based on learning data including a learning image stored in the storage section 54 and information indicating a position of a workpiece W as a target object associated with the learning image, and in a case where it is determined that image detection processing using the learning model does not satisfy a performance requirement, the learning image and the information indicating the position of the target object are re-acquired.

[0086] Thus, learning data is added in a case where the detection accuracy of the workpiece W as a target object is not improved, and therefore it is possible to reliably prevent a case where the detection accuracy is not sufficiently improved due to insufficient learning data.

[0087] The above describes an embodiment of the present disclosure, but the present disclosure is not limited to the foregoing embodiment. In addition, the effects of the present disclosure are not limited to the content described in the present embodiment.

[0088] In addition to the structure of the foregoing embodiment, processing of changing the illumination brightness and the like can be added in the processing of step S101. By learning the difference in illumination brightness, it is possible to further improve the detection accuracy of the workpiece W as a target object.

[0089] In the foregoing embodiment, the structure in which the robot 3 holds the vision sensor 50 is described, but the structure in which the vision sensor 50 side is fixed and the robot 3 moves the workpiece W into the field of view of the vision sensor 50 can also be provided.

[0090] In the foregoing embodiment, an example in which the machine control device 4 is provided separately from the image processing device 51 is described, but a separate control device having both the functions of the machine control device 4 and the image processing device 51 can also be provided as the annotation device. In this case, the annotation device can refer to the entire information processing device (computer). For example, a server, a PC, various control devices, and the like can also be provided as the annotation device.

[0091] In the foregoing embodiment, the industrial machine 1 is provided with the learning device 7, but the structure in which the learning device 7 is omitted and only annotation is performed and a data set is provided to another learning device can also be provided.

[0092] The industrial machine 1 can be a machine tool that performs positioning of a workpiece or a tool by a positioning mechanism and performs processing of the workpiece.

[0093] The annotation processing performed by the industrial machine 1 is implemented by software. In a case where it is implemented by software, programs constituting the software are installed in the image processing device 51. In addition, these programs can also be recorded in a removable medium and distributed to users, and can also be distributed by being downloaded to a computer of a user via a network.

Claims

1.An annotation apparatus comprising: an imaging device that captures an object to acquire an image; a robot that moves the imaging device or the object to bring the object into an imaging range of the imaging device; a control section that controls movement of the robot; a designation section that designates a position of the object in an image coordinate system of an image captured by the imaging device or a position of the object in a robot coordinate system based on the robot; a coordinate processing section that can convert the position of the object in the image coordinate system to the position of the object in the robot coordinate system and convert the position of the object in the robot coordinate system to the position of the object in the image coordinate system; and a storage section that stores the position of the object in the robot coordinate system acquired based on designation by the designation section, wherein the control section controls the robot to acquire a plurality of learning images of the object in which a positional relationship between the imaging device and the object is different, the designation section designates the position of the object in the image coordinate system by setting a model pattern designation region in which a contour of an image of the object in an image captured by the imaging device is included inside, and the storage section converts the position of the object in the robot coordinate system to the position of the object in the image coordinate system at the time of imaging and stores the position together with the learning image. 2.The annotation apparatus according to claim 1, wherein a holding section that holds the object is further included, the control section changes the positional relationship between the imaging device and the object by changing the position, posture, or both of the object held by the holding section to acquire the learning image. 3.The annotation apparatus according to claim 1 or 2, wherein a learning model is generated based on learning data including the learning image stored in the storage section and information indicating the position of the object associated with the learning image, the learning image and the information indicating the position of the object are re-acquired in a case where it is determined that image detection processing using the learning model does not satisfy a performance requirement. ​

Citation Information

Patent Citations

  • Information processing system, information processing method and program

    JP2018161692A

  • Teacher data generation device, teacher data generation method, teacher data generation program, and object detection system

    JP2018200531A

  • Information processing device, image recognition method and image recognition program

    JP2019056966A

  • Automated Collection And Labeling Of Object Data

    CN107428004A