Training device, image collecting device and control device

The training device enhances box detection accuracy in depalletizing by using both sharp and blurry training images to train a model, addressing issues with similar colors and adhesive materials, and improving detection precision.

DE112023006901T5Undetermined Publication Date: 2026-07-09FANUC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
FANUC LTD
Filing Date
2023-12-12
Publication Date
2026-07-09

AI Technical Summary

Technical Problem

Existing box detection techniques in depalletizing processes using machine learning struggle with accuracy due to similar background and box colors, close packing, and adhesive materials, and are sensitive to the number and quality of training images.

Method used

A training device captures both in-focus and out-of-focus training images to train a learning model for improved cardboard area detection, using a variety of training images based on specific conditions to enhance detection accuracy.

Benefits of technology

The trained model can accurately detect cardboard areas even in blurry images, improving detection precision and reducing noise from mismatched training conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A training device according to the present disclosure comprises: a storage unit designed to store a plurality of training images containing an area of ​​an object whose image was captured by an image acquisition unit; and a training unit designed to train a learning model to detect an area of ​​the object from an image using the plurality of training images, wherein the plurality of training images used to train the learning model includes an image that is blurred with respect to the object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA The present disclosure relates to a training device, an image acquisition device and a control device. BACKGROUND OF THE TECHNOLOGY In processes where stacked boxes are removed sequentially by a robot (depalletizing), a box detection technique is crucial for improving work efficiency. One such technique involves using threshold processing of a camera image to detect a specific area of ​​the box. However, this technique can sometimes fail to detect the box area, for example, because the background and box colors are similar, the boxes are packed closely together, or adhesive tape or similar materials are present on the box surface. A known learning-based technique can overcome these challenges (see, for example, patent literature 1).For example, in supervised machine learning with training images, a learning model is trained using previously captured training images to enable the detection of a cardboard area within an image. By inputting a captured image into a trained model that has completed the learning process, a cardboard area contained within a captured image can be detected. In such machine learning using training images, the detection accuracy of the trained model varies greatly depending on the number and quality of the training images. BIBLIOGRAPHY PATENT LITERATURE Patent literature 1: Japanese publication no. 11-272845 SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION There is a need for a technique that improves the accuracy of capturing an object area from a captured image using machine learning. SOLUTION TO THE PROBLEM A training device according to the present disclosure comprises: a storage unit configured to store a plurality of training images containing an area of ​​an object whose image was captured by an image acquisition unit; and a training unit configured to train a learning model to capture an area of ​​the object from an image using the plurality of training images, wherein the plurality of training images used to train the learning model includes an image that is blurred with respect to the object. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 shows a training system including a training device according to a first embodiment. Fig. 2 is a hardware configuration diagram of the training device shown in Fig. 1. Fig. 3 is a functional configuration diagram of the training device shown in Fig. 1. Fig. 4 shows an example of a training image management table stored in a memory unit shown in Fig. 3. Fig. 5 shows an example of a training model management table stored in the memory unit shown in Fig. 3. Fig. 6 is a flowchart showing an example of the image acquisition processing sequence by the training device shown in Fig. 1. Fig. 7 is a flowchart showing an example of the training sequence by the training device shown in Fig. 1. Fig. 8 shows a robot system including a control device according to a second embodiment.Figure 9 is a hardware configuration diagram of the control device shown in Figure 8. Figure 10 is a functional configuration diagram of the control device shown in Figure 8. Figure 11 is a flowchart showing an example of a robot control sequence using the control device shown in Figure 8. DETAILED DESCRIPTION OF THE INVENTION Embodiments of the present invention are described with reference to the drawings. In the following description, components that have essentially the same function and configuration are designated by the same reference numeral, and repeated descriptions are made only to the extent necessary. A training device according to a first embodiment is described below with reference to Figures 1, 2, 3, 4, 5, 6 to 7. The training device according to the first embodiment (hereinafter simply referred to as the training device) is an information processing device with an image acquisition function for capturing a multitude of training images by controlling a camera and a training function for using a multitude of training images to train a learning model to recognize a cardboard area from an image. A feature of the training device is that not only training images that are in focus with respect to the cardboard, but also training images that are out of focus with respect to the cardboard, are captured and used for training a learning model. In the present embodiment, the terms are defined as follows. Cardboard box: The cardboard box is an example of the object. The object is not limited to the cardboard box and can be any object that is to be captured by image recognition. Camera: The camera is an example of an image acquisition unit. In the first embodiment, the camera is assumed to be a device capable of changing its focus, focal length, and aperture based on externally input control signals. Focus: "In focus" means that the box appears sharp in the image, while "out of focus" means that the box appears blurry and not sharp. Optically, "in focus" means that the distance (defocus amount) from the image position to the image-taking position of the box is zero or close to zero, while "out of focus" means that the distance from the image position to the image-taking position of the box is large. Accordingly, the focus can be changed by altering the distance between the image position and the image-taking position of the box. For example, the focus can be changed by shifting the image-taking position along the optical axis. Note that the depth of field can be defined accordingly.For example, “in focus” can be defined as a defocus amount of zero – meaning that the imaging position matches the image capture position – or as a defocus amount that is less than a specified distance. Fig. 1 shows a configuration of a training system comprising the training device according to the first embodiment. As shown in Fig. 1, a training system 1 is configured such that a camera 4 and an information processing terminal 5 are connected to a training device 2, which acts as the center of the system. The training device 2 is a computer device that includes both a training function for training a learning model to recognize a cardboard area from an image using a large number of training images captured by camera 4, and an image acquisition function for controlling camera 4 to capture the large number of training images. It is implemented by a PC or similar device. It should be noted that the concept of the training device can include an information processing device that incorporates both the training function and the image acquisition function, as well as a camera. Camera 4 is connected to training device 2 via a cable, enabling data communication. Camera 4 is positioned to view stacked cardboard boxes W. It sets image capture parameters such as focus, aperture, and focal length based on a control signal from the training device and then records the images of the boxes W. The image captured by camera 4 is then input into training device 2 as a training image. It should be noted that camera 4 can be mounted on the hand of a robot positioned near the stacked cardboard boxes W. The information processing terminal 5 is connected to the training device 2 via a network 3, such as the internet, enabling data communication. The information processing terminal 5 is an input device for user input into the training device 2 and is implemented using a smartphone, tablet, or similar device. It should be noted that the step of acquiring a training image and the step of training a learning model are independent of each other, so these steps do not necessarily have to be performed by a single system. For example, an image acquisition device 2', which only has the image acquisition function, and a training device 2, which only has the training function, can be configured as independent devices. Furthermore, the function of the information processing terminal 5 for inputting user input into the training device 2 can be implemented within the training device 2 itself. That is, the training device 2 can include, as functions required for receiving user input, a display device that shows a screen created by the training device 2, and an operating device for receiving user input on the screen displayed on the display device. Fig. 2 is a hardware configuration diagram of the training device 2 according to the first embodiment. As shown in Fig. 2, the training device 2 is configured such that a RAM 12, a ROM 13, a storage device 14, a camera interface 15, and a communication device 16 are connected to a processor 11 via a data / control bus 10. The processor 11 is implemented by a CPU, a GPU, or the like. The RAM 12 functions as the main memory, workspace, and the like of the processor 11. A BIOS, an operating system, or the like is stored in the ROM 13. A training program is stored in the storage device 14.Various types of data, such as the training program, stored in the storage device 14, can be distributed to the user by recording them on removable media such as a USB stick, or they can be distributed to the training device 2 by downloading them over the network 3. The camera 4 is connected to the camera interface 15. The camera 4 defines the image capture conditions and performs image capture processes under the control of the processor 11. The communication device 16 is implemented by a communication module that conforms to any communication standard and sends and receives data to and from the information processing terminal 5 under the control of the processor 11. Fig. 3 is a functional configuration diagram of the training device 2. The training program loaded into the RAM 12 by the storage device 14 is executed by the processor 11, whereby the training device 2 functions as a transmission unit 21, as an operating information input unit 22, as a training image input unit 23, a storage unit 24, a screen generation unit 25, an image acquisition condition setting unit 26, a camera control unit 27, a training condition setting unit 28 and a training unit 29. The transmission unit 21 transmits data from various screens, relating to the training program generated by the screen generation unit 25, to the information processing terminal. The operating information input unit 22 provides user input via the information processing terminal. Specifically, object information, layout information, image capture conditions, training conditions, instructions for starting image capture, and the like are entered by the operating information input unit 22. The training image input unit 23 inputs data from a training image captured by the camera 4. The training image input data is stored in the memory unit 24.In addition, the object information, layout information and image capture conditions relating to the training image stored in memory unit 24 are registered in a training image management table stored in memory unit 24. Memory Unit 24 stores various types of information related to the processing of the training program. Specifically, Memory Unit 24 stores data from training images captured by Camera 4, data from training models trained by Training Unit 29, a training image management table for managing information about training images, and a trained model management table for managing information about trained models. Details of these management tables are described later. The screen generation unit 25 generates various screens related to the training program, according to a predefined format. These screens include an input screen for entering object and layout information, a screen for setting image capture conditions, a screen for setting training conditions, an instruction screen for instructing the start of image capture, and so on. The Image Capture Condition Setting Unit 26 sets the number of training images to be captured and the image capture conditions for each training image, based on the image capture conditions entered via the Operating Information Input Unit 22. The image capture conditions include a focus range and distance, an aperture range and distance, and a focal length range and distance. Based on these image capture conditions, the Image Capture Condition Setting Unit 26 defines a variety of image capture conditions in which at least one of the parameters—focus, aperture, and focal length—differs.For example, if the image capture condition setting unit 26 specifies three focus patterns, two aperture value patterns, and four focal length patterns based on the image capture conditions, the total number of permutations for these patterns - 24 - is specified as the number of training images to be captured, and 24 image capture condition patterns are specified. Triggered by the input of a setting instruction for starting image capture via the operating information input unit 22, the camera control unit 27 controls the camera 4 so that training images are captured under the image capture conditions defined by the image capture condition setting unit 26. The training condition setting unit 28 defines the training conditions entered via the operator information input unit 22. The training conditions are conditions for limiting the training images to be used for training a learning model from the multitude of training images stored in the memory unit 24. The training conditions comprise an object condition, a layout condition, and an image acquisition condition range (image acquisition condition set). Details of the training conditions are described later. Training Unit 29 trains a learning model to detect a cardboard area from an image using a variety of training images that satisfy the training conditions defined by Training Condition Setting Unit 28. For example, Training Unit 29 trains a learning model using machine learning with training images and a cardboard area in each training image as training data. This results in a trained model that, when given an image captured by Camera 4, detects a cardboard area contained within the image. The trained model created by Training Unit 29 is stored in Memory Unit 24, and the training conditions of the trained model stored in Memory Unit 24 are recorded in the Training Model Management Table.This does not preclude training the learning model using all training images, and if the training condition includes all training images, training unit 29 produces a trained model that has been trained using all training images. The training image management table stored in memory unit 24 is described below with reference to Fig. 4. Fig. 4 shows an example of the training image management table stored in memory unit 24. As shown in Fig. 4, the training image management table links the "Filename," "Object Information," "Layout Information," and "Image Acquisition Conditions" to the "Image No.", which uniquely identifies a training image. The object information, layout information, and image acquisition conditions are used as training conditions to extract the desired training images from a large number of training images. Object information is information about the object whose image was captured. For example, object information includes details about the type of object, such as a box or a pallet; detailed information about the object, such as its shape, size, color, and material; and supplementary information, such as a label, tape, or similar item attached to the object's surface. The layout information indicates the type of arrangement used for image capture. For example, the layout information includes image capture distance, stacking information, and environmental information. The image capture distance indicates the physical distance from camera 4 to the object. The stacking information indicates how the boxes W were stacked. For example, the stacking information includes 3 rows by 3 columns, 4 rows by 4 columns, and so on. The specification "3 rows by 3 columns" shows that one layer consists of a total of nine boxes W, three in a vertical direction and three in a horizontal direction. The environmental information shows the surroundings at the time the image was captured.For example, the environmental information includes information that can influence the detection of the object area based on the training image, such as the color of the lighting, the color of the background of the boxes W contained in the field of view of camera 4 (the color of the ground surface on which the boxes W are stacked), and the presence or absence of the reflection of a structure and the shadow of the structure within the field of view of camera 4. The image capture conditions indicate the optical settings of camera 4 at the time of image capture. These include, for example, the focus (focus deviation), the aperture value, and the focal length. The training model management table stored in memory unit 24 is described below with reference to Fig. 5. Fig. 5 shows an example of the training model management table stored in memory unit 24. As shown in Fig. 5, the "Model Name" and "Training Conditions" in the training model management table are linked to the "Model No.", which uniquely identifies a trained model. The Model Name displays the name of the model being trained. The Training Conditions are conditions used to limit the training images used for training the model. The Training Conditions include "Object Conditions," "Layout Conditions," and "Image Acquisition Conditions." The object condition is a condition used to limit the training images to be learned, based on object information. The object condition is determined from the object information managed in the training image management table. For example, the object condition is used to limit the training images to those assigned the object type "Cardboard" and the shape "Square". Layout conditions are used to limit the training images to be learned, based on layout information. Layout conditions include an image capture condition, a stacking condition, and an environmental condition, where the image capture condition specifies a range of image capture distances. The stacking condition is selected from the stacking information managed in the training image management table. The environmental condition is selected from the environmental information managed in the training image management table. For example, layout conditions are used to limit the training images to those associated with the image capture condition "200 cm to 300 cm," the stacking conditions "3 rows by 3 columns" and "4 rows by 4 columns," and the environmental condition "lighting color 'white' and floor surface color 'green'." Image capture conditions are a variety of image capture conditions used to limit the training images to be used for learning. These conditions include a focus range, an aperture range, and a focal length range. For example, image capture conditions are used to limit the training images to those corresponding to the focus range (blur range) "0.0 mm to 0.2 mm", the aperture range "f / 1.4 to f / 2.8", and the focal length "10 mm to 35 mm". The following describes the image acquisition process performed by the training device 2 with reference to Fig. 6. Fig. 6 is a flowchart illustrating an example of the image acquisition process performed by the training device 2. As shown in Fig. 6, the training device 2 first determines the object and layout information input from the information processing terminal (S11). Subsequently, the training device 2 determines the number of training images to be acquired and the image acquisition conditions for each training image, based on the image acquisition conditions input from the information processing terminal (S12).Triggered by a command to start image acquisition from the information processing terminal, the training device 2 selects one of the defined image acquisition conditions (focus, aperture, and focal length) (S13) and transmits a control signal to camera 4 (S14) to instruct the start of image acquisition. During step S14, camera 4 sets the image acquisition condition selected in step S13 and performs the image acquisition process. The training device 2 receives data from camera 4 of a training image captured under the image acquisition conditions defined in step S13 (S15), saves the training image (S16), assigns an image number to the saved training image, and records the object information, layout information, and image acquisition conditions in the training image management table (S17).The processes of steps S13 to S17 are repeated until the image acquisition of all training images to be recorded is complete (S18; No), and the image acquisition process is terminated when the image acquisition of all training images to be recorded is complete (S18; Yes). Following the image acquisition procedure in steps S11 to S18, the training device 2 can acquire a variety of training images, including a training image that is in focus with respect to the carton W and a training image that is out of focus with respect to the carton W, by controlling the camera 4 based on the input image acquisition condition. The following describes the training process performed by the training device 2 with reference to Fig. 7. Fig. 7 is a flowchart illustrating an example of the training process by the training device 2. As shown in Fig. 7, the training device 2 first defines the training conditions input by the information processing terminal (S21) and selects from a large number of training images those that meet the training conditions (S22). The training device 2 then trains a learning model using the selected training images (S23), saves the trained model (S24), and records the training conditions of the trained model in the training model management table (S25).Following the training process in steps S21 to S25, the training device 2 can generate a trained model using a variety of training images, including a training image that is in focus with respect to the cardboard W and a training image that is out of focus with respect to the cardboard W. The training device 2 according to the first embodiment, described with reference to Figs. 1, 2, 3, 4, 5, 6 to 7, can acquire a plurality of training images, comprising a sharp training image with respect to the cardboard W and a blurred training image with respect to the cardboard W, and create a trained model for acquiring a cardboard area from an image using the plurality of training images comprising a sharp training image with respect to the cardboard W and a blurred training image with respect to the cardboard W. The trained model, which was specifically trained using a blurred training image with respect to the cardboard W, performs the following: This means that in the system for removing stacked cartons W using a robot, as the removal process progresses and the number of stacked cartons W decreases, the distance of camera 4 from the carton W to be removed changes, which can cause camera 4 to lose focus on the carton W being removed. Specifically, for example, when removing cartons W stacked in nine layers: if the focus is set on the carton W in the ninth layer, the focus begins to shift as the number of layers decreases, which can result in the focus no longer being on the carton W stacked in the first layer.In such a case, by using a trained model that has been specifically trained using a blurry training image, it is possible to improve the probability of detecting a cardboard area even from an image that is blurry with respect to the cardboard W, compared to using a trained model trained only with sharp training images. In other words, it is possible to improve the accuracy of detecting a cardboard area from an image that is blurry with respect to the cardboard W. Furthermore, the training device 2 according to the first embodiment can generate a trained model that has been trained using training images selected from a large number of training images based on training conditions (object conditions, layout conditions, image acquisition conditions). For example, a trained model trained using training images selected under layout conditions that closely approximate the actual layout or under image acquisition conditions that closely approximate the image acquisition conditions actually used can reduce the time required for training while maintaining the accuracy of capturing a cardboard area compared to trained models trained using all training images.Furthermore, training images captured with layouts that differ significantly from the actual layout, or under image capture conditions that differ significantly from those actually used, can introduce noise, which can reduce the accuracy of detecting a cardboard area. In such a case, a trained model trained using training images restricted to layout conditions closely resembling the actual layout, or to image capture conditions closely resembling those actually used, can improve the accuracy of detecting a cardboard area compared to models trained using all training images. A control device according to a second embodiment is described below with reference to Figures 8, 9, 10 to 11. It should be noted that structures with the same functions as those of the first embodiment are designated with the same reference numerals and their detailed description is omitted here. The control device according to the second embodiment (hereinafter simply referred to as the control device) is an information processing device with a control function that detects a cardboard box area from an image captured by a camera using a trained model that was trained by machine learning using a variety of training images as training data, and controls a robot based on the detection result. A feature of the control device is that, during the operation in which the robot performs a carton removal operation, it detects a cardboard box area from an image using a trained model that was trained using a variety of training images, including both a sharp image and a blurred image with respect to the carton. Fig. 8 shows a configuration of a robot system comprising the control device according to the second embodiment. As shown in Fig. 8, a robot system 6 is configured such that a camera 4, an information processing terminal 5, a robot 8, and a three-dimensional sensor 9 are connected to a control device 7, which acts as the center of the system. The robot 8 is connected to the control device 7 via a cable, enabling data communication. The robot 8 is positioned near the stacked cartons W and performs a process to remove the cartons W after receiving a control signal from the control device 7. The three-dimensional sensor 9 is connected to the control device 7 via a cable, enabling data communication. The three-dimensional sensor 9 is mounted in a position from which it oversees the stacked cartons W and transmits data about the distance to the surface of the carton W to be removed to the control device 7 after receiving a control signal from the control device 7. Fig. 9 is a hardware configuration diagram of the control device 7 according to the second embodiment. As shown in Fig. 9, the control device 7 is configured such that a RAM 12, a ROM 13, a storage device 14, a camera interface 15, a communication device 16, a robot interface 17, and a sensor interface 18 are connected to a processor 11 via a data / control bus 10. The storage device 14 stores a control program for implementing a function to control the robot 8. The various types of data, such as a control program, stored in the storage device 14 can be distributed to the user by recording them on a removable storage medium (non-volatile storage medium) such as a USB flash drive, or they can be distributed by downloading them to the control device 7 via a network 3. The robot 8 is connected to the robot interface 17.The robot 8 operates under the control of the processor 11. The three-dimensional sensor 9 is connected to the sensor interface 18. Under the control of the processor 11, the three-dimensional sensor 9 measures the distance from a predefined reference position of the three-dimensional sensor 9 to the cardboard box W (measurement target). Fig. 10 shows a functional diagram of the control device 7. The control program loaded into the working memory 12 by the storage device 14 is executed by the processor 11, whereby the control device 7 functions as a transmission unit 71, as an operating information input unit 72, as a captured image input unit 73, as a distance input unit 74, a storage unit 75, a screen generation unit 76, an image acquisition condition setting unit 77, a camera control unit 78, a training model setting unit 79, a carton detection unit 80, a three-dimensional sensor control unit 81 and a robot control unit 82. The transmission unit 71 transmits data from various screens, relating to the control program generated by the screen creation unit 76, to the information processing terminal 5. The operating information input unit 72 inputs user commands via the information processing terminal 5. Specifically, the operating information input unit 72 inputs settings for the trained model, image capture conditions, an operation start command, and the like. The captured image input unit 73 inputs data from an image captured by the camera 4. The distance input unit 74 inputs data about the distance to the carton W to be removed, measured by the three-dimensional sensor 9. The storage unit 75 stores various types of information for processing the control program.In particular, the storage unit 75 stores a training model management table to manage data from a large number of trained models, where at least some of the training images used for training differ, as well as information regarding the trained models. Since the training model management table is identical to that of the first embodiment, a description of it is omitted. The screen generation unit 76 creates various screens related to the control program according to a predefined format. These screens include a screen for setting the image capture conditions, a screen for specifying the trained model, an instruction screen for instructing the start of a sampling operation, and the like. The image capture condition setting unit 77 sets the image capture conditions entered via the operator information input unit 72. These conditions include focus, aperture, and focal length. Triggered by the input of a start command via the operator information input unit 72, the camera control unit 78 controls the camera 4 to capture an image under the imaging conditions specified by the image capture condition setting unit 77. The training model setting unit 79 selects a training model from a variety of training models based on the training model setting conditions entered via the operating information input unit 72 and assigns the selected training model to the carton detection unit 80. Note that the training model setting unit 79 can narrow down a variety of trained models based on the trained model setting conditions and assign a trained model selected by the user from at least two narrowed-down trained models to the carton detection unit 80. The carton detection unit 80 uses the trained model defined by the training model setting unit 79 to detect a carton area from an image captured by camera 4 and to determine the position of the detected carton area. The carton detection unit 80 defines the detected carton W as the target to be removed. The three-dimensional sensor control unit 81 controls the 3D sensor 9 to obtain the distance from the 3D sensor 9 to the carton W to be removed. The robot control unit 82 controls the robot 8 to remove the carton W, based on the position of the carton area detected by the carton detection unit 80 and the distance detected by the 3D sensor 9. The following describes, with reference to Fig. 11, the robot control process by the control device 7 to instruct the robot 8 to perform the operation of removing the carton W. Fig. 11 is a flowchart showing an example of the robot control process by the control device 7. As shown in Fig. 11, the control device 7 first receives training model setting conditions (object information, layout information, image acquisition conditions) (S31) input from the information processing terminal 5 and, based on these conditions, selects a training model from a plurality of training models (S32). Next, in response to a start command input from the information processing terminal 5, the control device 7 transmits an image acquisition command, along with the image acquisition conditions input from the information processing terminal 5, to the camera 4 (S33). The operation in step S33 sets the user-defined image acquisition conditions in the camera 4, and the image acquisition process is executed.The control device 7 receives data from the captured image by camera 4 (S34) and performs the process of detecting a carton area from the captured image using the trained model (S35). If the carton area cannot be detected (S36; No), but stacked cartons W remain (S41; Yes), the control device 7 modifies the trained model (S42), returns to the process from step S33, and performs the carton W detection process again using the modified trained model. Whether stacked cartons W remain can be determined based on the image captured by camera 4, a three-dimensional measurement result from the three-dimensional sensor 9, and the like. If, on the other hand, the carton area can be detected (S36; Yes), a command to measure the distance from the three-dimensional sensor 9 to the carton W to be removed is sent to the three-dimensional sensor 9 (S37).In step S37, the distance from the three-dimensional sensor 9 to the carton W to be removed is entered into the control device 7. Based on the position of the carton area detected in step S35 and the distance from the three-dimensional sensor 9 to the carton W detected in step S37, the control device 7 sends an operating command to the robot 8 to remove the carton W (S38). In step S38, the robot 8 removes the carton W. The control device 7 repeatedly executes the operations of steps S37 and S38 until all cartons W detected in step S35 have been removed (S39; No). The control device 7 repeatedly executes the operations of steps S33 to S39 until all stacked cartons W have been removed (S40; Yes), and terminates control of the robot 8 when all stacked cartons W have been removed (S40; No). The control device 7 according to the second embodiment can first define a trained model according to a user instruction, execute the process of detecting the carton W using the initially defined trained model, and control the robot 8 to remove the detected carton W. The trained model is a model that has been trained using a variety of training images, including a training image that is in focus with respect to the carton W and a training image that is out of focus with respect to the carton W.As described in the first embodiment, the trained model, which was also trained using a blurry training image, can detect a cardboard area even from a blurry image. Therefore, even if an image is captured during actual operation that is blurry with respect to the cardboard W to be detected, the accuracy of detecting the cardboard W from the blurry image can be improved compared to the trained model trained only with sharp training images. This makes it possible to prevent the robot 8 from stopping when picking up the cardboard W due to its inability to detect it. Furthermore, according to control device 7 of the second embodiment, if the carton W cannot be detected using a trained model, the process of detecting the carton W can be repeated while switching between a multitude of pre-stored trained models. Since the multitude of pre-stored trained models were also trained using a training image that is blurred with respect to the carton W, and their training conditions differ from one another, the probability of detecting the carton W can be improved by switching between the trained models. In the present embodiment, the camera 4 is described as a device capable of automatically changing image acquisition conditions such as focus, focal length, and aperture value through internal processing of the camera 4 based on a control signal input from an external source; however, the imaging conditions of the camera 4, such as focus, focal length, and aperture value, can also be manually set by an operator operating the camera 4. In this case, the training device 2, in the sequence of image acquisition processes described with reference to Fig. 6, can display the image acquisition conditions (focus, aperture value, and focal length) of the next training image to be acquired instead of step S14. Furthermore, in the present embodiment it has been described that the camera 4 is configured to be able to change its focus, focal length and aperture value; however, it is sufficient if at least the focus can be changed, and the camera 4 only needs to be able to adjust at least the focus. The following annexes are further disclosed in connection with the present embodiments and modifications. (Annex 1) A training device 2 comprises: a storage unit 24 configured to store a plurality of training images containing an area of ​​an object whose image was captured by an image acquisition unit 4; and a training unit 29 configured to train a learning model to detect an area of ​​the object from an image using the plurality of training images. The plurality of training images includes an image that is blurred with respect to the object. (Annex 2) In the training device 2 according to Annex 1, the multitude of training images includes at least two training images with different focal lengths of the image acquisition unit 4. (Annex 3) In the training device 2 according to Annex 1 or 2, the multitude of training images includes at least two training images with different aperture values ​​of the image acquisition unit 4. (Annex 4) In the training device 2 according to one of Annexes 1 to 3, the plurality of training images comprises at least two training images with different physical distances from the image acquisition unit 4 to the object. (Annex 5) The training device 2 according to one of Annexes 1 to 4 further comprises an input unit 22 configured to input a training condition containing at least one physical distance from the image acquisition unit 4 to the object, a focus of the image acquisition unit 4, an aperture value of the image acquisition unit 4, and a focal length of the image acquisition unit 4. The training unit 29 trains the learning model using a training image that satisfies the training condition from among the plurality of training images. (Annex 6) An image acquisition device 2' comprises an image acquisition unit 4 configured to capture images of an object in order to acquire a plurality of training images to be used for training a learning model to detect an area of ​​the object from an image; and a control unit 27 configured to control the image acquisition unit to acquire the plurality of training images. The control unit 27 controls the image acquisition unit 4 to acquire a training image that is sharp with respect to the object and a training image that is blurred with respect to the object. (Annex 7) In the image acquisition device 2' according to Annex 6, the control unit 27 controls the image acquisition unit 4 so that at least two training images with different focal lengths of the image acquisition unit 4 are captured. (Annex 8) In the image acquisition device 2' according to Annex 7, the control unit 27 controls the image acquisition unit 4 so that at least two training images with different aperture values ​​of the image acquisition unit 4 are captured. (Annex 9) A control device 7 comprises an image acquisition unit 4 configured to capture an image of an object; a detection unit 80 configured to detect an area of ​​the object from an image captured by the image acquisition unit 4 using a trained model trained by machine learning using a variety of training images as training data, including a sharp image with respect to the object and a blurred image with respect to the object; and a control unit 82 configured to control a robot 8 based on the position of the area of ​​the object detected by the detection unit 80. While embodiments of the present disclosure have been described in detail, the present disclosure is not limited to the individual embodiments described above. These embodiments may be subjected to various additions, substitutions, modifications, partial deletions, etc., without departing from the core of the invention or the idea and spirit of the present invention as derived from the content described in the claims and their equivalents. For example, the embodiment described above shows the sequence of operations and the sequence of processes as examples, and the sequences are not limited to these. The same applies where numerical values ​​or formulas are used in the description of the embodiments described above. EXPLANATION OF THE REFERENCE SYMBOLS 1: Training system, 2: Training device, 2': Image acquisition device, 3: Network, 4: Camera, 5: Information processing terminal, 6: Robot system, 7: Control device, 8: Robot, 9: Three-dimensional sensor, 10: Data / control bus, 11: Processor, 12: RAM, 13: ROM, 14: Storage device, 15: Camera interface, 16: Communication device, 17: Robot interface, 18: Sensor interface, 21: Transmission unit, 22: Operating information input unit, 23: Training image input unit, 24: Storage unit, 25: Screen generation unit, 26: Image acquisition condition setting unit, 27: Camera control unit, 28: Training condition setting unit, 29: Training unit, 71: Transmission unit, 72: Operating information input unit, 73: 74: Captured image input unit, 75: Distance input unit, 76: Storage unit, 77: Screen creation unit, 78: Image capture condition setting unit, 79: Camera control unit,79: Training model setting unit, 80: Carton detection unit, 81: Three-dimensional sensor control unit, 82: Robot control unit.

Claims

A training device comprising: a storage unit comprising a plurality of training images containing an area of ​​an object whose image was captured by an image acquisition unit; and a training unit configured to train a learning model to capture an area of ​​the object from an image using the plurality of training images, wherein the plurality of training images includes an image that is blurred with respect to the object. The training device according to claim 1, wherein the plurality of training images comprises at least two training images with different focal lengths of the image acquisition unit. The training device according to claim 1 or 2, wherein the plurality of training images comprises at least two training images with different aperture values ​​of the image acquisition unit. The training device according to one of claims 1 to 3, wherein the plurality of training images comprises at least two training images with different physical distances from the image acquisition unit to the object. The training device according to one of claims 1 to 4, further comprising: an input unit configured to input a training condition comprising at least a physical distance from the image acquisition unit to the object, a focus of the image acquisition unit, an aperture value of the image acquisition unit and a focal length of the image acquisition unit, wherein the training unit trains the learning model using a training image that satisfies the training condition among the plurality of training images. An image acquisition device comprising: an image acquisition unit configured to capture images of an object in order to acquire a plurality of training images to be used for training a learning model to recognize an area of ​​the object from an image; and a control unit configured to control the image acquisition unit to acquire the plurality of training images, wherein the control unit controls the image acquisition unit to acquire a training image that is sharp with respect to the object and a training image that is blurred with respect to the object. The image acquisition device according to claim 6, wherein the control unit controls the image acquisition unit such that at least two training images with different focal lengths of the image acquisition unit are captured. The image acquisition device according to claim 6 or 7, wherein the control unit controls the image acquisition unit such that it captures at least two training images with different aperture values ​​of the image acquisition unit. A control device comprising: an image acquisition unit configured to capture an image of an object; a capture unit configured to capture an area of ​​the object from an image captured by the image acquisition unit using a trained model trained by machine learning using a variety of training images as training data, wherein the training images include a sharp image with respect to the object and a blurred image with respect to the object; and a control unit configured to control a robot based on the position of the area of ​​the object captured by the capture unit.