Learning device
The learning device automates and corrects annotation errors in waste sorting systems, improving efficiency and accuracy by using a replicated trained model and operator input, addressing the labor-intensive issue of manual training data annotation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-15
- Publication Date
- 2026-04-08
AI Technical Summary
Manual annotation of training data for machine learning in waste sorting devices is labor-intensive, hindering the efficiency of waste type identification.
A learning device comprising a first annotation unit, a learning unit, and a model management unit that automates the annotation process by using a replicated trained model and updates the model through machine learning with operator input to correct errors.
Improves the efficiency and accuracy of annotation work, reducing the effort required for manual correction of failed images and enhancing the overall recognition process.
Smart Images

Figure 0007842774000001 
Figure 0007842774000002 
Figure 0007842774000003
Abstract
Description
Technical Field
[0004] , , ,
[0005] , , ,
[0001] The present disclosure relates to a learning device.
Background Art
[0002] At waste treatment plants, a large amount of waste flows on a belt conveyor every day and is being processed. At the site where the waste is processed, the waste is sorted by hand. Although the waste sorting work is a simple task, the burden on the workers who sort the waste (hereinafter sometimes referred to as "sorting workers") is large, so a device that automatically sorts the waste (hereinafter sometimes referred to as a "waste sorting device") has been developed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] When the waste sorting device performs the work that the sorting worker used to do instead of the sorting worker, the waste sorting device recognizes each waste flowing on the belt conveyor and, based on the recognition result, uses a robot hand or a suction pad to extract a desired waste (hereinafter sometimes referred to as "desired waste") from the group of wastes flowing on the belt conveyor. Therefore, when a plurality of types of waste flow mixed on the belt conveyor, it is necessary for the waste sorting device to identify the types of waste. When the waste sorting device recognizes various types of waste, recognition using a learned model generated by machine learning is effective.
[0005] However, since machine learning requires a vast amount of training data, if the annotation work performed on image data—which involves visually inspecting images—is done manually when generating training data from image data of various types of waste, then generating the training data becomes extremely labor-intensive.
[0006] Therefore, this disclosure proposes a technology that can improve the efficiency of annotation work. [Means for solving the problem]
[0007] The learning device of this disclosure comprises a first annotation unit, a learning unit, and a model management unit. The first annotation unit performs a first annotation process to assign feature information, which is information indicating the features of the object image, to the object image using a second trained model replicated from a first trained model used by an object selection device that selects a desired object based on the recognition results of the object image. The learning unit updates the second trained model by performing machine learning using the object image and the feature information assigned by the first annotation process as training data. The model management unit updates the first trained model using the updated second trained model. [Effects of the Invention]
[0008] Disclosure technology can improve the efficiency of annotation work. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 shows an example of the configuration of an object sorting system according to Embodiment 1 of this disclosure. [Figure 2] Figure 2 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 3] Figure 3 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 4] Figure 4 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 5]Figure 5 shows an example of the operation of the learning device according to Embodiment 1 of this disclosure. [Figure 6] Figure 6 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 7] Figure 7 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 8] Figure 8 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 9] Figure 9 shows an example of the operation of the learning device of Embodiment 1 of this disclosure. [Figure 10] Figure 10 shows an example of the operation of the learning device according to Embodiment 1 of this disclosure. [Figure 11] Figure 11 shows an example of the operation of the learning device according to Embodiment 1 of this disclosure. [Figure 12] Figure 12 shows an example of the operation of the learning device according to Embodiment 1 of this disclosure. [Modes for carrying out the invention]
[0010] The embodiments of this disclosure will be described below with reference to the drawings. In the following embodiments, identical components will be denoted by the same reference numerals.
[0011] [Example 1] <Configuration of the object sorting system> Figure 1 shows an example of the configuration of an object sorting system according to Embodiment 1 of this disclosure.
[0012] In Figure 1, the object sorting system 1 includes a learning device 10, a camera 20, an object sorting device 30, a display 40, and an input device 50. The display 40 and the input device 50 are connected to the learning device 10. Examples of the input device 50 include a pointing device such as a mouse and a keyboard. The learning device 10, the camera 20, and the object sorting device 30 are connected to each other via a network.
[0013] The learning device 10 includes a first annotation unit 11, a teacher data storage unit 12, a machine learning unit 13, a learned model storage unit 14 for annotation, a model management unit 15, a replication unit 16, and a second annotation unit 17.
[0014] The object sorting device 30 includes a learned model storage unit 31 for recognition, an image recognition unit 32, and a desired object extraction unit 33.
[0015] Hereinafter, as an example, a case where the object sorting system 1 shown in FIG. 1 is installed in a waste treatment plant where a group of wastes flows on a belt conveyor will be described. That is, hereinafter, a case where the object to be sorted by the object sorting system 1 is waste will be described as an example. However, the object sorting system 1 may be installed in an assembly plant or the like where a group of parts flows on a belt conveyor. That is, the object to be sorted by the object sorting system 1 is not limited to waste, and the object sorting system 1 can be used for various objects.
[0016] The camera 20 is disposed above the belt conveyor on which the group of wastes is conveyed, and photographs the group of wastes conveyed by the belt conveyor. The image captured by the camera 20 (hereinafter sometimes referred to as "captured image") is transmitted from the camera 20 to both the learning device 10 and the object sorting device 30. The first annotation unit 11 and the image recognition unit 32 receive the captured image transmitted from the camera 20. <The desired object extraction unit 33 extracts the desired waste from the group of waste being transported by the belt conveyor, according to the recognition results from the image recognition unit 32. Examples of the desired object extraction unit 33 include a robot hand and a suction pad.
[0019] On the other hand, in the learning device 10, the duplication unit 16 operates only when the annotation-trained model is initialized. When the annotation-trained model is initialized, the duplication unit 16 duplicates the recognition-trained model stored in the recognition-trained model storage unit 31, and stores the duplicated recognition-trained model in the annotation-trained model storage unit 14, thereby initializing the annotation-trained model with the recognition-trained model.
[0020] The first annotation unit 11 uses the pre-trained annotation model stored in the pre-trained annotation model storage unit 14 to recognize each waste image present in the captured image and performs a process (hereinafter sometimes referred to as "first annotation process") to attach information indicating the features of each waste image (hereinafter sometimes referred to as "feature information") to each waste image. The first annotation unit 11 attaches feature information to waste images whose image recognition score is equal to or greater than the threshold TH1. The feature information includes information indicating the contour of the waste image (hereinafter sometimes referred to as "contour information") and information indicating the color of the waste image (hereinafter sometimes referred to as "color information"). The first annotation unit 11 associates the waste images with the feature information already attached to the waste images and generates training data (hereinafter sometimes referred to as "first training data") that includes the waste images and feature information. Therefore, the data in which feature information has been attached to each waste image in the captured image becomes the first training data. The first annotation unit 11 stores the generated first training data in the training data storage unit 12. The first annotation unit 11 recognizes waste images, for example, by performing instance segmentation on the captured image.
[0021] The second annotation unit 17 acquires the first training data stored in the training data storage unit 12 and displays the captured images to which feature information has been added for each waste image on the display 40. The input device 50 is operated by an operator, who can specify any part of the captured image displayed on the display 40 using the input device 50. The second annotation unit 17 uses the trained annotation model stored in the trained annotation model storage unit 14 to recognize an image of an arbitrarily specified part of the captured image (hereinafter sometimes referred to as the "specified area image") and performs a process (hereinafter sometimes referred to as the "second annotation process") to add feature information to the specified area image. The second annotation unit 17 adds feature information to specified area images whose image recognition score is greater than or equal to a threshold TH2 which has a value smaller than the threshold TH1 used by the first annotation unit 11. The second annotation unit 17 associates the specified location image with the feature information already assigned to the specified location image, and generates training data (hereinafter sometimes referred to as "second training data") that includes the specified location image and the feature information. Therefore, the data on which feature information has been assigned to the specified location image in the captured image becomes the second training data. The second annotation unit 17 stores the generated second training data in the training data storage unit 12. The second annotation unit 17 recognizes waste images by, for example, performing instance segmentation on arbitrarily specified locations in the captured image.
[0022] The machine learning unit 13 performs machine learning using the first and second training data stored in the training data storage unit 12, and updates the annotation-trained model stored in the annotation-trained model storage unit 14 with the trained model after machine learning.
[0023] The model management unit 15 updates the recognition trained model stored in the recognition trained model storage unit 31 using the updated annotation trained model stored in the annotation trained model storage unit 14 (hereinafter sometimes referred to as the "updated annotation trained model"). The model management unit 15 performs a recognition test similar to the recognition process performed by the image recognition unit 32, for example, using the updated annotation trained model, and updates the recognition trained model with the updated annotation trained model when the recognition accuracy reaches the target accuracy.
[0024] <Operation of the learning device> Figures 2 to 12 show examples of the operation of the learning device according to Embodiment 1 of this disclosure. In the following, we assume that empty bottles are used as waste, and that the object sorting device 30 sorts each empty bottle flowing on the conveyor belt into brown empty bottles (hereinafter sometimes referred to as "brown bottles") and empty bottles of other colors (hereinafter sometimes referred to as "other-colored bottles") (hereinafter sometimes referred to as "other-colored bottles"). Furthermore, in the following, we assume that the desired waste is brown bottles, and that first and second training data for brown bottles are generated.
[0025] As shown in Figure 2, for example, the captured image I1 includes images of empty bottles (hereinafter sometimes referred to as "empty bottle images") B11 and B12 as waste images. Empty bottle image B11 is an image of a brown bottle (hereinafter sometimes referred to as "brown bottle image"), and empty bottle image B12 is an image of a bottle of another color (hereinafter sometimes referred to as "other color bottle image").
[0026] The first annotation unit 11 assigns contour information INb1 to the empty bottle image B11 and contour information INb2 to the empty bottle image B12 in the captured image I1. Contour information INb1 and INb2 are formed by a plurality of coordinate points [x1, y1], [x2, y2], ... Furthermore, the first annotation unit 11 assigns label information INa1 and image information INc1 to the empty bottle image B11 in the captured image I1, and assigns label information INa2 and image information INc2 to the empty bottle image B12. Label information INa1 contains the color information "brown bottle" for the empty bottle image B11, and label information INa2 contains the color information "other color bottle" for the empty bottle image B12. In addition, image information INc1 and INc2 contain the image file name "XXX.bmp" indicating the captured image I1, and height information "1536" and width information "2048" indicating the pixel size of the captured image I1. Then, annotation data AD1 is formed for the empty bottle image B11 using label information INa1, contour information INb1, and image information INc1, and annotation data AD2 is formed for the empty bottle image B12 using label information INa2, contour information INb2, and image information INc2. The first annotation unit 11 associates the captured image I1 with the annotation data AD1 and AD2 with each other and generates first training data including the captured image I1 and the annotation data AD1 and AD2. The annotation data AD1 and AD2 are generated by the first annotation unit 11, for example, as a JSON format file.
[0027] For example, as shown in Figure 3, the captured image I2 includes empty bottle images B21, B22, B23, B24, and B25 as waste images. The first annotation unit 11, which performs the first annotation process on the captured image I2, assigns label information "brown bottle" and contour information CO1 to empty bottle image B21, label information "other color bottle" and contour information CO2 to empty bottle image B22, label information "brown bottle" and contour information CO4A to empty bottle image B24, and label information "brown bottle" and contour information CO5 to empty bottle image B25.
[0028] Here, we assume that the correct color of empty bottle image B21 is brown, the correct color of empty bottle image B22 is another color, the correct color of empty bottle image B24 is brown, and the correct color of empty bottle image B25 is another color. Therefore, the label information assigned to empty bottle images B21, B22, and B24 is correct, while the label information assigned to empty bottle image B25 is incorrect, as shown in Figure 3. Thus, empty bottle image B25 is a waste image (hereinafter sometimes referred to as a "failed image") in which the assignment of feature information by the first annotation process failed.
[0029] Furthermore, as shown in Figure 3, the contour indicated by the contour information CO4A attached to the empty bottle image B24 is deviated from the correct contour of the empty bottle image B24. Therefore, the empty bottle image B24, like the empty bottle image B25, is a failed image.
[0030] Furthermore, although the captured image I2 includes the empty bottle image B23, as shown in Figure 3, label information and contour information have not been assigned to the empty bottle image B23. In other words, the first annotation unit 11 has not detected the empty bottle image B23. Therefore, the empty bottle image B23, like the empty bottle images B24 and B25, is a failed image.
[0031] Furthermore, as shown in Figure 3, even though region R6 in the captured image I2 does not contain an empty bottle image with a contour indicated by contour information CO6, the label information "brown bottle" and contour information CO6 are assigned to region R6. In other words, the first annotation unit 11 mistakenly detected region R6 as an empty bottle image.
[0032] As described above, in the captured image I2, the empty bottle images B23, B24, and B25 are failure images, and region R6 is incorrectly detected as an empty bottle image.
[0033] Therefore, for example, the operator uses the pointer PO displayed on the captured image I2 by operating the input device 50 to specify the empty bottle image B23 by clicking on the empty bottle image B23 in the captured image I2, for example as shown in Figure 4. Alternatively, the operator uses the pointer PO to specify the empty bottle image B23 by drawing a region RS surrounding the location of the empty bottle image B23 in the captured image I2, for example as shown in Figure 5. Once the empty bottle image B23 is specified by operating the input device 50, the second annotation unit 17 uses the trained annotation model stored in the trained annotation model storage unit 14 to assign contour information CO3 to the empty bottle image B23, as shown in Figure 6.
[0034] For example, the operator uses the pointer PO to specify region R6 by clicking on the area enclosed by the contour indicated by the contour information CO6 in the captured image I2, as shown in Figure 7. Once region R6 is specified by the operation of the input device 50, the second annotation unit 17 deletes the contour information CO6 that was assigned to region R6, as shown in Figure 8.
[0035] For example, the operator uses the pointer PO to click on the label information "brown bottle" attached to the empty bottle image B25 in the captured image I2, as shown in Figure 9, thereby specifying the label information for the empty bottle image B25. Once the label information is specified by the operation of the input device 50, the second annotation unit 17 enables the modification of the specified label information. The operator then uses, for example, the keyboard to modify the label information attached to the empty bottle image B25 from "brown bottle" to "other color bottle," as shown in Figure 10. In accordance with the operator's modification, the second annotation unit 17 modifies the label information "brown bottle" attached to the empty bottle image B25 to the label information "other color bottle" which indicates the correct color of the empty bottle image B25.
[0036] For example, the operator uses the pointer PO to specify the contour indicated by the contour information CO4A in the captured image I2, as shown in Figure 11. Once the contour is specified by the operation of the input device 50, the second annotation unit 17 enables the modification of the specified contour. The operator then uses the pointer PO to modify the contour of the empty bottle image B24 to the correct contour, as shown in Figure 12. For example, the operator modifies the contour by moving the vertices of the contour indicated by the contour information CO4A using the pointer PO. In accordance with the operator's modification, the second annotation unit 17 modifies the contour information CO4A attached to the empty bottle image B24 to contour information CO4B that indicates the correct contour of the empty bottle image B24.
[0037] The above describes Example 1.
[0038] [Example 2] The training data storage unit 12, the annotation-trained model storage unit 14, and the recognition-trained model storage unit 31 are implemented as hardware, for example, by memory or storage. The first annotation unit 11, the machine learning unit 13, the model management unit 15, the replication unit 16, the second annotation unit 17, and the image recognition unit 32 are implemented as hardware, for example, by a processor such as a CPU (Central Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).
[0039] The above describes Example 2.
[0040] As described above, the learning device of this disclosure (learning device 10 in the embodiment) comprises a first annotation unit (first annotation unit 11 in the embodiment), a learning unit (machine learning unit 13 in the embodiment), and a model management unit (model management unit 15 in the embodiment). The first annotation unit performs a first annotation process to attach feature information, which is information indicating the features of an object image, to an object image using a second trained model (trained model for annotation in the embodiment) which is duplicated from a first trained model (trained model for recognition in the embodiment) used by an object selection device (object selection device 30 in the embodiment) that selects a desired object based on the recognition result of an object image. The learning unit updates the second trained model by performing machine learning using the object image and the feature information attached by the first annotation process as training data. The model management unit updates the first trained model using the updated second trained model.
[0041] This allows the first annotation unit to automatically perform annotation on object images, thereby increasing the efficiency of the annotation process. Furthermore, by sharing the trained model used for object image recognition in the object sorting device with the first annotation process, it becomes unnecessary to prepare a separate trained model for the first annotation process, further improving the efficiency of annotation. Additionally, since the second trained model is updated as needed, the accuracy of the first annotation process improves with each update to the second trained model. Moreover, since the first trained model is updated as needed with each update to the second trained model, the accuracy of object image recognition in the object sorting device also improves with each update to the second trained model.
[0042] Furthermore, the learning device of this disclosure (learning device 10 in the embodiment) has a second annotation unit (second annotation unit 17 in the embodiment). The second annotation unit performs a second annotation process to arbitrarily designated failure images, which are object images for which the first annotation process failed to assign feature information, by using a second trained model to assign feature information. The learning unit updates the second trained model by performing machine learning using the object image, the feature information assigned by the first annotation process, and the feature information assigned by the second annotation process as training data. For example, the feature information includes contour information that shows the contour of the object image.
[0043] This improves the efficiency of annotation work performed by operators on failed images. In particular, it reduces the effort required to add contour information to failed images. Furthermore, the accuracy of the second pre-trained model is improved by the second annotation process, which further improves the accuracy of the first annotation process.
[0044] Furthermore, the score threshold used when feature information is assigned by the second annotation process (threshold TH2 in the example) is smaller than the score threshold used when feature information is assigned by the first annotation process (threshold TH1 in the example).
[0045] By doing this, the success rate of recognizing failed images in the second annotation process can be increased compared to the success rate of recognizing object images in the first annotation process. [Explanation of Symbols]
[0046] 1. Object sorting system 10 Learning device 11. First Annotation Section 12. Training data storage unit 13 Machine Learning Department 14. Memory for pre-trained models used for annotation. 15 Model Management Department 16 Reproduction Department 17 Second Annotation Section 20 cameras 30 Object sorting device 31. Memory unit for pre-trained models for recognition. 32 Image Recognition Unit 33 Desired object extraction section
Claims
[Claim 1] A first annotation unit uses a second trained model, which is a replica of a first trained model used by an object sorting device that sorts desired objects based on the recognition results of an object image, to add feature information, which is information indicating the features of the object image, to the object image. A second annotation unit modifies the feature information based on the specified image, targeting the object image in which the first annotation unit failed to assign the feature information. A learning unit updates the first trained model by performing machine learning using the object image, the feature information assigned by the first annotation unit, and the feature information modified by the second annotation unit as training data. It is equipped with, The score threshold used when the feature information is modified by the second annotation unit is smaller than the score threshold used when the feature information is assigned by the first annotation unit. Learning device.
Citation Information
Patent Citations
Waste screening system and screening method therefor
JP2017109161A
Apparatus and method for online recognition and setting screen used therefor
JP2019074945A
Robot system and robot control method
JP2021010970A
Information processing apparatus, information processing method, and information processing program
JP2021043881A
Information processing apparatus, information processing method, and program
JP2021099582A