Annotation Device

The annotation device automates the generation of training data for waste sorting devices by comparing pre-and post-sorting images to determine difference regions and apply contour information, improving efficiency and accuracy.

JP7761658B2Active Publication Date: 2025-10-28PFU LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023548074
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-17
Publication Date
2025-10-28
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Manual annotation of image data for machine learning in waste sorting devices is labor-intensive and inefficient.

Method used

An annotation device that uses upstream and downstream cameras to capture images before and after a sorting operation, determines difference regions, assigns contour information, and selectively generates candidate data for training based on region area or feature similarity thresholds.

Benefits of technology

Improves annotation efficiency and accuracy by automating the process, reducing labor and preventing erroneous data from being used for training, thereby enhancing the performance of waste sorting devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007761658000001
    Figure 0007761658000001
  • Figure 0007761658000002
    Figure 0007761658000002
  • Figure 0007761658000003
    Figure 0007761658000003
Patent Text Reader

Abstract

In an annotation device 10a, an image acquisition unit 11 acquires an upstream image I1 that has been taken at a first position on an upstream side of a working position at a first timing before a work for selecting a desired waste from a waste group conveyed on a belt conveyor is performed and a downstream image I2 that has been taken at a second position on a downstream side of the working position at a second timing after the work is performed. A differential region determination unit 13 determines a differential region that is a region where a difference is present between the upstream image I1 and the downstream image I2. An annotation unit 14 provides a differential region in the upstream image I1 with contour information that is information indicating the contour of the differential region, and generates candidate data including the upstream image I1 and the contour information. An adoption determination unit 16 determines whether to adopt the candidate data as training data for machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an annotation device. [Background technology]

[0002] At waste disposal sites, large amounts of waste are processed daily on conveyor belts. At the waste disposal sites, the waste is sorted by hand. While sorting waste is a simple task, it places a heavy burden on the workers who sort the waste (hereinafter referred to as "sorters"). Therefore, devices that automatically sort waste (hereinafter referred to as "waste sorting devices") have been developed. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-049780 Summary of the Invention [Problem to be solved by the invention]

[0004] When a waste sorting device performs the tasks previously performed by a sorter, the waste sorting device can recognize each piece of waste flowing on the conveyor belt and use a robotic hand to pick up the desired waste based on the recognition results. Therefore, when multiple types of waste are mixed and flowing on the conveyor belt, the waste sorting device needs to identify the type of waste. For the waste sorting device to recognize various types of waste, recognition using a trained model generated by machine learning is effective.

[0005] However, machine learning requires a huge amount of training data, and when generating training data from image data of various types of waste, if the annotation work on the image data is done manually by visually inspecting the images, generating the training data will require a huge amount of effort.

[0006] Therefore, the present disclosure proposes a technique that can improve the efficiency of annotation work. [Means for solving the problem]

[0007] The annotation device disclosed herein includes an acquisition unit, a first determination unit, an annotation unit, and a second determination unit. The acquisition unit acquires a first image captured at a first timing before an operation of selecting a desired object from a group of objects on a transport path along which the group of objects is transported is performed, at a first position upstream of the position where the operation is performed, and a second image captured at a second timing after the operation is performed, at a second position downstream of the position where the operation is performed. The first determination unit compares the first image with the second image and determines a difference region, which is an area where there is a difference between the first image and the second image. The annotation unit assigns contour information, which is information indicating the contour of the difference region, to the difference region in the first image, and generates candidate data including the first image and the contour information. The second determination unit determines whether to use the candidate data as training data for machine learning. [Effects of the Invention]

[0008] According to the disclosed technology, the efficiency of annotation work can be improved. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating a configuration example of an annotation system according to a first embodiment of the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating a configuration example of an annotation device according to a first embodiment of the present disclosure. [Figure 3]FIG. 3 is a flowchart illustrating an example of a processing procedure in the annotation device according to the first embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of the operation of the annotation device according to the first embodiment of the present disclosure. [Figure 5] FIG. 5 is a diagram illustrating an example of the operation of the annotation device according to the first embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an example of the operation of the annotation device according to the first embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating a configuration example of an annotation device according to a second embodiment of the present disclosure. [Figure 8] FIG. 8 is a flowchart illustrating an example of a processing procedure in the annotation device according to the second embodiment of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating an example of the operation of the annotation device according to the second embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating an example of the operation of the annotation device according to the third embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. In the following embodiments, the same configurations and steps performing the same processes are denoted by the same reference numerals.

[0011] [Example 1] <Configuration of the annotation system> FIG. 1 is a diagram illustrating a configuration example of an annotation system according to a first embodiment of the present disclosure.

[0012] 1, the annotation system 1 includes an annotation device 10, an upstream camera 20, a downstream camera 30, and a belt conveyor 40. The annotation device 10 is connected to the upstream camera 20 and the downstream camera 30 via a network. The upstream camera 20 and the downstream camera 30 are disposed above the belt conveyor 40 and capture images of the upper surface of the belt conveyor 40 from above the belt conveyor 40.

[0013] 1 will be described as an example where the annotation system 1 is installed in a waste disposal site where waste materials flow on a belt conveyor 40. However, the annotation system 1 may also be installed in an assembly plant or the like where parts flow on a belt conveyor 40. In other words, the objects that can be targeted by the annotation system 1 are not limited to waste materials, and the annotation system 1 can be used for a variety of objects.

[0014] The belt conveyor 40 transports the waste mass placed on the belt conveyor 40 in the transport direction CD. In other words, the belt conveyor 40 forms a transport path along which the waste mass is transported in the transport direction CD.

[0015] The upstream camera 20 is disposed at a first position P1 on the upstream side of the belt conveyor 40, and the downstream camera 30 is disposed at a second position P2 on the downstream side of the belt conveyor 40. A sorter M is disposed at a third position P3 between the first position P1 and the second position P2, where the sorter M sorts desired waste from a group of waste transported by the belt conveyor 40 (hereinafter, sometimes referred to as "waste sorting work"). Thus, the upstream camera 20 photographs the waste on the belt conveyor 40 from above the belt conveyor 40 at a first position P1 upstream of the third position P3 where the sorter M performs the waste sorting work and at a first timing T1 before the waste sorting work is performed. The downstream camera 30 photographs the waste on the belt conveyor 40 from above the belt conveyor 40 at a second position P2 downstream of the third position P3 where the sorter M performs the waste sorting work and at a second timing T2 after the waste sorting work is performed. An image I1 captured by the upstream camera 20 (hereinafter sometimes referred to as the "upstream image") is transmitted from the upstream camera 20 to the annotation device 10, and an image I2 captured by the downstream camera 30 (hereinafter sometimes referred to as the "downstream image") is transmitted from the downstream camera 30 to the annotation device 10. The upstream image I1 and the downstream image I2 form a pair of images (hereinafter sometimes referred to as the "image set"). In order for the sorter M to perform the waste sorting work, the conveying speed of the belt conveyor 40 is preferably set to, for example, about 170 mm per second.

[0016] The annotation device 10 performs annotation using the upstream image I1 and the downstream image I2.

[0017] <Configuration of annotation device> Fig. 2 is a diagram illustrating a configuration example of an annotation device according to a first embodiment of the present disclosure. An annotation device 10a illustrated in Fig. 2 corresponds to the annotation device 10 illustrated in Fig. 1. In Fig. 2, the annotation device 10a includes an image acquisition unit 11, an image storage unit 12, a differential region determination unit 13, an annotation unit 14, a candidate data storage unit 15, an adoption determination unit 16, and a teacher data storage unit 17.

[0018] The image acquisition unit 11 acquires an upstream image I1 transmitted from the upstream camera 20, and acquires a downstream image I2 transmitted from the downstream camera 30. The image acquisition unit 11 stores the acquired upstream image I1 and downstream image I2 in the image storage unit 12.

[0019] The differential region determination unit 13 reads the upstream image I1 and the downstream image I2 stored in the image storage unit 12 from the image storage unit 12, compares the upstream image I1 with the downstream image I2, and determines a region where there is a difference between the upstream image I1 and the downstream image I2 (hereinafter, sometimes referred to as a "difference region"). The differential region determination unit 13 determines the differential region using a matching method such as SAD (Sum of Absolute Difference), SSD (Sum of Squared Difference), or NCC (Normalized Cross-Correlation). Since a differential region is determined as a region connecting multiple coordinate points, the differential region determination unit 13 outputs coordinate information indicating the differential region (hereinafter, sometimes referred to as "difference region information") to the annotation unit 14.

[0020] The annotation unit 14 reads the upstream image I1 stored in the image storage unit 12 from the image storage unit 12. Then, based on the differential region information, the annotation unit 14 assigns information indicating the contour of the differential region (hereinafter, sometimes referred to as "contour information") to the differential region in the upstream image I1, and generates candidate data including the upstream image I1 and the contour information. The annotation unit 14 stores the generated candidate data in the candidate data storage unit 15. The annotation unit 14 also outputs the contour information assigned to the differential region in the upstream image I1 to the adoption determination unit 16.

[0021] The adoption determination unit 16 determines whether or not to adopt the candidate data stored in the candidate data storage unit 15 as training data for machine learning based on the contour information (hereinafter, this may be referred to as "training data adoption determination"). The adoption determination unit 16 reads the candidate data that it has determined to adopt as training data (hereinafter, this may be referred to as "adopted data") from the candidate data storage unit 15 and stores it in the training data storage unit 17, and deletes the adopted data from the candidate data storage unit 15. On the other hand, the adoption determination unit 16 deletes the rejected data from the candidate data storage unit 15 without storing the candidate data that it has determined not to adopt as training data (hereinafter, this may be referred to as "rejected data") in the training data storage unit 17.

[0022] The training data stored in the training data storage unit 17 is used for machine learning to generate a trained model for the waste sorting device to recognize various types of waste.

[0023] <Processing procedure in annotation device> FIG. 3 is a flowchart illustrating an example of a processing procedure in the annotation device according to the first embodiment of the present disclosure.

[0024] As shown in FIG. 3, in step S100, at a first timing T1 before the waste sorting work is performed, an upstream image I1 is captured by the upstream camera 20 and transmitted to the annotation device 10a.

[0025] Next, in step S105, the sorter M performs the waste sorting work.

[0026] Next, in step S110, at a second timing T2 after the waste sorting work has been performed, a downstream image I2 is captured by the downstream camera 30 and transmitted to the annotation device 10a.

[0027] Next, in step S115, the differential region determining unit 13 determines the differential region.

[0028] Next, in step S120, the annotation unit 14 performs annotation to add contour information to the difference region in the upstream image I1. Then, in step S125, the annotation unit 14 generates candidate data including the upstream image I1 and the contour information. Then, in step S130, the annotation unit 14 stores the candidate data in the candidate data storage unit 15.

[0029] Next, in step S135, the adoption determination unit 16 calculates the area of ​​the differential region (hereinafter sometimes referred to as the "differential region area") based on the contour information, and determines whether to adopt the training data based on the calculated differential region area.

[0030] If the differential region area is less than the first threshold TH1 (step S135: Yes), the adoption determination unit 16 adopts the candidate data as training data and stores the candidate data as adopted data in the training data storage unit 17 (step S140).

[0031] On the other hand, when the differential region area is equal to or greater than the first threshold value TH1 (step S135: No), the adoption determination unit 16 does not adopt the candidate data as training data, and deletes the unadopted candidate data from the candidate data storage unit 15.

[0032] <Operation of the annotation device> 4, 5, and 6 are diagrams illustrating an example of the operation of the annotation device according to the first embodiment of the present disclosure. Hereinafter, it is assumed that empty bottles are waste materials, and that the waste sorting device sorts each empty bottle traveling on a belt conveyor into three types of bottles: brown empty bottles (hereinafter sometimes referred to as "brown bottles"), transparent empty bottles (hereinafter sometimes referred to as "transparent bottles"), and empty bottles of a color other than brown or transparent (hereinafter sometimes referred to as "other-colored bottles"). Hereinafter, it is assumed that a sorter M sorts brown bottles as desired waste materials from a group of waste materials transported by a belt conveyor 40 in order to generate training data for brown bottles.

[0033] As shown in FIG. 4, the first image set IS1 includes a first upstream image I1a and a first downstream image I2a. The first upstream image I1a also includes an image of a brown bottle (hereinafter sometimes referred to as a “brown bottle image”) B1 and an image of bottles of other colors (hereinafter sometimes referred to as “other-color bottle images”) B2. After the first upstream image I1a is captured by the upstream camera 20, a sorter M sorts the brown bottles. After the sorter M sorts the brown bottles, the downstream camera 30 captures the first downstream image I2a. Therefore, the first downstream image I2a includes the other-color bottle image B2 but does not include the brown bottle image B1. Therefore, the differential region determination unit 13 determines the first differential region DRa by comparing the first upstream image I1a and the first downstream image I2a. The outline of the first differential region DRa in the first upstream image I1a corresponds to the outline of the brown bottle image B1 in the first upstream image I1a. Therefore, the differential region determination unit 13 outputs differential region information formed by a plurality of coordinate points indicating the first differential region DRa to the annotation unit 14.

[0034] Next, as shown in FIG. 5, the annotation unit 14 assigns contour information INb to the first differential region DRa in the first upstream image I1a. Similar to the differential region information, the contour information INb is formed by multiple coordinate points [x1, y1], [x2, y2], .... Therefore, the contour information INb can be used to determine the position and area of ​​the first differential region DRa. Furthermore, the annotation unit 14 assigns label information INa and image information INc to the first differential region DRa. The label information INa includes information indicating the color of the empty bottle, such as "Brown Bottle," and the image information INc includes the image file name "XXX.bmp" indicating the first upstream image I1a, and height information "1536" and width information "2048" indicating the pixel size of the first upstream image I1a. The label information INa, contour information INb, and image information INc then form annotation data AD1. The annotation unit 14 associates the first upstream image I1a with the annotation data AD1, and generates candidate data including the first upstream image I1a and the annotation data AD1. The annotation data AD1 including the label information INa, the contour information INb, and the image information INc is generated by the annotation unit 14 as a file in, for example, a JSON format.

[0035] Here, waste materials transported by the belt conveyor 40 are often densely packed, for example, piled up. Therefore, when sorting waste materials, the sorter M may shuffle the waste materials while sorting the desired waste materials. If the waste materials are shuffled, the positions of each waste material in the waste materials change significantly before and after the waste sorting process, resulting in significant differences between the upstream and downstream images. For example, as shown in FIG. 6 , the area of ​​the second difference region DRb increases between the second upstream image I1b and the second downstream image I2b included in the second image set IS2. If the area of ​​the second difference region DRb increases in this way, the contour of the second difference region DRb no longer represents the shape of an empty bottle, increasing the likelihood of errors in the annotation data generated by the annotation unit 14.

[0036] Therefore, the adoption determination unit 16 adopts the candidate data as training data when the differential region area is less than the first threshold TH1, but does not adopt the candidate data as training data when the differential region area is equal to or greater than the first threshold TH1. This prevents an upstream image associated with erroneous annotation data from being adopted as training data.

[0037] In the above description of Example 1, the teacher data adoption determination was made based on the differential region area. However, instead of making the teacher data adoption determination based on the differential region area, the differential region determination unit 13 may calculate the differential region area, and the annotation unit 14 may perform annotation when the differential region area is less than the first threshold TH1, but may not perform annotation when the differential region area is equal to or greater than the first threshold TH1. Alternatively, when the differential region area is equal to or greater than the first threshold TH1, the differential region determination unit 13 may delete the image set used to calculate the differential region area (hereinafter sometimes referred to as the "calculation source image set") from the image storage unit 12, thereby preventing the annotation unit 14 from annotating the upstream image I1 included in the calculation source image set.

[0038] The first embodiment has been described above.

[0039] [Example 2] The second embodiment differs from the first embodiment in that the training data adoption determination is performed based on the image characteristics of the differential region. The differences from the first embodiment will be described below.

[0040] <Configuration of annotation device> Fig. 7 is a diagram illustrating a configuration example of an annotation device according to a second embodiment of the present disclosure. The annotation device 10b illustrated in Fig. 2 corresponds to the annotation device 10 illustrated in Fig. 1. In Fig. 7, the annotation device 10b includes an image acquisition unit 11, an image storage unit 12, a differential region determination unit 13, an annotation unit 14, a candidate data storage unit 15, an adoption determination unit 16, a teacher data storage unit 17, and a feature detection unit 18.

[0041] The feature detection unit 18 reads the upstream image I1 and the downstream image I2 stored in the image storage unit 12 from the image storage unit 12. In addition, the feature detection unit 18 receives input of differential region information from the differential region determination unit 13.

[0042] The feature detection unit 18 detects image features of the difference region in the upstream image I1 (hereinafter, sometimes referred to as "upstream difference region features") from the upstream image I1. The feature detection unit 18 also detects image features of the difference region in the downstream image I2 (hereinafter, sometimes referred to as "downstream difference region features") from the downstream image I2. The feature detection unit 18 then outputs information including information indicating the upstream difference region features and information indicating the downstream difference region features (hereinafter, sometimes referred to as "feature information") to the adoption determination unit 16. The feature detection unit 18 detects the upstream difference region features and the downstream difference region features using, for example, a method such as SIFT (Scaled Invariance Feature Transform) or HOG (Histograms of Oriented Gradients).

[0043] The adoption determination unit 16 determines whether to adopt the training data based on the characteristic information.

[0044] <Processing procedure in annotation device> Fig. 8 is a flowchart showing an example of a processing procedure in an annotation device according to Example 2 of the present disclosure. In Fig. 8, the processing of steps S100 to S115 and the processing of steps S120 to S130 are the same as those in Example 1 (Fig. 3), and therefore a description thereof will be omitted.

[0045] As shown in FIG. 8, in step S200, the feature detection unit 18 detects upstream differential region features from the upstream image I1, and detects downstream differential region features from the downstream image I2.

[0046] In step S205, the adoption determination unit 16 calculates the degree of difference between the upstream differential region feature and the downstream differential region feature (hereinafter, sometimes referred to as "feature difference") based on the feature information.

[0047] Next, in step S210, the adoption determination unit 16 determines whether to adopt the training data based on the feature difference.

[0048] If the feature dissimilarity is equal to or greater than the second threshold TH2 (step S210: Yes), the adoption determination unit 16 adopts the candidate data as training data and stores the candidate data as the adopted data in the training data storage unit 17 (step S140).

[0049] On the other hand, when the feature dissimilarity is less than the second threshold TH2 (step S210: No), the adoption determination unit 16 does not adopt the candidate data as training data, and deletes the unadopted candidate data from the candidate data storage unit 15.

[0050] <Operation of the annotation device> FIG. 9 is a diagram illustrating an example of the operation of the annotation device according to the second embodiment of the present disclosure.

[0051] During waste sorting, the sorter M may touch waste other than the desired waste (hereinafter referred to as "undesired waste") with his / her hands, causing the undesired waste to shift in position relative to the belt conveyor 40. Examples of undesired waste that may shift in position during waste sorting include waste adjacent to the desired waste and waste that is in the path of the sorter M's hands. For example, if the desired waste in the waste sorting is brown bottles, and the sorter M touches other-colored bottles while sorting the brown bottles, the position of the other-colored bottle image B2 may differ between the first upstream image I1a and the third downstream image I2c included in the third image set IS3, as shown in FIG. 9. If there is a difference in the position of the other-colored bottle image B2 between the first upstream image I1a and the third downstream image I2c, the differential region determination unit 13 will determine a third differential region DRc, and the annotation unit 14 will annotate the third differential region DRc in the first upstream image I1a. However, since the other color bottles are unwanted waste, annotation should not be performed on the third difference region DRc.

[0052] 9, the feature detection unit 18 detects, from the first upstream image I1a, image features of the first differential region DRa in the first upstream image I1a (hereinafter sometimes referred to as "first upstream differential region features") and image features of the third differential region DRc in the first upstream image I1a (hereinafter sometimes referred to as "second upstream differential region features"). Also, in FIG. 9, the feature detection unit 18 detects, from the third downstream image I2c, image features of the first differential region DRa in the third downstream image I2c (hereinafter sometimes referred to as "first downstream differential region features") and image features of the third differential region DRc in the third downstream image I2c (hereinafter sometimes referred to as "second downstream differential region features").

[0053] Next, the adoption determination unit 16 calculates the dissimilarity FD1 between the first upstream differential region feature and the first downstream differential region feature (hereinafter sometimes referred to as the “first feature dissimilarity”), and the dissimilarity FD2 between the second upstream differential region feature and the second downstream differential region feature (hereinafter sometimes referred to as the “second feature dissimilarity”). If both the first feature dissimilarity FD1 and the second feature dissimilarity FD2 are equal to or greater than the second threshold value TH2, the adoption determination unit 16 adopts the candidate data as training data. On the other hand, if either the first feature dissimilarity FD1 or the second feature dissimilarity FD2 is less than the second threshold value TH2, the adoption determination unit 16 does not adopt the candidate data as training data. This prevents an upstream image associated with erroneous annotation data from being adopted as training data.

[0054] In the above description of Example 2, the teacher data adoption determination is made based on the feature dissimilarity. However, instead of making the teacher data adoption determination based on the feature dissimilarity, the annotation unit 14 may annotate difference regions whose feature dissimilarity is equal to or greater than the second threshold value TH2, while not annotating difference regions whose feature dissimilarity is less than the second threshold value TH2.

[0055] The second embodiment has been described above.

[0056] [Example 3] In Example 3, the determination of whether training data is adopted is made based on the image features of the difference region, as in Example 2, but the difference from Example 2 is that the feature dissimilarity between two different image sets is calculated. The differences from Example 2 will be described below.

[0057] <Configuration of annotation device> The configuration example of the annotation device of the third embodiment is the same as that of the second embodiment (FIG. 7), and therefore the description thereof will be omitted.

[0058] <Processing procedure in annotation device> In the third embodiment, the processes of steps S100 to S110 in Fig. 8 are repeated to acquire two image sets, and the processes of steps S115, S200, and S120 to S130 in Fig. 8 are performed on the two consecutive image sets, after which the processes of step S205 and onwards are performed. Hereinafter, of the two consecutive image sets, the former set may be referred to as the "former image set" and the latter set may be referred to as the "latter image set".

[0059] By repeating the processes of steps S100 to S110, the image acquisition unit 11 acquires the former image set and then the latter image set. Therefore, in step S130, the candidate data generated based on the former image set (hereinafter may be referred to as "former candidate data") and the candidate data generated based on the latter image set (hereinafter may be referred to as "latter candidate data") are stored in the candidate data storage unit 15.

[0060] In step S200, the feature detection unit 18 detects upstream difference region features from the upstream image I1 included in the former image set, and detects downstream difference region features from the downstream image I2 included in the latter image set.

[0061] In step S205, the adoption determination unit 16 calculates the feature dissimilarity between the upstream differential region feature in the former image set and the downstream differential region feature in the latter image set.

[0062] In step S210, the adoption determination unit 16 determines whether to adopt the training data based on the feature dissimilarity. If the feature dissimilarity is equal to or greater than the threshold TH2 (step S210: Yes), the process proceeds to step S140. If the feature dissimilarity is less than the threshold TH2 (step S210: No), the process proceeds to step S145.

[0063] In step S140, the adoption determination unit 16 adopts both the former candidate data and the latter candidate data as training data, and stores the former candidate data and the latter candidate data, which are the adopted data, in the training data storage unit 17.

[0064] On the other hand, in step S145, the adoption determination unit 16 does not adopt either the former candidate data or the latter candidate data as training data, and deletes the former candidate data and the latter candidate data, which are unadopted data, from the candidate data storage unit 15.

[0065] <Operation of the annotation device> FIG. 10 is a diagram illustrating an example of the operation of the annotation device according to the third embodiment of the present disclosure.

[0066] When sorting waste, the sorter M may accidentally pick up undesired waste. Furthermore, the sorter M may realize several seconds (e.g., three seconds) after picking up the undesired waste and return the undesired waste to the belt conveyor 40. For example, if the conveying speed of the belt conveyor 40 is 170 mm per second, the belt conveyor 40 will move approximately 50 cm in three seconds. Therefore, if the angle of view of the downstream camera 30 is small, the downstream image I2 of the former image set may not include an image of the undesired waste (hereinafter, sometimes referred to as a "undesired waste image"), and the downstream image I2 of the latter image set may include an undesired waste image.

[0067] For example, it is assumed that brown bottles are desired waste, and the sorter M mistakenly picks up bottles of other colors, which are undesired waste, and then returns them to the belt conveyor 40.

[0068] In this case, if sorter M accidentally picks up a bottle of a different color, as shown in Figure 10, of the third upstream image I1c and the fourth downstream image I2d included in the former image set IS4, the third upstream image I1c includes the bottle image B3 of the different color and the bottle image B4 of the brown color, while the fourth downstream image I2d includes the bottle image B4 of the brown color but not the bottle image B3 of the different color. In other words, in the former image set IS4, the bottle image B3 of the different color is included in the third upstream image I1c but not in the fourth downstream image I2d. Therefore, the differential region determination unit 13 determines the fourth differential region DRd, and the annotation unit 14 annotates the fourth differential region DRd in the third upstream image I1c.

[0069] Furthermore, when sorter M returns the other-colored bottle he mistakenly picked up to the conveyor belt 40, as shown in FIG. 10 , of the fourth upstream image I1d and the fifth downstream image I2e included in the latter image set IS5, the fourth upstream image I1d includes the other-colored bottle image B5 and the clear bottle image (hereinafter sometimes referred to as the “clear bottle image”) B6, but does not include the other-colored bottle image B7, while the fifth downstream image I2e includes the other-colored bottle image B5, the clear bottle image B6, and the other-colored bottle image B7. In other words, in the latter image set IS5, the other-colored bottle image B7 is not included in the fourth upstream image I1d but is included in the fifth downstream image I2e. Therefore, the differential region determination unit 13 determines the fifth differential region DRe, and the annotation unit 14 annotates the fifth differential region DRe in the fourth upstream image I1d.

[0070] However, since the bottles of other colors are undesired waste that were mistakenly picked up by the sorter M and then returned to the belt conveyor 40, annotation should not be made to the fifth difference region DRe.

[0071] 10, the feature detection unit 18 detects image features of the fourth differential region DRd in the third upstream image I1c of the former image set IS4 (hereinafter, sometimes referred to as the "third upstream differential region feature") from the third upstream image I1c. Also, in FIG. 10, the feature detection unit 18 detects image features of the fifth differential region DRe in the fifth downstream image I2e of the latter image set IS5 (hereinafter, sometimes referred to as the "third downstream differential region feature") from the fifth downstream image I2e.

[0072] Next, the adoption determination unit 16 calculates the dissimilarity FD3 between the third upstream differential region feature and the third downstream differential region feature (hereinafter sometimes referred to as the "third feature dissimilarity") FD3. When the third feature dissimilarity FD3 is equal to or greater than the second threshold TH2, the adoption determination unit 16 adopts both the former candidate data and the latter candidate data as training data. On the other hand, when the third feature dissimilarity FD3 is less than the second threshold TH2, the adoption determination unit 16 does not adopt the former candidate data or the latter candidate data as training data. This makes it possible to prevent an upstream image associated with erroneous annotation data from being adopted as training data.

[0073] In the above description of Example 3, the teacher data adoption determination is made based on the feature dissimilarity. However, instead of making the teacher data adoption determination based on the feature dissimilarity, the annotation unit 14 may annotate difference regions whose feature dissimilarity is equal to or greater than the second threshold value TH2, while not annotating difference regions whose feature dissimilarity is less than the second threshold value TH2.

[0074] Furthermore, the threshold value TH2 used in the third embodiment may be set to a value greater than the threshold value TH2 used in the second embodiment. By doing so, the third embodiment can reduce the rate at which candidate data is adopted as training data compared to the second embodiment.

[0075] The third embodiment has been described above.

[0076] The above-described first, second, and third embodiments may be combined as appropriate and implemented in parallel.

[0077] [Example 4] The image acquisition unit 11 is realized as hardware, for example, by a network interface module. The image storage unit 12, candidate data storage unit 15, and teacher data storage unit 17 are realized as hardware, for example, by memory or storage. The difference region determination unit 13, annotation unit 14, adoption determination unit 16, and feature detection unit 18 are realized as hardware, for example, by a processor such as a CPU (Central Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).

[0078] The fourth embodiment has been described above.

[0079] As described above, the annotation device (annotation devices 10, 10a, and 10b in the embodiments) of the present disclosure includes an acquisition unit (image acquisition unit 11 in the embodiments), a first determination unit (difference region determination unit 13 in the embodiments), an annotation unit (annotation unit 14 in the embodiments), and a second determination unit (adoption determination unit 16 in the embodiments). The acquisition unit acquires a first image (upstream image in the embodiments) captured at a first timing before an operation of selecting a desired object from a group of objects on a conveyance path along which the group of objects is conveyed is performed, at a first position upstream of the position where the operation is performed, and a second image (downstream image in the embodiments) captured at a second timing after the operation is performed, at a second position downstream of the position where the operation is performed. The first determination unit compares the first image with the second image and determines a difference region, which is a region where there is a difference between the first image and the second image. The annotation unit assigns contour information, which is information indicating the contour of the difference region, to the difference region in the first image, and generates candidate data including the first image and the contour information. The second determination unit determines whether or not the candidate data is to be adopted as training data for machine learning.

[0080] For example, the second determination unit adopts the candidate data as training data when the area of ​​the differential region is less than a threshold, and does not adopt the candidate data as training data when the area of ​​the differential region is equal to or greater than the threshold.

[0081] For example, the second judgment unit adopts the candidate data as training data when the degree of difference between the features of the first region image, which is an image of the difference region in the first image, and the features of the second region image, which is an image of the difference region in the second image, is greater than or equal to a threshold, and does not adopt the candidate data as training data when the degree of difference between the features is less than the threshold.

[0082] For example, the acquisition unit acquires a first set of a first image and a second image, and then acquires a second set of the first image and the second image. When a degree of difference between a feature of a first region image, which is an image of a difference region in a first image included in the first set, and a feature of a second region image, which is an image of a difference region in a second image included in the second set, is equal to or greater than a threshold, the second determination unit adopts, as training data, first candidate data, which is candidate data generated based on the first set, and second candidate data, which is candidate data generated based on the second set. On the other hand, when the degree of difference between the features is less than the threshold, the second determination unit does not adopt the first candidate data or the second candidate data as training data.

[0083] This allows annotation to be performed automatically through an object sorting task, which requires less labor than manual annotation based on visual inspection of images, thereby improving the efficiency of annotation work. Furthermore, when annotation is performed manually based on visual inspection of images, the accuracy of the annotation varies depending on the ability of the worker. In contrast, automatic annotation through a simple task such as object sorting, which does not vary significantly in quality, allows for highly accurate annotation regardless of the ability of the worker performing the object sorting task. Furthermore, it prevents images with erroneous annotations from being used as training data. In particular, when annotation is performed automatically through an operation in which desired objects are selected from a group of objects transported in a dense state, highly accurate annotation can be performed, thereby preventing the generation of erroneous training data. [Explanation of symbols]

[0084] 1. Annotation System 10, 10a, 10b Annotation device 20 Upstream Camera 30 Downstream Camera 40 Conveyor Belt 11 Image acquisition unit 12 Image storage unit 13 Difference area determination section 14 Annotation section 15 Candidate data storage unit 16 Recruitment Judgment Department 17 Teacher data storage unit 18 Feature detection unit

Claims

1. an acquisition unit that acquires a first image taken at a first position upstream of a position where a desired object is selected from a group of objects on a conveyance path along which the group of objects is conveyed, at a first timing before the operation is performed, and a second image taken at a second position downstream of the position where the operation is performed, at a second timing after the operation is performed; a first determination unit that compares the first image with the second image and determines a difference region that is a region where there is a difference between the first image and the second image; an annotation unit that assigns contour information, which is information indicating a contour of the difference region, to the difference region in the first image and generates candidate data including the first image and the contour information; a second determination unit that determines whether or not to adopt the candidate data as training data for machine learning; An annotation device comprising:

2. The second determination unit When the area of ​​the difference region is less than a threshold, the candidate data is adopted as the training data; When the area is equal to or larger than the threshold, the candidate data is not adopted as the training data. The annotation device according to claim 1 .

3. The second determination unit adopting the candidate data as the training data when a degree of difference between a feature of a first region image that is an image of the difference region in the first image and a feature of a second region image that is an image of the difference region in the second image is equal to or greater than a threshold; When the degree of difference is less than the threshold, the candidate data is not adopted as the training data. The annotation device according to claim 1 .

4. the acquisition unit acquires a second set of the first image and the second image after acquiring a first set of the first image and the second image; The second determination unit When a degree of difference between a feature of a first region image, which is an image of the difference region in the first image included in the first set, and a feature of a second region image, which is an image of the difference region in the second image included in the second set, is equal to or greater than a threshold, first candidate data, which is the candidate data generated based on the first set, and second candidate data, which is the candidate data generated based on the second set, are adopted as the teacher data; When the degree of difference is less than the threshold, the first candidate data and the second candidate data are not adopted as the training data. The annotation device according to claim 1 .

Citation Information

Patent Citations

  • Waste screening system and screening method therefor

    JP2017109197A

  • Teacher data creation program, teacher data creation device and teacher data creation method

    JP2019049780A

  • Article sorting apparatus and article sorting method

    JP2021030219A