Teacher data determination device, teacher data determination method and teacher data determination program

JP2024133879A5Pending Publication Date: 2026-02-05PFU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023043880
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing waste sorting devices require large amounts of training data for accurate machine learning, which is challenging to obtain.

Method used

A teacher data determination device that includes a recognition section, calculation section, and determination section to automatically collect training data by recognizing images, calculating variations in recognition results, and determining which results to use for machine learning.

Benefits of technology

Enables the automatic collection of useful training data for machine learning, improving the accuracy of waste sorting devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To automatically collect teacher data useful for machine learning.SOLUTION: In a teacher data determination device 100, an image recognition unit 11 recognizes an image of an object using a learned model, and assigns attribute information to the recognized image, which is information indicating attributes of the image; a variance calculation unit 13 calculates variance among a plurality of recognition results for the image by the image recognition unit 11, each of which includes the image of the object and the attribute information; and a recognition result determination unit 14 determines which recognition result to adopt as teacher data for machine learning based on the variance among the plurality of recognition results.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a teacher data determination device, a teacher data determination method, and a teacher data determination program. [Background technology]

[0002] At waste disposal sites, large amounts of waste are transported on conveyer belts and processed every day. At the waste disposal sites, waste is sorted by hand. While sorting waste is a simple task, it places a heavy burden on the workers who sort the waste (hereinafter referred to as "sorters"). Therefore, devices that automatically sort waste (hereinafter referred to as "waste sorting devices") have been developed. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2016-191973 A Summary of the Invention [Problem to be solved by the invention]

[0004] When the waste sorting device performs the tasks that sorters used to perform, it is possible for the waste sorting device to recognize each waste flowing on the conveyor belt and, based on the recognition results, use a robot hand or a suction pad to extract the desired waste (hereinafter, sometimes referred to as "desired waste") from the waste flowing on the conveyor belt. Therefore, when multiple types of waste are mixed and flowing on the conveyor belt, it is necessary for the waste sorting device to identify the type of waste. For the waste sorting device to recognize various types of waste, image recognition using a trained model generated by machine learning is effective.

[0005] However, generating an accurate trained model using machine learning requires a large amount of useful training data.

[0006] Therefore, this disclosure proposes a technology that can automatically collect teacher data useful for machine learning. [Means for solving the problem]

[0007] The teacher data determination device of the present disclosure includes a recognition unit, a calculation unit, and a determination unit. The recognition unit recognizes an image of an object using a trained model, and assigns attribute information to the recognized image, the attribute information being information indicating attributes of the image. The calculation unit calculates a variance among a plurality of recognition results for the image by the recognition unit, each of the recognition results including the image and the attribute information. The determination unit determines a recognition result to be adopted as teacher data for machine learning based on the variance among the plurality of recognition results. Effect of the Invention

[0008] According to the present disclosure, training data useful for machine learning can be automatically collected. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating a configuration example of an object sorting system according to a first embodiment of the present disclosure. [Diagram 2] FIG. 2 is a diagram illustrating a configuration example of the control device according to the first embodiment of the present disclosure. [Diagram 3] FIG. 3 is a diagram illustrating an example of attribute information according to the first embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of a trace of the centroid coordinates according to the first embodiment of the present disclosure. [Diagram 5] FIG. 5 is a diagram illustrating an example of a trace of the centroid coordinates according to the first embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an example of a trace of the centroid coordinates according to the first embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating an example of a trace of the centroid coordinates according to the first embodiment of the present disclosure. [Figure 8]FIG. 8 is a diagram showing an example of extraction of desired waste according to the first embodiment of the present disclosure. [Figure 9] FIG. 9 is a diagram illustrating an example of an image recognition result according to the first embodiment of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating an operation example 1-1 of the teacher data determination device according to the first embodiment of the present disclosure. [Figure 11] FIG. 11 is a diagram illustrating an operation example 1-2 of the teacher data determination device according to the first embodiment of the present disclosure. [Figure 12] FIG. 12 is a diagram illustrating an operation example 1-3 of the teacher data determination device according to the first embodiment of the present disclosure. [Figure 13] FIG. 13 is a diagram illustrating a second operation example of the teacher data determination device according to the second embodiment of the present disclosure. [Figure 14] FIG. 14 is a diagram illustrating a third operation example of the teacher data determination device according to the third embodiment of the present disclosure. [Figure 15] FIG. 15 is a diagram illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. [Figure 16] FIG. 16 is a diagram illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. [Figure 17] FIG. 17 is a diagram illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. [Figure 18] FIG. 18 is a diagram illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. [Figure 19] FIG. 19 is a diagram illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. [Figure 20] FIG. 20 is a diagram illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. [Figure 21] FIG. 21 is a diagram illustrating a fourth operation example of the teacher data determination device according to the fourth embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the following embodiments, the same components are denoted by the same reference numerals.

[0011] [Example 1] <Configuration of object sorting system> FIG. 1 is a diagram illustrating a configuration example of an object sorting system according to a first embodiment of the present disclosure.

[0012] 1, the object sorting system 1 includes a control device 10, a camera 20, an object sorting device 30, and a belt conveyor 40. The control device 10, the camera 20, and the object sorting device 30 are connected to each other via a network.

[0013] The following description will be given, as an example, of a case where the object sorting system 1 shown in Fig. 1 is installed in a waste disposal site where waste materials flow on a belt conveyor 40. That is, the following description will be given, as an example, of a case where the objects to be sorted by the object sorting system 1 are waste materials. However, the object sorting system 1 may also be installed in an assembly plant or the like where parts flow on a belt conveyor. That is, the objects to be sorted by the object sorting system 1 are not limited to waste materials, and the object sorting system 1 can be used for various objects.

[0014] The belt conveyor 40 transports the waste mass placed on the belt conveyor 40 in a transport direction CD. In other words, the belt conveyor 40 forms a transport path along which the waste mass is transported in the transport direction CD.

[0015] The camera 20 is disposed above the belt conveyor 40 along which the waste mass is transported, has a predetermined angle of view, and continuously captures a predetermined area on the upper surface of the belt conveyor 40 from above the belt conveyor 40 at a constant frame rate. Thus, the image captured by the camera 20 (hereinafter sometimes referred to as the "captured image") is an image of the waste mass. The captured image is transmitted from the camera 20 to the control device 10.

[0016] <Control device configuration> Fig. 2 is a diagram illustrating a configuration example of a control device according to the first embodiment of the present disclosure. In Fig. 2, the control device 10 includes a teacher data determination device 100, a teacher data storage unit 15, a machine learning unit 16, and a learned model storage unit 17. The teacher data determination device 100 includes an image recognition unit 11, a recognition result storage unit 12, a variation calculation unit 13, and a recognition result determination unit 14.

[0017] The image recognition unit 11 recognizes each waste image (hereinafter may be referred to as "waste image") present in the captured image using the learned model stored in the learned model storage unit 17, and assigns information indicating the attribute of each waste image (hereinafter may be referred to as "attribute information") to each recognized waste image. The image recognition unit 11 assigns attribute information to a waste image for which the reliability score of image recognition using the learned model is equal to or greater than a threshold. The reliability score is output for each waste image from the learned model stored in the learned model storage unit 17. The image recognition unit 11 recognizes the waste image, for example, by performing instance segmentation on the captured image. The image recognition unit 11 associates the waste image with the attribute information already assigned to the waste image, generates a recognition result (hereinafter may be referred to as "image recognition result") in which attribute information is assigned to each waste image, and outputs the generated image recognition result to the recognition result storage unit 12 and the object sorting device 30.

[0018] The recognition result storage unit 12 stores in chronological order the image recognition results sequentially output from the image recognition unit 11. In other words, the recognition result storage unit 12 stores a plurality of image recognition results for waste images by the image recognition unit 11, each of which includes a waste image and attribute information.

[0019] The object sorting device 30 sorts by extracting desired waste from a mass of waste transported on the belt conveyor 40 based on the image recognition results (i.e., waste images and attribute information) output from the image recognition unit 11. The object sorting device 30 extracts the desired waste using, for example, a robot hand, a suction pad, or the like.

[0020] The variation calculation unit 13 calculates the variation among the multiple image recognition results stored in the recognition result memory unit 12 (hereinafter sometimes referred to as "recognition result variation"), and outputs the calculated recognition result variation to the recognition result judgment unit 14.

[0021] Based on the recognition result variability, the recognition result determination unit 14 determines an image recognition result to be adopted as training data for machine learning (hereinafter, may be referred to as an “adopted recognition result”) from the multiple image recognition results stored in the recognition result storage unit 12. Then, the recognition result determination unit 14 extracts an adopted recognition result from the multiple image recognition results stored in the recognition result storage unit 12, and outputs the extracted adopted recognition result to the training data storage unit 15.

[0022] Many waste images with attribute information added thereto are pre-stored as teacher data in the teacher data storage unit 15. The teacher data storage unit 15 also newly stores the adopted recognition result output from the recognition result determination unit 14 as additional teacher data.

[0023] The machine learning unit 16 periodically performs machine learning using the training data stored in the training data storage unit 15, and updates the trained model stored in the trained model storage unit 17 with the trained model after the machine learning. Therefore, the recognition of waste images in the image recognition unit 11 is performed using the trained model after it has been updated by additional machine learning performed by the machine learning unit 16 using the adopted recognition result.

[0024] <Attribute information> 3 is a diagram showing an example of attribute information according to the first embodiment of the present disclosure. In the following, it is assumed that the waste is an empty bottle, and that the object sorting device 30 sorts each empty bottle flowing on the belt conveyor 40 into three types of empty bottles: a brown empty bottle (hereinafter sometimes referred to as a "brown bottle"), a colorless and transparent empty bottle (hereinafter sometimes referred to as a "transparent bottle"), and an empty bottle having a color other than brown (hereinafter sometimes referred to as an "other color") (hereinafter sometimes referred to as an "other color bottle"). In the following, it is assumed that the desired waste is a brown bottle, and training data for brown bottles is stored in the training data storage unit 15.

[0025] As shown in Figure 3, when the captured image includes an image of an empty bottle (hereinafter sometimes referred to as an "empty bottle image") BI as a waste image, the image recognition unit 11 uses the learned model stored in the learned model memory unit 17 to recognize the empty bottle image BI present in the captured image, and assigns attribute information to the recognized empty bottle image BI, including the shooting time, individual ID, label information LA, contour information CO, center of gravity coordinates DG, contour area area, confidence score, center of gravity prediction error, and color statistics.

[0026] The photographing time is the time when the photographed image including the empty bottle image BI is captured by the camera 20.

[0027] The individual ID is an ID for uniquely identifying the same empty bottle among a plurality of empty bottle images. The image recognition unit 11 assigns an individual ID to each empty bottle image BI.

[0028] The label information LA is information indicating the type of empty bottle. If the empty bottle image BI is an image of a brown bottle (hereinafter may be referred to as a "brown bottle image"), the image recognition unit 11 assigns the label information LA of "brown bottle" to the empty bottle image BI, if the empty bottle image BI is an image of a transparent bottle (hereinafter may be referred to as a "transparent bottle image"), the image recognition unit 11 assigns the label information LA of "transparent bottle" to the empty bottle image BI, and if the empty bottle image BI is an image of a bottle of another color (hereinafter may be referred to as a "other color bottle image"), the image recognition unit 11 assigns the label information LA of "other color bottle" to the empty bottle image BI.

[0029] The image recognition unit 11 also assigns contour information CO, which indicates the contour of the empty bottle image BI, to the empty bottle image BI. The contour information CO is formed by a number of coordinate points (x0, y0), (x1, y1), ..., (xn, yn) with the long side of the rectangular captured image as the X-axis and the short side as the Y-axis.

[0030] The image recognition unit 11 also calculates the area center of gravity of the empty bottle image BI based on the contour information CO, and assigns the center of gravity coordinates DG indicating the position of the calculated area center of gravity to the empty bottle image BI. The center of gravity coordinates DG are formed by a single coordinate point (X, Y).

[0031] Here, the multiple coordinate points (x0, y0), (x1, y1), ..., (xn, yn) that form the contour information CO, and the single coordinate point (X, Y) that forms the center of gravity coordinates, are coordinate points in the coordinate system of the captured image, i.e., the coordinate system of the camera 20 (hereinafter sometimes referred to as the "camera coordinate system").

[0032] The image recognition unit 11 also calculates the area of ​​the region surrounded by the multiple coordinate points (x0, y0), (x1, y1), ..., (xn, yn) that form the contour information CO (hereinafter, sometimes referred to as the "contour region") (i.e., the contour region area), and assigns the calculated contour region area to the empty bottle image BI. The image recognition unit 11 calculates the total number of pixels in the contour region as the contour region area.

[0033] In addition, when the image recognition unit 11 recognizes the empty bottle image BI using the trained model, the image recognition unit 11 assigns a confidence score output from the trained model to the empty bottle image BI. The confidence score indicates the accuracy of the image recognition result of the empty bottle image BI.

[0034] Moreover, the image recognition unit 11 assigns a center of gravity prediction error [mm], which will be described later, to the empty bottle image BI.

[0035] The image recognition unit 11 also assigns color statistics of the empty bottle image BI to the empty bottle image BI. Based on the color information of each pixel in the contour region, the image recognition unit 11 assigns the following color statistics to the empty bottle image BI: an R component (red component) statistic (hereinafter sometimes referred to as "R component amount") for all pixels in the contour region, a G component (green component) statistic (hereinafter sometimes referred to as "G component amount") for all pixels in the contour region, and a B component (blue component) statistic (hereinafter sometimes referred to as "B component amount") for all pixels in the contour region. The R component amount, G component amount, and B component amount are calculated, for example, by arithmetic averaging, and are expressed using gradation values ​​of 0 to 255.

[0036] <Tracing the center of gravity coordinates> 4, 5, 6, and 7 are diagrams showing examples of traces of barycentric coordinates according to the first embodiment of the present disclosure.

[0037] First, as shown in FIG. 4, the image recognition unit 11 recognizes a brown bottle image BO in a captured image acquired at time t11, assigns an individual ID to the brown bottle image BO at time t11, and calculates the center of gravity coordinate DG1 of the brown bottle image BO at time t11.

[0038] 5, the image recognition unit 11 predicts the center of gravity EG2 of the brown bottle image BO in the image captured at time t12 following the image captured at time t11 based on the frame rate FR [frame / sec] of the camera 20 and the conveying speed CS [mm / sec] of the belt conveyor 40. The center of gravity EG2 at time t12 is predicted to move by CS / FR in the conveying direction CD of the belt conveyor 40 (i.e., the X direction in FIG. 1) per frame relative to the center of gravity DG1 at time t11. The image recognition unit 11 also sets a circular allowable range TR of a predetermined size centered on the predicted center of gravity EG2.

[0039] 6, the image recognition unit 11 recognizes the brown bottle image BO in the captured image acquired at time t12, and calculates the barycentric coordinate DG2 of the brown bottle image BO at time t12. The image recognition unit 11 also determines whether the barycentric coordinate DG2 is within the allowable range TR. The image recognition unit 11 also calculates the distance between the barycentric coordinate EG2 and the barycentric coordinate DG2 as a barycentric prediction error PE, which is a prediction error of the barycentric position.

[0040] When the center of gravity coordinate DG2 is within the allowable range TR, the image recognition unit 11 determines that the subjects of both the brown bottle images BO at time t11 and time t12 are the same brown bottle, assigns the same individual ID to the brown bottle image BO at time t12 as the individual ID assigned to the brown bottle image BO at time t11, and sets a straight trace line TL connecting the center of gravity coordinate DG1 to the center of gravity coordinate DG2, as shown in Fig. 7. On the other hand, when the center of gravity coordinate DG2 is not within the allowable range TR, the image recognition unit 11 stops tracing the center of gravity coordinate.

[0041] The image recognition unit 11 sequentially executes tracing of the center of gravity coordinates as described above for the captured images sequentially acquired at the frame rate FR.

[0042] <Extraction of desired waste> Fig. 8 is a diagram showing an example of extraction of desired waste according to the first embodiment of the present disclosure. In Fig. 8, the frame rate FR of the camera 20 is set to a frame rate that allows the same empty bottle conveyed in the conveying direction CD to be photographed a maximum of three times based on the conveying speed CS of the belt conveyor 40 and the angle of view of the camera 20. In addition, the object sorting device 30 uses the suction pad 31 to extract brown bottles, which are desired waste, from the group of empty bottles.

[0043] In FIG. 8, the image recognition unit 11 recognizes the brown bottle image BO in the captured image acquired at time t11, and calculates the barycentric coordinates DG1 (X1, Y1) of the brown bottle image BO at time t11.

[0044] Next, the image recognition unit 11 recognizes the brown bottle image BO in the captured image acquired at time t12, and calculates the barycentric coordinates DG2 (X2, Y2) of the brown bottle image BO at time t12.

[0045] Next, the image recognition unit 11 determines that the subject of both brown bin images BO at times t11 and t12 is the same brown bin, and sets a straight trace line TL connecting the center of gravity coordinates DG1 (X1, Y1) to the center of gravity coordinates DG2 (X2, Y2).

[0046] Next, the image recognition unit 11 recognizes the brown bottle image BO in the captured image acquired at time t13, and calculates the barycentric coordinates DG3 (X3, Y3) of the brown bottle image BO at time t13.

[0047] Next, the image recognition unit 11 determines that the subject of both brown bin images BO at times t12 and t13 is the same brown bin, and sets a straight trace line TL connecting the center of gravity coordinates DG1 (X1, Y1) to the center of gravity coordinates DG3 (X3, Y3).

[0048] Next, the image recognition unit 11 extends the trace line TL connecting the center of gravity coordinate DG1 (X1, Y1) to the center of gravity coordinate DG3 (X3, Y3) by linear approximation to the installation position of the suction pad 31 in the X direction, and calculates the extraction target coordinate TC (X, Y) of the brown bottle by converting the camera coordinate system to the coordinate system of the object sorting device 30. In addition, the image recognition unit 11 calculates the time (hereinafter sometimes referred to as the "extraction target time") tH at which the brown bottle, which is the subject of the brown bottle image BO recognized at times t11, t12, and t13, reaches the extraction target coordinate TC (X, Y), based on the conveying speed CS. The image recognition unit 11 transmits a control signal including the calculated extraction target coordinate TC (X, Y) and extraction target time tH to the object sorting device 30.

[0049] In the object sorting device 30, the suction pad 31 moves to the extraction target coordinates TC(X,Y) according to the extraction target coordinates TC(X,Y) and extraction target time tH received from the image recognition unit 11, and extracts an empty bottle located at the extraction target coordinates TC(X,Y) at the extraction target time tH from the group of empty bottles. As a result, in the object sorting device 30, the brown bottle that is the subject of the brown bottle image BO recognized by the image recognition unit 11 at times t11, t12, and t13 is extracted from the group of empty bottles.

[0050] <Image recognition results> Fig. 9 is a diagram illustrating an example of an image recognition result according to the first embodiment of the present disclosure. Fig. 9 illustrates, as an example, an image recognition result for an empty bottle image included in a captured image acquired at time t12.

[0051] As shown in Figure 9, the image recognition result includes the shooting time, the empty bottle image, the individual ID, the contour area area, the center of gravity prediction error, the R component amount, the G component amount, the B component amount, the confidence score, and the type of empty bottle indicated by the label information LA.

[0052] <Operation of the teacher data judgment device> 10, 11, and 12 are diagrams illustrating operation examples 1-1, 1-2, and 1-3 of the teacher data determination device according to the first embodiment of the present disclosure.

[0053] In the following, as an example, the threshold for the variation in the area of ​​the contour region (hereinafter sometimes referred to as "area variation") is set to "5", the threshold for the variation in the center of gravity prediction error (hereinafter sometimes referred to as "error variation") is set to "3", the threshold for the variation in the amount of R component (hereinafter sometimes referred to as "R component variation") is set to "5", the threshold for the variation in the amount of G component (hereinafter sometimes referred to as "G component variation") is set to "5", the threshold for the variation in the amount of B component (hereinafter sometimes referred to as "B component variation") is set to "5", the threshold for the variation in the confidence score (hereinafter sometimes referred to as "confidence score variation") is set to "3", and the threshold for the variation in the type of empty bottle (hereinafter sometimes referred to as "type variation") is set to "0.3".

[0054] The variation calculation unit 13 calculates the variation of each of the area variation, the error variation, the R component variation, the G component variation, the B component variation, and the confidence score variation by using the standard deviation. The variation calculation unit 13 also calculates the type variation by using the Simpson's diversity index. Each of the variations of the area variation, the error variation, the R component variation, the G component variation, the B component variation, the confidence score variation, and the type variation corresponds to the recognition result variation.

[0055] The recognition result determination unit 14 determines an image recognition result in which at least one of the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation is equal to or greater than a threshold value as an adopted recognition result.

[0056] <Operation example 1-1 (Fig. 10)> As shown in Fig. 10, when the area variation and error variation among three image recognition results obtained in time series by the image recognition unit 11 at times t11, t12, and t13 and having the same individual ID are equal to or greater than a threshold value, the recognition result determination unit 14 determines that the cause of the variation is insufficient recognition accuracy of the image recognition unit 11, and determines these three image recognition results as adopted recognition results. On the other hand, the image recognition unit 11 determines that the subjects of multiple empty bottle images assigned the same individual ID are the same empty bottle. As an example of a case where the area variation and error variation among three image recognition results obtained in time series and having the same individual ID are equal to or greater than a threshold value, a case where a part of the empty bottle image at time t12 is missing is assumed.

[0057] <Operation example 1-2 (Fig. 11)> As shown in FIG. 11, when the B component variation and type variation among three image recognition results having the same individual ID obtained by the image recognition unit 11 in chronological order at times t11, t12, and t13 are equal to or greater than a threshold value, the recognition result determination unit 14 determines that the cause of the variation is insufficient recognition accuracy of the image recognition unit 11, and determines these three image recognition results as adopted recognition results.

[0058] <Operation example 1-3 (Fig. 12)> 12, when the error variation is equal to or greater than a threshold value while the area variation is less than a threshold value among three image recognition results acquired in time series by the image recognition unit 11 at times t11, t12, and t13 and having the same individual ID, the recognition result determination unit 14 determines that the variation is not caused by insufficient recognition accuracy of the image recognition unit 11, and does not adopt these three image recognition results as teacher data for machine learning and extract them from the recognition result storage unit 12. As an example of a case where the error variation is equal to or greater than a threshold value while the area variation is less than a threshold value among three image recognition results acquired in time series and having the same individual ID, a case where the same empty bottle rolls on the belt conveyor 40 while being transported on the belt conveyor 40 is assumed.

[0059] The first embodiment has been described above.

[0060] [Example 2] <Operation of the teacher data judgment device> The image recognition unit 11 obtains a plurality of image recognition results for a single empty bottle image using a plurality of different trained models for the single empty bottle image. The plurality of different trained models are stored in advance in the trained model storage unit 17.

[0061] The variation calculation unit 13 calculates each of the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation between multiple image recognition results for a single empty bottle image.

[0062] <Example of operation 2 (Fig. 13)> Fig. 13 is a diagram showing an operation example 2 of the teacher data determination device according to the second embodiment of the present disclosure. In the following, as an example, a case will be described in which the image recognition unit 11 obtains three image recognition results for a single empty bottle image using three different learned models, a first learned model, a second learned model, and a third learned model, for a single empty bottle image. In Fig. 13, the image recognition result in the first row is obtained using the first learned model, the image recognition result in the second row is obtained using the second learned model, and the image recognition result in the third row is obtained using the third learned model.

[0063] As shown in FIG. 13, when the confidence score variation among three image recognition results having the same individual ID obtained by the image recognition unit 11 at the same time t12 using each of the learned models, namely the first learned model, the second learned model, and the third learned model, is equal to or greater than a threshold value, the recognition result determination unit 14 determines that the cause of the variation is insufficient recognition accuracy of the image recognition unit 11, and determines these three image recognition results to be adopted recognition results.

[0064] The second embodiment has been described above.

[0065] [Example 3] <Operation of the teacher data judgment device> The image recognition unit 11 processes a single empty bottle image to generate a plurality of processed images. In addition, the image recognition unit 11 uses the same trained model for each of the plurality of processed images to obtain a plurality of image recognition results for each of the plurality of processed images as a plurality of image recognition results for the single empty bottle image.

[0066] The variation calculation unit 13 calculates each of the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation between multiple image recognition results for each of the multiple processed images.

[0067] <Operation example 3 (Fig. 14)> Fig. 14 is a diagram showing an operation example 3 of the teacher data determination device according to the third embodiment of the present disclosure. In the following, as an example, a case will be described in which the image recognition unit 11 processes a single empty bottle image to generate three processed images, a first processed image, a second processed image, and a third processed image. In Fig. 14, the image recognition result in the first row is the image recognition result for the first processed image, the image recognition result in the second row is the image recognition result for the second processed image, and the image recognition result in the third row is the image recognition result for the third processed image.

[0068] As shown in FIG. 14, when the type variation among three image recognition results obtained at the same time t12 by the image recognition unit 11 using the same trained model and having the same individual ID is equal to or greater than a threshold, the recognition result determination unit 14 determines that the cause of the variation is insufficient recognition accuracy of the image recognition unit 11, and determines these three image recognition results to be adopted recognition results.

[0069] The third embodiment has been described above.

[0070] [Example 4] <Calculation of recognition result variance> In the first to third embodiments, the outline region area, the center of gravity prediction error, the R component amount, the G component amount, the B component amount, and the reliability score are continuous values, and the variation calculation unit 13 calculates the area variation, the error variation, the R component variation, the G component variation, the B component variation, and the reliability score variation using standard deviation. In contrast, in the fourth embodiment, the variation calculation unit 13 converts the outline region area, the center of gravity prediction error, the R component amount, the G component amount, the B component amount, and the reliability score into discrete values ​​using a histogram, and calculates the area variation, the error variation, the R component variation, the G component variation, the B component variation, and the reliability score variation using Simpson's diversity index.

[0071] 15 to 20 are diagrams illustrating an example of calculation of the recognition result variation according to the fourth embodiment of the present disclosure. For example, when the contour area is a histogram as shown in FIG. 15, the variation calculation unit 13 calculates the area variation shown as the diversity index to be 0.00. For example, when the centroid prediction error is a histogram as shown in FIG. 16, the variation calculation unit 13 calculates the error variation shown as the diversity index to be 0.00. For example, when the R component amount is a histogram as shown in FIG. 17, the variation calculation unit 13 calculates the R component variation shown as the diversity index to be 0.00. For example, when the G component amount is a histogram as shown in FIG. 18, the variation calculation unit 13 calculates the G component variation shown as the diversity index to be 0.00. Furthermore, for example, when the B component amount is a histogram as shown in Fig. 19, the variation calculation unit 13 calculates the B component variation shown as a diversity index to be 0.44. Furthermore, for example, when the confidence score is a histogram as shown in Fig. 20, the variation calculation unit 13 calculates the confidence score variation shown as a diversity index to be 0.00.

[0072] <Operation of the teacher data judgment device> <Example 4 (Fig. 21)> FIG. 21 is a diagram illustrating a fourth operation example of the teacher data determination device according to the fourth embodiment of the present disclosure.

[0073] 15 to 20, the variation calculation unit 13 calculates the area variation, error variation, R component variation, G component variation, and confidence score variation as diversity indexes of 0.00, while calculating the B component variation as diversity index of 0.44 (FIG. 21). For example, as shown in FIG. 21, the variation calculation unit 13 calculates the type variation as diversity index of 0.44 in the same manner as in the first embodiment.

[0074] As shown in FIG. 21, weights are set for each of the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation. For example, the variation calculation unit 13 calculates the sum of the multiplication results of each recognition result variation and each weight (hereinafter, may be referred to as "weighted sum"). When the weighted sum is equal to or greater than a threshold, the recognition result determination unit 14 determines that the cause of the variation is insufficient recognition accuracy of the image recognition unit 11, and determines that three image recognition results obtained in time series by the image recognition unit 11 at each of the times t11, t12, and t13 and having the same individual ID are adopted recognition results. In the example shown in FIG. 21, the weighted sum is calculated as (0.44×1)+(0.44)×2=1.32, and the threshold for the weighted sum is set to 1.0.

[0075] Also, for example, when the maximum value among the recognition result variations shown in Fig. 21 is equal to or greater than a threshold value, the recognition result determination unit 14 may determine that the cause of the variation is insufficient recognition accuracy of the image recognition unit 11, and determine that three image recognition results acquired in time series by the image recognition unit 11 at each of times t11, t12, and t13 and having the same individual ID are to be adopted recognition results. In the example shown in Fig. 21, the maximum value of the recognition result variation (hereinafter sometimes referred to as "maximum variation") is 0.44, and the threshold value for the maximum variation is set to 0.3.

[0076] For example, the variation calculation unit 13 calculates the average value of each recognition result variation (hereinafter sometimes referred to as "average variation"). Then, when the average variation is equal to or greater than the threshold value, the recognition result determination unit 14 determines that the cause of the variation lies in the insufficient recognition accuracy of the image recognition unit 11. It is also possible to determine that three image recognition results obtained in time series by the image recognition unit 11 at each of the times t11, t12, and t13 and having the same individual ID are adopted recognition results. In the example shown in FIG. 21, the average variation is calculated as (0.44 + 0.44) ÷ 7 = 0.13 (rounded to the third decimal place), and the threshold value for the average variation is set to 0.2.

[0077] The above has described Example 4.

[0078] [Example 5] In the above Examples 1 to 4, the recognition result determination unit 14 determined that all of the image recognition results with a recognition result variation equal to or greater than the threshold value were adopted recognition results. However, among all N image recognition results with a recognition result variation equal to or greater than the threshold value, for example, only the top M (M < N) image recognition results with a large variation or only randomly selected M image recognition results may be determined as the adopted recognition results.

[0079] [Example 6] The recognition result storage unit 12, the teacher data storage unit 15, and the learned model storage unit 17 are realized as hardware by, for example, a memory or a storage. The image recognition unit 11, the variation calculation unit 13, the recognition result determination unit 14, and the machine learning unit 16 are realized as hardware by, for example, a processor such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0080] In addition, all or part of each process in the above description in the control device 10 may be realized by having a processor of the control device 10 execute a program corresponding to each process. For example, a program corresponding to each process in the above description may be stored in a memory or storage of the control device 10, and the program may be read from the memory or storage by the processor and executed. In addition, the program may be stored in a program server connected to the control device 10 via an arbitrary network, downloaded from the program server to the control device 10 and executed, or stored in a recording medium readable by the control device 10, read from the recording medium and executed. Examples of recording media readable by the control device 10 include portable storage media such as memory cards, USB memories, SD cards, flexible disks, magneto-optical disks, CD-ROMs, and DVDs.

[0081] The sixth embodiment has been described above.

[0082] In the above-mentioned first to fifth embodiments, the case where all of the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation are used as the recognition result variation is given as an example. However, the technology of the present disclosure is applicable to the case where at least one of the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation is used. In other words, the technology of the present disclosure is applicable to the case where at least one of the contour area, the center of gravity prediction error, the color statistics, the confidence score, and the label information LA is included in the attribute information. In addition, the technology of the present disclosure is also applicable to the case where the variation other than the area variation, error variation, R component variation, G component variation, B component variation, confidence score variation, and type variation is used as the recognition result variation.

[0083] As described above, the teacher data determination device (teacher data determination device 100 in the embodiment) of the present disclosure includes a recognition unit (image recognition unit 11 in the embodiment), a calculation unit (variation calculation unit 13 in the embodiment), and a determination unit (recognition result determination unit 14 in the embodiment). The recognition unit recognizes an image of an object using a trained model, and assigns attribute information, which is information indicating attributes of the image, to the recognized image. The calculation unit calculates the variation among multiple recognition results for an image by the recognition unit, each of which includes an image of an object and attribute information. The determination unit determines which recognition result to adopt as teacher data for machine learning based on the variation among the multiple recognition results.

[0084] This makes it possible to use recognition results that are expected to have low accuracy as training data for machine learning, thereby automatically collecting training data that is useful for machine learning.

[0085] For example, the recognition unit obtains a plurality of recognition results for a plurality of images of the same object transported on the transport path over time, the images being captured at different times, and the calculation unit calculates a variation among the plurality of recognition results for the plurality of images of the same object captured at different times.

[0086] Also, for example, the recognition unit obtains multiple recognition results for an image of a single object using multiple different trained models for the image of a single object, and the calculation unit calculates the variation among the multiple recognition results.

[0087] For example, the recognition unit generates a plurality of processed images by processing an image of a single object, and obtains a plurality of recognition results for each of the plurality of processed images using the same trained model for each of the plurality of processed images, and the calculation unit calculates the variation among the plurality of recognition results. [Explanation of symbols]

[0088] 1. Object sorting system 10 Control device 20 Camera 30 Object sorting device 100 Teacher data judgment device 11 Image Recognition Unit 12 Recognition result storage unit 13 Variation calculation section 14 Recognition result judgment section 15 Teacher data storage unit 16 Machine Learning Department 17 Trained model memory

Claims

1. a recognition unit that recognizes an image of an object using a trained model and assigns attribute information to the recognized image, the attribute information being information indicating the attributes of the image; a selection unit that selects a recognition result to be adopted as training data for machine learning based on a variation among a plurality of recognition results for the image by the recognition unit, each of the plurality of recognition results including the image and the attribute information; A teacher data determination device comprising:

2. The attribute information is at least one of the area of ​​a region surrounded by the contour of the image of the object, the type of the object, a confidence score output from the trained model, a prediction error of the center of gravity position of the region, and a statistic of the color of the image of the object. The teacher data determination device according to claim 1 .

3. the recognition unit acquires the plurality of recognition results for each of a plurality of images of the same object conveyed on a conveyance path over time, the images being captured at different times; the selection unit calculates the variation among the plurality of recognition results for each of the plurality of images of the same object captured at different times. The teacher data determination device according to claim 1 .

4. the recognition unit acquires the plurality of recognition results for the image of the single object using each of a plurality of trained models that are different from one another for the image of the single object; the selection unit calculates the variation among the plurality of recognition results. The teacher data determination device according to claim 1 .

5. the recognition unit generates a plurality of processed images by processing an image of the single object, and acquires a plurality of recognition results for each of the plurality of processed images using the same trained model for each of the plurality of processed images; the selection unit calculates the variation among the plurality of recognition results. The teacher data determination device according to claim 1 .

6. Recognizes images of objects using a trained model, assigning attribute information to the recognized image, the attribute information being information indicating the attribute of the image; selecting a plurality of recognition results for the image, each of which includes the image and the attribute information, based on variations among the plurality of recognition results to be adopted as training data for machine learning; Method for determining teacher data.

7. Recognizes images of objects using a trained model, assigning attribute information to the recognized image, the attribute information being information indicating the attribute of the image; selecting a plurality of recognition results for the image, each of which includes the image and the attribute information, based on variations among the plurality of recognition results to be adopted as training data for machine learning; A training data judgment program for causing a processor to execute processing.