ANNOTATION DEVICE, ANNOTATION METHOD, AND ANNOTATION PROGRAM
The annotation device and method address the challenges of high cost and low accuracy in existing annotation techniques by classifying and generating learning data based on annotator reliability, resulting in more efficient and accurate annotation processes for supervised learning in machine learning.
Patent Information
- Application Number
- JP2022569380
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-12-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-12-15
AI Technical Summary
Existing annotation techniques for supervised learning in machine learning struggle to achieve high accuracy and low cost, particularly when tasks are difficult to understand in a short time, leading to large variations in labels assigned by annotators and reduced reliability of annotation results.
The proposed annotation device and method involve acquiring first learning data, distributing it to multiple annotators, classifying the data based on the reliability of the correct answer labels, and generating second learning data that includes reference points and data that is easy or difficult to annotate accurately, allowing for more accurate and efficient annotation.
This approach enables annotations to be performed at a lower cost and with higher accuracy in supervised learning, by improving the reliability of annotation results and reducing the time and effort required for annotation tasks.
Smart Images

Figure 0007673757000001 
Figure 0007673757000002 
Figure 0007673757000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an annotation device, an annotation method, and an annotation program. [Background technology]
[0002] Traditionally, supervised learning in machine learning requires training data and corresponding correct labels. In many studies, multiple people view the data and assign metadata (annotation).
[0003] For example, in the case of annotation for audio or video, a worker (referred to as "annotator" as appropriate) listens to the presented audio or video for several to several tens of seconds and adds metadata to meet the specifications. Specifically, in the case of annotation for research and development of emotion recognition from audio, the most appropriate emotion for the audio heard is selected, and in the case of object detection or object recognition for images, an area within the image of the object is selected and an explanation of the object is added.
[0004] Conventional annotation methods can be divided into those that have a comparison target for the task and those that do not. When there is no comparison target, the annotator watches still images or a few seconds of audio or video and assigns metadata. This method has a low time cost because the number of times data is viewed = the total number of samples N. In addition, if the task can be understood by anyone even in a short time (e.g., transcription, tagging objects, or a state where the person is clearly angry no matter who sees or hears it), annotation can be performed accurately.
[0005] On the other hand, when there is a comparison target, annotators listen to audio or video for a long period of time (tens of seconds to several minutes) and assign metadata about continuous and relative changes in events (see, for example, Non-Patent Document 3), or listen to multiple audio or video and assign relative rankings or scores (see, for example, Non-Patent Document 4). This method reduces variation between annotators and allows for more accurate metadata to be assigned because there is a comparison target. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Mohammad Soleymani, and Martha Larson, “Crowdsourcing for Affective Annotation of Video: Development of a Viewer-reported Boredom Corpus”, 2010. [Non-Patent Document 2] Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C. Alexander, and Nathan Silberman, “Learning From Noisy Labels By Regularized Estimation Of Annotator Confusion”, 2019. [Non-Patent Document 3] David Melhart, Antonios Liapis, and Georgios N. Yannakakis “PAGAN: Video Affect Annotation Made Easy”, 2019. [Non-Patent Document 4] Lifang Yang, and Rui Zhu, “Subjective Evaluation of Cooling Fan Sound based on Grade Scoring and Paired Comparison”, 2016. Summary of the Invention [Problem to be solved by the invention]
[0007] However, the above-mentioned conventional techniques cannot perform annotation at a lower cost and with higher accuracy in supervised learning in machine learning because annotation methods without a comparison target cannot perform accurate annotation for tasks that are difficult to understand in a short viewing period, and there are problems such as large variation in the labels assigned by annotators, resulting in low reliability of annotation results.
[0008] To address this issue, there are measures such as taking into consideration the quality of the annotator, answer trends, and reducing the effects of noise by using multiple people to annotate, but these do not fundamentally solve the problem of increasing the reliability of the annotation itself (see, for example, Non-Patent Documents 1 and 2). For example, when annotating the degree of concentration, if the person appears to be very focused or not focused at all, that is, if it is clearly visible to anyone who sees or hears it, the votes of multiple annotators are likely to match, but if it is difficult to tell whether they are focused or not, accurate annotation is difficult, and as a result, subtle differences cannot be expressed.
[0009] On the other hand, annotation methods without a comparison target require viewing a long time or a large amount of data, which requires huge costs for annotation. For example, when viewing several combinations of data at the same time, selecting n items from the total N sample data can result in a maximum of N C n It may be possible to reduce the number of combinations while maintaining the quality of annotation by referring to psychological experiments, but careful consideration is required to decide which combinations to exclude. [Means for solving the problem]
[0010] In order to solve the above-mentioned problems and achieve the object, the annotation device of the present invention is characterized in comprising an acquisition unit that acquires first learning data to be used for machine learning, a first distribution unit that distributes the first learning data acquired by the acquisition unit to a plurality of annotators, a classification unit that classifies the first learning data based on the reliability of first correct labels respectively assigned to the first learning data by each annotator, and a second distribution unit that distributes the classification result of the first learning data classified by the classification unit.
[0011] Furthermore, an annotation method according to the present invention is an annotation method executed by an annotation device, and is characterized in that it includes an acquisition step of acquiring first learning data to be used for machine learning, a first distribution step of distributing the first learning data acquired by the acquisition step to a plurality of annotators, a classification step of classifying the first learning data based on the reliability of first correct labels respectively assigned to the first learning data by each annotator, and a second distribution step of distributing the classification result of the first learning data classified by the classification step.
[0012] In addition, the annotation program of the present invention is characterized in that it causes a computer to execute an acquisition step of acquiring first learning data to be used in machine learning, a first distribution step of distributing the first learning data acquired by the acquisition step to a plurality of annotators, a classification step of classifying the first learning data based on the reliability of first correct labels respectively assigned to the first learning data by each annotator, and a second distribution step of distributing the classification result of the first learning data classified by the classification step. Effect of the Invention
[0013] The present invention makes it possible to perform annotation at lower cost and with higher accuracy in supervised learning in machine learning. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an annotation system according to the first embodiment. [Diagram 2] FIG. 2 is a block diagram showing an example of the configuration of the annotation device according to the first embodiment. [Diagram 3] FIG. 3 is a diagram illustrating an example of learning data according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the first learning data and the first correct label according to the first embodiment. [Diagram 5] FIG. 5 is a diagram illustrating an example of second learning data and a second correct label according to the first embodiment. [Figure 6] FIG. 6 is a flowchart showing an example of the flow of the annotation process according to the first embodiment. [Figure 7] FIG. 7 is a flowchart showing an example of the flow of the first learning data classification process according to the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating a computer that executes a program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An annotation device, an annotation method, and an annotation program according to the present invention will be described in detail below with reference to the accompanying drawings. Note that the present invention is not limited to the following embodiments.
[0016] [First embodiment] The configuration of the annotation system according to this embodiment, the configuration of the annotation device, a specific example of the annotation process, the flow of the annotation process, and the flow of the data classification process will be described below in that order, and finally the effects of this embodiment will be described.
[0017] [Annotation system configuration] The configuration of an annotation system (or "this system" as appropriate) 100 according to this embodiment will be described in detail with reference to Fig. 1. Fig. 1 is a diagram showing an example of an annotation system according to a first embodiment. The annotation system 100 has an annotation device 10 such as a server, annotators 20 (20A, 20B, 20C) such as various terminals, and various databases 30 (30A, 30B, 30C).
[0018] Here, the annotation device 10, the annotator 20, and the database 30 are connected to each other via a predetermined communication network (not shown) in a wired or wireless manner so as to be able to communicate with each other. Note that the annotation system 100 shown in Fig. 1 may include a plurality of annotation devices 10.
[0019] First, the annotation device 10 acquires learning data necessary for research and development as first learning data from various databases 30 (step S1). Here, the acquired learning data is data such as audio, images, and videos, and is acquired in a medium and on a scale appropriate for the purpose of the research and development.
[0020] Next, the annotation device 10 distributes the acquired first learning data to the annotator 20 (step S2). Here, the annotator 20 is a terminal and a user of the terminal that respectively assign correct answer labels to the distributed learning data, but is not particularly limited thereto. The annotator 20 may be a machine learning model that can assign a specific correct answer label that is created separately.
[0021] Next, the annotator 20 assigns a correct label (first correct label) to the distributed first learning data (step S3). In addition, the annotation device 10 acquires the first learning data to which the correct label has been assigned (step S4).
[0022] Thereafter, the annotation device 10 classifies the first learning data based on the first correct answer label (step S5). At this time, the annotation device 10 selects the learning data to which reliable correct answer data has been assigned as a reference point (appropriately, "reference data") S based on the answer acquired from the annotator 20. The annotation device 10 further classifies the learning data other than the reference point S into data to which it is easy to assign an accurate correct answer label (appropriately, "data D") and data to which it is difficult to assign an accurate correct answer label (appropriately, "data E").
[0023] Furthermore, the annotation device 10 generates second learning data from the classified first learning data (step S6). At this time, the annotation device 10 generates a data group including the reference point S, data E, and data D that have the same transmission source. Note that the classification of the first learning data and the generation of the second learning data will be described later.
[0024] Then, the annotation device 10 distributes the generated second learning data to the annotator 20 (step S7). At this time, when distributing the data group of the second learning data, the annotation device 10 distributes each data to the annotator 20 so that the user views the reference point S, followed by data E and data D. The annotator 20 also assigns a correct answer label (second correct answer label) to the distributed second learning data (step S8). Finally, the annotation device 10 acquires the second learning data to which the correct answer label has been assigned (step S9).
[0025] In the annotation system 100 according to the present embodiment, the annotation device 10 includes reliable data on an event to which a correct label is to be assigned in a data group and indicates the data group. This allows the annotator 20 to use the data as a comparison target, thereby realizing more accurate annotation.
[0026] [Configuration of annotation device] The configuration of the annotation device 10 according to this embodiment will be described in detail with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the configuration of the annotation device according to this embodiment. The annotation device 10 has an input unit 11, an output unit 12, a communication unit 13, a storage unit 14, and a control unit 15.
[0027] The input unit 11 controls input of various information to the annotation device 10. The input unit 11 is, for example, a mouse or a keyboard, and accepts input of setting information and the like to the annotation device 10. The output unit 12 controls output of various information from the annotation device 10. The output unit 12 is, for example, a display, and outputs setting information and the like stored in the annotation device 10.
[0028] The communication unit 13 is responsible for data communication with other devices. For example, the communication unit 13 performs data communication with each communication device. The communication unit 13 can also perform data communication with an operator's terminal (not shown).
[0029] The storage unit 14 stores various information referenced when the control unit 15 operates and various information acquired when the control unit 15 operates. Here, the storage unit 14 is, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. Note that, in the example of Fig. 2, the storage unit 14 is installed inside the annotation device 10, but it may be installed outside the annotation device 10, or multiple storage units may be installed.
[0030] The memory unit 14 stores the first learning data obtained from the database 30 described later, the first learning data to which a first correct label has been assigned obtained from the annotator 20, the classification result classified by the classification unit 15c of the control unit 15, the second learning data generated by the generation unit 15d, the second learning data to which a second correct label has been assigned obtained from the annotator 20, etc., as well as information about the annotator 20, such as a user name and an identification number of the machine learning model.
[0031] The control unit 15 controls the entire annotation device 10. The control unit 15 includes an acquisition unit 15a, a first distribution unit 15b, a classification unit 15c, a generation unit 15d, and a second distribution unit 15e. Here, the control unit 15 is, for example, an electronic circuit such as a central processing unit (CPU) or a micro processing unit (MPU), or an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0032] The acquiring unit 15a acquires first learning data used for machine learning. For example, the acquiring unit 15a acquires the first learning data including audio, images, or videos. The acquiring unit 15a also acquires the first learning data from the database 30. The acquiring unit 15a also acquires learning data to which a correct answer label has been assigned from the annotator 20. The acquiring unit 15a further stores the first learning data, the learning data to which a correct answer label has been assigned, and the like in the storage unit 14.
[0033] The first distribution unit 15b distributes the first learning data acquired by the acquisition unit 15a to a plurality of annotators 20. For example, the first distribution unit 15b distributes the first learning data in a format in which a predetermined number is assigned as a first correct answer label. The first distribution unit 15b also distributes the first learning data to a machine learning model as the annotator 20. Note that detailed processing of the first learning data and the first correct answer label will be described later.
[0034] The classification unit 15c classifies the first learning data based on the reliability of the first correct label assigned to the first learning data by each annotator. For example, the classification unit 15c classifies the first learning data into reference data, data to which the correct label is easily assigned accurately, and data to which the correct label is difficult to assign accurately, based on the variance of the first correct label as the reliability. The classification unit 15c also classifies the first learning data based on the posterior probability of the first correct label as the reliability. Furthermore, the classification unit 15a stores the calculation result of the reliability of the first correct label and the classification result based on the reliability in the storage unit 14.
[0035] Here, when the annotator 20 is a human, the reliability is the variance of the numerical values of the correct labels of each annotator for certain learning data, but is not particularly limited thereto. The index used for the reliability may be any index that represents the dispersion of the numerical values, and the smaller the dispersion of the numerical values, the higher the reliability of the correct labels. Also, when the annotator 20 is a machine learning model, the reliability is the posterior probability of the numerical values that are the estimation results of the machine learning model for certain learning data, but is not particularly limited thereto. The index used for the reliability may be any index that represents the accuracy of the estimation results of the machine learning model, and the higher the accuracy of the estimation results, the higher the reliability of the correct labels.
[0036] The generating unit 15d generates, as a classification result, second learning data that includes a plurality of reference data having different extreme values, data that is easy to accurately label, and data that is difficult to accurately label, and the data is a data group that originates from the same source. Furthermore, the generating unit 15d stores the classification result of the second learning data, etc., in the storage unit 14.
[0037] Here, the extreme value is, for example, the minimum number "1" and the maximum number "5" in the case of learning data in a format in which the degree of a specific state such as concentration is judged as a correct label using a number on a 5-level scale {1, 2, 3, 4, 5}, but is not particularly limited to this. The extreme value may be any number that indicates that the annotator 20 can clearly judge that the state is extreme, and is not limited to the minimum or maximum value of the range of numbers set in advance as the correct label.
[0038] The second distribution unit 15e distributes the classification result of the first learning data classified by the classification unit 15c. For example, the second distribution unit 15e first distributes a plurality of reference data with different extreme values. In addition, the second distribution unit 15e distributes the classification result to the plurality of annotators who distributed the first learning data, or to a predetermined annotator other than the plurality of annotators who distributed the first learning data.
[0039] Here, the classification result is the first learning data classified by the classification unit 15c based on the reliability of the assigned correct label, and is, for example, learning data labeled with three categories of reference point S (reference data), data E (data to which a correct label is easily assigned accurately), and data D (data to which a correct label is difficult to assign accurately), but is not particularly limited. The classification result may be learning data labeled with the reliability of the correct label, or may be learning data selected by the generation unit 15d.
[0040] [Example of annotation processing] A specific example of the annotation process of the annotation device 10 according to the present embodiment will be described with reference to Figs. 3 to 5. Fig. 3 is a diagram showing an example of learning data according to the first embodiment. Fig. 4 is a diagram showing an example of first learning data and a first correct label according to the first embodiment. Fig. 5 is a diagram showing an example of second learning data and a second correct label according to the first embodiment.
[0041] (First annotation process) First, the first annotation process from acquisition of the first learning data to acquisition of the first learning data to which the first correct answer label is assigned will be described. First, the first learning data acquired by the annotation device 10 from the database 30 or the like will be described with reference to FIG. 3. Here, the acquired first learning data is data such as audio, image, video, etc., and is data acquired in a medium and on a scale according to the purpose of research and development. For example, when performing annotation to realize concentration level estimation from audio, the annotation device 10 acquires audio data from an audio database that stores audio data in the database 30.
[0042] FIG. 3 shows a dataset X that holds voice data. The dataset X contains {x0, x1, x2, x3, x4, . . . x N 3 includes the voice data of the following: The voice data shown in Fig. 3 is a representation of the voice waveform as a relationship between the passage of time and the voice signal strength.
[0043] In the following description, annotation processing using voice data as the first learning data will be described, but the type of learning data is not particularly limited. The first learning data may be image data, video data, or a combination thereof, in addition to voice data. Furthermore, the first learning data may be data obtained by converting the above voice data, etc., into numerical values or text.
[0044] Next, the first learning data that the annotation device 10 distributes to the annotator 20 and the first correct label assigned to the first learning data that the annotation device 10 acquires from the annotator 20 will be described with reference to Fig. 4. For example, when performing annotation to realize concentration level estimation from voice, the annotation device 10 sets five levels of concentration levels ("1": not concentrated, "2": slightly not concentrated, "3": neither concentrated nor concentrated (flat), "4": slightly concentrated, "5": concentrated) in advance, and distributes the first learning data to the annotator 20, which causes a correct label to be assigned according to which concentration level is most suitable for each voice data held in the dataset X.
[0045] When performing annotation to realize concentration level estimation from voice, the annotation device 10 distributes learning data for judging, for example, from voice data regarding questions and dialogue between a teacher and a student during a class, whether the student to whom a question was asked was concentrating on the class, or not, and how concentrated he or she was, on a five-point scale. When assigning a correct answer label regarding the concentration level from image data or video data, the annotation device 10 may have the annotator 20 read the facial expressions of the students from images or videos during the class, and distribute learning data for judging the concentration level.
[0046] Then, the annotation device 10 acquires the first learning data to which the correct answer labels have been assigned by the annotator 20. In FIG. 3, the speech data “x0” to “x N The correct labels for "ANNOT1(X)" are shown (see Figure 4 "ANNOT1(X)"). For example, the correct labels for the audio data x0 assigned by "Annotator 01" to "Annotator 03" are "2", "1", and "1", respectively.
[0047] (Second annotation process) Secondly, the second annotation process from classification of the first learning data assigned with the first correct label to acquisition of the second learning data assigned with the second correct label will be described. First, a specific example of classification process of the first learning data based on the reliability of the correct label acquired from the annotator 20 will be described with reference to Fig. 4. The annotation device 10 calculates the average and variance values of the correct label assigned to each voice data.
[0048] In Figure 4, x0 has a mean of 1.3 and a variance of 0.3 (small variance). Similarly, x1 has a mean of 5 and a variance of 0 (all annotators' answers match), x2 has a mean of 1 and a variance of 0 (all annotators' answers match), x3 has a mean of 3.3 and a variance of 0.3 (small variance), x4 has a mean of 4.0 and a variance of 1.0 (large variance), and xN The mean is 1.6 and the variance is 1.3 (large variance).
[0049] At this time, the annotation device 10 selects, as the reference point S, the learning data to which reliable correct answer data has been assigned from the answers acquired from the annotator 20. In the example of Fig. 4, the annotation device 10 selects x1 (extreme value "5") and x2 (extreme value "1") as the cases where the answers of all the annotators are consistent and the numerical values of the assigned correct answer labels are extreme values.
[0050] The annotation device 10 further classifies the learning data other than the reference point S into data D, which is easy to accurately label as data D, and data E, which is difficult to accurately label as data E. For example, the annotation device 10 sets a reliability threshold and classifies data D if the variance is 1.0 or more, and classifies data E if not. In the example of FIG. 4, the annotation device 10 classifies x0 and x3 into data E because the variance is less than 1.0, and x4 and x5 into data E. N Since the variance is greater than 1.0, it is classified as data D.
[0051] When the annotation device 10 uses an estimation result by a separately created machine learning model as a correct label, the annotation device 10 classifies, for example, learning data in which the posterior probability of the numerical value that is the estimation result is 80% or more as a reference point S, learning data in which the posterior probability is 50% or more and less than 80% as data E, and learning data in which the posterior probability is less than 50% as data D. Furthermore, the annotation device 10 can statically or dynamically change the classification method, such as the number of classifications of the learning data and the threshold value.
[0052] Next, a specific example of the generation and distribution process of the second learning data based on the classification of the first learning data will be described with reference to Fig. 5. The annotation device 10 generates a data group including three types of data, namely, a reference point S, data E, and data D, as the second learning data. At this time, the reference point S is always made to include data of "1" and "5", which are the extreme values of each. In addition, the source of the data of each data group (speaker, person appearing in a video, object, etc.) is assumed to be the same.
[0053] For example, the annotation device 10 may generate a set of data {p0, p1, . . . p M A data set P is generated with elements {x0, x1, x2, x3, x4, x N} (see Figures 3 and 4) as elements, and the data group p M In a ,x b ,x c ,x d ,x e ,x f} (not shown in FIG. 3 and FIG. 4) as elements. In the example of FIG. 5, the annotation device 10 includes reference points S{x1, x2}, data E{x0, x3}, data D{x4, x N} and the data set p M Then, the reference point S{x a ,x b}, data E{x c ,x d}, data D{x e ,x f} is selected.
[0054] The number of data in each data group and the selection method can be changed as desired as long as they include reference points S that are different extreme values and have the same source, as described above. For example, the number of data may be a random number within a certain range. Also, the included data may be two each of data E and data D, and the average value of each annotation result may be closer to "1" or "5."
[0055] Thereafter, the annotation device 10 distributes the data group selected as described above to the annotator 20 as the second learning data. At this time, when distributing the data group of the second learning data, the annotation device 10 first distributes the reference point S for each data group, and then distributes data E and data D to the annotator 20. In the example of FIG. 5, when distributing the data group p0, the annotation device 10 distributes the reference point S{x1, x2}, data E{x0, x3}, data D{x4, x N}, and the data set p M When distributing, the reference point S{x a ,x b}, data E{x c ,x d}, data D{x e ,x f} in that order.
[0056] The annotation device 10 may instruct the annotator 20 to view the data in the order of data E and data D. The annotation device 10 may also distribute data E and data D to be viewed randomly. Furthermore, the annotation device 10 may present a correct label for the reference point S that is distributed first, at the same time as distribution, and instruct the annotator 20 not to assign a correct label, or may instruct the annotator 20 to assign a correct label to all learning data regardless of the classification of the learning data.
[0057] Finally, the annotation device 10 acquires the second learning data to which the correct answer labels have been assigned by the annotator 20. In the example of Fig. 5, the annotation device 10 acquires the correct answer labels of the data E and the data D excluding the reference point S for each data group for each annotator. For example, for "annotator 01", 3, x4,x N} training data, in order {1,4 , 3,2}, and the correct labels are obtained for the data set p M {x c ,x d ,x e ,x f} training data, in order {2,4 , The correct label {X,3,3} is obtained (see Figure 5 “ANNOT2(X)”).
[0058] The final processing of the correct labels assigned to the second learning data is not particularly limited. The annotation device 10 may take a majority vote for each learning data and determine the most common correct label as the final correct label, or may calculate the average score of the numerical values and determine the numerical value as the final correct label.
[0059] [Annotation process flow] The flow of annotation processing according to this embodiment will be described in detail with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the flow of annotation processing according to the first embodiment.
[0060] First, the acquisition unit 15a of the annotation device 10 acquires first learning data including voice, image, video, etc. from the database 30 or the like (step S101). At this time, the acquisition unit 15a may acquire the first learning data from the storage unit 14. The acquisition unit 15a may process the original data such as voice data acquired from the database 30 or the storage unit 14, divide the data into appropriate sizes as learning data, or classify the data appropriately. Furthermore, the acquisition unit 15a may acquire voice data, etc. from the outside via the input unit 11.
[0061] Next, the first distribution unit 15b distributes the first learning data to the annotator 20 (step S102). At this time, the first distribution unit 15b may select an annotator 20 to distribute the first learning data according to the first learning data. In addition, the acquisition unit 15a acquires the first learning data to which the first correct answer label has been assigned by the annotator 20 (step S103).
[0062] Then, the classification unit 15c classifies the first learning data based on the reliability of the first correct label (step S104). The generation unit 15d generates second learning data from the classified first learning data (step S105). Then, the second distribution unit 15e distributes the second learning data to the annotator 20 (step S106).
[0063] In addition, the second distribution unit 15e can also distribute the second learning data to an annotator other than the annotator 20 that distributed the first learning data. For example, the second distribution unit 15e can also distribute the first learning data to a human annotator and distribute the second learning data to an annotator that is a machine learning model.
[0064] Finally, the acquiring unit 15a acquires the second learning data to which the second correct label has been assigned by the annotator 20 (step S107), and the process ends. Note that if the accuracy of the acquired second correct label is not sufficient, the processes of steps S104 to S107 may be performed again.
[0065] [Flow of classification process for the first training data] The flow of the classification process of the first learning data according to the present embodiment will be described in detail with reference to Fig. 7. Fig. 7 is a flowchart showing an example of the flow of the classification process of the first learning data according to the first embodiment. First, the acquisition unit 15a of the annotation device 10 acquires the first correct label assigned to the first learning data from the annotator 20 (step S201). Next, when the annotator 20 is a human (step S202: annotator is a human), the classification unit 15c performs classification processes of steps S208 to S210 based on the processes of steps S203 to S205.
[0066] If the answers of all the annotators match (step S203: YES) and the answers are extreme values (step S204: YES), the classification unit 15c classifies the first learning data to which the correct label has been assigned as a reference point S (step S208). If the answers of the annotators 20 include ones that do not match (step S203: NO) or if the answers of the annotators 20 are not extreme values (step S204: NO), the classification unit 15c performs the process of step S205.
[0067] If the variance of the answer of the annotator 20 is 1.0 or more (step S205: Yes), the classification unit 15c classifies the first learning data to which the correct label has been assigned as data E (step S209). If the variance of the answer of the annotator 20 is less than 1.0 (step S205: No), the classification unit 15c classifies the first learning data to which the correct label has been assigned as data D (step S210). When the classification process of steps S208 to S210 is completed, the classification unit 15c ends the process.
[0068] On the other hand, when the annotator 20 is a machine learning model (step S202: annotator is a machine learning model), the classification unit 15c performs classification processing of steps S208 to S210 based on the processing of steps S206 to S207. When the posterior probability of the value that is the estimation result of the annotator 20 is 80% or more (step S206: Yes), the classification unit 15c classifies the first learning data to which the correct label has been assigned as the reference point S (step S208).
[0069] Furthermore, if the posterior probability of the value that is the estimation result of the annotator 20 is less than 80% (step S206: No) and the posterior probability is 50% or more (step S207: Yes), the classification unit 15c classifies the first learning data to which the correct label has been assigned as data E (step S209). Furthermore, if the posterior probability of the value that is the estimation result of the annotator 20 is less than 50% (step S207: No), the classification unit 15c classifies the first learning data to which the correct label has been assigned as data D (step S210). When the classification process of steps S208 to S210 is completed, the classification unit 15c ends the process.
[0070] [Advantages of the First Embodiment] First, in the annotation process according to the present embodiment described above, first learning data used in machine learning is acquired, the acquired first learning data is distributed to a plurality of annotators, and the first learning data is classified based on the reliability of the first correct answer labels respectively assigned to the first learning data by each annotator, and the classification result of the classified first learning data is distributed. Therefore, in this process, annotation can be performed at a lower cost and with higher accuracy in supervised learning in machine learning.
[0071] Secondly, in the annotation process according to the present embodiment described above, first learning data including audio, images or videos is acquired, the first learning data is distributed in a format in which a predetermined number is assigned as a first correct label, and the first learning data is classified into reference data, data that is easy to accurately assign a correct label to, and data that is difficult to accurately assign a correct label to based on the variance of the first correct label as the reliability. Therefore, in the present process, in supervised learning in machine learning, it is possible to assign a highly reliable correct label even if there is no comparison target, and it is possible to perform annotation at a lower cost and with higher accuracy.
[0072] Thirdly, in the annotation process according to the present embodiment described above, the annotator distributes first learning data to the machine learning model and classifies the first learning data based on the posterior probability of the first correct label as the reliability. Therefore, in this process, in supervised learning in machine learning, even if the annotator is not human, it is possible to assign a highly reliable correct label, and it is possible to perform annotation at a lower cost and with higher accuracy.
[0073] Fourth, in the annotation process according to the present embodiment described above, as a classification result, second learning data is generated, which is a data group including multiple reference data with different extreme values, data that is easy to accurately label, and data that is difficult to accurately label, and each data is generated from the same source, and the multiple reference data are distributed first. Therefore, in supervised learning in machine learning, this process enables reliable and efficient labeling of correct answers even when there is no comparison target, and can perform annotation at a lower cost and with higher accuracy.
[0074] Fifth, in the annotation process according to the present embodiment described above, the classification result is distributed to the multiple annotators who distributed the first learning data, or to a predetermined annotator other than the multiple annotators who distributed the first learning data. Therefore, in the present process, in supervised learning in machine learning, even if there is no comparison target, it is possible to assign a reliable, efficient, and more flexible correct answer label, and to perform annotation at a lower cost and with higher accuracy.
[0075] [System configuration, etc.] Each component of each device shown in the figures according to the above embodiment is a functional concept, and does not necessarily have to be physically configured as shown in the figures. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figures, and all or a part of them can be functionally or physically distributed and integrated in any unit according to various loads, usage conditions, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.
[0076] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed arbitrarily unless otherwise specified.
[0077] 〔program〕 It is also possible to create a program in which the processing executed by the annotation device 10 described in the above embodiment is written in a language executable by a computer. In this case, the same effect as in the above embodiment can be obtained by the computer executing the program. Furthermore, such a program may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read and executed by a computer to realize the same processing as in the above embodiment.
[0078] Fig. 8 is a diagram showing a computer that executes a program. As shown in Fig. 8, the computer 1000 includes, for example, a memory 1010, a CPU (Central Processing Unit) 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070, and these components are connected by a bus 1080.
[0079] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012, as exemplified in FIG. 8. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090, as exemplified in FIG. 8. The disk drive interface 1040 is connected to a disk drive 1100, as exemplified in FIG. 8. For example, a removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120, as exemplified in FIG. The video adapter 1060 is connected to, for example, a display 1130, as exemplified in FIG. 8.
[0080] 8, the hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the above programs are stored in, for example, the hard disk drive 1090 as program modules in which instructions to be executed by the computer 1000 are written.
[0081] Furthermore, the various data described in the above embodiment are stored as program data, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes various processing procedures.
[0082] Note that the program module 1093 and program data 1094 relating to the program are not limited to being stored in the hard disk drive 1090, and may be stored in, for example, a removable storage medium and read by the CPU 1020 via a disk drive or the like. Alternatively, the program module 1093 and program data 1094 relating to the program may be stored in another computer connected via a network (such as a local area network (LAN) or wide area network (WAN)) and read by the CPU 1020 via the network interface 1070.
[0083] The above-described embodiments and their modifications are included in the technology disclosed in this application, as well as in the scope of the invention described in the claims and their equivalents. [Explanation of symbols]
[0084] 10 Annotation device 11 Input section 12 Output section 13. Communications Department 14 Storage section 15 Control section 15a Acquisition part 15b 1st Distribution Section 15c Classification section 15d Generator 15e 2nd Distribution Section 20, 20A, 20B, 20C Annotator 30, 30A, 30B, 30C Database 100 Annotation System
Claims
1. an acquisition unit that acquires first learning data to be used in machine learning; a first distribution unit that distributes the first learning data acquired by the acquisition unit to a plurality of annotators; a classification unit that classifies the first training data based on a reliability of a first correct label that is assigned to each of the first training data by each annotator; a second distribution unit that distributes a classification result of the first learning data classified by the classification unit to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Equipped with The acquisition unit acquires the first learning data including audio, images, or videos, The first distribution unit distributes the first learning data in a format in which a predetermined number is assigned as the first correct label; The classification unit classifies the first learning data into reference data, data that is easy to accurately assign a correct label, or data that is difficult to accurately assign a correct label based on the variance of the first correct label as the reliability.
2. The first distribution unit distributes the first training data to a machine learning model as the annotator; The annotation device according to claim 1 , wherein the classification unit classifies the first learning data based on a posterior probability of the first correct label as the reliability.
3. a generation unit that generates second learning data as the classification result, the second learning data including a plurality of reference data having different extreme values, the data to which the correct label is easily assigned accurately, and the data to which the correct label is difficult to assign accurately, the second learning data being generated from a same data group; The annotation device according to claim 1 , wherein the second distribution unit distributes the plurality of reference data first.
4. An annotation method performed by an annotation device, comprising: An acquisition step of acquiring first learning data to be used in machine learning; a first distribution step of distributing the first training data acquired by the acquisition step to a plurality of annotators; a classification step of classifying the first training data based on the reliability of a first correct label assigned to each of the first training data by each annotator; a second distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Including, The acquiring step acquires the first learning data including audio, an image, or a video; The first distribution step distributes the first learning data in a format in which a predetermined number is assigned as the first correct label; The annotation method is characterized in that the classification step classifies the first learning data into reference data, data that is easy to accurately assign a correct label, or data that is difficult to accurately assign a correct label based on the variance of the first correct label as the reliability.
5. An acquisition step of acquiring first learning data used in machine learning; a first distribution step of distributing the first training data acquired by the acquisition step to a plurality of annotators; a classification step of classifying the first training data based on the reliability of a first correct label assigned to each of the first training data by each annotator; a second distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Run the following on your computer: The acquiring step acquires the first learning data including audio, an image, or a video; The first distribution step distributes the first learning data in a format in which a predetermined number is assigned as the first correct label; The classification step classifies the first learning data into reference data, data that is easy to accurately assign a correct label, or data that is difficult to accurately assign a correct label based on the variance of the first correct label as the reliability.
6. a classification unit that classifies first learning data used in machine learning based on the reliability of first correct labels that are respectively assigned by a plurality of annotators to the first learning data; a distribution unit that distributes a classification result of the first learning data classified by the classification unit to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Equipped with The annotation device, wherein the reliability represents a variance of the first correct label.
7. a classification unit that classifies first learning data used in machine learning based on the reliability of first correct labels that are respectively assigned by a plurality of annotators to the first learning data; a distribution unit that distributes a classification result of the first learning data classified by the classification unit to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Equipped with An annotation device characterized in that the classification result includes reference data selected based on the reliability.
8. a classification unit that classifies first learning data used in machine learning based on the reliability of first correct labels that are respectively assigned by a plurality of annotators to the first learning data; a distribution unit that distributes a classification result of the first learning data classified by the classification unit to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Equipped with The annotation device, characterized in that the classification result includes classification of the first learning data into data that is easy to accurately assign a correct label to, or data that is difficult to accurately assign a correct label to, based on the reliability.
9. An annotation method performed by an annotation device, comprising: a classification step of classifying first learning data used for machine learning based on the reliability of first correct labels respectively assigned by a plurality of annotators to the first learning data; a distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Including, The annotation method, wherein the reliability represents a variance of the first correct label.
10. An annotation method performed by an annotation device, comprising: a classification step of classifying first learning data used for machine learning based on the reliability of first correct labels respectively assigned by a plurality of annotators to the first learning data; a distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Including, An annotation method, characterized in that the classification result includes reference data selected based on the reliability.
11. An annotation method performed by an annotation device, comprising: a classification step of classifying first learning data used for machine learning based on the reliability of first correct labels respectively assigned by a plurality of annotators to the first learning data; a distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; Equipped with The annotation method, characterized in that the classification result includes classification of the first learning data into data that is easy to accurately assign a correct label to, or data that is difficult to accurately assign a correct label to, based on the reliability.
12. a classification step of classifying first learning data used for machine learning based on the reliability of first correct labels respectively assigned by a plurality of annotators to the first learning data; a distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; on the computer, The annotation program, wherein the reliability represents a variance of the first correct label.
13. a classification step of classifying first learning data used for machine learning based on the reliability of first correct labels respectively assigned by a plurality of annotators to the first learning data; a distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; on the computer, The annotation program, wherein the classification result includes reference data selected based on the reliability.
14. a classification step of classifying first learning data used for machine learning based on the reliability of first correct labels respectively assigned by a plurality of annotators to the first learning data; a distribution step of distributing the classification result of the first learning data classified by the classification step to the plurality of annotators or a predetermined annotator other than the plurality of annotators; on the computer, The annotation program, characterized in that the classification result includes classification of the first learning data into data that is easy to accurately assign a correct label to, or data that is difficult to accurately assign a correct label to, based on the reliability.
Citation Information
Patent Citations
Cell annotation method and system using adaptive incremental learning
JP2019521443A
Operation device
JP2020144755A
Systems and Methods for the Determining Annotator Performance in the Distributed Annotation of Source Data
US20130346409A1
Method for managing annotation job, apparatus and system supporting the same
US20200152316A1
Management of annotation jobs
US20200342165A1