Annotation-based learning data storage system, and method for storing annotated learning data.

The annotation-based learning data storage system efficiently distributes annotation tasks using pre-trained models on mobile devices and servers, reducing time and effort while enhancing accuracy.

JP7839541B2Active Publication Date: 2026-04-02THE RITSUMEIKAN TRUST
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Creating annotations for machine learning datasets is time-consuming and labor-intensive, often inaccurate due to differences in lighting and imaging methods, and scaling large-scale data collection is difficult.

Method used

An annotation-based learning data storage system comprising a mobile terminal and a server, utilizing pre-trained machine learning models to perform initial classification on the mobile device and detailed checks on the server, ensuring accurate annotation distribution and storage.

Benefits of technology

Reduces working time and effort while improving annotation accuracy by distributing the annotation work across mobile devices and cloud servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839541000001
    Figure 0007839541000001
  • Figure 0007839541000002
    Figure 0007839541000002
  • Figure 0007839541000003
    Figure 0007839541000003
Patent Text Reader

Abstract

To provide a system that achieves reduction in work time and improvement in annotation accuracy.SOLUTION: A mobile terminal transmits image data to an annotation-added learning data storage server if classification data indicates that an image of an object of data collection is included. The annotation-added learning data storage server stores the image data from the mobile terminal in a storage device as annotation-added learning data when out of confidence data corresponding to each piece of class data acquired for each piece of the image data of an image portion, the maximum confidence data is greater than a predetermined threshold, and for each piece of the image data of the image portion, class data corresponding to the maximum confidence data matches the class data associated with the image data of the image portion.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an annotation learning data accumulation system for creating teacher data for machine learning, and a method for accumulating annotation learning data.

Background Art

[0002] Currently, the work of creating annotations required for creating a dataset (i.e., a training dataset) for training a machine learning model (AI model) takes time and effort.

[0003] Conventionally, in a machine learning model that uses image data as an explanatory variable, in many cases, a person who was not at the shooting site visually performs the annotation work on the captured image data. When the work is performed without seeing the actual object, a difference in the accuracy of the annotation appears depending on the way the light hits during shooting and the shooting method.

[0004] Therefore, in such a method, not only does it take time and effort, but it is also likely to lack accuracy in creating training data. Furthermore, creating large-scale data becomes increasingly difficult.

[0005] However, since a certain amount of labor is indispensable for the annotation work, the realization of a system that implements the annotation work simultaneously with the acquisition work of image data is currently difficult.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0007] If the imager can perform annotation work while comparing the target object with the image data at each site where data, especially image data, is acquired, it will save time spent solely on creating a dataset of ground truth data, and furthermore, increase the accuracy of the annotation work.

[0008] This disclosure aims to provide an annotation training data storage system that enables the essential parts of annotation work to be performed at the data collection site, thereby distributing the implementation of annotation work, reducing working time and effort, and improving the accuracy of annotation. [Means for solving the problem]

[0009] The annotation-based learning data storage system disclosed herein is comprised of a mobile terminal and an annotation-based learning data storage server, connected via a predetermined network. The mobile terminal includes a first interface device and a first processing circuit. The annotation-assigning learning data storage server includes a second interface device, a storage device, and a second processing circuit. Mobile devices include: A pre-trained first machine learning model has been constructed, which has been trained using image data and classification data indicating whether or not the image data contains an image of the object being collected as training data. The annotation-assigned training data storage server includes: A pre-trained second machine learning model has been constructed, which was trained using image data and class data of the objects included in the image data as training data. The first processing circuit of the mobile terminal is Image data is acquired via a first interface device, which includes an image portion of one or more data collection targets, the coordinate data of the image portion of the data collection targets, and the associated class data of the targets. Image data acquired via the first interface device is input to a first trained machine learning model, and classification data output from the first trained machine learning model is acquired. If the acquired classification data indicates that it includes an image of the object being collected, the image data acquired via the first interface device is further transmitted via the first interface device to the annotation training data storage server. The second processing circuit of the annotation-assigned learning data storage server is: Image data from a mobile device is received via a second interface device. From the received image data, obtain image data of one or more image portions specified by coordinate data. Each image data of the image portion is input into a pre-trained second machine learning model, and for each image data of the image portion, one or more class data and confidence data corresponding to each of those class data are obtained. For each image data of the image portion, if the maximum confidence value data among the confidence value data corresponding to one or more class data obtained is greater than a predetermined threshold, For each image data in the image portion, if the class data corresponding to the confidence data of the maximum value matches the class data associated with the image data in the image portion, Image data from mobile devices is stored in a memory device as annotation-attached training data. [Effects of the Invention]

[0010] By utilizing the annotation-based learning data storage system disclosed herein, the annotation work can be distributed, further reducing working time and effort while improving the accuracy of annotations. [Brief explanation of the drawing]

[0011] [Figure 1]FIG. 1 is a system configuration diagram of an annotation learning data accumulation system according to an embodiment. [Figure 2A] FIG. 2A is an example of image data captured by a mobile terminal with a bounding box and a label attached to an object of data collection. [Figure 2B] FIG. 2B is an example of image data collected by a mobile terminal including prank data and misdata. [Figure 3] FIG. 3 is a schematic flowchart of a prank data check process (first check process) and a transmission process in a mobile annotation system that constitutes an annotation learning data accumulation system according to an embodiment. [Figure 4] FIG. 4 is a schematic flowchart of a misdata check process (second check process) in an annotation learning data accumulation server that constitutes an annotation learning data accumulation system according to an embodiment. [Figure 5] FIG. 5 is an example of a process of a misdata check process (second check process) in an annotation learning data accumulation server that constitutes an annotation learning data accumulation system according to an embodiment. [Figure 6] FIG. 6 is an example of image data output and accumulated by an annotation learning data accumulation server that constitutes an annotation learning data accumulation system according to an embodiment.

MODE FOR CARRYING OUT THE INVENTION

[0012] Hereinafter, embodiments will be described in detail with reference to the drawings as appropriate. However, a more detailed description than necessary may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially the same configurations may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate understanding by those skilled in the art.

[0013] The inventors provide the accompanying drawings and the following description to enable those skilled in the art to fully understand the present disclosure, but they are not intended to limit the subject matter recited in the claims thereby.

[0014] 1. Background of the Present Disclosure In constructing a machine learning model using image data as an explanatory variable, a certain amount of working time and labor are indispensable for the annotation work required to create a training dataset.

[0015] Moreover, conventionally, the work of collecting image data and the work of annotating the collected image data have been performed by different workers on different occasions. For example, in many cases, for image data collected by various data collectors, only those who perform the annotation work visually annotate them collectively. Since those who only perform the annotation work work intensively at a place completely different from the site of collecting image data, they perform annotations for various image data, that is, various image data with different light intake methods and imaging methods. Then, differences will inevitably occur in the accuracy of the annotations in the created training dataset. In addition, since the annotation work is performed for various image data, the labor and time in the annotation work tend to increase instead.

[0016] If, at the site of acquiring image data, while comparing the object of actual data collection with the image, the important part of the annotation work can be carried out, the time to be spent for creating correct data can be greatly reduced, and furthermore, the accuracy of the correct data can be improved. The present disclosure is obtained from such an idea, and the present disclosure provides a system that realizes highly accurate annotation in two stages of a distributed type and a centralized type.

[0017] 2. Embodiments Hereinafter, preferred embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0018] 2.1. System Configuration The annotation-based learning data storage system and method for storing annotation-based learning data according to this embodiment are systems and methods that perform annotation, particularly of image data, through a two-stage check between a mobile terminal and an annotation-based learning data storage server. Figure 1 is a system configuration diagram of the annotation-based learning data storage system 2 according to this embodiment.

[0019] The annotation-based learning data storage system 2 includes at least one mobile terminal 4, a computer device 16, and a storage device 12. The mobile terminal 4 has a mobile annotation system 3 implemented as an application system. The computer device 16 and the storage device 12 constitute the annotation-based learning data storage server 5.

[0020] The computer device 16 and the storage device 12 are connected by wire or wireless means, and can send and receive data from each other. The computer device 16 and the storage device 12 are connected to an external network 18, and can exchange data with other computer systems connected to the external network 18, such as the mobile terminal 4. Furthermore, the computer device 16 of the annotation-assigning learning data storage server 5 may be connected to the learning server 14.

[0021] The computer device 16 is a server machine or workstation computer equipped with one or more processors.

[0022] The storage device 12 is a storage device located outside the computer device 16, such as a disk drive or flash memory, and stores various databases, various datasets (e.g., training datasets), and various computer programs used by the computer device 16. For example, the storage device 12 stores training datasets as image data transmitted from the mobile terminal 4, which will be described later, or from an external source.

[0023] The mobile terminal 4 is a portable computing device, such as a tablet or smartphone. The mobile terminal 4 includes an image sensor and imaging control means for controlling the image sensor. The mobile terminal 4 transmits and receives image data to and from a computer device 16, etc.

[0024] The learning server 14 uses image data and other data recorded in the storage device 12 to train various machine learning models. Note that the training of machine learning models may also be performed by the computer device 16 of the annotation-adding training data storage server 5, in which case the learning server 14 does not need to be used.

[0025] The external network 18 is, for example, the internet, and is connected to the mobile terminal 4 via the first interface device 6 (described below) and to the computer device 16 via the second interface device 22 (described later), such as a network terminal.

[0026] Furthermore, the mobile terminal 4 includes a first interface device 6 and a first processing circuit 8.

[0027] The first interface device 6 is an interface unit capable of exchanging data with the outside world, including a network terminal, video input terminal, USB terminal, keyboard, mouse, etc. The data here is, for example, image data of the object to be collected, which will be described later. After acquisition, this data can be recorded in a storage unit (not shown) inside the mobile terminal 4. The data recorded in the storage unit can be acquired into the computer device 16 via the first interface device 6 as appropriate.

[0028] The first processing circuit 8 is comprised of a processor. Here, the processor encompasses a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). In this embodiment, the various processes of the portable annotation system 3 in the annotation-based learning data storage system 2 are realized by the execution of various programs by the first processing circuit 8. These various processes may be implemented by an ASIC (Application Specific Integrated Circuit) or a combination thereof.

[0029] The portable annotation system 3 in the annotation-based learning data storage system 2 according to this embodiment is constructed using a computer language such as Python. The computer languages ​​that can be used to construct the portable annotation system 3 according to this embodiment are not limited to these, and of course, other computer languages ​​may be used.

[0030] Furthermore, in the portable annotation system 3 according to this embodiment, a pre-trained first machine learning model is constructed. The first machine learning model according to this embodiment is constructed using a lightweight deep learning model so that it can operate on smartphones and small tablet devices. As will be explained later, the pre-trained first machine learning model only classifies whether or not the input image data contains images of the objects to be collected, so it is constructed appropriately using a lightweight deep learning model.

[0031] Furthermore, the aforementioned computer device 16 includes a second interface device 22, a second processing circuit 24, and a memory 10.

[0032] The second interface device 22, like the first interface device 6, is an interface unit capable of exchanging data with the outside world, including a network terminal, video input terminal, USB terminal, keyboard, mouse, etc. The data here includes, for example, image data of the data collection target received from the mobile terminal 4, which will be explained later, and annotated image data of the data collection target. After acquisition, this data can be recorded in the storage device 12. The data recorded in the storage device 12 can be acquired into the computer device 16 via the second interface device 22 as appropriate.

[0033] Various types of data generated by the annotation-assigned learning data storage server 5 are appropriately recorded in the storage device 12. These types of data include, for example, image data of the objects that have been annotated for data collection.

[0034] The second processing circuit 24, like the first processing circuit 8, is composed of a processor. Here, the processor encompasses a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). In the annotation-based learning data storage system 2 according to this embodiment, the various processes of the annotation-based learning data storage server 5 are realized by the execution of various programs by the second processing circuit 24. These various processes may be realized by an ASIC (Application Specific Integrated Circuit) or a combination thereof.

[0035] The second processing circuit 24 in this embodiment may be composed of multiple signal processing circuits. Each signal processing circuit may be, for example, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), and may be called a "processor". In the annotation-based learning data storage system 2 according to this embodiment, one processor (for example, a GPU) may perform some of the various processes of the annotation-based learning data storage server 5, and another processor (for example, a CPU) may perform some of the other processes.

[0036] Memory 10 is a data rewritable storage unit inside the computer device 16, and is composed of, for example, RAM (Random Access Memory) containing a large number of semiconductor memory elements. Memory 10 temporarily stores specific computer programs, variable values, parameter values, etc., when the second processing circuit 24 performs various processes. Memory 10 may also include so-called ROM (Read Only Memory). The ROM pre-stores computer programs that realize various processes of the annotation learning data storage server 5 in the annotation learning data storage system 2 described below. The second processing circuit 24 reads the computer programs from the ROM and expands them into RAM, enabling the second processing circuit 24 to execute the computer programs.

[0037] The various processing programs of the annotation-based learning data storage server 5 in the annotation-based learning data storage system 2 according to this embodiment are constructed using a computer language such as Python. The computer languages ​​that can be used to construct the various processes in the annotation-based learning data storage server 5 according to this embodiment are not limited to these, and of course, other computer languages ​​may be used.

[0038] Furthermore, in the annotation-adding training data storage server 5 according to this embodiment, a trained second machine learning model is constructed. The second machine learning model according to this embodiment takes image data as input and outputs image data and class data, and can be implemented with a network structure of a convolutional neural network (CNN) including, for example, U-Net, Res-net, skipped connection, batch normalization, Max pooling, and autoencoder.

[0039] 2.2. [System Operation] The objects to which image data is collected by the annotation-assigning learning data storage system 2 according to this embodiment are not particularly limited. They may be objects to be identified, including tableware, fruits, and / or people, or they may be objects to be inspected, including metal, printed materials, food, and / or fruits, and may include visual inspection items such as scratches and stains. In this embodiment, tableware, in particular, tableware used in a dish return system operating in a school or workplace cafeteria, is used as the object to which image data is collected.

[0040] Therefore, in an image in which one or more dishes are captured, for example, after a meal, an example of ground truth data would be to surround each dish with a bounding box and label each bounding box with the name of the dish (i.e., class data or label). Creating such ground truth data for an image in which one or more dishes are captured is the annotation in this embodiment. Note that the bounding box is one example of data for indicating the coordinates of the dishes. The data indicating the coordinates of the dishes may also be a circle surrounding each dish.

[0041] 2.2.1. [Operation of the mobile annotation system] The main function of the portable annotation system 3 in the annotation-adding learning data storage system 2 according to this embodiment is to exclude image data of objects that are clearly not tableware. This is intended to pre-delete (pre-remove) so-called "prank data". For example, the image of a flower shown in Figure 2B (B1) is an image of an object that is clearly not tableware and can be said to be prank data that should be pre-deleted. In order to realize this main function, a pre-trained first machine learning model is constructed in the portable annotation system 3.

[0042] This first machine learning model only classifies whether the input image data contains the image portion of tableware. Therefore, the first machine learning model is trained using image data containing various tableware portions and classification data indicating the presence of tableware as training data. Furthermore, since the trained first machine learning model only classifies whether the input image data contains the image portion of tableware, and must be deployable on various smartphones and tablet devices, it is desirable that the first machine learning model be a lightweight deep learning model.

[0043] Each of the mobile devices 4 is pre-equipped with a pre-trained first machine learning model.

[0044] Figure 3 is a flowchart illustrating the prank data check process (first check process) and transmission process in the mobile annotation system 3, which is configured using a mobile terminal 4 and constitutes the annotation-adding learning data storage system 2 according to the embodiment. The operation of the mobile annotation system 3 will be described below with reference to the flowchart shown in Figure 3.

[0045] As a preprocessing step (step S02) before the first processing circuit 8 of the mobile terminal 4 acquires image data, an image (see Figure 2A) of the object to be collected, such as tableware in the dish return system, is captured by the operator of the mobile terminal 4, for example, a user of the cafeteria. The operator of the mobile terminal 4 operates the mobile terminal 4 to add bounding boxes and labels to the image portion of the tableware. The labels are selected from a predetermined class. When the object to be collected is tableware in the dish return system, the class includes various tableware names, such as [Rice-Bowl, Soup-Bowl, Main-Dish, Dessert-Dish, Square-bowl, Fish-dish, Cup, Water-cup, Tea-cup, Wine-cup, Spoon, Chopsticks-two].

[0046] Figure 2A shows an example of image data captured by the mobile terminal 4, where a bounding box and label have been attached to the tableware, which is the object of data collection, by the operator of the mobile terminal 4. The bounding box indicates the coordinate data of the image portion of the tableware. After the start (step S01), the first processing circuit 8 of the mobile terminal 4 first acquires the image data in this state (step S02). That is, the first processing circuit 8 of the mobile terminal 4 acquires image data via the first interface device 6 that includes the image portions of one or more tableware objects, includes the coordinate data of those image portions of the tableware, and further associates the class data of each of those tableware items with its own individual tableware class data. Note that the bounding box is not the only way to indicate the coordinates of the image portion of the tableware object. For example, the coordinates of the center point and the length of the radius, i.e., a circle, can also be used to indicate the coordinates of the image portion of the tableware object.

[0047] Next, the first processing circuit 8 of the mobile terminal 4 inputs the image data acquired via the first interface device 6 into the trained first machine learning model and obtains classification data output from the trained first machine learning model (step S04).

[0048] Next, the first processing circuit 8 of the mobile terminal 4 determines whether or not the acquired classification data indicates that it contains an image of tableware, which is the object of data collection (step S06). If the classification data indicates that it contains an image of tableware (step S06 - YES), the first processing circuit 8 of the mobile terminal 4 transmits the image data acquired via the first interface device to the annotation-assigned learning data storage server 5 via the first interface device (step S08). The image data shown in Figure 2A and the image data shown in Figures 2B (B2-1) and (B2-2) are examples of images that indicate that the classification data contains an image of tableware.

[0049] On the other hand, if the classification data does not indicate that an image of tableware is included (step S06, NO), the first processing circuit 8 of the mobile terminal 4 deletes the image data acquired via the first interface device without transmitting it externally (step S10). This removes the malicious data. The image data shown in Figure 2B (B1) is an example image that indicates the classification data does not include an image of tableware.

[0050] The malicious data check process (first check process) and transmission process for one image data are now complete (step S12).

[0051] 2.2.2. [Operation of the Annotation-Adding Training Data Storage Server] In the annotation-based learning data storage system 2 according to this embodiment, the main processing of the annotation-based learning data storage server 5 is to perform detailed checks and remove image data with errors. For example, this involves removing image data with errors such as the bounding box being outside the image portion of the tableware, as shown in Figure 2B(B2-1), or image data with errors such as the labeling of "rice bowl" being incorrectly done as "chopsticks," as shown in Figure 2B(B2-2). To achieve this main processing, a trained second machine learning model is constructed in the annotation-based learning data storage server 5.

[0052] This second machine learning model outputs class data (the name of the tableware) and confidence data for the tableware in the input image data of the tableware. Therefore, the second machine learning model is trained using image data of various tableware and the class data (the name of the tableware) attached to each of those image data as training data. Since the second machine learning model recognizes the specific name of the tableware in the input image data, it is desirable to implement it using a CNN, for example.

[0053] The annotation-assigned training data storage server 5 is pre-equipped with a second, pre-trained machine learning model.

[0054] Figure 4 is a flowchart illustrating the misdata check process (second check process) and storage process in the annotation-assigned learning data storage server 5, which constitutes the annotation-assigned learning data storage system according to the embodiment. The operation of the annotation-assigned learning data storage server 5 will be described below with reference to the flowchart shown in Figure 4.

[0055] After the start (step S22), the second processing circuit 24 of the annotation-assigned learning data storage server 5 receives image data from the mobile terminal 4 via the second interface device 22 (step S24).

[0056] Next, the second processing circuit 24 of the annotation-assigned learning data storage server 5 acquires image data of one or more image portions specified by coordinate data, that is, by bounding boxes, from the received image data (step S26). If the received image data contains multiple image portions of tableware, multiple image portions will also be acquired.

[0057] Next, the second processing circuit 24 of the annotation-assigned learning data storage server 5 inputs each image data of the image portion into the trained second machine learning model, and for each image data of the image portion, it obtains one or more class data and confidence data corresponding to each of the said class data (step S28). The class data here are specific tableware names, for example, Rice-Bowl, Soup-Bowl, Fish-dish, Water-cup, Spoon, etc.

[0058] Next, the second processing circuit 24 of the annotation-assigned learning data storage server 5 determines whether the maximum confidence value data among the confidence value data corresponding to one or more class data acquired for each image data of the image portion is greater than a predetermined threshold (step S30).

[0059] If the maximum value of the confidence data is lower than a predetermined threshold, it indicates that the second machine learning model cannot definitively recognize what the tableware in the image portion is (i.e., what the class data is). In other words, in many cases, the maximum value of the confidence data will be lower than the predetermined threshold for image data with errors, such as the bounding box being outside the image portion of the tableware, as shown in Figure 2B (B2-1). Therefore, step S30 here is provided to remove image data that contains image portions with errors, such as the bounding box being outside the image portion of the tableware.

[0060] If the confidence data for the maximum value is greater than a predetermined threshold (step S30, YES), the second processing circuit 24 of the annotation-assigned learning data storage server 5 further determines whether the class data (i.e., the dish name) corresponding to the confidence data for the maximum value of the image portion matches the class data associated with the image portion of the image (step S32).

[0061] Step S32 here is provided to remove image data that has errors, such as the incorrect class data (label) attached to the image portion of the tableware, as shown in Figure 2B (B2-2).

[0062] If a single received image data contains multiple image portions of tableware, steps S28, S30, and S32 described above are performed for all image data of the multiple tableware portions. For all image data of the multiple tableware portions, if the confidence data of the maximum value is greater than a predetermined threshold (step S30 - YES), and the class data corresponding to the confidence data of the maximum value matches the class data associated with the image data of that portion (step S32 - YES), then the second processing circuit 24 of the annotation training data storage server 5 stores the image data received from the mobile terminal 4 in the storage device 12 as annotation training data (step S34).

[0063] Furthermore, the second processing circuit 24 of the annotation-assigned learning data storage server 5 deletes the image data received from the mobile terminal 4 without storing it in the storage device 12 if the maximum confidence data obtained for each image data portion of the tableware is below a predetermined threshold (step S30-NO). If a single received image data contains multiple image portions of tableware, the image data received from the mobile terminal 4 is deleted if the maximum confidence data for even one of those image portions is below a predetermined threshold (step S36).

[0064] Furthermore, the second processing circuit 24 of the annotation-assigned learning data storage server 5 deletes the image data received from the mobile terminal 4 without storing it in the storage device 12 if the class data corresponding to the maximum confidence score data for each image data portion of the tableware does not match the class data associated with the image data of that portion (step S32-NO). If a single received image data contains multiple image portions of tableware, the image data received from the mobile terminal 4 is deleted if the class data corresponding to the maximum confidence score data for even one of those image portions does not match the class data associated with the image data of that portion (step S36).

[0065] The misdata check process (second check process) and storage process for one image data are now complete (step S12).

[0066] Figure 5 shows an example of the misdata check process (second check process) in the annotation-assigned learning data storage server 5, which constitutes the annotation-assigned learning data storage system 2 according to the embodiment. In the example shown in Figure 5, a bounding box is added to the central portion of the image data of the tableware from the mobile terminal 4 by the operator of the mobile terminal 4, and further class data (label) "chopsticks" is added. The central portion of the image of the tableware is cropped (see Figure 4, step S26) and input into the trained second machine learning model, and the recognition function of the trained second machine learning model yields the output result that the tableware in that portion of the image is a "bowl" (see Figure 4, step S28).

[0067] Furthermore, in the processing example shown in Figure 5, the output result "bowl" is compared with the class data "chopsticks" assigned by the operator of the mobile terminal 4 (see Figure 4, step S32), and a mismatch is found. As a result, the image data from this mobile terminal 4 is discarded (deleted) (see Figure 4, step S36).

[0068] Figure 6 shows an example of image data output and stored by the annotation-attached learning data storage server 5, which constitutes the annotation-attached learning data storage system 2 according to this embodiment. Bounding boxes are attached to all the tableware in the image data, and specific class data is attached to each portion of the image of the tableware separated by the bounding boxes. The numbers to the right of each class data (label) are confidence score data, but these confidence score data are not required.

[0069] 2.3. [Summary of Embodiments] The annotation-based learning data storage system 2 according to this embodiment consists of a mobile terminal 4 and an annotation-based learning data storage server 5, which are connected via a predetermined network. The mobile terminal 4 includes a first interface device 6 and a first processing circuit 8. The annotation-based learning data storage server 5 includes a second interface device 22, a storage device 12, and a second processing circuit 24. The mobile terminal 4 has a trained first machine learning model built on it, which has been trained using image data and classification data indicating whether or not the image data contains an image of the object to be collected as training data. The annotation-based learning data storage server 5 has a trained second machine learning model built on it, which has been trained using image data and class data of the object to be collected contained in the image data as training data. The first processing circuit 8 of the mobile terminal 4 acquires image data via the first interface device 6, which includes an image portion of one or more objects to be collected, includes coordinate data of the image portion of the object to be collected, and further associates the class data of the object. The first processing circuit 8 of the mobile terminal 4 inputs the image data acquired via the first interface device 6 to the trained first machine learning model and acquires classification data output from the trained first machine learning model. If the acquired classification data indicates that it contains an image of the object to be collected, the first processing circuit 8 of the mobile terminal 4 transmits the image data acquired via the first interface device 6 to the annotation-assigned learning data storage server 5 via the first interface device 6. The second processing circuit 24 of the annotation-assigned learning data storage server 5 receives the image data from the mobile terminal 4 via the second interface device 22. The second processing circuit 24 of the annotation-assigned learning data storage server 5 acquires image data of one or more image portions specified by coordinate data from the received image data.The second processing circuit 24 of the annotation-assigned learning data storage server 5 inputs each image data of the image portion into the trained second machine learning model and obtains one or more class data and confidence data corresponding to each of the class data for each of the image data of the image portion. If the maximum confidence data among the confidence data corresponding to each of the one or more class data obtained for each of the image data of the image portion is greater than a predetermined threshold, and the class data corresponding to the maximum confidence data for each of the image data of the image portion matches the class data associated with the image data of the image portion, the second processing circuit 24 of the annotation-assigned learning data storage server 5 stores the image data from the mobile terminal 4 in the storage device 12 as annotation-assigned learning data.

[0070] By using the annotation-based learning data storage system 2 according to this embodiment, annotation work can be performed in a distributed and centralized manner, such as performing simple but important checks on a mobile device while performing complex and resource-intensive checks on a cloud server. Furthermore, it is possible to reduce working time and effort and improve the accuracy of annotations.

[0071] 3. [Other Embodiments] As described above, embodiments have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited thereto and can be applied to embodiments that are modified, replaced, added, or omitted as appropriate.

[0072] In the above-described embodiment, the second processing circuit 24 of the annotation-assigned learning data storage server 5 deletes the image data from the mobile terminal 4 if the maximum confidence value data among the confidence value data corresponding to one or more class data acquired for each image data of the image portion is below a predetermined threshold. For example, in this case, instead of deleting the image data from the mobile terminal 4, the second processing circuit 24 of the annotation-assigned learning data storage server 5 may be configured to send the image data back to the mobile terminal 4. Similarly, if the class data corresponding to the maximum confidence value data for each image data of the image portion does not match the class data associated with the image data of the image portion, the second processing circuit 24 of the annotation-assigned learning data storage server 5 may be configured to send the image data back to the mobile terminal 4 instead of deleting the image data from the mobile terminal 4. As a result, the operator of the mobile terminal 4 can perform correction work such as adding bounding boxes to the image data and adding class data.

[0073] Furthermore, attached drawings and a detailed description are provided to illustrate the embodiments. Therefore, among the components described in the attached drawings and detailed description, there may be not only components essential for solving the problem, but also components that are not essential for solving the problem, provided that they illustrate the above technology. For this reason, the mere presence of such non-essential components in the attached drawings and detailed description should not be immediately assumed to mean that those non-essential components are essential.

[0074] Furthermore, since the embodiments described above are for illustrative purposes of the technology described herein, various modifications, substitutions, additions, omissions, etc., can be made within the claims or their equivalents. [Explanation of symbols]

[0075] 2...Annotation-based learning data storage system, 3...Mobile annotation system, 4...Mobile terminal, 5...Annotation-based learning data storage server, 6...First interface device, 8...First processing circuit, 10...Memory, 12...Storage device, 14...Learning server, 16...Computer device, 18...External network, 22...Second interface device, 24...Second processing circuit.

Claims

1. In an annotation-based learning data storage system, The annotation-based learning data storage system is comprised of a mobile terminal and an annotation-based learning data storage server, both connected via a predetermined network. The aforementioned mobile terminal includes a first interface device and a first processing circuit, The annotation-assigning learning data storage server includes a second interface device, a storage device, and a second processing circuit. The aforementioned mobile device includes: A pre-trained first machine learning model has been constructed, which has been trained using image data and classification data indicating whether or not the image data contains an image of the object being collected as training data. The annotation-assigned learning data storage server includes: A pre-trained second machine learning model has been constructed, which was trained using image data and class data of the objects included in the image data as training data. The first processing circuit of the aforementioned mobile terminal is Image data is acquired via the first interface device, which includes an image portion of one or more data collection targets, includes coordinate data of the image portion of the data collection targets, and is associated with the class data of the targets. Image data acquired via the first interface device is input to a trained first machine learning model, and classification data output from the trained first machine learning model is acquired. If the acquired classification data indicates that it includes an image of the object to be collected, the image data acquired via the first interface device is further transmitted via the first interface device to the annotation-assigned learning data storage server. The second processing circuit of the annotation-assigned learning data storage server is: Image data from the aforementioned mobile terminal is received via the second interface device. From the received image data, obtain image data of one or more image portions specified by coordinate data. Each of the image data of the aforementioned image portion is input into a trained second machine learning model, and for each of the image data of the aforementioned image portion, one or more class data and confidence data corresponding to each of the said class data are obtained. Of the confidence data corresponding to one or more class data obtained for each of the image data of the aforementioned image portion, the maximum confidence data is greater than a predetermined threshold, and For each of the image data of the aforementioned image portion, if the class data corresponding to the confidence data of the maximum value matches the class data associated with the image data of the aforementioned image portion, Image data from the aforementioned mobile terminal is stored in the storage device as annotation-assigned training data. An annotation-based learning data storage system.

2. The aforementioned trained first machine learning model is a lightweight deep learning model. The annotation-assigned learning data storage system according to claim 1.

3. The first processing circuit of the aforementioned mobile terminal is If the acquired classification data does not indicate that it includes an image of the object being collected, the image data acquired via the first interface device will not be transmitted externally. The annotation-assigned learning data storage system according to claim 1.

4. The second processing circuit of the annotation-assigned learning data storage server is: If, among the confidence data corresponding to one or more class data obtained for each of the image data of the aforementioned image portion, the maximum confidence data is less than or equal to a predetermined threshold, If, for each of the image data of the aforementioned image portion, the class data corresponding to the confidence data of the maximum value does not match the class data associated with the image data of the aforementioned image portion, Delete the image data from the aforementioned mobile device. The annotation-assigned learning data storage system according to claim 1.

5. The object of data collection is an object identification target including one of the following: tableware, fruit, or human being, or an object inspection target including one of the following: metal, printed material, food, or fruit, and the object is an object inspection item including defects or stains. The annotation-assigned learning data storage system according to claim 1.

6. The first step of preparing a trained first machine learning model, which is trained using image data and classification data indicating whether or not the image data contains an image of the object to be collected, as training data by a first processing circuit constituting a mobile terminal, The second processing circuit, which constitutes the annotation-assigning learning data storage server, prepares a trained second machine learning model, which has been trained using image data and class data of the objects to be collected included in the image data as training data. The first processing circuit acquires image data in which the image data includes an image portion of one or more objects to be collected, the coordinate data of the image portion of the objects to be collected, and the class data of the objects to be collected is associated with the image data. The first processing circuit inputs the acquired image data to the trained first machine learning model and obtains classification data output from the trained first machine learning model. If the acquired classification data indicates that it includes an image of the object to be collected, the first processing circuit transmits the acquired image data to the annotation-assigned learning data storage server. The second processing circuit performs the steps of receiving image data from the mobile terminal, The second processing circuit performs the step of obtaining image data of one or more image portions specified by coordinate data from the received image data, The second processing circuit inputs each of the image data of the image portion into a trained second machine learning model, and for each of the image data of the image portion, it obtains one or more class data and confidence data corresponding to each of the class data. Of the confidence data corresponding to one or more class data obtained for each of the image data of the aforementioned image portion, the maximum confidence data is greater than a predetermined threshold, and For each of the image data of the aforementioned image portion, if the class data corresponding to the confidence data of the maximum value matches the class data associated with the image data of the aforementioned image portion, The second processing circuit performs the steps of storing the image data from the mobile terminal in a storage device as annotation-assigned learning data. A method for accumulating annotated training data, including [specific data].

7. The aforementioned trained first machine learning model is a lightweight deep learning model. The method according to claim 6.

8. Furthermore, If, among the confidence data corresponding to one or more class data obtained for each of the image data of the aforementioned image portion, the maximum confidence data is less than or equal to a predetermined threshold, If, for each of the image data of the aforementioned image portion, the class data corresponding to the confidence data of the maximum value does not match the class data associated with the image data of the aforementioned image portion, The second processing circuit performs the following steps: deleting image data from the mobile terminal including, The method according to claim 6.

9. The object of data collection is an object identification target including one of the following: tableware, fruit, or human being, or an object inspection target including one of the following: metal, printed material, food, or fruit, and the object is an object inspection item including defects or stains. The method according to claim 6.

Citation Information

Patent Citations

  • Trade information interchange device and method

    JP2018195328A

  • Machine learning method and device

    JP2019101740A

  • Learning data preparation device, learning model preparation system, learning data preparation method and program

    JP2019159499A

  • Annotation device, annotation method, and program

    JP2020126311A

  • Purchase and sales information exchange device and method

    JP2020144912A