Information processing apparatus and control method thereof
The information processing device automates the annotation of training data by using metadata to identify and label the main subject in images, addressing the heavy workload issue in existing systems.
Patent Information
- Application Number
- JP2024112651
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-01-23
AI Technical Summary
Existing systems require manual user intervention to determine the main subject in images, leading to a heavy workload in annotating training data for machine learning.
An information processing device that utilizes metadata such as AF information, AF frame coordinates, touch AF information, and defocus maps to automate the selection of the main subject in images, reducing the user's workload in creating training data.
Automates the annotation process by identifying and labeling the main subject in images, significantly reducing the burden on users in generating training data for machine learning.
Smart Images

Figure 2026011777000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technology for supporting annotation. [Background technology]
[0002] In recent years, the use of artificial intelligence (AI) has been progressing in various fields. One example is "supervised learning," which generates an inference model by performing machine learning based on training data containing correct labels. In supervised learning, high-quality training data is required to obtain a machine learning model with high generalization performance. Training data consists of various input data determined by the task to be solved and corresponding correct labels that have been appropriately assigned. In general, public datasets created for purposes such as competitions are often used as training data.
[0003] If there is no training data suitable for the purpose, it is necessary to create a dataset by oneself. However, creating training data, especially annotation such as assigning correct labels, requires specialized knowledge and is a heavy workload. Patent Document 1 discloses a device that annotates images using a machine learning model for image recognition. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 7,055,259 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the device disclosed in Patent Document 1 does not determine which object in an image is the main subject, which means that the user must determine the main subject in the image themselves before annotating it, which creates a heavy workload.
[0006] The present invention has been made in view of the above problems, and aims to provide a technique for supporting annotation. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems, an information processing device according to the present invention has the following configuration. That is, the information processing device for creating training data used in machine learning comprises: a first acquisition means for acquiring image data; a second acquisition means for acquiring metadata, which is information obtained when the image represented by the image data is captured; a detection means for detecting an object included in the image; an assigning means for selecting a specific object from the objects detected by the detecting means based on the metadata and assigning a label as the training data; Equipped with. [Effects of the Invention]
[0008] According to the present invention, a technique for supporting annotation can be provided. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating the overall configuration of a system. [Figure 2] FIG. 1 is a diagram illustrating a configuration of an imaging device. [Figure 3] FIG. 1 is a diagram illustrating a hardware configuration of an information processing device. [Figure 4] FIG. 2 is a diagram illustrating various types of metadata. [Figure 5] 10A and 10B are diagrams illustrating examples of various metadata for image data. [Figure 6] FIG. 2 is a diagram illustrating a functional configuration of an annotation program. [Figure 7] FIG. 10 is a diagram illustrating an example of object detection. [Figure 8] FIG. 10 is a diagram illustrating an example of training data. [Figure 9]1 is an overall flowchart of annotation processing. [Figure 10] FIG. 2 is a diagram illustrating the positional relationship between the AF frame and the BB. [Figure 11] 10 is a detailed flowchart of the process (S504). [Figure 12] 10 is a diagram illustrating the positional relationship between touch coordinates and BB. FIG. [Figure 13] 10 is a detailed flowchart of the process (S504). [Figure 14] FIG. 10 is a diagram illustrating a GUI for checking and correcting labels. [Figure 15] FIG. 10 is a diagram illustrating a defocus map. [Figure 16] 10 is a detailed flowchart of the process (S504). [Figure 17] FIG. 2 is a diagram illustrating an AF frame and a defocus map. [Figure 18] 10 is a detailed flowchart of the process (S504). [Figure 19] 10A and 10B are diagrams illustrating gaze coordinates and a defocus map. [Figure 20] 10 is a detailed flowchart of the process (S504). DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0011] (First embodiment) As a first embodiment of an information processing device according to the present invention, an information processing device that performs annotation on a captured image will be described below as an example.
[0012] <System configuration and device configuration> 1 is a diagram showing the overall configuration of the system. An imaging device 100 associates and records image data captured using the imaging device 100 with metadata of the image. An information processing device 200 acquires the image data and metadata and performs annotation.
[0013] FIG. 2 shows the configuration of the imaging device 100. The photographing lens 101 forms an optical image of a subject on the image sensor 103. The lens control unit 102 controls the photographing lens 101 for zoom control, focus control, aperture control, image stabilization, and other functions. The image sensor 103 is a CMOS image sensor in which pixels, each having two pairs of pupil-divided photoelectric conversion units, are arranged in a matrix, and converts the optical image of the subject into an electrical image signal. The signal processing unit 104 performs various corrections and compressions on the image signal output from the image sensor 103. The signal processing unit 104 temporarily holds the image signal output from the image sensor 103 and has an arithmetic and logic unit (ALU) for calculating a corrected image signal by performing differential operations and the like. The timing generation unit 105 outputs a timing signal for driving the image sensor 103.
[0014] The overall control unit 106 performs various calculations and controls the overall operation of the imaging device 100, including the operation of the image sensor 103. Based on signals acquired by the image sensor 103, the overall control unit 106 also performs focus detection using a phase difference detection method and calculates the defocus amount. The overall control unit 106 is generally configured with a central processing unit (CPU). However, some of the calculations can be delegated to an FPGA (Field Programming Gate Array). Alternatively, these can be combined into a single package called an application-specific integrated circuit (ASIC). The memory 107 is a work area where the overall control unit 106 temporarily stores image data output by the signal processing unit 104. The non-volatile memory 108 stores programs, various thresholds, adjustment values that differ for each imaging device 100, and the like.
[0015] The display unit 109 displays various information and captured images. The operation unit 110 includes a group of input devices such as switches, buttons, and a touch panel, and accepts user instructions for the imaging device 100. The display unit 109 may also include a viewfinder and a user's line of sight detector. In this case, the operation unit 110 can also accept the user's line of sight as a user instruction. The recording unit 111 records or reads image data and metadata associated with the image. The recording unit 111 is a circuit that reads and writes data from a removable recording medium such as a semiconductor memory. An external storage device may also serve a similar purpose. Here, the external storage device can be realized, for example, by a medium (recording medium) and an external storage drive for accessing the medium. Known examples of such media include a flexible disk (FD), CD-ROM, DVD, USB memory, MO, and flash memory. The external storage device may also be a server device connected via a network. The communication unit 112 can communicate with other devices via one or more networks, such as the Internet, a wide area network (WAN), or a local area network (LAN).
[0016] In this embodiment, the imaging device 100 acquires a defocus map, which is depth information of an image, based on a signal acquired by the imaging element 103. In this embodiment, the imaging device 100 also has an autofocus (AF) function as an automatic focus adjustment function of the photographing lens 101. In this embodiment, the imaging device 100 also has a touch panel as the operation unit 110, and is equipped with a touch shooting function that allows a user to operate the AF function by touching the touch panel. In this embodiment, the recording unit 111 associates image data captured using the imaging device 100 with metadata of the image, and records the associated data in a flash memory as an external storage device. Note that the specific configuration for acquiring depth information is not limited to the above-described embodiment. For example, the imaging device 100 may acquire depth information using a dedicated distance sensor such as LiDAR (Light Detection and Ranging, Laser Imaging Detection) or ToF (Time of Flight).
[0017] FIG. 3 is a diagram showing the hardware configuration of the information processing device 200. The information processing device 200 can be configured as a general-purpose computer (such as a PC). The CPU 201 executes various processes by executing programs. The GPU 202 is a graphics processing unit (GPU) that can perform efficient calculations by processing a larger amount of data in parallel. Therefore, it is effective to use the GPU 202 to perform processing when performing learning and inference multiple times using a machine learning model such as deep learning. Specifically, when executing a learning program including a machine learning model, the CPU 201 and GPU 202 cooperate to perform calculations, thereby performing learning. Furthermore, the GPU 202 may also be used when executing an inference program including a trained machine learning model, as in the case of learning.
[0018] A control program is stored in the read-only memory (ROM) 203. The random access memory (RAM) 204 is used as a temporary storage area such as the main memory or work area of the CPU 201. The HDD 205 is a hard disk for storing electronic data and programs according to this embodiment. An external storage device may also be used to perform a similar function. Here, the external storage device can be realized, for example, by a medium (recording medium) and an external storage drive for accessing the medium. Examples of such media include a flexible disk (FD), CD-ROM, DVD, USB memory, MO, and flash memory. The external storage device may also be a server device connected via a network. The input unit 206 is composed of a mouse, keyboard, touch panel, etc., and accepts input from the user. The display unit 207 is composed of an LCD display, etc., and can display various data and processing results to the user. The communication unit 208 can communicate with other devices via one or more of various networks, such as the Internet, a wide area network (WAN), and a local area network.
[0019] In this embodiment, the information processing device 200 acquires image data and metadata recorded by the imaging device 100 via flash memory and stores them in the HDD 205. Note that the specific configuration for acquiring data from the imaging device 100 is not limited to the above-described embodiment. For example, the communication unit 208 of the information processing device 200 may receive data by communicating with the communication unit 112 of the imaging device 100. Furthermore, both the imaging device 100 and the information processing device 200 may be connected to various networks, and data may be received via a server device or the like serving as an external storage device.
[0020] FIG. 4 is a diagram illustrating various types of metadata recorded by the imaging device 100. Here, the metadata is data related to information when an image is captured using the imaging device 100. Here, it is assumed that the metadata includes AF information, AF frame coordinates, touch AF information, touch coordinates, gaze information, gaze coordinates, and a defocus map. Note that the AF information and AF frame coordinates relate to the AF function of the imaging device 100. The touch AF information and touch coordinates relate to the touch shooting function of the imaging device 100. The gaze information and gaze coordinates relate to the user's gaze detection function in the imaging device 100. The defocus map is information related to depth measured by the imaging device 100.
[0021] 5 is a diagram showing examples of various types of metadata for image data recorded by the imaging device 100. In this embodiment, it is assumed that the metadata is embedded in the image data as Exif (Exchangeable image file format). Below, the correspondence between image data and metadata will be described in detail using two pieces of image data ("IMG_0010.JPG" and "IMG_0011.JPG" shown in FIG. 5) as an example.
[0022] IMG_0010.JPG: This is image data captured when the AF function of the imaging device 100 is enabled, and obtained when the user touches the operation unit 110 to specify the coordinates of the AF frame while shooting in live view. Therefore, the AF information and touch AF information are "TRUE" in the metadata, and the touch coordinates and AF frame coordinates are recorded. Furthermore, because the user performed live view shooting, the gaze detector in the viewfinder of the display unit 109 was unable to detect the user's gaze, so the gaze information is recorded as "FALSE" and the gaze coordinates are recorded as "N / A." Additionally, because depth measurement was enabled, the metadata also records a defocus map calculated by the overall control unit 106 based on the signal acquired by the image sensor 103.
[0023] IMG_0011.JPG: This is image data captured by a user adjusting the focus using manual focus (MF) while looking through the viewfinder (display unit 109) and then pressing the release button (operation unit 110). Therefore, the AF information and AF frame coordinates, which are metadata related to the AF function, and the touch AF information and touch coordinates, which are metadata related to the touch shooting function, are all recorded as "FALSE" or "N / A." The viewfinder's gaze detector detects the user's gaze, and the gaze information is recorded as "TRUE," while the gaze coordinates indicate the user's gaze position on the image. Even in MF mode, the defocus map can be calculated correctly based on the signal acquired by the image sensor 103, so the defocus map is recorded in the metadata.
[0024] In this way, image data and metadata, which is information about when the image data was captured, are recorded in a configuration in which they are associated with each other. Note that the specific configuration for associating image data with metadata is not limited to the above-mentioned form. For example, a unique identifier may be assigned to image data, table data may be created using the identifier as a primary key, and metadata may be recorded therein as individual records.
[0025] 6 is a diagram showing the functional configuration realized by the information processing device 200 executing the annotation program 300. The annotation program 300 is stored in the HDD 205. The CPU 201 acquires the annotation program 300 from the HDD 205 and loads it on the RAM 204. The annotation program 300 thus loaded is calculated by the CPU 201 and the GPU 202, whereby each process constituting the annotation program 300 is executed. Details of the processing of the annotation program 300 will be described later with reference to FIG. 9 etc.
[0026] FIG. 7 is a diagram illustrating an example of object detection by the object detection unit 303. The input image 400 is an image captured by the imaging device 100. The object detection unit 303 receives the input image 400 as input and outputs a detection result 450. In the detection result 450, a bounding box (BB) 451 and a class name 452 are assigned to each object detected by the object detection unit 303. The BB 451 is a rectangular area indicating the area in the input image 400 in which each object is captured. The class name 452 is text indicating the class into which the object has been classified in a predetermined classification determined by the object detection unit 303.
[0027] 8 is a diagram illustrating an example of training data output by the annotation unit 304. The annotation unit 304 selects one or more objects in the detection result 450 based on metadata, and outputs the BB 451 and class name 452 assigned to the selected object as labels, linking the input image 400 with the labels, as training data. However, if object selection based on metadata cannot be performed normally, the label is exceptionally set to "N / A."
[0028] In this embodiment, the annotation unit 304 selects one object (there may be multiple objects of the same type) as the main subject (specific object) for one piece of image data, and labels the BB 451 and class name 452 corresponding to the main subject. In particular, the annotation unit 304 selects one object to be the main subject from the detection result 450 (which generally includes one or more objects) based on the above-mentioned metadata. Hereinafter, an example will be described in which one object is selected as the main subject and a label is assigned, but labels can also be assigned to multiple objects that satisfy a predetermined threshold condition.
[0029] <Device Operation> 9 is a flowchart of the annotation process. The annotation process is realized by the CPU 201 expanding the annotation program 300 stored in the HDD 205 onto the RAM 204 and performing calculations. The following steps are executed by the CPU 201 and the GPU 202 performing calculations.
[0030] In step S501, the image data acquisition unit 301 acquires image data stored in the HDD 205 and expands it in the RAM 204. In step S502, the metadata acquisition unit 302 acquires image data stored in the HDD 205 and expands it in the RAM 204. In this embodiment, since the metadata is embedded in the image data as Exif, S501 and S502 are executed in parallel.
[0031] In step S503, the object detection unit 303 receives the image data expanded in the RAM 204 as the input image 400, detects objects captured in the image, and outputs the detection results 450 to the RAM 204. In this embodiment, the object detection unit 303 assigns one BB 451 and one class name 452 to each object.
[0032] In this embodiment, the object detection unit 303 uses a single shot multibox detector (SSD) described in the following document A. The SSD is a known machine learning model capable of object detection, and is suitable for this embodiment in that it can detect objects of relatively small size. Literature A: W. Liu et al., "SSD: Single Shot MultiBox Detector", ECCV 2016
[0033] In step S504, the annotation unit 304 receives as input the image data, metadata, and detection result 450 expanded in the RAM 204. Then, for one or more objects detected in the detection result, the annotation unit 304 assigns a label to an object selected based on the metadata, and outputs the label to the HDD 205 as training data.
[0034] By performing the above steps, the processing of the annotation program 300 is completed. Note that the specific configuration for detecting an object appearing in the input image 400 in S503 is not limited to the above-described configuration. For example, a known machine learning model capable of object detection, such as Faster R-CNN (Region-based Convolutional Neural network) or YOLO (You Only Look Once), may be used. A rule-based algorithm based on image binarization, edge extraction, or the like may also be used. In addition, the detection results of multiple machine learning models may be used in an integrated manner. Furthermore, if the detection result 450 also includes an index such as the likelihood of detection or the reliability of each BB assigned to the object, and the index is below a predetermined threshold, an exception process such as that shown in S709, which will be described later, may be executed.
[0035] <Details of processing (S504)> The following describes in detail the process of selecting and annotating a main subject based on image data captured by a user and metadata such as AF information and AF frame coordinates. In particular, the process of selecting a main subject (main subject selection process) will be described in detail.
[0036] FIG. 10 is a diagram illustrating the positional relationship between an AF frame 600 and a BB 451 near the center of the detection result 450. The AF frame 600 shown in FIG. 10(a) is an area on an image designated by a user for the AF function of the imaging device 100. The AF frame 600 is information calculated from the AF frame coordinates, which are metadata. The center-to-center distance A 601 shown in FIG. 10(c) is the distance on the image between the center coordinates of the BB of object A and the center coordinates of the AF frame 600. Similarly, the center-to-center distance B 602 is the distance on the image between the center coordinates of the BB of object B and the center coordinates of the AF frame 600. The distance threshold 603 shown in FIG. 10(e) is the radius of a circle centered on the center coordinates of the AF frame 600.
[0037] 10(a) to 10(d), the AF frame 600 overlaps with the BB of one of the objects. Specifically, in FIG. 10(a), the AF frame 600 is contained only in the BB of object B. In FIG. 10(b), the AF frame 600 is contained in the BB of object B and also partially overlaps with the BB of object A. In FIG. 10(c), the AF frame 600 partially overlaps with the BB of both object A and object B. In FIG. 10(d), the AF frame 600 is contained in the BB of both object A and object B.
[0038] 10(e) and 10(f), the AF frame 600 does not overlap the BB of either object. Specifically, in FIG. 10(e), the center coordinates of the BB of both object A and object B are located inside a circle having a radius equal to the distance threshold 603. In FIG. 10(f), the center coordinates of the BB of both object A and object B are located outside the circle having a radius equal to the distance threshold 603.
[0039] 11 is a detailed flowchart of the process (S504) in the first embodiment. In the main subject selection process in the first embodiment, the main subject is selected based on the AF information and the AF frame coordinates.
[0040] In step S701, the annotation unit 304 determines how many objects are detected in the detection result 450. If one or more objects are detected, the process proceeds to S702. Otherwise (if no objects are detected), the process proceeds to S709.
[0041] In step S702, the annotation unit 304 determines where the AF frame 600 is located in the detection result 450. In this embodiment, as illustrated in Fig. 5, the AF frame coordinates indicate the coordinates on the image of the upper left and lower right of the rectangular AF frame 600. Therefore, it is possible to determine where the AF frame 600 is located in the detection result 450 based on the AF frame coordinates.
[0042] In step S703, the annotation unit 304 determines whether there are one or more objects where the AF frame 600 and the BB overlap, based on the position of the AF frame 600 determined in S702 and the BB of each object detected in the detection result 450. If there are one or more overlapping objects, the process proceeds to S704. If not, the process proceeds to S707.
[0043] In step S704, the annotation unit 304 again determines the number of overlapping objects calculated in S703. If there are two or more overlapping objects, the process proceeds to S705 to select a main subject from among them. If not (if there is only one object where the AF frame 600 and BB overlap), the process selects that one object as the main subject and proceeds to S706.
[0044] In step S705, the annotation unit 304 selects a main subject from the detected objects based on a predetermined index.
[0045] The main subject is the subject that the user was trying to photograph when taking a photograph. Generally, when photographing with the AF function of the imaging device 100 enabled, it is considered that the main subject is often photographed with the AF frame 600 overlapping it. Therefore, in this embodiment, the area (intersection) of the overlapping region between the AF frame 600 and the BB of each object is used as the first index for selecting the main subject. Specifically, the object with the largest intersection is selected. For example, in FIGS. 10(a) to (c), the intersection with object B is larger than the intersection with object A, so object B is selected as the main subject.
[0046] On the other hand, in FIG. 10(d), the AF frame 600 is contained within the BB of both object A and object B, resulting in the same intersection value. Also, in FIG. 10(e), the AF frame 600 does not overlap with either object, so intersection cannot be calculated. Such an image is likely to have been captured when another object approaches the main subject while the user is overlaying the AF frame 600 on the main subject, or when the AF frame 600 fails to completely overlay the moving main subject. Therefore, in this embodiment, the center-to-center distance, which is the distance on the image between the center coordinates of the object BB and the center coordinates of the AF frame 600, is used as a second index for selecting the main subject, and the object with the shortest center-to-center distance is selected. For example, in FIG. 10(d), center-to-center distance A601 is shorter than center-to-center distance B602, so object A is selected as the main subject. Also, in FIG. 10(e), center-to-center distance B602 is shorter than center-to-center distance A601, so object B is selected as the main subject.
[0047] In addition, if there is no object whose center coordinates of BB are within the circle indicated by the distance threshold 603 as shown in Figure 10(f), the process will proceed to S709 based on the judgment of S707 described below, and the processing of S705 will not be performed.
[0048] In step S706, the annotation unit 304 assigns the BB 451 and class name 452 assigned to the object selected as the main subject as labels to the input image 400, and outputs the result as training data.
[0049] In step S707, the annotation unit 304 determines the number of objects whose center coordinates of BB exist within the circle indicated by the distance threshold 603. If one or more objects exist within the circle, the process proceeds to S708. If not (if no objects exist around the AF frame 600), the process proceeds to S709.
[0050] In step S708, the annotation unit 304 again determines the number of objects within the circle calculated in step S707. If there are two or more objects within the circle, the process proceeds to step S705 to select a main subject from among them. If not (if there is only one object whose center coordinates are BB within the circle indicated by the distance threshold 603), the object is selected as the main subject, and the process proceeds to step S706.
[0051] In step S709, the annotation unit 304 performs exception processing when an object to be selected as the main subject is not found. Specifically, "N / A" is assigned to both the BB 451 and the class name 452 of the input image 400, and the input image 400 is output as training data. Note that in this embodiment, if an image is captured with the AF function of the imaging device 100 disabled, the metadata related to the function is also disabled and annotation cannot be performed, so exception processing (S709) is performed.
[0052] Note that the index for selecting the main subject in S705 is not limited to the above-described form. For example, it is also possible to select an object corresponding to the largest or smallest BB among objects positioned to overlap with the AF frame 600. It is also possible to select an object with the largest IoU (Intersection over Union) based on the AF frame 600 and the BB corresponding to the object.
[0053] As described above, according to the first embodiment, a main subject included in an image is identified and annotation is performed using metadata. In particular, the main subject is selected and annotation is performed based on the AF information and AF frame coordinates, which are metadata. This enables the automation / semi-automation of annotation, significantly reducing the burden on the user in creating training data for machine learning.
[0054] (Variation 1) In the first embodiment, the AF information and AF frame coordinates, which are metadata, are used to select the main subject. However, other metadata may also be used to select the main subject. In Modification 1, an example is shown in which the main subject is selected based on touch AF information and touch coordinates.
[0055] 12 is a diagram illustrating the positional relationship between touch coordinates 800 and BB 451 near the center of the detection result 450. The touch coordinates 800 are coordinates on an image specified by a user touching the operation unit 110 for the AF function in the imaging device 100. The center-to-center distance A 801 is the distance on the image between the center coordinates of BB of object A and the touch coordinates 800. Similarly, the center-to-center distance B 802 is the distance on the image between the center coordinates of BB of object B and the touch coordinates 800. The distance threshold 803 is the radius of a circle centered on the touch coordinates 800.
[0056] 12(a) and 12(b), the touch coordinate 800 overlaps with the BB of one of the objects. Specifically, in FIG. 12(a), the touch coordinate 800 is included only in the BB of the object B. In FIG. 12(b), the touch coordinate 800 is included in the BB of both the object A and the object B.
[0057] 12(c) and 12(d), the touch coordinates 800 do not overlap with the BB of either object. Specifically, in FIG. 12(c), the center coordinates of the BB of both object A and object B exist within a circle having a radius of the distance threshold 803. In FIG. 12(d), the center coordinates of the BB of both object A and object B exist outside the circle having a radius of the distance threshold 803.
[0058] 13 is a detailed flowchart of the process (S504) in Modification 1. In the main subject selection process in Modification 1, the main subject is selected based on the touch AF information and touch coordinates. Note that S901, S905, and S906 are the same processes as S701, S706, and S709 in the first embodiment (FIG. 11), and therefore descriptions thereof will be omitted.
[0059] In step S902, the annotation unit 304 determines whether there are one or more objects where the touch coordinates 800 and BB overlap, based on the position of the touch coordinates 800 and the BB of each object detected in the detection result 450. If there are one or more overlapping objects, the process proceeds to S903. If not, the process proceeds to S906.
[0060] In step S903, the annotation unit 304 again determines the number of overlapping objects calculated in S902. If there are two or more overlapping objects, the process proceeds to S904 to select a main subject from among them. Otherwise (if there is only one object), the process selects that one object as the main subject and proceeds to S905.
[0061] In step S904, the annotation unit 304 selects a main subject from the detected plurality of objects based on a predetermined index. In this modification, the inclusion relationship on the image between the object BB and the touch coordinate 800 (the touch position in the image) is used as the first index for selecting the main subject, and only one object that includes the touch coordinate 800 is selected. For example, in FIG. 12(a), only the BB of the object B includes the touch coordinate 800, so the object B is selected as the main subject.
[0062] On the other hand, in FIG. 12(b), both object A and object B's BB contain touch coordinate 800. Furthermore, in FIGS. 12(c) and 12(d), neither object contains touch coordinate 800. Such images are considered to be images in which multiple objects are crowded around the main subject, or in which the user's touch instruction position is slightly shifted from the main subject. Therefore, in this modification, the center-to-center distance, which is the distance on the image between the center coordinate of object BB and touch coordinate 800, is used as a second index for selecting the main subject, and the object with the smallest center-to-center distance is selected. For example, in FIGS. 12(b) and 12(c), center-to-center distance A 801 is shorter than center-to-center distance B 802, so object A is selected as the main subject.
[0063] In addition, if there is no object whose center coordinates are BB within the circle indicated by the distance threshold 803 as shown in Figure 12 (d), the process proceeds to S906 based on the judgment of S902, and the processing of S904 is not performed.
[0064] In this embodiment, if an image is captured with the touch capture function of the imaging device 100 disabled, the metadata related to the function is also disabled and annotation cannot be performed, so exception processing (S906) is performed. Note that the index for selecting the main subject in S904 is not limited to the above-described form. For example, the object corresponding to the largest or smallest BB among the objects positioned overlapping the touch coordinate 800 may be selected.
[0065] Furthermore, even when the user's gaze information is used as other metadata, it is possible to select the main subject and perform annotation using the same process as in Modification 1. This is because the user's gaze information is also information indicating coordinates on an image, and has the same characteristics as touch coordinates 800 in that it is metadata that tends to be located around the main subject on an image.
[0066] (Variation 2) In the second modification, an example of a user interface is shown that allows the user to confirm and modify the label after the label has been assigned to the selected object based on the metadata in S504.
[0067] FIG. 14 is a diagram illustrating a graphical user interface (GUI) that allows a user to check and correct the labels added by the annotation program 300. As shown in FIG.
[0068] Window 1000 is a window that allows the user to check and correct the labels assigned by the annotation program 300, and is displayed on the display unit 207. The user can check the labels by looking at window 1000. Furthermore, as will be described below, label corrections from the user can be accepted via various GUI components arranged in window 1000.
[0069] The image display unit 1001 is a GUI component that displays the input image 400 and its corresponding label. In the initial state, the image display unit 1001 displays only the BB and class name corresponding to the object selected by the annotation program 300 as labels, and does not display the BB and class names corresponding to other objects. The class name input unit 1002 is a GUI component that receives the class name of the main subject corresponding to the image data being referenced. In the initial state, the class name corresponding to the object selected by the annotation program 300 is input to the class name input unit 1002.
[0070] The BB start point input unit 1003 and the BB end point input unit 1004 are GUI components for inputting the coordinates on the image of the BB of the main subject corresponding to the image data being referenced. In the initial state, the coordinates of the BB corresponding to the object selected by the annotation program 300 are input to the BB start point input unit 1003 and the BB end point input unit 1004. In this modified example, the upper left vertex of the rectangular BB is set as the BB start point, and the lower right vertex is set as the BB end point. Furthermore, the upper left vertex of the image data is set as the origin, and the coordinates of a position x pixels to the right and y pixels downward are expressed as x, y.
[0071] For example, the value of the BB start point input section 1003 in FIG. 14(a) indicates that the BB start point is located 3000 pixels to the right and 2000 pixels below the top left vertex of the image data. The edit button 1005 is a GUI component that the user presses (touches / clicks) when modifying the input class name or BB. The processing when the user presses the edit button 1005 will be described later. The image file name display section 1006 is a GUI component that displays the file name of the image data being referenced. The backward image transition button 1007 and the forward image transition button 1008 are GUI components for changing the image data being referenced.
[0072] FIG. 14(b) is a diagram showing the window 1000 when the user presses the Modify button 1005 in FIG. 14(a). Pressing the Modify button 1005 enters the user modification state, allowing the user to modify the label. In addition, in the user modification state, the image display section 1001 displays the BBs and class names of all objects based on the detection results 450. At this time, the user can modify the label by selecting another object displayed in the image display section 1001 or by entering a character string and a numerical value in the class name input section 1002, BB start point input section 1003, and BB end point input section 1004. After modifying the label, the user can confirm the modification by pressing the Modify button 1005 again.
[0073] (Second embodiment) In the second embodiment, a description will be given of a mode in which a main subject is selected and annotated based on image data captured by a user and defocus information, which is metadata. Note that the system configuration, device configuration, various data, and overall processing are the same as those in the first embodiment (FIGS. 1 to 6 and 9), and therefore descriptions thereof will be omitted.
[0074] <Details of processing (S504)> FIG. 15 is a diagram illustrating a defocus map. FIG. 15(a) shows a defocus map 1100, in which the defocus amount of each region of the image is associated with a region on the reference image data. For example, in this embodiment, the image data is divided into nine regions horizontally and eight regions vertically, and the defocus amount is calculated for each region. Also, in FIG. 15, regions where the defocus amount falls within a range of ±0.5 (the absolute value of the defocus amount is equal to or less than a predetermined value) are displayed with shading. Of these shaded regions, the region with the largest area when adjacent regions are combined is defined as the in-focus region 1101. FIG. 15(b) shows regions where the defocus amount falls within a range of ±0.5 superimposed on the image represented by the image data.
[0075] In this embodiment, the defocus amount is a value indicating how far a region on an image is displaced from the focus position at the time of image capture. Specifically, a defocus amount of zero indicates that the region is located at the same position as the focus position at the time of image capture. Furthermore, a positive defocus amount indicates that the region is located closer to the photographer than the focus position at the time of image capture. Similarly, a negative defocus amount indicates that the region on an image is located farther from the photographer than the focus position at the time of image capture. The larger the absolute value of the defocus amount, the farther the region is from the focus position (closer or farther from the photographer).
[0076] The method for determining the in-focus area 1101 is not limited to the above-described method, as long as it satisfies the condition that it is possible to determine in which area of the image the subject intentionally focused on by the user is located. For example, among areas where the defocus amount falls within the range of ±0.5, the area where the average defocus amount is closest to zero when adjacent areas are combined may be determined as the in-focus area 1101.
[0077] 16 is a detailed flowchart of the process (S504) in the second embodiment. In the main subject selection process in the second embodiment, the main subject is selected based on defocus information. Note that S1201, S1207, and S1208 are the same processes as S701, S706, and S709 in the first embodiment (FIG. 11), and therefore descriptions thereof will be omitted.
[0078] In step S1202, the annotation unit 304 determines where the in-focus region 1101 is located in the detection result 450. In this embodiment, as described above, the in-focus region 1101 is determined to be the largest continuous region among regions where the defocus amount falls within the range of ±0.5.
[0079] In step S1203, the annotation unit 304 determines whether or not the in-focus region 1101 exists (whether or not the in-focus region 1101 was found in S1202). If the in-focus region 1101 exists, the process proceeds to S1204. If not, the process proceeds to S1208, since it means that no region on the image is in focus.
[0080] In step S1204, the annotation unit 304 determines whether there are one or more objects where the focused area 1101 and the BB overlap, based on the position of the focused area 1101 and the BB of each object detected in the detection result 450. If there are one or more overlapping objects, the process proceeds to S1205. If not, the process proceeds to S1208.
[0081] In step S1205, the annotation unit 304 again determines the number of overlapping objects calculated in S1204. If there are two or more overlapping objects, the process proceeds to S1206 to select a main subject from among them. Otherwise (if there is only one object), the process selects that one object as the main subject and proceeds to S1207.
[0082] In step S1206, annotation unit 304 selects a main subject from the detected objects based on a predetermined index. In this embodiment, the area (intersection) of the overlapping region between in-focus region 1101 and BB of each object is used as the first index for selecting the main subject. Specifically, the object with the largest intersection is selected.
[0083] In this embodiment, if an image is captured while the depth measurement function of the image capture device 100 is disabled, the metadata related to the function is also disabled and annotation cannot be performed, so exceptional processing is performed. Furthermore, in S1206, the index for selecting the main subject is not limited to the above-described form. For example, the barycentric coordinates of the focus area 1101 may be calculated, and an object with the smallest distance between the barycentric coordinates and the central coordinates of BB may be selected.
[0084] As described above, according to the second embodiment, a main subject included in an image is identified and annotation is performed using metadata. In particular, the main subject is selected and annotation is performed based on defocus information, which is metadata. This enables automation / semi-automation of annotation, significantly reducing the burden on the user in creating training data for machine learning.
[0085] (Third embodiment) In the third embodiment, a description will be given of a mode in which a main subject is selected and annotated based on image data captured by a user and metadata such as defocus information, AF information, and AF frame coordinates. Note that the system configuration, device configuration, various data, and overall processing are the same as those in the first embodiment (FIGS. 1 to 6 and 9), and therefore a description thereof will be omitted.
[0086] <Details of processing (S504)> As shown in Fig. 5, the metadata recorded differs depending on the shooting conditions. Also, depending on the shooting conditions, using multiple pieces of metadata can enable more accurate selection of the main subject. Therefore, in the third embodiment, a form will be described in which both the metadata used in the first embodiment (AF information and AF frame coordinates) and the metadata used in the second embodiment (defocus information) are used.
[0087] FIG. 17 is a diagram illustrating an AF frame and a defocus map. In the defocus map shown in the upper part of FIG. 17(a), the in-focus region 1101 is located in the upper left of the image. On the other hand, the AF frame 600 shown in the lower part of FIG. 17(a) is located at the center of the image. Furthermore, the defocus amount of the region overlapping with the AF frame 600 falls within the range of ±0.5. In the defocus map shown in the upper part of FIG. 17(b), the in-focus region 1101 is located on the left side of the image. On the other hand, the AF frame 600 shown in the lower part of FIG. 17(b) is located at the center of the image. Furthermore, the defocus amount of the region overlapping with the AF frame 600 does not fall within the range of ±0.5.
[0088] 18 is a detailed flowchart of the process (S504) in the third embodiment. In the main subject selection process in the third embodiment, the main subject is selected based on defocus information, AF information, and AF frame coordinates. Note that S1301 and S1306 are the same processes as S701 and S709 in the first embodiment (FIG. 11), and therefore a description thereof will be omitted.
[0089] In step S1302, the annotation unit 304 determines where the AF frame 600 and the in-focus region 1101 are located in the detection result 450. This process is implemented by the same process as S702 in the first embodiment and S1202 in the second embodiment.
[0090] In step S1303, the annotation unit 304 determines the defocus amount at the position overlapping with the AF frame 600. If the defocus amount at the position overlapping with the AF frame 600 falls within a range of ±0.5, it is considered that the user captured the image with the AF frame 600 overlapping the main subject, and the process proceeds to S1304. On the other hand, if the defocus amount at the position overlapping with the AF frame 600 does not fall within a predetermined range near zero, it is considered that the user focused on the main subject, then changed the angle of view, and then captured the image, and the process proceeds to S1305.
[0091] In step S1304, the annotation unit 304 performs annotation based on the AF frame. Specifically, the same processes as S703 to S709 in the first embodiment are executed. On the other hand, in step S1305, the annotation unit 304 performs annotation based on the defocus amount. Specifically, the same processes as S1203 to S1208 in the second embodiment are executed.
[0092] For example, in FIG. 17(a), if the main subject is selected based only on the defocus information, a person overlapping the in-focus area 1101 is selected. However, considering the amount of defocus in the periphery and the depth relationship in the image, the in-focus area 1101 in FIG. 17(a) is located farther away from the focus point, and therefore the correct value is thought to be approximately -5 to -20. According to the method of this embodiment, the amount of defocus at the position overlapping with the AF frame 600 is within the range of ±0.5, so the person in the center of the screen is selected as the main subject. In this way, even if part of the defocus map 1100 is not calculated correctly due to the influence of noise caused by the subject contrast or the image sensor 103, the main subject can be correctly selected.
[0093] Furthermore, in FIG. 17(b), if the main subject is selected based solely on the AF frame 600, the car on the right side of the image is selected, or exception processing is performed depending on the setting of the distance threshold 603. However, the defocus amount at the position overlapping with the AF frame 600 in FIG. 17(b) is not within the range of ±0.5, and the same is true for the defocus amount at the position overlapping with the car. Based on the method of this embodiment, the person overlapping with the in-focus area 1101 in FIG. 17(b) is selected as the main subject. In this way, even if the user focuses on the main subject using the AF frame 600 and then changes the angle of view before capturing an image, the main subject can be correctly selected.
[0094] As described above, according to the third embodiment, the main subject is selected based on the defocus information, AF information, and AF frame coordinates, which allows for more accurate selection of the main subject than the first and second embodiments (which use less metadata).
[0095] (Variation 3) In Modification 3, a form will be described in which line-of-sight information and line-of-sight coordinates are used instead of AF information and AF frame coordinates. That is, an example will be shown in which a main subject is selected based on defocus information, line-of-sight information, and line-of-sight coordinates.
[0096] FIG. 19 is a diagram illustrating gaze coordinates and a defocus map. Gaze coordinates 1401 are the user's gaze position at the time of image capture, acquired by a user gaze detector provided in the display unit 109. In FIG. 19(a), the in-focus area 1101 is located in the upper left of the image, while the gaze coordinates 1401 are located in the center of the image. Furthermore, the defocus amount of the area overlapping with the in-focus area 1101 falls within the range of ±0.5. In FIG. 19(b), the in-focus area 1101 is located on the left side of the image, while the gaze coordinates 1401 are located in the center of the image.
[0097] 20 is a detailed flowchart of the process (S504) in Modification 3. In the main subject selection process in Modification 3, the main subject is selected based on defocus information and user line-of-sight information. Note that S1501 and S1506 are the same processes as S701 and S709 in the first embodiment (FIG. 11), and therefore a description thereof will be omitted.
[0098] In step S1502, the annotation unit 304 determines where the in-focus region 1101 is located in the detection result 450. This process is realized by the same process as S1202 in the second embodiment.
[0099] In step S1503, the annotation unit 304 determines the amount of defocus at the position overlapping with the gaze coordinates 1401. If the amount of defocus at the position overlapping with the gaze coordinates 1401 is within the range of ±0.5, it is considered that the user captured the image with their gaze focused on the main subject, and the process proceeds to S1504. On the other hand, if the amount of defocus at the position overlapping with the gaze coordinates 1401 is not within a predetermined range near zero, there is a possibility of erroneous measurement by the user gaze detector, and the process proceeds to S1505.
[0100] In step S1504, the annotation unit 304 performs annotation based on gaze coordinates. Specifically, it executes the same processes as S902 to S906 in Modification 1 (except that "gaze coordinates" are used instead of "touch coordinates"). On the other hand, in step S1505, the annotation unit 304 performs annotation based on the defocus amount. Specifically, it executes the same processes as S1203 to S1208 in the second embodiment.
[0101] The disclosure of this specification includes the following information processing device, control method, and program. (Item 1) An information processing device that creates teacher data used in machine learning, a first acquisition means for acquiring image data; a second acquisition means for acquiring metadata, which is information obtained when the image represented by the image data is captured; a detection means for detecting an object included in the image; an assigning means for selecting a specific object from the objects detected by the detecting means based on the metadata and assigning a label as the training data; An information processing device comprising: (Item 2) The image and the label are displayed on a graphical user interface (GUI), and a receiving unit is further provided for receiving a correction to the label. 2. The information processing device according to item 1, (Item 3) the metadata includes information about an AF frame of an autofocus (AF) function of the imaging device when the imaging device captured the image, The assigning means selects the one object based on the AF frame. 3. The information processing device according to item 1 or 2. (Item 4) The detection means outputs the bounding (BB) and label of the detected object, The assigning means selects one object based on the area of an overlapping region between the AF frame and the BB of each detected object, and assigns a label corresponding to the selected one object. 4. The information processing device according to item 3, (Item 5) The assigning means selects the one object further based on the distance between the center coordinates of the AF frame and the center coordinates of the BB of each object. 5. The information processing device according to item 4. (Item 6) the metadata includes information about a touch position of a user obtained by a touch shooting function of the imaging device when the imaging device captured the image, The assigning means selects the one object based on the touch position. 3. The information processing device according to item 1 or 2. (Item 7) The detection means outputs the BB and label of the detected object, The assigning means selects an object corresponding to a BB that includes the touch position as the one object, and assigns a label corresponding to the one selected object. 7. The information processing device according to item 6, (Item 8) the metadata includes depth information corresponding to the image; The assigning means selects the one object based on the depth information. 3. The information processing device according to item 1 or 2. (Item 9) the depth information is information about a defocus amount obtained when an imaging device captures the image, The assigning means selects an object that exists in an area where the absolute value of the defocus amount is equal to or less than a predetermined value as the one object. 9. The information processing device according to item 8, (Item 10) the metadata includes user's line of sight information obtained by a line of sight detection function of the imaging device when the imaging device captured the image, The assigning means selects the one object based on the line-of-sight information. 3. The information processing device according to item 1 or 2. (Item 11) A control method for an information processing device that creates teacher data used in machine learning, comprising: a first acquisition step of acquiring image data; a second acquisition step of acquiring metadata, which is information obtained when the image represented by the image data is captured; a detection step of detecting an object included in the image; an assignment step of selecting a specific object from the objects detected by the detection step based on the metadata and assigning a label as the training data; A control method comprising: (Item 12) Item 12. A program for causing a computer to execute the control method according to Item 11.
[0102] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0103] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0104] 200 Information processing device; 301 Image data acquisition unit; 302 Metadata acquisition unit; 303 Object detection unit; 304 Annotation unit
Claims
1. An information processing device that creates teacher data used in machine learning, a first acquisition means for acquiring image data; a second acquisition means for acquiring metadata, which is information obtained when the image represented by the image data is captured; a detection means for detecting an object included in the image; an assigning means for selecting a specific object from the objects detected by the detecting means based on the metadata and assigning a label as the training data; An information processing device comprising:
2. The image and the label are displayed on a graphical user interface (GUI), and a receiving unit is further provided for receiving a correction of the label.
2. The information processing apparatus according to claim 1, wherein:
3. the metadata includes information about an AF frame of an autofocus (AF) function of the imaging device when the imaging device captured the image, The assigning means selects the one object based on the AF frame.
2. The information processing apparatus according to claim 1, wherein:
4. The detection means outputs the bounding (BB) and label of the detected object, The assigning means selects the one object based on the area of an overlapping region between the AF frame and the BB of each detected object, and assigns a label corresponding to the one selected object.
4. The information processing apparatus according to claim 3,
5. The assigning means selects the one object further based on the distance between the center coordinate of the AF frame and the center coordinate of BB of each object.
5. The information processing apparatus according to claim 4,
6. the metadata includes information about a touch position of a user obtained by a touch shooting function of the imaging device when the imaging device captured the image, The assigning means selects the one object based on the touch position.
2. The information processing apparatus according to claim 1, wherein:
7. The detection means outputs the BB and label of the detected object, The assigning means selects an object corresponding to the BB that includes the touch position as the one object, and assigns a label corresponding to the one selected object.
7. The information processing apparatus according to claim 6,
8. the metadata includes depth information corresponding to the image; The assigning means selects the one object based on the depth information.
2. The information processing apparatus according to claim 1, wherein:
9. the depth information is information about a defocus amount obtained when an imaging device captures the image, The assigning means selects an object that exists in an area where the absolute value of the defocus amount is equal to or less than a predetermined value as the one object.
9. The information processing apparatus according to claim 8,
10. the metadata includes user's line of sight information obtained by a line of sight detection function of the imaging device when the imaging device captured the image, The assigning means selects the one object based on the line-of-sight information.
2. The information processing apparatus according to claim 1, wherein:
11. A control method for an information processing device that creates teacher data used in machine learning, comprising: a first acquisition step of acquiring image data; a second acquisition step of acquiring metadata, which is information obtained when the image represented by the image data is captured; a detection step of detecting an object included in the image; an assignment step of selecting a specific object from the objects detected by the detection step based on the metadata and assigning a label as the training data; A control method comprising:
12. A program for causing a computer to execute the control method according to claim 11.
Citation Information
Patent Citations
Labeling device and learning device
JP7055259B2