Image capturing apparatus, method for controlling image capturing apparatus, and storage medium

The imaging device addresses the shortage of training data and annotators by displaying annotation information with live view images, enabling efficient image collection and annotation data generation.

JP2026013085APending Publication Date: 2026-01-28CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024113265
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2026-01-28

AI Technical Summary

Technical Problem

The shortage of training data and annotators for AI image recognition poses a significant challenge, necessitating a reduction in the workload associated with acquiring and annotating images.

Method used

An imaging device that receives annotation information, displays it with a live view image, allows user selection, and generates annotation data by associating selected information with captured images, thereby facilitating efficient image collection.

Benefits of technology

Reduces the workload associated with acquiring and annotating images, addressing the shortage of training data and annotators for AI image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026013085000001_ABST
    Figure 2026013085000001_ABST
Patent Text Reader

Abstract

It is possible to provide a technique of reducing the shortage of annotation data and the workload associated with the acquisition of annotation data.SOLUTION: An imaging apparatus includes a reception unit configured to receive annotation information related to an image, a display unit configured to display a live view image, a display control unit configured to display the annotation information received by the reception unit on the display unit in association with the live view image, an imaging unit configured to generate a captured image, an input unit configured to receive a selection operation for selecting the annotation information displayed on the display unit, and a data generation unit configured to generate annotation data in which the annotation information selected by the selection operation is associated with the captured image generated by the imaging unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an imaging device, a control method for an imaging device, and a program. [Background technology]

[0002] Generally, machine learning is required when performing processes such as image recognition using AI (Artificial Intelligence), and training data to be used for training must be prepared.In addition, training data may be created by adding information that gives meaning to images, called annotation information, to the image data that forms the basis of the training data.

[0003] In recent years, with the advancement of AI, the demand for training data used in AI training has also been on the rise. However, there is a chronic shortage of training data used in AI training, due to a shortage of image data that forms the basis of training data and a shortage of annotators, who are the workers who work on annotating information.

[0004] Therefore, Patent Document 1 proposes creating training data for AI learning by adding annotation information to a composite image, while Patent Document 2 proposes providing a method for limiting the complexity of adding annotation information to promote the creation of training data for AI learning. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent Publication No. 2021-157404 [Patent Document 2] Patent Publication No. 2021-68450 Summary of the Invention [Problem to be solved by the invention]

[0006] However, this does not fundamentally solve the problems of a shortage of image data (images for annotation) that form the basis of the training data mentioned above, and a shortage of annotators, and given the recent spread of AI technology, there is an urgent need to resolve these issues.

[0007] Generally, annotation information providers receive annotation requests from AI developers and retrieve images related to the annotation information that best suits their needs from the internet or data servers. It is believed that tens of thousands to hundreds of thousands of images for annotation are required for AI training. Therefore, if there are insufficient images for annotation, they must secure images that meet the requirements by taking new photographs, etc., which raises concerns about the burden of retrieving images for annotation.

[0008] The present invention has been made in view of the above, and aims to provide a technique for reducing the workload associated with the shortage of annotation data and the acquisition of annotation data. [Means for solving the problem]

[0009] The imaging device according to the present invention includes a receiving unit that receives annotation information related to an image, a display unit that displays a live view image, a display control unit that displays the annotation information received by the receiving unit on the display unit in association with the live view image, an imaging unit that generates a captured image, an input unit that accepts a selection operation to select the annotation information to be displayed on the display unit, and annotation data that associates the annotation information selected by the selection operation with the captured image generated by the imaging unit. and a data generating unit that generates the image data.

[0010] In addition, a control method for an imaging device according to the present invention is a control method for an imaging device, characterized by including: a receiving step of receiving annotation information related to an image; a display step of displaying a live view image on a display unit; a display control step of displaying the annotation information received by the receiving step in association with the live view image on the display unit; an imaging step of generating a captured image; a step of accepting a selection operation to select the annotation information to be displayed on the display unit; and a data generation step of generating annotation data in which the annotation information selected by the selection operation is associated with the captured image generated by the imaging step. [Effects of the Invention]

[0011] According to the present invention, it is possible to provide a technique for reducing the workload associated with the shortage of annotation data and the acquisition of annotation data. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 2 is a diagram for explaining functions of the information processing system according to the first embodiment. [Figure 2] 1 is a block diagram schematically showing the configuration of an imaging device according to a first embodiment. [Figure 3] 4 is a flowchart of processing executed by the imaging device according to the first embodiment. [Figure 4] 2A to 2C are diagrams schematically showing examples of displays in the imaging device according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the embodiments described below are examples of means for realizing the present invention, and may be modified or changed as appropriate depending on the configuration of the device to which the present invention is applied and various conditions. Furthermore, the embodiments may be combined as appropriate.

[0014] First Embodiment An imaging device according to a first embodiment of the present invention will be described. Fig. 1 is a block diagram showing an example of the configuration of an information processing system 1 including an imaging device 102 according to this embodiment. Note that, although a digital camera is assumed as an example of the imaging device here, the imaging device is not limited to this. The imaging device according to this embodiment may also be a communication device such as a portable tablet device or a personal computer.

[0015] 1 is a diagram showing a schematic configuration of an information processing system 1 according to this embodiment. As shown in FIG. 1, the information processing system 1 includes a management device 101, an imaging device 102, and a server 103.

[0016] The management device 101 functions as a server that manages annotation requests and information related to the processing of images (annotation images) that satisfy the annotation requests in this embodiment. The management device 101 acquires information related to annotation requests from a server 103 managed by an AI development company that provides annotation requests related to annotation images. This annotation request information includes various conditions for identifying objects appearing in images requested by the AI ​​development company for use in AI training. The annotation request information also includes feature information related to the characteristics of the subject appearing in the image and the imaging conditions of the image used for AI training. For example, if the subject is a person, the feature information related to the subject's features includes information such as the subject's gender, age, height, physique, and facial expression.

[0017] The management device 101 transmits the annotation request information acquired from the server 103 to the imaging device 102. The management device 101 also receives annotation data including captured images (annotation images) and tag information (annotation information), which will be described later, from the imaging device 102, and manages the received annotation data. When the management device 101 meets the conditions for a dataset for AI learning required by the AI ​​development company, such as when the number of annotation images received from the imaging device 102 reaches a predetermined number, the management device 101 generates a dataset using the annotation images. Then, the management device 101 transmits the generated dataset to the server 103.

[0018] The imaging device 102 receives information about the annotation request of the AI ​​development company, which the management device 101 has acquired from the server 103, and displays the received information as tag information together with a live view image on the display unit. The imaging device 102 also accepts a selection of tag information displayed on the display unit from the user of the imaging device 102, and associates the selected tag information with the captured image. The imaging device 102 then transmits the associated tag information and captured image to the management device 101 as annotation data.

[0019] In this embodiment, the imaging device 102 receives annotation request information managed by the management device 101 and displays the information as tag information along with the live view image. This allows the user of the imaging device 102 to grasp the annotation request from the AI ​​development company managing the server 103 in real time as tag information for the live view image. Knowing the annotation request from the AI ​​development company when capturing an image allows the user to capture the subject in accordance with the annotation request or recapture the image in accordance with the request. As a result, the imaging device 102 of this embodiment can collect images for annotation more efficiently than conventional methods, reducing the possibility of a shortage of images for annotation and a shortage of annotators. It is assumed that the captured images captured by the imaging device 102 in this embodiment are acquired using still image data or video data.

[0020] In FIG. 1, the imaging device 102 and the server 103 are depicted as standalone devices, but multiple imaging devices and / or servers may be connected to the management device 101 in the information processing system 1. The management device 101 also determines whether an image received from the imaging device 102 satisfies the annotation request received from the server 103. This determination may be made by a user of the management device 101 visually checking the image, instead of being performed by the management device 101. The management device 101 may also make the determination using AI. This allows the management device 101 to prevent images that do not satisfy the annotation request from being mixed into the annotation data.

[0021] FIG. 2 is a block diagram showing the schematic configuration of the image capture device 102 of FIG. 1. Note that the configuration shown in FIG. 2 is merely an example and does not limit the configuration of the image capture device 102 of this embodiment. As shown in FIG. 2, the image capture device 102 includes an MPU (Micro Processor Unit) The imaging device 102 includes a timing signal generating unit 201, a timing signal generating circuit 202, an imaging unit 203, an A / D converter, and a memory controller 205. The imaging device 102 also includes a buffer memory 206, a display unit 207, a storage medium I / F 208, a storage medium 209, an operation unit 210, an information processing unit 211, a receiving unit 212, and a transmitting unit 213.

[0022] In this embodiment, the image capturing device 102 is a camera such as a digital camera or a digital video camera, or an electronic device such as a mobile phone or a computer with a camera function.

[0023] The MPU 201 is a microcontroller for controlling the processing of each part of the image capture device 102, such as the image capture sequence. It controls processes such as receiving information on annotation requests from development companies, displaying live view images and tag information on the display unit 207, associating captured images with tag information, and transmitting annotation data. Note that instead of the MPU 201 controlling the entire device, the entire device may be controlled by having multiple hardware devices share the processing load.

[0024] The timing signal generation circuit 202 generates timing signals required to operate the imaging unit 203. The imaging unit 203 includes, for example, an optical lens unit, an optical system that performs optical control such as aperture, zoom, and focus, and an imaging element that converts light (image) entering the device via the optical lens unit into an electrical video signal. A CMOS (Complementary Metal Oxide Semiconductor) or a CCD (Charge Coupled Device) is used as the imaging element. Under the control of the MPU 201, the imaging unit 203 converts subject light formed by a lens of the imaging unit 203 into an electrical signal using the imaging element, and outputs the resulting digital data as image data after performing noise reduction processing and the like.

[0025] The A / D converter 204 converts analog image data read from the imaging unit 203 into digital image data. The memory controller 205 temporarily stores tag information to be displayed on the display unit 207 and controls memory read / write and the refresh operation of the buffer memory 206. The buffer memory 206 stores captured image data. The display unit 207 displays the tag information and the image data stored in the buffer memory 206. The storage medium I / F 208 is an interface for controlling reading and writing of data from and to the storage medium 209. The storage medium 209 is a memory card, a hard disk, or the like, and stores programs for processing executed in the imaging device 102 and data required for the processing.

[0026] The operation unit 210 is an input unit that accepts user instructions regarding annotation information selection when associating tag information with a captured image. The operation unit 210 may be, for example, a physical button or a button displayed on a touch panel. The information processing unit 211 associates the captured image with tag information. The tag information may be associated with the captured image as meta information using, for example, EXIF ​​(Exchangeable Image File Format), or may be directly embedded in the image data of the captured image. The receiving unit 212 and the transmitting unit 213 are connected to the Internet and transmit and receive data to and from external devices such as the management device 101. The annotation request information acquired by the receiving unit 212 is displayed together with the image as tag information on the display unit 207 via the MPU 201. The receiving unit 212 may receive tag information from the management device 101. The captured image associated with tag information by the information processing unit 211 is transmitted as annotation data from the transmitting unit 213 to the management device 101.

[0027] Next, an example of processing executed by the imaging device 102 of this embodiment will be described. Fig. 3 is a flowchart showing an example of the procedure of processing executed by the imaging device 102. The processing in Fig. 3 is realized by the MPU 201 of the imaging device 102 expanding and executing a program stored in the storage medium 209. The processing shown in Fig. 3 is executed when the imaging device 102 receives an operation related to the start of imaging, such as a user pressing the imaging button on the imaging device 102.

[0028] In step S301, the receiving unit 212 of the MPU 201 acquires tag information from the management device 101. Note that the acquisition of tag information by the MPU 201 may be performed continuously, or may be performed in response to a processing start instruction based on an input from the operation unit 210.

[0029] Next, in step S302, the MPU 201 detects the light received by the imaging unit 203 from the subject. The captured light is converted into an electrical signal, image capturing processing is performed, a live view image is generated, and the live view image is displayed on the display unit 207.

[0030] Next, in step S303, the MPU 201 uses the tag information related to the annotation request acquired by the receiving unit 212 in step S301 to display the tag information together with the live view image on the display unit 207. Here, the MPU 201 is a display control unit that associates the annotation information received by the receiving unit with the live view image and displays it on the display unit.

[0031] Next, in step S304, the MPU 201 controls the image capturing unit 203 in accordance with the user's operation on the operation unit 210, and performs image capturing processing to generate a captured image.

[0032] Next, in step S305, the MPU 201 accepts a user selection of tag information displayed on the display unit 207 in step S303. Note that the user's selection of tag information in step S304 may be performed before capturing the image in step S304. In this case, the flowchart in FIG. 3 is changed so that the processing in step S304 is performed after step S305.

[0033] Next, in step S306, the MPU 201 associates the tag information selected in step S304 with the captured image using the information processing unit 211. Then, the MPU 201 generates annotation data including the image and tag information that meets the annotation request received by the management device 101 from the server 103. Here, the MPU 201 is a data generation unit that generates annotation data in which the annotation information selected by the selection operation is associated with the captured image generated by the imaging unit.

[0034] Then, in step S307, the MPU 201 causes the transmission unit 213 to transmit the annotation data generated in step S306 to the management device 101. The MPU 201 then ends the processing of this flowchart. Note that the MPU 201 may encrypt the data transmitted in step S307, or may additionally perform processing to associate the data to be transmitted with a blockchain.

[0035] In this manner, in this embodiment, the image capture device 102 displays the annotation request as tag information together with the live view image, allowing the user to know the annotation request in real time when capturing an image.

[0036] 4A to 4C schematically show specific examples of images and tag information displayed on the display unit 207 by the above-described processing of the imaging device 102 in this embodiment.

[0037] In the display example shown in FIG. 4A, tag information 401 and a subject 402 in a live view image are displayed on the display unit 207. Here, it is assumed that the subject 402 is a man. As shown in FIG. 4A, the tag information 401 is displayed superimposed on the subject 402 displayed in live view on the display unit 207, allowing the user to check the tag information in real time without interfering with the user's shooting. Furthermore, in the example of FIG. 4A, the background of the display area for the tag information 401 is displayed transparently, so that the subject 402 is not obstructed by the display area for the tag information 401, which is expected to improve usability for the user's shooting. Note that the transparency of the background of the display area for the tag information 401 may be set as appropriate.

[0038] In the example of FIG. 4A, the user can recognize that the tag information 401 contains the information "tag information 2: male" when capturing an image of the subject 402. As a result, the user can operate the operation unit 210 to select "tag information 2: male" after capturing an image of the subject 402. By doing so, it is possible to associate the captured image of the captured subject 402 with the selected tag information. In this way, the user can recognize the annotation request of the AI ​​development company that provided the tag information from the tag information in real time when capturing an image. As a result, for example, when a male subject and a female subject are displayed on the live view screen, the user can select the male subject based on the tag information. Therefore, by using the processing of the flowchart of this embodiment, which cannot be realized with conventional technology, the user can capture an image that meets the annotation request.

[0039] Furthermore, in this embodiment, the order of the individual pieces of tag information displayed on the display unit 207 is changed depending on the subject appearing in the live view image. The MPU 201 acquires subject information about the subject displayed on the display unit 207, and rearranges the multiple pieces of tag information in descending order of relevance to the subject based on the acquired subject information. As a specific example, the MPU 201 can use conventional image recognition technology to estimate characteristics (gender, age, height, build, facial expression, etc.) of the subject displayed on the display unit 207, and acquire the estimation results as subject information. Here, the MPU 201 is an identification unit that identifies the subject appearing in the live view image.

[0040] FIG. 4B shows a display example in which the MPU 201 rearranges the tag information displayed on the display unit 207 in FIG. 4A and displays the rearranged tag information on the display unit 207. In FIG. 4B, the MPU 201 performs image recognition using a trained model stored in the storage medium 209 of the imaging device 102 and estimates that the subject 402 displayed on the display unit 207 is a male. Then, based on the estimation result that the subject 402 is a male, the MPU 201 rearranges the tag information 401, among the tag information 1, 2, 3, ..., that constitutes the tag information 401, so that the tag information "tag information 2: male" takes priority over the other tag information. The MPU 201 then changes the tag information "tag information 2: male" to the tag information "tag information 1: male." As a result, the display unit 207 displays the changed tag information 501. As a result, the tag information "male," which has a stronger association with the subject 402, is displayed preferentially on the display unit 207.

[0041] Therefore, according to this embodiment, even when a list of various tag information is displayed on the display unit 207, the user can more easily select tag information that is easily associated with the subject. Note that in this embodiment, the MPU 201 may rearrange the tag information so that the newest tag information, i.e., the tag information with the most recent acquisition date and time, is displayed preferentially. Furthermore, the MPU 201 may rearrange the tag information using location information such as the user's current location so that tag information that is more relevant to the current location is displayed preferentially.

[0042] In this embodiment, the MPU 201 generates training data including, for example, input images, which are images of subjects with different characteristics captured by an arbitrary imaging device, and ground truth data, which is information indicating the characteristics of the subjects, to generate a training data set consisting of multiple sets of training data. The MPU 201 then uses the training data set to train a learning model that estimates the subject from a live view image displayed on the display unit 207. The MPU 201 stores the trained model obtained by training the learning model in the storage medium 209. In this embodiment, a convolutional neural network (CNN) is used to build the training model. For example, a well-known network such as VGG16 or Dense Net may be used. Note that the MPU 201 may externally acquire a trained model in advance, which has been trained using the above learning model, and store it in the storage medium 209 before starting the flowchart of FIG. 3 .

[0043] In this embodiment, it is assumed that the trained model is obtained by learning a learning model that estimates the characteristics of a person as a subject. However, the subject is not limited to a person, and various objects such as animals, buildings, mountains, forests, and the sea can be used as subjects, and the trained model can estimate the characteristics of each subject. The trained model may be obtained by training a training model.

[0044] Furthermore, in this embodiment, it is assumed that an estimator based on deep learning such as CNN is used as an estimator that realizes the learning model. However, the learning model according to this embodiment may use an estimator based on deep learning other than CNN, such as a vision transformer. Alternatively, the learning model according to this embodiment may use an estimator based on a known machine learning method other than deep learning, such as regression using Random Forest or Adaboost.

[0045] Furthermore, in this embodiment, the tag information displayed on the display unit 207 may display information related to the image of the subject when the user captures the image. In FIG. 4C, details of the annotation request of the AI ​​development company, which the management device 101 has acquired from the server 103, are displayed as tag information related to the subject 402 displayed on the display unit 207 in FIG. 4A. In the example of FIG. 4C, just displaying "male" as tag information, as in FIGS. 4A and 4B, does not reveal the angle of view and imaging conditions at the time of image capture, so even if the image is captured by the user, it may not be usable as the image desired by the AI ​​development company.

[0046] Therefore, for example, when the user operates the operation unit 210 to select "tag information 2: male" in Fig. 4A, or when the user operates the operation unit 210 to select "tag information 1: male" in Fig. 4B, the MPU 201 displays details of the tag information. 4C shows an example of detailed display of tag information displayed on the display unit 207. As shown in FIG. 4C, the display unit 207 displays tag information 601 as details of the tag information for "male," including the shooting condition that "an image capturing the upper body of a male from the front is preferable. Images in which other subjects appear at the same angle of view are not acceptable." The detailed information displayed in the tag information 601 can be generated based on the annotation request information acquired by the management device 101 from the server 103. In this way, in this embodiment, the display unit 207 displays the detailed shooting conditions and requests for the subject as tag information, allowing the user to understand what angle of view and conditions should be used to shoot the subject 402.

[0047] While the present invention has been described in detail above based on preferred embodiments, it is not limited to these specific embodiments, and various modifications within the spirit and scope of the present invention are also encompassed by the present invention. Parts of the above-described embodiments may be combined as appropriate. Furthermore, the present invention also encompasses a case in which a software program implementing the functions of the above-described embodiments is supplied to a system or device having a computer capable of executing the program, either directly from a storage medium or via wired or wireless communication, and the program is then executed. Therefore, the program code itself supplied to and installed on a computer to implement the functional processing of the present invention also embodies the present invention. In other words, the computer program itself for implementing the functional processing of the present invention is also encompassed by the present invention. In this case, the program may take any form, such as object code, a program executed by an interpreter, or script data supplied to an OS, as long as it has the program functionality. Examples of storage media for providing the program include magnetic storage media such as hard disks and magnetic tapes, optical / magneto-optical storage media, and nonvolatile semiconductor memory. Another conceivable method of providing the program is to store the computer program implementing the present invention on a server on a computer network, and then download and install the computer program on a connected client computer.

[0048] The various controls described above may or may not be performed by a single piece of hardware (e.g., a processor or circuit). The entire device may be controlled by multiple pieces of hardware (e.g., multiple processors, multiple circuits, or a combination of one or more processors and one or more circuits) sharing the processing.

[0049] The above processor is a processor in the broad sense, and includes general-purpose processors and dedicated processors. General-purpose processors include, for example, CPUs (Central Processing Units), MPUs (Micro Processing Units), and DSPs (Digital Signal Processors). Dedicated processors include, for example, GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and PLDs (Programmable Logic Devices). Programmable logic devices include, for example, FPGAs (Field Programmable Gate Arrays) and CPLDs (Complex Programmable Logic Devices).

[0050] Although the embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Furthermore, each of the above-described embodiments merely represents one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0051] (Other embodiments) The present invention can also be realized by a process in which a program that realizes one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in the computer of the system or device read and execute the program, or by a circuit that realizes one or more functions. [Explanation of symbols]

[0052] 102 imaging device, 201 MPU, 207 display unit, 211 information processing unit

Claims

1. a receiving unit for receiving annotation information related to an image; a display unit that displays a live view image; a display control unit that displays the annotation information received by the receiving unit on the display unit in association with the live view image; an imaging unit that generates a captured image; an input unit that accepts a selection operation to select the annotation information to be displayed on the display unit; and a data generating unit that generates annotation data in which the annotation information selected by the selection operation and the captured image generated by the imaging unit are associated with each other.

2. an identifying unit that identifies a subject appearing in the live view image; The display control unit sorts the plurality of pieces of annotation information in descending order of relevance to the subject identified by the identification unit and displays the sorted pieces of annotation information on the display unit.

2. The imaging device according to claim 1.

3. 3. The imaging device according to claim 2, wherein the identification unit is configured to identify the subject appearing in the live view image based on an estimation result of the subject using a trained model obtained by training a learning model that estimates the subject appearing in the live view image using teacher data including an input image and ground truth data that is information about the subject in the input image.

4. the annotation information includes feature information relating to features of a subject captured in the live view image and imaging conditions; The display control unit displays the imaging conditions together with the feature information on the display unit.

2. The imaging device according to claim 1.

5. The imaging device according to any one of claims 1 to 4, a management device that receives the annotation data from the imaging device; An information processing system having the above.

6. a receiving step of receiving annotation information associated with the image; a display step of displaying a live view image on a display unit; a display control step of displaying the annotation information received in the receiving step on the display unit in association with the live view image; an imaging step of generating a captured image; receiving a selection operation for selecting the annotation information to be displayed on the display unit; a data generating step of generating annotation data in which the annotation information selected by the selection operation and the captured image generated by the imaging step are associated with each other; 11. A method for controlling an imaging device, comprising:

7. A program for causing a computer to function as each unit of the imaging device according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and device for providing annotated traffic area data

    JP2021068450A

  • Learning data generation method, learning data generation device, and program

    JP2021157404A