Information processing device, control method for information processing device, and program

JP7912035B2Active Publication Date: 2026-08-27CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024087244
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2026-08-27
Estimated Expiration
2044-05-29

AI Technical Summary

Benefits of technology

【0010】 本発明によれば、学習データを生成する手間を低減することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007912035000001
    Figure 0007912035000001
  • Figure 0007912035000002
    Figure 0007912035000002
  • Figure 0007912035000003
    Figure 0007912035000003
Patent Text Reader

Abstract

To provide a technique for reducing the time and effort required to generate training data.SOLUTION: An information processing device for generating training data comprises: position information acquisition means for acquiring focus position information of imaging means; image acquisition means for acquiring one or more images on the basis of the timing at which the focus position information is acquired; and generation means for generating the training data by adding annotation information to the one or more images acquired by the image acquisition means on the basis of the focus position information.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , , , , , , , , , , , ,

[0005] ,

[0001] The present invention relates to an information processing apparatus, a control method for the information processing apparatus, and a program.

Background Art

[0002] In recent years, the use of AI (Artificial Intelligence) has been promoted in various fields. Among them, there is supervised learning in which machine learning is performed based on training data including correct data to generate an inference model.

[0003] In supervised learning, in order to obtain a machine learning model with high generalization performance, various inputs determined by the task to be solved and training data with high-quality annotation information added thereto are required. Generally, high-quality training data can utilize publicly available datasets created for purposes such as competitions. However, if there is no training data suitable for the purpose, it is necessary to create a dataset by oneself. When creating a dataset by oneself, an annotation operation for creating training data by adding annotation information manually by a person or the like is required. Since a huge amount of training data is required to create a machine learning model with high generalization performance, a lot of time is required for annotation.

[0004] In Patent Document 1, a method of performing annotation semi-automatically or automatically using a learned object detector to generate training data is disclosed.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Non-Patent Documents

[0006] <00项目00030>[[ID=4I]]

Non-Patent Document 1

[0007] However, the technology disclosed in Patent Document 1 requires the user to prepare image data themselves in order to generate training data, which necessitates taking photographs and collecting image data, thus presenting the challenge of time-consuming data generation.

[0008] This invention has been made in view of the above-mentioned problems, and aims to provide a technology for reducing the effort required to generate training data. [Means for solving the problem]

[0009] An information processing apparatus according to one aspect of the present invention, which achieves the above objective, An information processing device that generates training data, A location information acquisition means for acquiring focus position information of a shooting means, Image acquisition means that acquires one or more images based on the time when the focus position information is acquired, The time and period from when the aforementioned focus position information was acquired A synchronized defocus map, wherein the amount of defocus is mapped on the input image data. Get Defocus Map means of acquisition, The aforementioned defocus map and the focus position information and Based on this, a generation means determines the amount of defocus of the main subject and generates training data in which the amount of defocus of the main subject is added as annotation information to one or more images acquired by the image acquisition means, It is characterized by having the following features. [Effects of the Invention]

[0010] According to the present invention, the effort required to generate training data can be reduced. [Brief explanation of the drawing]

[0011] [Figure 1] Hardware configuration diagram of a learning data generation device according to an embodiment. [Figure 2] Diagram for explaining the functional configuration of the learning data generation device in the first embodiment. [Figure 3] Flowchart showing the procedure of the learning data generation process in the first embodiment. [Figure 4] Diagram for explaining the learning data generation process in the first embodiment. [Figure 5] Flowchart showing the procedure of the learning data generation process in Modification 1 of the first embodiment. [Figure 6] Diagram for explaining the learning data generation process in Modification 1 of the first embodiment. [Figure 7] Flowchart showing the procedure of the learning data generation process in the second embodiment. [Figure 8] Diagram for explaining the functional configuration of the learning data generation device in the second embodiment. [Figure 9] Diagram for explaining the learning data generation process in the second embodiment. [Figure 10] Diagram for explaining the functional configuration of the learning data generation device in the third embodiment. [Figure 11] Flowchart showing the procedure of the learning data generation process in the third embodiment. [Figure 12] Flowchart showing the procedure of the learning data generation process in the third embodiment. [Figure 13] Diagram for explaining the learning data generation process in the third embodiment.

Modes for Carrying Out the Invention

[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0013] (First Embodiment) In this embodiment, a case of generating learning data will be described by taking shooting with an interchangeable-lens digital camera as an example.

[0014] <Hardware Configuration> FIG. 1 shows an example of the hardware configuration of an information processing apparatus (learning data generation apparatus 200) according to this embodiment. The CPU 100 is a central processing unit that performs operations such as calculations and logical determinations for various processes. A control program is stored in the ROM (Read-Only-Memory) 110. The RAM (Random Access Memory) 120 is used as a temporary storage area such as the main memory and work area of the CPU 100. The HDD 130 is a hard disk for storing electronic data and programs according to this embodiment. An external storage device may be used to perform the same role. Here, the external storage device can be realized, for example, by a medium (recording medium) and an external storage drive for realizing access to the medium. Examples of such a medium include a flexible disk (FD), CD-ROM, DVD, USB memory, MO, flash memory, and the like. Further, the external storage device may be a server device connected via a network or the like.

[0015] The input unit 140 consists of a keyboard, touch panel, etc., and accepts input from the user. The display unit 150 consists of a liquid crystal display, etc., and can display various data and processing results to the user. The learning data generation device 200 can also communicate with other devices via the communication unit 160. It may receive instructions from the user via the communication unit 160 from other devices, or it may output processing results to other devices. The learning data generation device 200 can be configured using a general-purpose information processing device having the above configuration.

[0016] <Functional Configuration> Figure 2 is a diagram illustrating the learning data generation device 200 in this embodiment. The shooting unit 201 photographs the subject. For example, it can be composed of an image sensor such as a CMOS sensor. The focus position information acquisition unit 202 acquires the focus position information entered by the user, which is received from the input unit 140. The focus position information entered by the user is the coordinates indicating the position within the displayed image (within the field of view of the shooting unit 201) where the AF frame to be focused is selected by touch operation (touch AF). Other methods for selecting the AF frame include selecting it by dragging the AF frame on a touch panel or joystick (multicontroller) (touch and drag AF). Alternatively, the user may select it by operating a pointer with their gaze (gaze input).

[0017] The image acquisition unit 203 acquires images captured by the shooting unit 201. The acquired images include live view images, and the timing of image acquisition does not depend on the user's shutter release operation. Based on the focus position information acquired by the focus position information acquisition unit 202, the image acquisition unit 203 adds annotation information to the acquired images. The learning data storage unit 205 stores the learning data generated by the learning data generation unit 204.

[0018] <Processing> Figure 3 shows an overall flowchart of the learning data generation process according to this embodiment. In S301, the imaging unit 201 starts imaging. The display unit 150 displays the live view image captured by the imaging unit 201.

[0019] In S302, the shooting unit 201 activates the touch AF shooting mode, which is a method in which the user selects the AF frame during autofocus (AF). The input unit 140 waits for the user to input touch input.

[0020] In S303, the input unit 140 accepts user touch input. The user checks the live view image displayed on the display unit 150 and touches the position on the touch panel of the input unit 140 where they want to focus, thereby performing touch input.

[0021] In S304, the focus position information acquisition unit 202 acquires the coordinates (focus position information) of the touch input received in S303. The image acquisition unit 203 also saves the live view image of the shooting unit 201 at the moment the touch input was made in S303. The timing of saving the live view image may be the moment when the touch input is confirmed by any button operation, such as when the shutter button is half-pressed after the touch input.

[0022] In S305, the shooting unit 201 starts the autofocus process. The autofocus process starts when the touch input is completed in S303. In S306, the shooting unit 201 drives the focus lens based on the autofocus process in S305. In S307, the shooting unit 201 determines that focus has been achieved if it is determined that the focus has been set to the position specified in S303 based on the focus drive result in S306.

[0023] In S308, the input unit 140 determines whether the user has performed touch input again. If this step is Yes, the process returns to S303. On the other hand, if this step is No, the process proceeds to S309. The user checks the focusing result in S307 and determines that the desired subject is in focus, then the process proceeds to S309 without performing touch input again. Otherwise, the process returns to S303, and the user touches the position they want to focus on to perform touch input. After that, the processes from S304 to S307 are repeated.

[0024] In S309, the learning data generation device 200 determines whether the user has performed a shutter release operation. If a shutter release operation has been performed, the shooting unit 201 takes a picture and proceeds to the process in S310. If a shutter release operation has not been performed, the process returns to S308 and determines again whether the user has performed touch input again.

[0025] In S310, the training data generation unit 204 generates training data by adding annotation information to the image saved in S304 based on the focus position information saved in S304. For example, the saved focus position information is added to the image as annotation information to generate training data. This completes the series of processes shown in Figure 3.

[0026] Here, Figures 4(a) to 4(c) are explanatory diagrams of the learning data generation process in S310. Figure 4(a) shows image 400 saved in S304. 401 represents an object, and 402 represents a person. Figure 4(b) is a visualization of the touch input coordinates, which are the focus position information saved in S304. The marker 410 represents the touch input coordinates in touch AF saved in S304. Figure 4(c) is a diagram in which the touch input coordinates, which are the focus position information saved in S304, are superimposed on the image saved in S304. On the person 421 in image 420, the coordinates 422 of the position touched by the user are visualized as a marker. For example, Figure 4(c) shows that the coordinates indicating the pupil of person 421 are determined as annotation information by the user's touch input.

[0027] As described above, according to this embodiment, it is possible to collect image data for estimating focus position information and annotation information by performing normal shooting using the shooting unit. Therefore, since the user can collect image data and perform annotation without being aware of it, the effort required to generate training data can be reduced.

[0028] <Learning method in the first embodiment> The training data generated by the training data generation process in this embodiment can be used to train a machine learning model for estimating focus position information.

[0029] Specific machine learning algorithms include nearest neighbors, naive Bayes, decision trees, and support vector machines. Deep learning, which uses neural networks to generate features and connection weights for learning, is another example. Appropriately, any of the above algorithms can be applied to this embodiment.

[0030] This section explains learning using neural networks. The learning process uses images saved in S304 as input data. The learning process involves error detection and weight update.

[0031] In the error detection process, the error between the output data output from the output layer of the neural network, which is based on the input data input to the input layer, and the training data is obtained. At this time, the focus position information saved in S304 is used as the training data. The focus position information is, for example, the coordinates of the touch position in touch AF. In the error detection process, a loss function may also be used to calculate the error between the output data from the neural network and the training data.

[0032] In the weight update process, the connection weight coefficients between nodes of the neural network are updated based on the error obtained in the error detection process, so as to reduce the error. This weight update process can be performed, for example, by updating the connection weight coefficients using backpropagation. Backpropagation is a method that adjusts the connection weight coefficients between nodes of each neural network so as to reduce the above error.

[0033] The output data produced as a result of the learning process is a machine learning model that estimates focus position information. The machine learning model refers to parameters such as weighting coefficients obtained through the weight update process.

[0034] In this way, by using images as input data and focus position information as training data, it becomes possible to train a neural network that regresses focus position information corresponding to the input image.

[0035] (Inference method in the first embodiment) Focus position information can be inferred using a machine learning model trained using the learning method described above. Here, we will explain the case where the trained machine learning model is applied to a digital camera with interchangeable lenses for inference.

[0036] The input data used is, for example, a live view image captured by an image sensor such as a CMOS sensor. After acquisition, the live view image is used directly as input data for a pre-trained machine learning model.

[0037] The output data is the inference result, and an estimated value of the focus position information is output. For example, the output data may be the estimated coordinates of the touch position in touch AF, and the position will represent the coordinate information within the image, such as (310,452).

[0038] In this way, when a user takes a picture with an interchangeable-lens digital camera, it becomes possible to estimate the focus position information contained in the training data from within the live view image using a pre-trained machine learning model. If a subject included in the training data is present in the live view image, the AF frame in the image can be automatically selected without the need for input operations such as touch AF, thus reducing the effort required during shooting.

[0039] Furthermore, when photographing subjects included in the training data, even if the subject is moving quickly and it is difficult for the user to select an AF frame, the system can automatically select an AF frame from within the image, making it easier for the user to focus on the subject they want to focus on.

[0040] [Example 1] On smartphones, selecting an AF target is frequently done by touching the screen. Below, we will describe a modification of the first embodiment when applied to a smartphone. In this modification, we will describe a case where a smartphone, which has multiple imaging sensors and multiple lenses with different angles of view, generates multiple training data simultaneously in a single process.

[0041] For example, if a smartphone has three lenses—a telephoto lens, a standard lens, and a wide-angle lens—each lens has its own image sensor (such as a CMOS sensor), and the shooting operation is performed simultaneously through each lens. The telephoto lens has a longer focal length than the standard lens, making it possible to magnify and photograph subjects that are further away. Also, the wide-angle lens has a shorter focal length than the standard lens, so using the wide-angle lens allows you to photograph a wider area than when using the standard lens.

[0042] In other words, the focal length decreases in the order of telephoto lens, standard lens, and wide-angle lens, and the field of view widens accordingly. Furthermore, the telephoto lens, standard lens, and wide-angle lens are assumed to be lenses with a zoom function, allowing for a continuous change in the field of view between the telephoto and wide-angle ends. The telephoto lens, standard lens, and wide-angle lens may not only have a mechanism that optically achieves a predetermined magnification of 1:1, but may also have a mechanism that allows the user to change the magnification.

[0043] Furthermore, multiple live view images from each lens can be viewed on the smartphone's display. Users can take photos while checking the multiple live view images displayed on the screen.

[0044] The process in this modified example is the same as the process shown in Figure 3 of the first embodiment, so a basic explanation will be omitted. The process in S304 differs from that of the first embodiment, so it will be explained in detail using the flowchart in Figure 5.

[0045] Figure 5 is a flowchart showing the overall processing steps for the training data generation process in Modification Example 1. Since each of the processes S501 to S503 is the same as the processes S301 to S303 in Figure 3, their explanation is omitted.

[0046] In S304, the focus position information acquisition unit 202 acquires the coordinates (focus position information) of the touch input received in S303. The image acquisition unit 203 also saves the live view image from the shooting unit 201 at the moment the touch input was made in S303.

[0047] In S504, the input unit 140 determines whether the user has performed touch input on the live view image of the lens with the longest focal length. If this step is Yes, the process proceeds to S505. On the other hand, if this step is No, the process proceeds to S512.

[0048] Here, Figure 6 shows the results of displaying the live view image 604 of the telephoto lens, the live view image 603 of the standard lens, and the live view image 602 of the wide-angle lens on the display 601 of the smartphone 600. For example, if the user touches the live view image 604 of the telephoto lens in S504, the process proceeds to S505.

[0049] Furthermore, if, for example, the user touches the live view image 602 of the wide-angle lens, the process proceeds to S512. When touch input is performed on the live view image 602 of the wide-angle lens with a short focal length, if the touch position is at the edge of the screen, the field of view of the live view image of the lens with a longer focal length is narrower, so the focus position information may not fit in the image. If the focus position information does not fit in the image, it may not be possible to generate training data, so the judgment process in S504 is performed.

[0050] In S505, the focus position information acquisition unit 202 acquires the coordinates of the touch input entered in S503. The image acquisition unit 203 also saves two or more live view images that were simultaneously captured with two or more of the three lenses: telephoto lens, standard lens, and wide-angle lens.

[0051] Since the processes S506 to S511 are the same as the processes S305 to S310 in Figure 3, their explanation will be omitted.

[0052] In S512, the input unit 140 determines whether the touch position is within the image range of the telephoto lens when touch input is performed on the live view image of the wide-angle lens. If this step is Yes, the process proceeds to S505. On the other hand, if this step is No, the process ends.

[0053] This process takes into account the case where, if touch input is performed on a live view image of a wide-angle lens with a short focal length, and the touch position is at the edge of the screen, the field of view of the live view image of a long-focal-length lens is narrower, so the focus position information may not fit in the image. Even if touch input is performed on a live view image of a wide-angle lens, if the touch position is within the image range of the telephoto lens, multiple training data can be generated simultaneously, so this judgment process is performed.

[0054] Let f1 be the focal length of the wide-angle lens, f2 be the focal length of the telephoto lens, w1 be the horizontal resolution of the live view image of the wide-angle lens, and h1 be the vertical resolution. Then, in the live view image of the wide-angle lens, the resolution of the area corresponding to the field of view of the live view image of the telephoto lens, with respect to the center of the image, is Δw in the horizontal direction and Δh in the vertical direction. In this case, Δw and Δh are defined by the following equations (1) and (2).

[0055] Δw = w1·f1 / f2 ...(1) Δh = h1·f1 / f2 ...(2) Therefore, if we define the horizontal resolution w2 and the vertical resolution h2 as such that the touch input in the wide-angle lens's live view image fits within the range of the telephoto lens's live view image, then the range of values ​​for w2 and h2 is defined by equations (3) and (4) below.

[0056] (w1-Δw) / 2≦w2≦(w1+Δw) / 2 ...(3) (h1-Δh) / 2≦h2≦(h1+Δh) / 2 ...(4) For example, when f1=13, f2=26, w1=4032, and h1=3024, Δw=2016 and Δh=1512. Also, the range of w2 values ​​is from 1008 to 3024, and the range of h2 values ​​is from 756 to 2268. However, since the lenses of a smartphone are not mounted in exactly the same position, the center positions of the wide-angle lens and the telephoto lens may be shifted by the amount of the mounting position difference. Therefore, the ranges of w2 and h2 may be corrected by that amount of shift. If the coordinates touched in the live view image of the wide-angle lens are included in this range of w2 and h2, proceed to processing S505; otherwise, terminate the learning data generation process.

[0057] Thus, in the case of smartphones with multiple lenses having different angles of view, it is possible to collect multiple training data sets with different angles of view simultaneously in a single processing step, making it possible to efficiently increase the variety of training data.

[0058] [Differentiation 2] In the above embodiment, only the live view image at the moment of touch input in S304 was saved, but one or more additional live view images may be saved at other times. For example, the saving timing may be immediately after the focus determination in S307.

[0059] In this way, increasing the timing of saving live view images makes it possible to capture the movement and changes of a moving subject in the image, for example, when photographing a moving subject, in the short time between touch input and focusing. It also makes it possible to increase the variety of captured images and increase the amount of training data that can be acquired in a single processing step.

[0060] (Second embodiment) This embodiment describes an example of creating training data to detect a main subject located at a user-desired focal position.

[0061] <Functional Configuration> Figure 8 is a diagram illustrating the learning data generation device 200 according to this embodiment. The image acquisition unit 203 from the image capture unit 201 in this embodiment is the same as the image acquisition unit 203 from the image capture unit 201 in Embodiment 1, so its description is omitted.

[0062] The learning data generation unit 204 in this embodiment includes an object recognition unit 2041, which performs object recognition on images acquired by the image acquisition unit 203. Here, the object recognition process is a process that includes at least one of object detection and region segmentation. The region segmentation may be semantic region segmentation that performs class classification at the pixel level.

[0063] Object detection is the task of detecting specific objects (people, dogs, vehicles, etc.) or specific body parts (faces, eyes, heads, etc.) within an image. In object detection, the position and size of the object to be detected are estimated, and the estimated result is output as a rectangle (bounding box).

[0064] Region segmentation is the task of identifying a class for each pixel in an image and classifying the image into regions according to that class. Region segmentation includes semantic segmentation, which performs class classification for all pixels in an image, and instance segmentation, which not only performs class classification but also distinguishes and classifies each object within the image. Specific object detection and region segmentation methods are described in Non-Patent Document 1, but the explanation of the methods themselves is omitted in this embodiment.

[0065] In this embodiment, the object recognition unit 2041 is a pre-trained neural network. The training model is not particularly limited as long as it is capable of object recognition. The data input to the object recognition unit 2041 is an image acquired by the image acquisition unit 203.

[0066] <Processing> The processing in this embodiment is the same as the processing shown in Figure 3 described in the first embodiment, so a basic explanation will be omitted. The processing at S310 differs from that in the first embodiment. Figure 7 is a flowchart showing the procedure for generating training data in this embodiment, and shows the detailed processing procedure according to this embodiment corresponding to the processing at S310. Here, using instance segmentation as an example, we will explain while referring to Figures 9(a) to (d) as well.

[0067] In S701, the object recognition unit 2041 performs object recognition on the image acquired by the image acquisition unit 203. Figure 9(a) shows that in image 900 acquired by the image acquisition unit 203, an object 901 (tree) and a subject 902 (person) are captured. The object recognition unit 2041 performs instance segmentation on this image. Figure 9(b) shows the output result 910 after performing instance segmentation on image 900. Class A 911 and class B 912 are classified as pixel-level classes in the output result. Here, class A represents, for example, a tree, and class B represents, for example, a person.

[0068] In S702, the learning data generation unit 204 determines the main subject from the recognition results of the object recognition unit 2041 based on the coordinates acquired by the focus position information acquisition unit 202. Figure 9(c) shows the output result 920 after performing instance segmentation, with the coordinates 921 of the focus position information acquired by the focus position information acquisition unit 202 superimposed. In the output result 920, the class to which the coordinates indicated by the coordinates 921 of the focus position information belong is determined to be the main subject in the image. Here, class B922 is determined to be the main subject.

[0069] In S703, the learning data generation unit 204 recognizes the information of the main subject determined in S702 as the correct label for object detection, and generates learning data by adding the correct label as annotation information related to the main subject. Figure 9(d) shows the result of creating learning data for detecting the main subject in an image by surrounding the pixels of class B931, to which the main subject in the image determined in S702 belongs, with a bounding rectangle 932.

[0070] For the class labels of the created data, the class of the primary subject determined in S702 may be used. Alternatively, the class labels of the training data may be set by accumulating multiple sets of created training data and clustering them.

[0071] According to this embodiment, it is possible to collect image data and annotation information for detecting the main subject by performing normal shooting using the shooting unit. This makes it possible to reduce the effort required to create training data, including the collection of image data.

[0072] <Learning method in the second embodiment> The training data generated by the training data generation process in this embodiment can be used to train a machine learning model for object detection.

[0073] One specific machine learning algorithm is deep learning, which uses neural networks to generate features and connection weights for learning. This section will explain learning using neural networks.

[0074] The learning process uses images saved in S304 as input data. During learning, error detection and weight update processes are performed. In the error detection process, the error between the output data output from the output layer of the neural network, based on the input data input to the input layer, and the training data is obtained. At this time, the bounding box used as the training data is the bounding box that estimates the position and size of the object, which was saved in S703. In the error detection process, a loss function may also be used to calculate the error between the output data from the neural network and the training data.

[0075] In the weight update process, the connection weight coefficients between nodes of the neural network are updated based on the error obtained in the error detection process, so as to reduce the error. In this weight update process, for example, the backpropagation method is used to update the connection weight coefficients. The backpropagation method is a technique that adjusts the connection weight coefficients between nodes of each neural network so as to reduce the above error.

[0076] The output data produced as a result of the learning process is a machine learning model that estimates bounding boxes that define the position and size of objects. The machine learning model refers to parameters such as weighting coefficients obtained through the weight update process.

[0077] In this way, by using images as input data and bounding boxes that estimate the position and size of objects as training data, it becomes possible to train a neural network that can detect objects from input images.

[0078] <Inference method in the second embodiment 2> Using a machine learning model trained with the learning method described above, object detection inference can be performed. Here, we will explain how to apply the trained machine learning model to an interchangeable-lens digital camera and perform inference.

[0079] For input data, live view images captured by an image sensor such as a CMOS sensor are used. After acquisition, the live view images are used directly as input data for a pre-trained machine learning model.

[0080] The output data is the inference result, and it outputs a bounding box that estimates the position and size of the object. For example, the output data represents the position and size according to the spatial direction of the image, such as (310, 452, 105, 40) in the order of horizontal coordinates, vertical coordinates, horizontal size, and vertical size.

[0081] In this way, when a user takes a picture with an interchangeable-lens digital camera, it becomes possible to estimate the position and size of objects included in the training data from within the live view image using a pre-trained machine learning model. If the user wants to photograph a subject included in the training data, the system can automatically detect the subject and focus on it without the user having to select an AF frame using input operations such as touch AF in the image, thus reducing the effort required for input during shooting.

[0082] (Third embodiment) This embodiment describes an example of creating training data to estimate the amount of defocus at a user-desired focal position from a defocus map.

[0083] First, let's explain the defocus map. A defocus map is a map that maps the amount of defocus at multiple locations on the input image data. The amount of defocus is expressed in units of Fδ. A known pupil-splitting phase difference detection method can be used to generate the defocus map. For example, correlation calculations are performed on the image signals of two different pupil regions to calculate the phase difference, which is the amount of displacement between images with different disparities (hereinafter referred to as image A and image B), i.e., the image displacement between image A and image B. Then, the amount of defocus is calculated based on the calculated phase difference (image displacement) between image A and image B, thereby generating a defocus map.

[0084] <Functional Configuration> Figure 10 is a diagram illustrating the functional configuration of the learning data generation device 200 in this embodiment. The imaging unit 201 to the learning data storage unit 205 in this embodiment are the same as those in Figure 2, so a basic explanation is omitted.

[0085] The defocus map acquisition unit 206 acquires a defocus map that is time-synchronized with the image acquired by the image acquisition unit 203.

[0086] <Processing> Figure 11 is a flowchart showing the procedure for generating training data in this embodiment. The process performed in S1105 and the process in S1111, which corresponds to the process in S310, differ from those in the first embodiment, so the differences will be explained in detail.

[0087] The processes S1101 to S1104 are the same as the processes S301 to S304 in Figure 3, so their explanations are omitted. In S1105, the defocus map acquisition unit 206 acquires and saves a defocus map that is time-synchronized with the image acquired in S1104. The processes S1106 to S1110 are the same as the processes S305 to S309 in Figure 3, so their explanations are omitted. In S1111, the training data generation unit 204 generates training data from the defocus map acquired in S1105.

[0088] Now, with reference to Figure 12, the details of the process in S1111 will be explained. In S1201, the learning data generation unit 204 reads the defocus map acquired in S1105. In S1202, the learning data generation unit 204 determines the amount of defocus of the main subject from the defocus map acquired in S1105.

[0089] Figures 13(a) to 13(c) illustrate the process of determining the amount of defocus on the main subject based on the focus position information from the values ​​on the defocus map read in S1201, and generating training data. Figure 13(a) shows that the subject 1301 (person) is captured in image 1300 acquired in S1104.

[0090] Figure 13(b) is an image obtained by superimposing the defocus map 1311 read in S1201 onto the image 1310 acquired in S1104. The defocus map 1311 shows that the darker the color, the greater the amount of defocus, and the amount of defocus in the area where the subject 1312 is located is greater than that of the background.

[0091] Figure 13(c) shows the result of visualizing and superimposing the focus position information 1322 saved in S1104 onto the defocus map 1321. The amount of defocus at the coordinates indicated by the focus position information 1322 on the defocus map 1321 is determined as annotation information. At this time, the determined amount of defocus is, for example, 0.1Fδ.

[0092] In S1203, the learning data generation unit 204 generates learning data using the defocus amount determined in S1202 as annotation information. In this embodiment, learning data is generated by adding the defocus amount of the main subject determined in S1202 as annotation information to the image saved in S1104.

[0093] According to this embodiment, by performing imaging using the imaging unit, it becomes possible to collect image data and annotation information for estimating the amount of defocus. This makes it possible to reduce the effort required to create training data, including the collection of image data.

[0094] <Learning method in the third embodiment> The training data generated by the training data generation process in this embodiment can be used to train a machine learning model for estimating the amount of defocus.

[0095] One specific machine learning algorithm is deep learning, which uses neural networks to generate features and connection weights for learning. This section will explain learning using neural networks.

[0096] The learning process uses images saved in S304 as input data. The learning process consists of error detection and weight update. In the error detection process, the error between the output data output from the output layer of the neural network and the training data is obtained according to the input data input to the input layer. At this time, the defocus amount saved in S1203 is used as the training data. In the error detection process, a loss function may also be used to calculate the error between the output data from the neural network and the training data.

[0097] In the weight update process, the connection weight coefficients between nodes of the neural network are updated based on the error obtained in the error detection process, so as to reduce the error. In this weight update process, for example, the backpropagation method is used to update the connection weight coefficients. The backpropagation method is a technique that adjusts the connection weight coefficients between nodes of each neural network so as to reduce the above error.

[0098] The output data produced as a result of the learning process is a machine learning model that estimates the amount of defocus of the main subject. The machine learning model refers to parameters such as weighting coefficients obtained through the weight update process.

[0099] In this way, by using images as input data and defocus levels as training data, it becomes possible to train a neural network that estimates the defocus level corresponding to the input image.

[0100] <Inference method in the third embodiment> Using a machine learning model trained with the learning method described above, it is possible to infer the amount of defocus of the main subject. Here, we will explain the case where the trained machine learning model is applied to an interchangeable-lens digital camera to perform inference.

[0101] The input data is, for example, a live view image captured by an image sensor such as a CMOS sensor. After acquisition, the live view image is used directly as input data for a pre-trained machine learning model. The output data is the inference result, which is an estimated value of the amount of defocus of the main subject. The output data is, for example, 0.13Fδ.

[0102] In this way, when a user takes a picture with an interchangeable-lens digital camera, it becomes possible to estimate the amount of defocus of the main subject included in the training data from the live view image using a pre-trained machine learning model. If the user wants to photograph the main subject included in the training data, the system can automatically identify the main subject and estimate its defocus amount without the user having to select an AF frame using input operations such as touch AF in the image. As a result, the effort required during shooting is reduced, and it becomes easier to focus on the main subject included in the training data by performing AF using the estimated defocus amount.

[0103] The disclosures herein include the following information processing devices, methods for controlling the information processing devices, and programs. (Item 1) An information processing device that generates training data, A location information acquisition means for acquiring focus position information of a shooting means, Image acquisition means that acquires one or more images based on the time when the focus position information is acquired, A generation means that generates the training data by adding annotation information to one or more images acquired by the image acquisition means based on the focus position information, An information processing device characterized by comprising: (Item 2) The information processing device according to item 1, characterized in that the focus position information is a coordinate indicating the position of a distance measuring point selected within the field of view of the shooting means. (Item 3) The system further includes a display means for displaying an image within the field of view of the aforementioned shooting means, The information processing device according to item 2, characterized in that the position of the distance measuring point is a position selected by the user on the image within the field of view of the shooting means. (Item 4) The information processing device according to any one of items 1 to 3, characterized in that the one or more images are live view images. (Item 5) Further comprising object recognition means for recognizing pre-learned objects, The generating means is Based on the focus position information, the object to be the main subject is determined from the recognition result of the object recognition means, An information processing device according to any one of items 1 to 4, characterized in that it generates training data in which annotation information related to the main subject is attached to one or more images acquired by the image acquisition means. (Item 6) The information processing apparatus according to item 5, characterized in that the object recognition means performs at least one of the processes of object detection or region segmentation. (Item 7) The information processing apparatus according to item 6, characterized in that the domain partitioning includes instance segmentation. (Item 8) The information processing device according to any one of items 1 to 7, characterized in that the one or more images are images taken at a time that is temporally close to the time when the focus position information was acquired. (Item 9) The aforementioned shooting means is equipped with multiple lenses with different angles of view, The plurality of lenses include a lens with a first angle of view and a lens with a second angle of view that is wider than the first angle of view. A display means for displaying a first image acquired by a lens with a first field of view and a second image acquired by a lens with a second field of view, The system further comprises a determination means for determining whether or not the focus position information selected on the second image is included in the first image, The information processing apparatus according to item 1 or 2, characterized in that the generation means generates the learning data by adding the annotation information to the first image and the second image based on the focus position information when the focus position information selected on the second image is included in the first image. (Item 10) The information processing device according to any one of items 1 to 9, characterized in that the generation means assigns the focus position information as annotation information. (Item 11) The system further includes a map acquisition means for acquiring a defocus map at a time point that is temporally close to the time when the focus position information was acquired. The generating means is Based on the defocus map acquired by the map acquisition means, the amount of defocus on the main subject is determined. An information processing device according to any one of items 1 to 3, characterized in that it generates training data in which the amount of defocus of the main subject is added as annotation information to one or more images acquired by the image acquisition means. (Item 12) The information processing device according to item 11, characterized in that the generation means generates the learning data to which the amount of defocus at the coordinates indicated by the focus position information is added as annotation information. (Item 13) A control method for an information processing device that generates training data, A position information acquisition process to acquire focus position information of the shooting device, An image acquisition step in which one or more images are acquired based on the time when the focus position information is acquired, A generation step that generates the training data by adding annotation information to one or more images acquired in the image acquisition step based on the focus position information, A control method for an information processing device, characterized by having the following features. (Item 14) A program to cause a computer to function as an information processing device as described in any one of items 1 through 12.

[0104] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0105] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of symbols]

[0106] 140: Input unit, 150: Display unit, 200: Learning data generation device, 201: Shooting unit, 202: Focus position information acquisition unit, 203: Image acquisition unit, 204: Learning data generation unit, 205: Learning data storage unit

Claims

1. An information processing device that generates training data, A location information acquisition means for acquiring focus position information of a shooting means, Image acquisition means that acquires one or more images based on the time when the focus position information is acquired, A defocus map acquisition means for acquiring a defocus map that is time-synchronized with the time when the focus position information is acquired, wherein the amount of defocus is mapped on the input image data. A generation means that determines the amount of defocus of the main subject based on the defocus map and the focus position information, and generates learning data in which the amount of defocus of the main subject is added as annotation information to one or more images acquired by the image acquisition means, An information processing device characterized by comprising:

2. The information processing device according to claim 1, characterized in that the focus position information is a coordinate indicating the position of a distance measuring point selected within the field of view of the shooting means.

3. The system further includes a display means for displaying an image within the field of view of the aforementioned shooting means, The information processing device according to claim 2, characterized in that the position of the distance measuring point is a position selected by the user on the image within the field of view of the shooting means.

4. The information processing apparatus according to claim 1, characterized in that one or more of the aforementioned images are live view images.

5. Further comprising object recognition means for recognizing pre-learned objects, The generating means is Based on the focus position information, the object to be the main subject is determined from the recognition result of the object recognition means, The information processing apparatus according to claim 1, characterized in that it generates training data in which annotation information related to the main subject is added to one or more images acquired by the image acquisition means.

6. The information processing apparatus according to claim 5, characterized in that the object recognition means performs at least one of the processes of object detection or region segmentation.

7. The information processing apparatus according to claim 6, characterized in that the domain partitioning includes instance segmentation.

8. The aforementioned shooting means is equipped with multiple lenses with different angles of view, The plurality of lenses include a lens with a first angle of view and a lens with a second angle of view that is wider than the first angle of view. A display means for displaying a first image acquired by a lens with a first field of view and a second image acquired by a lens with a second field of view, The system further includes a determination means for determining whether the focus position information selected on the second image is included in the first image, The information processing apparatus according to claim 1, characterized in that the generation means generates the learning data by adding the annotation information to the first image and the second image based on the focus position information when the focus position information selected on the second image is included in the first image.

9. The information processing apparatus according to claim 1, characterized in that the generation means further adds the focus position information as annotation information.

10. The information processing apparatus according to claim 1, characterized in that the generation means generates the learning data to which the amount of defocus at the coordinates indicated by the focus position information is added as annotation information.

11. A control method for an information processing device that generates training data, A position information acquisition process to acquire focus position information of the shooting device, An image acquisition step in which one or more images are acquired based on the time when the focus position information is acquired, A defocus map acquisition step involves acquiring a defocus map that is time-synchronized with the time when the focus position information is acquired, wherein the amount of defocus is mapped on the input image data. A generation step of generating training data in which the amount of defocus of the main subject is determined based on the defocus map and the focus position information, and the amount of defocus of the main subject is added as annotation information to one or more acquired images. A control method for an information processing device, characterized by having the following features.

12. A program for causing a computer to function as an information processing device according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Distance calculation device, imaging apparatus, and distance calculation method

    JP2015034732A

  • Electronic device, moving body, distance calculation method, and computer program

    JP2023131720A

  • Labeling device and learning device

    JP7055259B2