Information processing device, method for controlling information processing device, and program
The information processing device automates the generation of learning data by acquiring focus position information and adding annotation to images, addressing the time-consuming manual processes in existing methods, thereby enhancing the efficiency of training data collection.
Patent Information
- Application Number
- JP2024087244
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2044-05-29
AI Technical Summary
Existing methods for generating training data for machine learning models require manual annotation and collection of image data, which is time-consuming.
An information processing device that automatically generates learning data by acquiring focus position information and adding annotation to images using a digital camera or smartphone, allowing for efficient collection of training data without manual input.
Reduces the effort required to generate training data by automating the annotation process, enabling faster and more efficient training data collection.
Smart Images

Figure 2025180116000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a control method for an information processing device, and a program. [Background technology]
[0002] In recent years, the use of AI (Artificial Intelligence) has been increasing in various fields. One example is supervised learning, which performs machine learning based on training data containing correct answers to generate inference models.
[0003] In supervised learning, obtaining a machine learning model with high generalization performance requires various inputs determined by the task to be solved, and training data with high-quality annotation information for those inputs. Generally, high-quality training data can be obtained from public datasets created for purposes such as competitions. However, if training data suitable for the purpose is not available, it is necessary to create a dataset yourself. Creating a dataset yourself requires annotation work, whereby annotation information is added manually to create training data. Creating a machine learning model with high generalization performance requires a huge amount of training data, and annotation takes a lot of time.
[0004] Patent Document 1 discloses a method for generating training data by semi-automatically or automatically annotating using a trained object detector. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 7055259 [Non-patent literature]
[0006] [Non-Patent Document 1] Kaiming He et al. "Mask R-CNN" Summary of the Invention [Problem to be solved by the invention]
[0007] However, the technology disclosed in Patent Document 1 requires users to prepare image data themselves in order to generate training data, which requires taking photographs and collecting image data, resulting in the problem that generating training data is time-consuming.
[0008] The present invention has been made in view of the above-mentioned problems, and aims to provide a technique for reducing the effort required to generate training data. [Means for solving the problem]
[0009] To achieve the above object, an information processing device according to one aspect of the present invention comprises: An information processing device for generating learning data, a position information acquisition means for acquiring focus position information of the photographing means; image acquisition means for acquiring one or more images based on the time when the focus position information is acquired; a generation means for generating the learning data by adding annotation information to the one or more images acquired by the image acquisition means based on the focus position information; The present invention is characterized by comprising: [Effects of the Invention]
[0010] According to the present invention, it is possible to reduce the effort required to generate training data. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a hardware configuration diagram of a training data generation device according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating the functional configuration of the training data generation device according to the first embodiment. [Figure 3]4 is a flowchart showing the procedure of a learning data generation process according to the first embodiment. [Figure 4] FIG. 3 is a diagram for explaining a learning data generation process according to the first embodiment. [Figure 5] 10 is a flowchart showing the procedure of a learning data generation process according to a first modification of the first embodiment. [Figure 6] FIG. 10 is a diagram for explaining a learning data generation process according to Modification 1 of the first embodiment. [Figure 7] 10 is a flowchart showing the procedure of a learning data generation process according to the second embodiment. [Figure 8] FIG. 10 is a diagram illustrating the functional configuration of a training data generation device according to a second embodiment. [Figure 9] FIG. 10 is a diagram for explaining a learning data generation process according to the second embodiment. [Figure 10] FIG. 10 is a diagram illustrating the functional configuration of a training data generation device according to a third embodiment. [Figure 11] 10 is a flowchart showing the procedure of a learning data generation process according to the third embodiment. [Figure 12] 10 is a flowchart showing the procedure of a learning data generation process according to the third embodiment. [Figure 13] FIG. 10 is a diagram for explaining a learning data generation process according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0013] (First embodiment) In this embodiment, a case where learning data is generated will be described using an example of photography with an interchangeable lens digital camera.
[0014] <Hardware configuration> FIG. 1 shows an example of the hardware configuration of an information processing device (learning data generation device 200) according to this embodiment. The CPU 100 is a central processing unit that performs calculations and logical decisions for various processes. A read-only memory (ROM) 110 stores a control program. The random access memory (RAM) 120 is used as a temporary storage area such as the main memory or work area of the CPU 100. The hard disk drive (HDD) 130 is a hard disk for storing electronic data and programs according to this embodiment. An external storage device may also be used to perform a similar function. Here, the external storage device can be realized, for example, by a medium (recording medium) and an external storage drive for accessing the medium. Known examples of such media include a flexible disk (FD), a CD-ROM, a DVD, a USB memory, an MO, and a flash memory. The external storage device may also be a server device connected via a network.
[0015] The input unit 140 is configured with a keyboard, touch panel, etc., and accepts input from a user. The display unit 150 is configured with a liquid crystal display, etc., and can display various data and processing results to a user. The training data generation device 200 can also communicate with other devices via the communication unit 160. It may receive instructions from a user from other devices via the communication unit 160, or output processing results to other devices. The training data generation device 200 can be configured using a general-purpose information processing device having the above configuration.
[0016] <Functional configuration> FIG. 2 is a diagram illustrating a training data generation device 200 according to this embodiment. The image capture unit 201 captures an image of a subject. For example, the image capture unit 201 can be configured with an image sensor such as a CMOS sensor. The focus position information acquisition unit 202 acquires focus position information input by the user received from the input unit 140. The focus position information input by the user is coordinates indicating the position within the displayed image (within the field of view of the image capture unit 201) of the ranging point (AF frame) to be focused on, selected by a touch operation (touch AF). Another method for selecting an AF frame may be a method in which the AF frame is selected by dragging it with a touch panel or joystick (multi-controller) and inputting it (touch-and-drag AF). Alternatively, the user may select the AF frame by operating a pointer with their line of sight and inputting it (gaze input).
[0017] The image acquisition unit 203 acquires images captured by the shooting unit 201. The acquired images include live view images, and the timing of image acquisition does not depend on the user's release operation. Annotation information is added to the images acquired by the image acquisition unit 203 based on focus position information acquired by the focus position information acquisition unit 202. The learning data storage unit 205 stores the learning data generated by the learning data generation unit 204.
[0018] <Processing> 3 is a flowchart showing the overall learning data generation process according to this embodiment. In S301, the image capturing unit 201 starts capturing an image. A live view image captured by the image capturing unit 201 is displayed on the display unit 150.
[0019] In S302, the photographing unit 201 activates a touch AF photographing mode, which is a method in which the user selects an AF frame during autofocus (AF). The input unit 140 waits for reception of a touch input from the user.
[0020] In S303, the input unit 140 receives a touch input from the user. The touch input is performed by the user touching a position on the touch panel of the input unit 140 at which the user wants to focus while checking the live view image displayed on the display unit 150.
[0021] In S304, the focus position information acquisition unit 202 acquires the coordinates (focus position information) of the touch input entered in S303. Furthermore, the image acquisition unit 203 saves a live view image of the imaging unit 201 at the moment when the touch input is made in S303. The timing for saving the live view image may be the moment when the touch input is confirmed by operating any button, such as when the shutter button is half-pressed after the touch input.
[0022] In S305, the image capturing unit 201 starts autofocus processing. The process of executing autofocus is started when the touch input is completed in S303. In S306, the image capturing unit 201 drives the focus lens by the autofocus processing of S305. In S307, the image capturing unit 201 determines that the focus has been achieved if it is determined that the focus has been achieved at the position specified in S303 based on the focus drive result of S306.
[0023] In S308, the input unit 140 determines whether the user has performed a touch input again. If this step is Yes, the process returns to S303. On the other hand, if this step is No, the process proceeds to S309. If the user checks the focus result of S307 and determines that the desired subject is in focus, the process proceeds to the processing of S309 without performing another touch input. If not, the process returns to the processing of S303 again, and the user performs a touch input by touching the position where the user wants to focus. Thereafter, the processes of S304 to S307 are repeated.
[0024] In S309, the learning data generation device 200 determines whether or not the user has performed a release operation. If a release operation has been performed, the image capturing unit 201 captures an image and the process proceeds to S310. If a release operation has not been performed, the process returns to S308, and it is again determined whether or not the user has performed a touch input again.
[0025] In S310, the learning data generation unit 204 generates learning data by adding annotation information to the image saved in S304 based on the focus position information saved in S304. For example, the learning data is generated by adding the saved focus position information to the image as annotation information. This completes the series of processes in FIG. 3.
[0026] 4(a) to 4(c) are explanatory diagrams of the learning data generation process in S310. FIG. 4(a) shows an image 400 saved in S304. Reference numeral 401 denotes an object, and 402 denotes a person. FIG. 4(b) is a diagram visualizing the coordinates of the touch input, which is the focus position information saved in S304. The mark 410 represents the coordinates of the touch input in touch AF saved in S304. FIG. 4(c) is a diagram superimposed on the image saved in S304, which is the focus position information saved in S304. Coordinates 422 of the position touched by the user are visualized as a mark on a person 421 in an image 420. For example, FIG. 4(c) shows that the coordinates indicating the pupils of the person 421 are determined as annotation information by the user's touch input.
[0027] As described above, according to this embodiment, it is possible to collect image data and annotation information for estimating focus position information by performing normal imaging using the imaging unit. Therefore, the user can collect image data and perform annotation without being aware of it, thereby reducing the effort required to generate learning data.
[0028] <Learning Method in the First Embodiment> The learning data generated by the learning data generation process in this embodiment can be used to train a machine learning model for estimating focus position information.
[0029] Specific examples of machine learning algorithms include nearest neighbor algorithms, naive Bayes algorithms, decision trees, and support vector machines. Deep learning, which uses a neural network to generate features and connection weighting coefficients for learning, is also an example. Any of the above algorithms that can be used can be used as appropriate and applied to this embodiment.
[0030] Here, learning using a neural network will be explained. Learning is performed using the images saved in S304 as input data. During learning, error detection processing and weight update processing are performed.
[0031] In the error detection process, the error between the output data output from the output layer of the neural network in response to the input data input to the input layer and the training data is obtained. At this time, the focus position information saved in S304 is used as the training data. The focus position information is, for example, the coordinates of the touch position in touch AF. In the error detection process, a loss function may be used to calculate the error between the output data from the neural network and the training data.
[0032] In the weight update process, the connection weighting coefficients between the nodes of the neural network are updated based on the error obtained in the error detection process so as to reduce the error. This weight update process can be performed by updating the connection weighting coefficients using, for example, the backpropagation method. The backpropagation method is a technique for adjusting the connection weighting coefficients between the nodes of each neural network so as to reduce the error.
[0033] The output data output as a result of learning is a machine learning model that estimates focus position information. The machine learning model refers to parameters such as weighting coefficients obtained by the weight update process.
[0034] In this way, by using an image as input data and focus position information as training data, it becomes possible to train a neural network that regresses focus position information according to the input image.
[0035] <Inference method in the first embodiment> Focus position information can be inferred using a machine learning model trained using the above-mentioned learning method. Here, we will explain the case where inference is performed by applying a trained machine learning model to an interchangeable lens digital camera.
[0036] The input data is a live view image captured by an image sensor such as a CMOS sensor. After capture, the live view image is used as input data for the trained machine learning model.
[0037] The output data is the inference result, and the estimated value of the focus position information is output. For example, the output data is the estimated coordinates of the touch position in touch AF, and represents coordinate information within the image, such as position (310, 452).
[0038] In this way, when a user takes a photo using an interchangeable-lens digital camera, the trained machine learning model can be used to estimate focus position information contained in the training data from within the live-view image. If a subject contained in the training data is present in the live-view image, the AF frame within the image can be automatically selected without requiring input operations such as touch AF, reducing the effort required when taking photos.
[0039] Furthermore, when photographing a subject included in the learning data, even if the subject moves quickly and it is difficult for the user to select an AF frame, the AF frame can be automatically selected from within the image, making it easier for the user to focus on the subject they want to focus on.
[0040] [Variation 1] On smartphones, the AF target is often selected by touching the screen. A modified example of the first embodiment will be described below, in which the first embodiment is applied to a smartphone. In this modified example, a case where multiple learning data are simultaneously generated in a single process for a smartphone having multiple lenses with different angles of view as a device with multiple image sensors will be described.
[0041] For example, if a smartphone has three lenses - a telephoto lens, a standard lens, and a wide-angle lens - each lens has its own image capture unit (such as an image sensor such as a CMOS sensor) and images are captured simultaneously through each lens. The focal length of a telephoto lens is longer than that of a standard lens, allowing for magnified capture of subjects further away. The focal length of a wide-angle lens is shorter than that of a standard lens, allowing for a wider range of images to be captured using the wide-angle lens than using the standard lens.
[0042] That is, the focal length decreases in the order of telephoto lens, standard lens, and wide-angle lens, and the photographic angle of view increases accordingly. Note that the telephoto lens, standard lens, and wide-angle lens are each assumed to have a zoom function and to be capable of continuously changing the photographic angle of view between the telephoto side and the wide-angle side. The telephoto lens, standard lens, and wide-angle lens may not only be a lens with a mechanism that optically multiplies the image to a predetermined magnification, but also a lens with a mechanism that allows the user to change the magnification.
[0043] In addition, multiple live view images captured by each lens can be viewed on the smartphone display, allowing users to capture images while checking the multiple live view images displayed on the display.
[0044] The processing in this modification is similar to the processing in Fig. 3 of the first embodiment, so a basic explanation will be omitted. The processing in S304 is different from that in the first embodiment, so it will be explained in detail using the flowchart in Fig. 5.
[0045] 5 is a flowchart showing the overall processing procedure of the learning data generation process in Modification 1. The processes in S501 to S503 are the same as the processes in S301 to S303 in FIG. 3, and therefore description thereof will be omitted.
[0046] In S304, the focus position information acquisition unit 202 acquires the coordinates (focus position information) of the touch input input in S303. In addition, the image acquisition unit 203 saves a live view image of the imaging unit 201 at the moment when the touch input was made in S303.
[0047] In S504, the input unit 140 determines whether the user has performed a touch input on the live view image of the lens with the longest focal length. If the result of this step is Yes, the process proceeds to S505. On the other hand, if the result of this step is No, the process proceeds to S512.
[0048] 6 shows the results of displaying a live view image 604 of the telephoto lens, a live view image 603 of the standard lens, and a live view image 602 of the wide-angle lens on a display 601 of a smartphone 600. For example, if the user touches the live view image 604 of the telephoto lens in S504, the process proceeds to S505.
[0049] Also, for example, if the user touches the live view image 602 of a wide-angle lens, the process proceeds to S512. When a touch input is performed on the live view image 602 of a wide-angle lens with a short focal length, if the touch position is at the edge of the screen, the focus position information may not fit within the image because the angle of view of the live view image of a lens with a long focal length is narrower. If the focus position information does not fit within the image, there is a possibility that learning data cannot be generated, so the determination process of S504 is performed.
[0050] In S505, the focus position information acquisition unit 202 acquires the coordinates of the touch input input in S503. In addition, the image acquisition unit 203 saves two or more live view images captured simultaneously with two or more of the telephoto lens, the standard lens, and the wide-angle lens.
[0051] The processes in S506 to S511 are similar to the processes in S305 to S310 in FIG. 3, and therefore will not be described here.
[0052] In S512, when a touch input is performed on a live view image of the wide-angle lens, the input unit 140 determines whether the touch position is within the image range of the telephoto lens. If this step is Yes, proceed to S505. On the other hand, if this step is No, end the processing.
[0053] This process takes into account the possibility that when a touch input is made on a live view image captured by a wide-angle lens with a short focal length, if the touch position is at the edge of the screen, the focus position information may not fit within the image because the angle of view of a live view image captured by a long-focal length lens is narrower.Even when a touch input is made on a live view image captured by a wide-angle lens, if the touch position is within the image range of the telephoto lens, multiple pieces of learning data can be generated simultaneously, so this determination process is performed.
[0054] Let the focal length of the wide-angle lens be f1, the focal length of the telephoto lens be f2, the horizontal resolution of the wide-angle lens's live view image be w1, and the vertical resolution be h1. In this case, in the live view image of the wide-angle lens, the resolution of the area corresponding to the angle of view of the live view image of the telephoto lens, based on the center of the image, is taken as Δw in the horizontal direction and Δh in the vertical direction. In this case, Δw and Δh are defined by the following equations (1) and (2).
[0055] Δw=w1·f1 / f2 ...(1) Δh=h1·f1 / f2 ...(2) Therefore, if the horizontal resolution of the wide-angle lens live view image is w2 and the vertical resolution is h2 so that touch input falls within the range of the telephoto lens live view image, the range of values for w2 and h2 is defined by the following equations (3) and (4).
[0056] (w1-Δw) / 2≦w2≦(w1+Δw) / 2 ...(3) (h1-Δh) / 2≦h2≦(h1+Δh) / 2 ...(4) For example, when f1 = 13, f2 = 26, w1 = 4032, and h1 = 3024, Δw = 2016 and Δh = 1512. The range of w2 values is from 1008 to 3024, and the range of h2 values is from 756 to 2268. However, because the lenses on a smartphone are not attached in exactly the same position, the center positions of the wide-angle lens and telephoto lens may be shifted by the amount of deviation in their attachment positions. Therefore, the range of w2 and h2 values may be corrected by the amount of deviation. If the coordinates touched on the live view image of the wide-angle lens are included in this range of w2 and h2, proceed to processing in S505; if not, end the learning data generation process.
[0057] In this way, in the case of a smartphone with multiple lenses with different angles of view, multiple pieces of training data with different angles of view can be collected simultaneously with a single process, making it possible to efficiently increase the variety of training data.
[0058] [Variation 2] In the above embodiment, only the live view image at the moment when the touch input is performed in S304 is saved, but one or more live view images may also be saved at other times, such as immediately after the focus determination in S307.
[0059] Increasing the timing at which live view images are saved in this way makes it possible, for example, when photographing a moving subject, to capture the subject's movements and changes that occur in the short time between touch input and focusing in the image. This also makes it possible to increase the variety of captured images, and increase the amount of learning data that can be acquired in a single process.
[0060] (Second embodiment) In this embodiment, an example will be described in which learning data is created to detect a main subject that exists at a focus position desired by a user.
[0061] <Functional configuration> 8 is a diagram illustrating a training data generation device 200 according to this embodiment. The components from the imaging unit 201 to the image acquisition unit 203 in this embodiment are the same as those from the imaging unit 201 to the image acquisition unit 203 in the first embodiment, and therefore will not be described further.
[0062] The learning data generation unit 204 according to this embodiment includes an object recognition unit 2041, and performs object recognition on the images acquired by the image acquisition unit 203. Here, the object recognition process is a process that includes at least one of object detection and region segmentation. The region segmentation may be semantic region segmentation that performs class classification on a pixel-by-pixel basis.
[0063] Object detection processing is the task of detecting specific objects (people, dogs, vehicles, etc.) or specific body parts (faces, eyes, heads, etc.) from within an image. Object detection processing estimates the position and size of the object to be detected, and outputs the estimation results as a rectangle (bounding box).
[0064] Region segmentation processing is a task of identifying a class for each pixel in an image and classifying the image into regions for each class. Region segmentation processing includes semantic segmentation, which classifies all pixels in an image, and instance segmentation, which not only classifies but also distinguishes and classifies each object in the image. Specific object detection and region segmentation techniques are described in Non-Patent Document 1, but a description of the techniques themselves will be omitted in this embodiment.
[0065] The object recognition unit 2041 in this embodiment is a neural network that has been trained in advance. The training model is not particularly limited as long as it can recognize objects. Data input to the object recognition unit 2041 is an image acquired by the image acquisition unit 203.
[0066] <Processing> The processing in this embodiment is similar to the processing in FIG. 3 described in the first embodiment, and therefore a basic description will be omitted. The processing in S310 differs from that in the first embodiment. FIG. 7 is a flowchart showing the procedure of the training data generation processing in this embodiment, and shows the detailed processing procedure according to this embodiment corresponding to the processing in S310. Here, the processing will be described using instance segmentation as an example, also with reference to FIGS. 9(a) to (d).
[0067] In S701, the object recognition unit 2041 performs object recognition on an image acquired by the image acquisition unit 203. FIG. 9(a) shows that an object 901 (tree) and a subject 902 (person) are captured in an image 900 acquired by the image acquisition unit 203. The object recognition unit 2041 performs instance segmentation on this image. FIG. 9(b) shows an output result 910 obtained by performing instance segmentation on the image 900. Class A 911 and class B 912 are classified on a pixel-by-pixel basis as the output result. Here, class A indicates, for example, a tree, and class B indicates, for example, a person.
[0068] In S702, the learning data generation unit 204 determines an object to be the main subject from the recognition result of the object recognition unit 2041, based on the coordinates acquired by the focus position information acquisition unit 202. FIG. 9(c) shows an output result 920 obtained by executing instance segmentation, with coordinates 921 of the focus position information acquired by the focus position information acquisition unit 202 superimposed on it. In the output result 920, the class to which the coordinates indicated by the coordinates 921 of the focus position information belong is determined as the main subject in the image. Here, class B 922 is determined as the main subject.
[0069] In S703, the training data generation unit 204 recognizes the information about the main subject determined in S702 as a correct label for object detection, and generates training data by assigning the correct label as annotation information related to the main subject. Figure 9(d) shows the result of creating training data for detecting the main subject in an image by surrounding, with a circumscribing rectangle 932, the pixels of class B 931 to which the main subject in the image determined in S702 belongs.
[0070] The class label of the created data may be the class of the main subject determined in S702. Alternatively, a plurality of created learning data may be accumulated and clustered to set the class label of the learning data.
[0071] According to this embodiment, it is possible to collect image data for detecting a main subject and annotation information by performing normal photographing using the photographing unit, thereby reducing the effort required for creating learning data, including the collection of image data.
[0072] <Learning Method in the Second Embodiment> The learning data generated by the learning data generation process in this embodiment can be used to train a machine learning model for object detection.
[0073] A specific example of a machine learning algorithm is deep learning, which uses a neural network to generate features and connection weighting coefficients for learning. Here, we will explain learning using a neural network.
[0074] Learning is performed using the image saved in S304 as input data. During learning, error detection processing and weight update processing are performed. During error detection processing, the error between the training data and the output data output from the output layer of the neural network according to the input data input to the input layer is obtained. At this time, the circumscribing rectangle 932 saved in S703 is used as the training data as a bounding box that estimates the position and size of the object. During error detection processing, a loss function may be used to calculate the error between the output data from the neural network and the training data.
[0075] In the weight update process, the connection weighting coefficients between the nodes of the neural network are updated based on the error obtained in the error detection process so as to reduce the error. In this weight update process, the connection weighting coefficients are updated using, for example, the backpropagation method. The backpropagation method is a technique for adjusting the connection weighting coefficients between the nodes of each neural network so as to reduce the error.
[0076] The output data from the learning process is a machine learning model that estimates a bounding box that defines the position and size of an object. The machine learning model refers to parameters such as weighting coefficients obtained by the weight update process.
[0077] In this way, by using an image as input data and a bounding box that estimates the position and size of an object as training data, it becomes possible to train a neural network to detect objects from input images.
[0078] <Inference Method in Second Embodiment 2> The machine learning model trained using the above-mentioned learning method can be used to perform inference for object detection. Here, we will explain the case where inference is performed by applying the trained machine learning model to an interchangeable lens digital camera.
[0079] As input data, live view images captured by an image sensor such as a CMOS sensor are used. After being acquired, the live view images are used as input data for the trained machine learning model.
[0080] The output data is the inference result, which is a bounding box that estimates the position and size of the object. The output data represents the position and size according to the spatial direction of the image, for example, (310, 452, 105, 40) in the order of horizontal coordinate, vertical coordinate, horizontal size, and vertical size.
[0081] In this way, when a user takes a photo using an interchangeable lens digital camera, the trained machine learning model can be used to estimate the position and size of objects included in the training data from the live view image.When a user wants to photograph a subject included in the training data, the subject can be automatically detected and focused on without the user having to select an AF frame using input operations such as touch AF in the image, reducing the effort required for input when shooting.
[0082] (Third embodiment) In this embodiment, an example will be described in which learning data is created to estimate the defocus amount at a user-desired focus position from a defocus map.
[0083] First, a defocus map will be described. A defocus map is a map in which the defocus amount is mapped at multiple locations on input image data. The defocus amount is expressed in units of Fδ. A known split-pupil phase difference detection method can be used to generate a defocus map. For example, a correlation calculation is performed on each of the image signals of two different pupil regions to calculate the phase difference, which is the amount of shift between images with different parallax (hereinafter referred to as image A and image B), i.e., the image shift amount between image A and image B. Then, the defocus map is generated by calculating the defocus amount based on the calculated phase difference (image shift amount) between image A and image B.
[0084] <Functional configuration> 10 is a diagram illustrating the functional configuration of the training data generation device 200 according to this embodiment. The components from the imaging unit 201 to the training data storage unit 205 according to this embodiment are the same as those from the imaging unit 201 to the training data storage unit 205 in FIG. 2, and therefore a basic description thereof will be omitted.
[0085] The defocus map acquisition unit 206 acquires a defocus map that is synchronized in time with the image acquired by the image acquisition unit 203 .
[0086] <Processing> 11 is a flowchart showing the procedure of the learning data generation process in this embodiment. The process performed in S1105 and the process performed in S1111, which corresponds to the process in S310, are different from those in the first embodiment, and therefore the differences will be mainly described.
[0087] The processes of S1101 to S1104 are similar to the processes of S301 to S304 in Fig. 3, and therefore descriptions thereof will be omitted. In S1105, the defocus map acquisition unit 206 acquires and stores a defocus map that is temporally synchronized with the image acquired in S1104. The processes of S1106 to S1110 are similar to the processes of S305 to S309 in Fig. 3, and therefore descriptions thereof will be omitted. In S1111, the learning data generation unit 204 generates learning data from the defocus map acquired in S1105.
[0088] The processing of S1111 will now be described in detail with reference to Fig. 12. In S1201, the learning data generation unit 204 reads the defocus map acquired in S1105. In S1202, the learning data generation unit 204 determines the defocus amount of the main subject from the defocus map acquired in S1105.
[0089] 13(a) to 13(c) are explanatory diagrams of the process of determining the defocus amount of the object to be the main subject based on the focus position information from the values on the defocus map read in S1201 and generating learning data. Fig. 13(a) shows that a subject 1301 (person) is photographed in an image 1300 acquired in S1104.
[0090] 13(b) is an image obtained by superimposing a defocus map 1311 read in S1201 on an image 1310 acquired in S1104. In the defocus map 1311, the darker the color, the greater the amount of defocus, and the defocus amount in the area where a subject 1312 is present is greater than that of the background.
[0091] 13(c) shows a result 1320 in which the focus position information 1322 saved in S1104 is visualized and superimposed on a defocus map 1321. The defocus amount at the coordinates indicated by the focus position information 1322 on the defocus map 1321 is determined as annotation information. At this time, the determined defocus amount is, for example, 0.1Fδ.
[0092] In S1203, the learning data generation unit 204 generates learning data using the defocus amount determined in S1202 as annotation information. In this embodiment, learning data is generated by adding the defocus amount of the main subject determined in S1202 as annotation information to the image saved in S1104.
[0093] According to this embodiment, it is possible to collect image data and annotation information for defocus amount estimation by capturing images using the capturing unit, thereby reducing the effort required for creating learning data, including the collection of image data.
[0094] <Learning Method in the Third Embodiment> The learning data generated by the learning data generation process in this embodiment can be used to train a machine learning model for estimating the defocus amount.
[0095] A specific example of a machine learning algorithm is deep learning, which uses a neural network to generate features and connection weighting coefficients for learning. Here, we will explain learning using a neural network.
[0096] Learning is performed using the image saved in S304 as input data. Learning involves error detection processing and weight update processing. In the error detection processing, the error between the training data and the output data output from the output layer of the neural network in response to the input data input to the input layer is obtained. At this time, the defocus amount saved in S1203 is used as the training data. In the error detection processing, a loss function may be used to calculate the error between the output data from the neural network and the training data.
[0097] In the weight update process, the connection weighting coefficients between the nodes of the neural network are updated based on the error obtained in the error detection process so as to reduce the error. In this weight update process, the connection weighting coefficients are updated using, for example, the backpropagation method. The backpropagation method is a technique for adjusting the connection weighting coefficients between the nodes of each neural network so as to reduce the error.
[0098] The output data output as a result of the learning is a machine learning model that estimates the defocus amount of the main subject. The machine learning model refers to parameters such as weighting coefficients obtained by the weight update process.
[0099] In this way, by using an image as input data and the defocus amount as training data, it becomes possible to train a neural network that estimates the defocus amount according to the input image.
[0100] <Inference Method in the Third Embodiment> The defocus amount of the main subject can be inferred using a machine learning model trained using the above-mentioned learning method. Here, we will explain the case where inference is performed by applying the trained machine learning model to an interchangeable lens digital camera.
[0101] The input data is a live view image captured by an image sensor such as a CMOS sensor. After the live view image is acquired, it is used as input data for the trained machine learning model. The output data is the inference result, and an estimated value of the defocus amount of the main subject is output. The output data is, for example, 0.13Fδ.
[0102] In this way, when a user takes a photograph using an interchangeable-lens digital camera, the trained machine learning model can be used to estimate the defocus amount of the main subject included in the training data from the live view image. When a user wants to photograph a main subject included in the training data, the system can automatically identify the main subject and estimate its defocus amount without the user having to select an AF frame by performing an input operation such as touch AF within the image. This reduces the effort required for shooting, and also makes it easier to focus on the main subject included in the training data by performing AF using the estimated defocus amount.
[0103] The disclosure of this specification includes the following information processing device, control method for an information processing device, and program. (Item 1) An information processing device for generating learning data, a position information acquisition means for acquiring focus position information of the photographing means; image acquisition means for acquiring one or more images based on the time when the focus position information is acquired; a generation means for generating the learning data by adding annotation information to the one or more images acquired by the image acquisition means based on the focus position information; An information processing device comprising: (Item 2) 2. The information processing device according to item 1, wherein the focus position information is coordinates indicating the position of a distance measurement point selected within the angle of view of the imaging means. (Item 3) further comprising a display means for displaying an image within the angle of view of the photographing means, 3. The information processing device according to item 2, wherein the position of the distance measurement point is a position selected by a user on an image within the angle of view of the imaging means. (Item 4) 4. The information processing device according to any one of items 1 to 3, wherein the one or more images are live view images. (Item 5) Further comprising an object recognition means for recognizing a learned object, The generating means determining an object to be a main subject from the recognition result of the object recognition means based on the focus position information; 5. The information processing device according to any one of items 1 to 4, wherein the learning data is generated by adding the annotation information related to the main subject to the one or more images acquired by the image acquisition means. (Item 6) 6. The information processing device according to item 5, wherein the object recognition means executes at least one of object detection and area division processing. (Item 7) 7. The information processing device according to item 6, wherein the region division includes instance segmentation. (Item 8) 8. The information processing device according to any one of items 1 to 7, wherein the one or more images are images taken at a time point close in time to the time point at which the focus position information is acquired. (Item 9) The photographing means includes a plurality of lenses with different angles of view, the plurality of lenses include a lens with a first angle of view and a lens with a second angle of view wider than the first angle of view, a display means for displaying a first image acquired by the lens with the first angle of view and a second image acquired by the lens with the second angle of view; a determination unit that determines whether the focus position information selected on the second image is included in the first image, 3. The information processing device according to item 1 or 2, wherein, when the focus position information selected on the second image is included in the first image, the generation means generates the learning data by adding the annotation information to the first image and the second image based on the focus position information. (Item 10) 10. The information processing device according to any one of items 1 to 9, wherein the generating means assigns the focus position information as the annotation information. (Item 11) further comprising a map acquisition means for acquiring a defocus map at a time point close in time to the time point at which the focus position information is acquired, The generating means determining a defocus amount of an object to be a main subject based on the defocus map acquired by the map acquisition means; 4. The information processing device according to any one of items 1 to 3, characterized in that the learning data is generated by adding the defocus amount of the main subject to the one or more images acquired by the image acquisition means as the annotation information. (Item 12) Item 12. The information processing device according to item 11, wherein the generating means generates the learning data to which the defocus amount at the coordinates indicated by the focus position information is added as the annotation information. (Item 13) A control method for an information processing device that generates learning data, comprising: a position information acquisition step of acquiring focus position information of an imaging means; an image acquisition step of acquiring one or more images based on the time when the focus position information was acquired; a generating step of generating the learning data by adding annotation information to the one or more images acquired by the image acquiring step based on the focus position information; 1. A method for controlling an information processing device, comprising: (Item 14) A program for causing a computer to function as the information processing device according to any one of items 1 to 12.
[0104] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0105] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0106] 140: Input unit, 150: Display unit, 200: Learning data generation device, 201: Photography unit, 202: Focus position information acquisition unit, 203: Image acquisition unit, 204: Learning data generation unit, 205: Learning data storage unit
Claims
1. An information processing device for generating learning data, a position information acquisition means for acquiring focus position information of the photographing means; image acquisition means for acquiring one or more images based on the time when the focus position information is acquired; a generation means for generating the learning data by adding annotation information to the one or more images acquired by the image acquisition means based on the focus position information; An information processing device comprising:
2. 2. The information processing apparatus according to claim 1, wherein the focus position information is coordinates indicating the position of a distance measurement point selected within the angle of view of the image capturing means.
3. further comprising a display means for displaying an image within the angle of view of the photographing means, 3. The information processing apparatus according to claim 2, wherein the position of the distance measurement point is a position selected by a user on an image within the angle of view of the image capturing means.
4. The information processing apparatus according to claim 1 , wherein the one or more images are live view images.
5. Further comprising an object recognition means for recognizing a learned object, The generating means determining an object to be a main subject from the recognition result of the object recognition means based on the focus position information; 2. The information processing apparatus according to claim 1, wherein the learning data is generated by adding the annotation information related to the main subject to the one or more images acquired by the image acquisition means.
6. 6. The information processing apparatus according to claim 5, wherein the object recognition means executes at least one of object detection and area segmentation.
7. The information processing apparatus according to claim 6 , wherein the region division includes instance segmentation.
8. The information processing apparatus according to claim 1 , wherein the one or more images are images taken at a time point close in time to the time point at which the focus position information was acquired.
9. The photographing means includes a plurality of lenses with different angles of view, the plurality of lenses include a lens having a first angle of view and a lens having a second angle of view wider than the first angle of view, a display means for displaying a first image acquired by the lens having the first angle of view and a second image acquired by the lens having the second angle of view; a determination unit that determines whether the focus position information selected on the second image is included in the first image, The information processing device according to claim 1, characterized in that, when the focus position information selected on the second image is included in the first image, the generation means generates the learning data by adding the annotation information to the first image and the second image based on the focus position information.
10. 2. The information processing apparatus according to claim 1, wherein the generating means assigns the focus position information as the annotation information.
11. further comprising a map acquisition means for acquiring a defocus map at a time point close in time to the time point at which the focus position information is acquired, The generating means determining a defocus amount of an object to be a main subject based on the defocus map acquired by the map acquisition means; 2. The information processing apparatus according to claim 1, wherein the learning data is generated by adding the defocus amount of the main subject to the one or more images acquired by the image acquisition means as the annotation information.
12. 12. The information processing apparatus according to claim 11, wherein the generating means generates the learning data to which the defocus amount at the coordinates indicated by the focus position information is added as the annotation information.
13. A control method for an information processing device that generates learning data, comprising: a position information acquisition step of acquiring focus position information of an imaging means; an image acquisition step of acquiring one or more images based on the time when the focus position information was acquired; a generating step of generating the learning data by adding annotation information to the one or more images acquired by the image acquiring step based on the focus position information; 1. A method for controlling an information processing device, comprising:
14. A program for causing a computer to function as the information processing device according to any one of claims 1 to 12.
Citation Information
Patent Citations
Distance calculation device, imaging apparatus, and distance calculation method
JP2015034732A
Electronic device, moving body, distance calculation method, and computer program
JP2023131720A
Labeling device and learning device
JP7055259B2