Image generation device and method, program, and storage medium
The image generating device addresses the lack of control in virtual photography by incorporating acquisition and correction means, enhancing the enjoyment of virtual photography experiences.
Patent Information
- Application Number
- JP2024139155
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Existing virtual photography technologies lack control over various corrections, making the shooting experience less enjoyable and disrupting the enjoyment of taking pictures in virtual space.
An image generating device that includes a generating means for generating a virtual space image, acquisition means for photographer operation information, camera information, and subject information, and a modification means for correcting the virtual space image based on acquired information.
Enables an enjoyable photography experience in virtual space by allowing appropriate control and correction adjustments based on real-time conditions.
Smart Images

Figure 2026036508000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image generation device that generates a virtual space image. [Background technology]
[0002] Patent Document 1 discloses a technology that allows a photographer to have a photography experience without going to the shooting location by combining an image of a virtual space created by a 3D model with an image of real space captured by a camera and displaying the composite image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-78908 Summary of the Invention [Problem to be solved by the invention]
[0004] However, while the technology described in Patent Document 1 allows the photographer to shoot a virtual image with freely selected composition, angle of view, etc., it does not provide control over various corrections to make the shooting experience more enjoyable.
[0005] The virtual space is entirely defined by data, and it is possible to know the characteristics of the subject itself, as well as its past and future movements. This makes it possible to realize advanced control such as prediction and correction, allowing users to take photos without mistakes, or even fully automatically.
[0006] However, achieving perfect control and correction at all times can feel strange compared to the experience of taking pictures in real space, and can even get in the way of enjoying the act of taking pictures. Therefore, in order to enjoy the experience of taking pictures in virtual space, it is necessary to appropriately judge and switch between whether to use correction and how much correction to use depending on the conditions.
[0007] The present invention has been made in view of the above-mentioned problems, and an object of the present invention is to provide an image generating device that allows an enjoyable photography experience when taking pictures in a virtual space. [Means for solving the problem]
[0008] The image generating device of the present invention is characterized by comprising a generating means for generating a virtual space image, a first acquisition means for acquiring operation information of the photographer on the camera, a second acquisition means for acquiring information about the camera, a third acquisition means for acquiring information about the subject, and a modification means for modifying corrections to the virtual space image captured by the camera based on at least one of the information acquired by the first to third acquisition means. [Effects of the Invention]
[0009] According to the present invention, it is possible to enjoy the experience of taking pictures in a virtual space. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing the configuration of an imaging system according to a first embodiment of the present invention. [Figure 2] FIG. 1 is a block diagram showing the configuration of a camera according to the present invention. [Figure 3] FIG. 2 is a diagram showing a pixel array in the camera of the first embodiment. [Figure 4] 3A and 3B are a plan view and a cross-sectional view of a pixel according to the first embodiment. [Figure 5] FIG. 2 is a view showing a focus detection area according to the first embodiment. [Figure 6] FIG. 2 is a block diagram showing the hardware configuration of an external computing device according to the first embodiment. [Figure 7] FIG. 2 is a block diagram showing the functional configuration of an external calculation device according to the first embodiment. [Figure 8] 5 is a flowchart for explaining the processing of real space shooting and virtual space shooting in the first embodiment. [Figure 9] 6 is a flowchart illustrating a real space imaging process according to the first embodiment. [Figure 10] 5 is a flowchart illustrating an imaging process according to the first embodiment. [Figure 11] 5 is a flowchart illustrating subject tracking AF processing in the first embodiment. [Figure 12] 5 is a flowchart illustrating subject detection and tracking processing in the first embodiment. [Figure 13] 6 is a flowchart illustrating virtual space shooting processing in the first embodiment. [Figure 14] 3A and 3B are diagrams for explaining information of a camera lens information storage device, a camera / lens, and an external computing device in the first embodiment. [Figure 15] 4 is a flowchart for explaining generation and output of an image of a virtual space in the first embodiment. [Figure 16] 5 is a flowchart of a virtual subject tracking process according to the first embodiment. [Figure 17] 5 is a flowchart for acquiring photographing difficulty level information in the first embodiment. [Figure 18] FIG. 10 is a diagram illustrating correction related to framing in the first embodiment. [Figure 19] 5A to 5C are diagrams illustrating correction related to zooming in the first embodiment. [Figure 20] 5A to 5C are diagrams illustrating correction related to focusing in the first embodiment. [Figure 21] 6 is a flowchart illustrating a defocus amount processing process according to the first embodiment. [Figure 22] FIG. 4 is a graph showing a virtual defocus amount calculation in the first embodiment. [Figure 23] FIG. 4 is a diagram showing an example of a virtual defocus map according to the first embodiment. [Figure 24] 6 is a flowchart illustrating a subroutine for virtual space photography in the first embodiment. [Figure 25] 5A to 5C are diagrams illustrating virtual space photography that reflects operation information according to the first embodiment. [Figure 26] 6 is a flowchart illustrating a subroutine of a viewpoint movement process in the first embodiment. [Figure 27] FIG. 3 is a diagram showing an example of viewpoint movement in the first embodiment. [Figure 28] 1 is a flowchart illustrating the operation of a camera when capturing a virtual space in the first embodiment. [Figure 29] 5 is a flowchart for explaining playback of a photographed result and evaluation of a photographed image in the first embodiment. [Figure 30] FIG. 4 is an explanatory diagram of a defocus map display in the first embodiment. [Figure 31] FIG. 4 is an explanatory diagram of a focus degree display of a series of captured images in the first embodiment. [Figure 32] FIG. 4 is an explanatory diagram of a setting change list according to the first embodiment. [Figure 33] FIG. 10 is an explanatory diagram of the best setting display after evaluation in the first embodiment. [Figure 34] 10 is a flowchart for explaining generation and output of an image of a virtual space in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.
[0012] First Embodiment FIG. 1 is a diagram showing the configuration of an imaging system 10 including an imaging device according to a first embodiment of the present invention, an external computing device (information processing device), and a camera / lens information storage device.
[0013] 1, an imaging device (camera) 100 has a function of capturing an image of a subject existing in real space, a function of instructing the capturing of an image of a subject existing in virtual space, and a function of displaying the captured image. Camera 100 also functions as a tactile reproduction device.
[0014] The external computing device 1000 is connected to the camera 100 via wire or wirelessly so as to be able to exchange information, and is configured to include a virtual space reproduction device 1100 and a virtual image generation device 1200. The virtual space reproduction device 1100 places a subject as an object in a virtual space whose position and shape change from moment to moment within a set virtual space (background space). The virtual image generation device 1200 acquires, from the camera 100, setting information for the camera and lens, control information, operation information for the operation members, position information including the shooting direction, and the like. Furthermore, using the information acquired from the camera 100, it acquires related information from a camera / lens information storage device 2000.
[0015] The camera / lens information storage device 2000 may be a server on a network such as a cloud, or may be provided within the external computing device 1000.
[0016] The virtual image generation device 1200 uses the information acquired as described above to generate (capture) an image from the virtual space constructed by the virtual space reproduction device 1100. The image generated here may be a two-dimensional image or a three-dimensional image containing information that allows for stereoscopic display. In FIG. 1, the external computing device 1000 is configured to reproduce the virtual space and generate the image, but the camera 100 may be configured to realize these functions internally.
[0017] Fig. 2 is a diagram showing the configuration of a camera 100 as an imaging apparatus according to a first embodiment of the present invention. In Fig. 2, a first lens group 101 is arranged closest to the subject (front side) in the imaging optical system as an imaging optical system, and is held so as to be movable in the direction of the optical axis. An aperture 102 adjusts the light amount by adjusting its aperture diameter. A second lens group 103 moves in the direction of the optical axis together with the aperture 102, and performs magnification change (zooming) together with the first lens group 101, which moves in the direction of the optical axis.
[0018] The third lens group (focus lens) 105 moves in the optical axis direction to adjust the focus. The optical low-pass filter 108 is an optical element for reducing false colors and moiré in captured images. The first lens group 101, the aperture 102, the second lens group 103, the third lens group 105, and the optical low-pass filter 108 constitute an imaging optical system.
[0019] The zoom actuator 111 rotates a cam barrel (not shown) around the optical axis, causing cams provided on the cam barrel to move the first lens group 101 and the second lens group 103 in the optical axis direction, thereby varying the magnification. The diaphragm actuator 112 drives a plurality of light-shielding blades (not shown) in opening and closing directions to adjust the amount of light from the diaphragm 102. The focus actuator 114 moves the third lens group 105 in the optical axis direction, thereby adjusting the focus.
[0020] The focus drive circuit 126 drives the focus actuator 114 in response to a focus drive command from the camera CPU 121, and moves the third lens group 105 in the optical axis direction. The aperture drive circuit 128 drives the aperture actuator 112 in response to an aperture drive command from the camera CPU 121. The zoom drive circuit 129 drives the zoom actuator 111 in response to a zoom operation by the user.
[0021] In this embodiment, the interchangeable lens having the photographic optical system, actuators 111, 112, 114, and drive circuits 126, 128, 129 is configured to be detachable from the camera body using a mount M that enables electrical and mechanical connection. However, the photographic optical system, actuators 111, 112, 114, and drive circuits 126, 128, 129 may be configured to be integral with the camera body including the image sensor 107.
[0022] The electronic flash 115 has a light-emitting element such as a xenon tube or an LED, and emits light to illuminate the subject. The AF assist light emitter 116 has a light-emitting element such as an LED, and projects an image of a mask with a predetermined aperture pattern onto the subject via a projection lens, thereby improving focus detection performance for dark or low-contrast subjects. The electronic flash control circuit 122 controls the electronic flash 115 to turn on in synchronization with the imaging operation. The assist light drive circuit 123 controls the AF assist light emitter 116 to turn on in synchronization with the focus detection operation.
[0023] The camera CPU 121 is responsible for various controls in the camera 100. The camera CPU 121 has a calculation unit, ROM, RAM, an A / D converter, a D / A converter, a communication interface circuit, etc. The camera CPU 121 drives various circuits within the camera 100 in accordance with a computer program stored in the ROM, and controls a series of operations such as AF, image capture, image processing, and recording. The camera CPU 121 functions as an image processing device.
[0024] The image sensor 107 is composed of a two-dimensional CMOS photosensor including multiple pixels and its peripheral circuitry, and is disposed on the imaging plane of the photographing optical system. The image sensor 107 photoelectrically converts the subject image formed by the photographing optical system. The image sensor drive circuit 124 controls the operation of the image sensor 107, and also A / D converts the analog signal generated by the photoelectric conversion and sends the digital signal to the camera CPU 121.
[0025] The shutter 106 has a focal plane shutter configuration and is driven by a shutter drive circuit built into the shutter 106 based on instructions from the camera CPU 121. The image sensor 107 is shielded from light while a signal from the image sensor 107 is being read out. Furthermore, when exposure is being performed, the focal plane shutter is opened and a photographing light beam is guided to the image sensor 107.
[0026] The image processing circuit 125 performs predetermined image processing on image data stored in the RAM in the camera CPU 121. The image processing performed by the image processing circuit 125 includes, but is not limited to, so-called development processing such as white balance adjustment processing, color interpolation (demosaic) processing, and gamma correction processing, as well as signal format conversion processing and scaling processing. Furthermore, the image processing circuit 125 determines the main subject based on the posture information of the subject and the position information of objects unique to the scene (hereinafter referred to as unique objects). The results of the determination processing may be used for other image processing (e.g., white balance adjustment processing). The image processing circuit 125 saves the processed image data, joint position information of each subject, position and size information of the unique objects, center of gravity position information of the subject determined to be the main subject, position information of the face and eyes, etc. in the RAM in the camera CPU 121.
[0027] The display (display means) 131 includes a display element such as an LCD, and displays information about the imaging mode of the camera 100, a preview image before imaging, a confirmation image after imaging, an index for the focus detection area, and an in-focus image. The operation switch group 132 includes a main (power) switch, a release (photography trigger) switch, a zoom operation switch, an imaging mode selection switch, and the like, and is operated by the user. The flash memory 133 records captured images. The flash memory 133 is detachable from the camera 100.
[0028] The subject detection unit 140, which serves as a subject detection means, performs subject detection based on dictionary data generated by machine learning. In this embodiment, the subject detection unit 140 uses dictionary data for each subject to detect multiple types of subjects. Each dictionary data is, for example, data in which the characteristics of the corresponding subject are registered. The subject detection unit 140 performs subject detection by sequentially switching between dictionary data for each subject. The dictionary data for each subject is stored in a dictionary data storage unit (a ROM in the camera CPU 121). Therefore, multiple dictionary data are stored in the dictionary data storage unit. The camera CPU 121 determines which dictionary data from the multiple dictionary data to use for subject detection based on pre-set subject priorities and settings of the imaging device.
[0029] The video input unit 141 inputs the generated image when shooting (generating an image) in a virtual space, and the camera CPU 121 performs processes such as displaying the input image on the display 131 and storing it in the flash memory 133. The information output unit 142 outputs various information to the external computing device 1000 when shooting in a virtual space. The camera operation information to be output includes a release operation that issues shooting instructions, lens zooming, focus operation, etc. The camera setting information to be output includes setting information related to the continuous shooting mode, autofocus, photometry, exposure condition settings, image generation, lens control, etc. The camera control information to be output includes information related to correction values and thresholds used in various algorithms used for shooting and image generation. The information output also indicates the camera position and shooting direction. Details will be described later.
[0030] Examples of dictionary data for subject detection include dictionary data for detecting "people" as subjects, dictionary data for detecting "animals," dictionary data for detecting "vehicles," etc. Furthermore, dictionary data for detecting "the whole person" and dictionary data for detecting "the person's face" may be stored separately in the dictionary data storage unit.
[0031] In this embodiment, the object detection unit 140 is configured by a machine-learned convolutional neural network (CNN) and estimates the position of an object included in image data, etc. The object detection unit 140 may be realized by a circuit specialized for estimation processing using a GPU (graphics processing unit) or a CNN.
[0032] The machine learning of the CNN may be performed by any method. For example, a predetermined computer such as a server may perform the machine learning of the CNN, and the camera 100 may acquire the trained CNN from the predetermined computer. For example, the CNN of the subject detection unit 140 may be trained by the predetermined computer performing supervised learning using training image data as input and the position of the subject corresponding to the training image data as training data. In this way, a trained CNN is generated. The training of the CNN may be performed by the camera 100 or the image processing device described above.
[0033] Next, the image array of the image sensor 107 will be described with reference to Fig. 3. Fig. 3 shows the pixel array of the image sensor 107, which is in a range of 4 pixel columns x 4 pixel rows, as viewed from the optical axis direction (z direction).
[0034] Each pixel unit 200 includes four imaging pixels arranged in two rows and two columns. Arranging a large number of pixel units 200 on the image sensor 107 enables photoelectric conversion of a two-dimensional subject image. An imaging pixel 200R having R (red) spectral sensitivity (hereinafter referred to as R pixel) is arranged in the upper left of each pixel unit 200, and imaging pixels 200G having G (green) spectral sensitivity (hereinafter referred to as G pixel) are arranged in the upper right and lower left. Furthermore, an imaging pixel 200B having B (blue) spectral sensitivity (hereinafter referred to as B pixel) is arranged in the lower right. Each imaging pixel includes a first focus detection pixel 201 and a second focus detection pixel 202, which are divided in the horizontal direction (x direction).
[0035] In the image sensor 107 of this embodiment, the pixel pitch P of the imaging pixels is 4 μm, the number of imaging pixels N is 5,575 horizontal (x) columns × 3,725 vertical (y) rows = approximately 20.75 million pixels, the pixel pitch PAF of the focus detection pixels is 2 μm, and the number of focus detection pixels NAF is 11,150 horizontal columns × 3,725 vertical rows = approximately 41.5 million pixels.
[0036] In this embodiment, the case where each imaging pixel is divided into two in the horizontal direction is described, but it may also be divided in the vertical direction. Furthermore, the image sensor 107 in this embodiment has a plurality of imaging pixels, each of which includes a first and second focus detection pixel, but the imaging pixel and the first and second focus detection pixels may be provided as separate pixels. For example, the first and second focus detection pixels may be discretely arranged among the plurality of imaging pixels.
[0037] Fig. 4(a) shows one imaging pixel (200R, 200G, 200B) as viewed from the light receiving surface side (+z direction) of the image sensor 107. Fig. 4(b) shows the aa cross section of the imaging pixel in Fig. 4(a) as viewed from the -y direction. As shown in Fig. 4(b), one imaging pixel is provided with one microlens 305 for collecting incident light.
[0038] Each imaging pixel is provided with photoelectric conversion units 301 and 302 that are divided into N parts in the x direction (two parts in this embodiment). The photoelectric conversion units 301 and 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively. The centers of gravity of the photoelectric conversion units 301 and 302 are decentered on the -x side and +x side, respectively, with respect to the optical axis of the microlens 305.
[0039] An R, G, or B color filter 306 is provided between the microlens 305 and the photoelectric conversion units 301 and 302 in each imaging pixel. The spectral transmittance of the color filter may be changed for each photoelectric conversion unit, or the color filter may be omitted.
[0040] Light incident on the imaging pixels from the imaging optical system is collected by the microlens 305, dispersed by the color filter 306, and then received by the photoelectric conversion units 301 and 302, where it is photoelectrically converted. The camera 100 having the image sensor 107 shown in FIGS. 3 and 4 can perform so-called phase-difference focus detection, which detects the phase difference from a pair of signal sequences obtained by splitting a light beam passing through the imaging optical system using known technology (e.g., JP 2023-95509 A). Phase-difference focus detection can detect the amount of defocus in a specified area within the imaging range, including the direction. A detailed description will be omitted.
[0041] Next, the focus detection area, which is an area of the image sensor 107 where a pair of signal sequences for detecting a phase difference is acquired, will be described with reference to Fig. 5. In Fig. 5, A(n,m) indicates the nth focus detection area in the x direction and the mth focus detection area in the y direction out of multiple focus detection areas (three in the x direction and three in the y direction, for a total of nine) set in the effective pixel area 300 of the image sensor 107. A pair of signal sequences is generated from multiple pixels included in the focus detection area A(n,m). I(n,m) indicates an index that indicates the position of the focus detection area A(n,m) on the display 131.
[0042] Note that the nine focus detection areas shown in FIG. 5 are merely examples, and the number, positions, and sizes of the focus detection areas are not limited. For example, one or more focus detection areas may be set within a predetermined range centered on a position specified by the user or the subject position detected by the subject detector. In this embodiment, the focus detection areas are arranged to obtain focus detection results with higher resolution when acquiring a defocus map, which will be described later. For example, a total of 9,600 focus detection areas are arranged on the image sensor, divided into 120 horizontal and 80 vertical sections.
[0043] 6 is a block diagram showing an example of the hardware configuration of the external calculation device 1000. The external calculation device 1000 includes a CPU 1001, a RAM 1003, a ROM 1002, a storage unit 1004, an input interface 1005, an output interface 1006, and a system bus 1007. The input interface 1005 is connected to the camera 100 and the camera / lens information storage device 2000. The output interface 1006 is connected to the camera 100.
[0044] The CPU 1001 is a processor that comprehensively controls each component of the external computing device 1000. The RAM 1003 is a memory that functions as the main memory and work area of the CPU 1001. The ROM 1002 is a memory that stores programs and the like used for processing within the external computing device 1000. The CPU 1001 uses the RAM 1003 as a work area and executes programs stored in the ROM 1002 to perform various processes, which will be described later.
[0045] The storage unit 1004 is a storage device that stores image data used for processing in the external computing device 1000, parameters (i.e., setting values) for the processing, etc. The storage unit 1004 may be an HDD, an optical disk drive, a flash memory, or the like.
[0046] The input interface 1005 is, for example, a serial bus interface such as USB or IEEE1394. The external computing device 1000 can acquire the above-mentioned various information from the camera 100 via the input interface 1005. The output interface 1006 is, for example, a video output terminal such as DVI or HDMI (registered trademark). The external computing device 1000 can output image data processed by the external computing device 1000 to the display 131 of the camera 100 via the output interface 1006. It can also output images to be recorded in the flash memory 133 of the camera 100. Note that the external computing device 1000 may include components other than those described above, but as these are not the main focus of the present invention, detailed description thereof will be omitted.
[0047] Next, a virtual image generation process performed by the external computing device 1000 using the hardware configuration of Fig. 6 will be described with reference to Fig. 7. Fig. 7 is a block diagram showing the functional configuration of the external computing device 1000. In this embodiment, the CPU 1001 executes a program stored in the ROM 1002 to realize each block shown in Fig. 7. However, it is not necessary for the CPU 1001 to execute all functions, and each part of the external computing device 1000 may be provided with a processing circuit that executes each function.
[0048] First, the virtual space reproduction device 1100 will be described. The foreground object acquisition unit 1102 acquires a 3D object of a person, such as a stage performer, as a foreground subject stored in the foreground object storage unit 1101. A 3D object is 3D shape data describing information indicating shape and color, and is composed of a textured mesh model, a 3D point cloud with colored points, or the like. Note that the 3D object does not have to be colored. The objects stored in the foreground object storage unit 1101 can be various 3D objects, such as people of different races, genders, and ages, various animals, and moving objects such as cars. The foreground object acquisition unit 1102 may acquire multiple foreground objects, rather than just one. The 3D object also has subject information such as velocity, acceleration, angular velocity, angular acceleration, size, and contrast. Alternatively, a pre-trained model that estimates a 3D model from an image may be used to generate a 3D object from a captured image of a subject the photographer wants to capture in virtual space photography, and size and contrast information may also be stored. Furthermore, by using a plurality of time-series images, it may be possible to generate and store information on the velocity, acceleration, angular velocity, and angular acceleration of a three-dimensional object.
[0049] As another method, a three-dimensional object may be generated and stored from captured images taken using multiple imaging devices with different viewpoints. The imaging area is captured from multiple directions using multiple imaging devices. This imaging area may be, for example, an indoor imaging studio or a theatrical stage. The multiple imaging devices are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple imaging devices do not need to be installed around the entire periphery of the imaging area; they may be installed only in certain directions of the imaging area depending on installation space restrictions, etc. The number of imaging devices can be set using various methods. For example, if the imaging area is a soccer stadium, approximately 30 imaging devices may be installed around the stadium. Furthermore, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0050] For each of the multiple image capture devices, a parameter set may be described, including parameters representing the three-dimensional position, parameters representing the orientation of the image capture device in the pan, tilt, and roll directions, and the size of the field of view (angle of view) and resolution of the image capture device. The information included in the parameter set is calculated in advance using a known camera calibration procedure and stored in an appropriate storage device (e.g., the foreground object storage unit 1101). That is, points in multiple images captured by the multiple image capture devices are associated with each other and calculated using geometric calculations. Note that the content of the information included in the parameter set is not limited to the above. For example, the parameter set may have multiple parameter sets corresponding to multiple frames constituting a video captured by the image capture device, and the information may indicate the position and orientation of the image capture device at each of multiple consecutive points in time.
[0051] The foreground object acquisition unit 1102 generates a 3D object of a person, such as a stage performer, who is the foreground subject, based on the multiple viewpoint images and parameter set received from the imaging device, for example, according to the method described in JP 2017-211827 A.
[0052] Similarly, a background object acquisition unit 1105 acquires a 3D object that serves as a space for placing a foreground object, such as a stage or stadium that serves as a background stored in a background object storage unit 1104. Possible background objects stored in the background object storage unit 1104 include 3D objects of various spaces, such as a large concert hall or soccer stadium, or a small indoor room. The background object may be made using design data such as CAD, or shape and color data scanned with a laser scanner or the like. Alternatively, the background object may be generated from a group of images from multiple viewpoints using computer vision techniques such as Structure from Motion.
[0053] The object composition unit 1103 places the foreground object within the space of the acquired background object. The information about the foreground object acquired by the object composition unit 1103 can include three-dimensional models of the subject at multiple times, corresponding to the shape and color of the subject at multiple times. When placing the foreground object, the foreground object is placed so that it does not float with respect to the ground contained in the background object, except for interference between objects or actions such as jumping. The foreground object may be placed according to the object placement information (position, orientation) held by the background object, or may be placed based on instructions from an external source such as a user.
[0054] Next, the virtual image generation device 1200 will be described.
[0055] The viewpoint information acquisition unit 1201 acquires virtual viewpoint parameters including the position and direction (pan, tilt, roll) of the virtual viewpoint in the virtual space. The virtual viewpoint parameters may be set to initial values, registered values, previous history positions, etc. in the virtual space, or may be set by user instructions.
[0056] The camera lens information acquisition unit 1202 acquires information about the camera and lens used for virtual space photography from the camera / lens information storage device 2000 or the camera 100. Details of the information will be described later. The camera lens information update unit 1203 acquires and updates the camera lens information, which is updated over time, each time.
[0057] An operation information acquisition unit 1205 acquires camera and lens operation information from the camera 100. Details of the information will be described later.
[0058] The viewpoint information, camera lens information, and operation information are input to image correction amount calculation unit 1206, which calculates the amount of image correction. Image correction amount calculation unit 1206 calculates the amount of image correction using information obtained from shooting difficulty calculation unit 1261 and photographer's intention extraction unit 1262. Details of the processing will be described later.
[0059] The display image generation unit 1204 performs rendering and generates a virtual image using the foreground and background object information, virtual viewpoint information, and camera lens information acquired from the object synthesis unit 1103. The generated virtual image is output to the camera 100 and displayed on the display 131 of the camera 100. It is also recorded in the flash memory 133 of the camera 100 and the storage unit 1004 of the external computing device 1000.
[0060] (Photography processing) The flowchart in Figure 8 shows the process of having the camera 100 of this embodiment capture real space and virtual space images. Specifically, it shows the process from the pre-capture operation of displaying an image on the display 131 of the camera 100 to capturing a still image. The camera CPU 121, which is a computer, executes this process in accordance with a computer program. In the following description, S means step.
[0061] First, in S1, the camera CPU 121 causes the display device 131 to start displaying a menu for setting, live view images of real space and virtual space, and the like. The generation of live view images to be played back will be described later. At the first startup or in response to a user operation, a menu setting screen is displayed on the display device 131, allowing the user to select whether to capture real space or virtual space. The display content may be determined based on the history from the previous startup. If capture of real space or virtual space has already been set, the live view display that started beforehand is continued.
[0062] In S2, the camera CPU 121 determines whether or not to perform virtual space shooting based on a user instruction or the previous history. If the answer is Yes in S2, the process proceeds to S1000, where virtual space shooting processing is performed. On the other hand, if the answer is No in S2, the process proceeds to S10, where real space shooting processing is performed. After the processing of S10 or S1000 is completed, the process proceeds to S3.
[0063] In S3, the camera CPU 121 determines whether or not the main switch included in the operation switch group 132 has been turned off. If the main switch has been turned off, the camera CPU 121 ends this process, and if the main switch has not been turned off, the process returns to S1.
[0064] (Real space photography processing) The flowchart in Fig. 9 shows the real space imaging process shown in S10 in Fig. 8. Specifically, it shows the process from the pre-imaging operation of displaying a live view image on the display 131 of the camera 100 to the operation of capturing a still image. The camera CPU 121, which is a computer, executes this process in accordance with a computer program. In the following description, S means step.
[0065] First, in S11, the camera CPU 121 causes the image sensor drive circuit 124 to drive the image sensor 107 and acquires image data from the image sensor 107. Thereafter, the camera CPU 121 acquires, from the acquired image data, paired focus detection signals from paired focus detection pixels included in each focus detection area shown in FIG. 5. The camera CPU 121 also adds the paired focus detection signals of all effective pixels of the image sensor 107 to generate an image signal, and causes the image processing circuit 125 to perform image processing on the image signal (image data) to acquire image data. Note that if the image sensor pixels and the focus detection pixels are provided separately, the camera CPU 121 acquires image data by performing interpolation processing on the focus detection pixels.
[0066] In S12, the camera CPU 121 causes the image processing circuit 125 to generate a live view image from the image data obtained in S11 and displays it on the display 131. The live view image is a reduced image matched to the resolution of the display 131, and the user can adjust the image composition, exposure conditions, etc. while viewing this image. Therefore, the camera CPU 121 performs exposure adjustment based on the photometric value obtained from the image data and displays it on the display 131. Exposure adjustment is achieved by appropriately adjusting the exposure time, opening and closing the aperture of the photographing lens, and adjusting the gain of the image sensor output.
[0067] Next, in S13, the camera CPU 121 determines whether or not a switch Sw1, which instructs the start of an image capture preparation operation, has been turned on by half-pressing a release switch included in the operation switch group 132. If Sw1 is not turned on, the camera CPU 121 repeats the determination of S13 to monitor the timing at which Sw1 is turned on. On the other hand, if Sw1 is turned on, the camera CPU 121 proceeds to S400 and performs subject tracking autofocus (AF) processing. Here, it performs predictive AF processing to detect the subject area from the focus detection signal obtained, set the focus detection area, and reduce the effect of the time lag between the focus detection processing and the image capture processing for recording. Details will be described later.
[0068] In S15, the camera CPU 121 determines whether or not the switch Sw2, which instructs the start of an imaging operation, has been turned on by fully pressing the release switch. If Sw2 is not turned on, the camera CPU 121 returns to S13. On the other hand, if Sw2 is turned on, the camera CPU 121 proceeds to S300 and executes an imaging subroutine. Details of the imaging subroutine will be described later. When the imaging subroutine ends, this process ends.
[0069] In this embodiment, the subject detection process and AF process are performed after it is detected in S3 that Sw1 is turned on, but the timing of these processes is not limited to this. By performing the subject tracking AF process in S400 before Sw1 is turned on, it is possible to eliminate the need for the photographer to take preparatory actions before shooting.
[0070] Next, the imaging subroutine executed by the camera CPU 121 in S300 of FIG. 9 will be described with reference to the flowchart shown in FIG.
[0071] In S301, the camera CPU 121 performs exposure control processing and determines the imaging conditions (shutter speed, aperture value, imaging sensitivity, etc.) This exposure control processing can be performed using brightness information acquired from image data of the live view image.
[0072] Then, the camera CPU 121 transmits the determined aperture value to the aperture drive circuit 128 to drive the aperture 102. The camera CPU 121 also transmits the determined shutter speed to the shutter 106 to open the focal plane shutter. Furthermore, the camera CPU 121 causes the image sensor 107 to accumulate charge during the exposure period via the image sensor drive circuit 124.
[0073] In S302, the camera CPU 121, which has performed the exposure control processing, causes the image sensor drive circuit 124 to read out all pixels of the image signal resulting from still image capture from the image sensor 107. The camera CPU 121 also causes the image sensor drive circuit 124 to read out one of a pair of focus detection signals from a focus detection area (focus target area) within the image sensor 107. The focus detection signal read out at this time is used to detect the focus state of the image during image playback, which will be described later. One of the pair of focus detection signals can be subtracted from the image capture signal to obtain the other focus detection signal.
[0074] In S303, the camera CPU 121 causes the image processing circuit 125 to perform defective pixel correction processing on the imaging data that was read out and A / D converted in S302.
[0075] In S304, the camera CPU 121 causes the image processing circuit 125 to perform image processing and encoding processes such as demosaic (color interpolation), white balance, gamma correction (tone correction), color conversion, and edge enhancement on the image data after the defective pixel correction process.
[0076] In S305, the camera CPU 121 records the still image data obtained as image data by the image processing and encoding processing in S304 and one of the focus detection signals read out in S302 in the memory 133 as an image data file.
[0077] In S306, the camera CPU 121 associates the camera characteristic information as characteristic information of the camera 100 with the still image data recorded in S305 and records it in the memory 133 and in the memory within the camera CPU 121. The camera characteristic information includes, for example, the following information. Imaging conditions (aperture value, shutter speed, imaging sensitivity, etc.) Information about image processing performed by the image processing circuit 125 Information about the light-receiving sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 107 Information about vignetting of the imaging light beam within the camera 100 Information on the distance from the mounting surface of the photographing optical system in the camera 100 to the image sensor 107 Information about the manufacturing tolerances of the Camera 100 Information regarding the light sensitivity distribution of the imaging pixels and focus detection pixels (hereinafter simply referred to as light sensitivity distribution information) is information regarding the sensitivity of the image sensor 107 according to the distance (position) from the optical axis. This light sensitivity distribution information depends on the microlens 305 and the photoelectric conversion units 301 and 302, and therefore may be information regarding these. Furthermore, the light sensitivity distribution information may be information regarding changes in sensitivity with respect to the angle of incidence of light.
[0078] In S307, the camera CPU 121 associates the lens characteristic information as characteristic information of the photographing optical system with the still image data recorded in S305 and records it in the memory 133 and in a memory within the camera CPU 121. The lens characteristic information includes, for example, information on the exit pupil, information on a frame such as a lens barrel that blocks light beams, information on the focal length and F-number at the time of image capture, information on aberrations of the photographing optical system, information on manufacturing errors of the photographing optical system, and information on the position of the focus lens 105 at the time of image capture (subject distance).
[0079] In S308, the camera CPU 121 records image-related information as information related to the still image data in the memory 133 and in a memory within the camera CPU 121. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.
[0080] In S309, the camera CPU 121 displays a preview of the captured image on the display 131. This allows the user to easily check the captured image.
[0081] When the process of S309 is completed, the camera CPU 121 ends this imaging subroutine.
[0082] Next, the subject tracking AF processing subroutine executed by the camera CPU 121 in S400 of FIG. 9 will be described with reference to the flowchart shown in FIG.
[0083] In S401, the camera CPU 121 calculates the amount of image shift between pairs of focus detection signals obtained in each of the multiple focus detection areas acquired in S11, calculates the amount of defocus for each focus detection area from the image shift amount, and acquires a defocus map. As described above, in this embodiment, the group of focus detection results obtained from the focus detection areas arranged on the image sensor in 120 horizontal and 80 vertical divisions, for a total of 9,600 points, is called a defocus map.
[0084] In S402, the camera CPU 121 performs subject detection and tracking processing. The subject detection processing is performed by the above-mentioned subject detection unit 140. Since subject detection may not be possible depending on the state of the obtained image, in such cases, tracking processing is performed using other means such as template matching to estimate the position of the subject. Details will be described later.
[0085] In S403, the camera CPU 121, which serves as a local area selection means, sets a focus detection area using the subject detection area information obtained in S402. The camera CPU 121 acquires information such as the subject's position, size, and reliability as subject detection area information obtained as the output of the subject detection and tracking process performed in S402. The focus detection area can be set by selecting a focus detection result that is highly reliable and indicates a subject that is relatively close, based on the results of the focus detection area within the area set as the subject detection area. Alternatively, the focus detection area can be set by relocating a focus detection area within the area set as the subject detection area, acquiring image data and focus detection signals again, and similarly selecting a focus detection result.
[0086] In S404, the camera CPU 121 acquires the focus detection result of the set focus detection area. The focus detection result acquired here may be a focus detection result closest to the desired area selected from the focus detection results calculated in S401, or a focus detection signal corresponding to a newly set focus detection area may be used to calculate the defocus amount. The focus detection area for calculating the defocus amount is not limited to one, and multiple focus detection areas may be arranged around the area to calculate the defocus amount.
[0087] In S405, the camera CPU 121 performs predictive AF processing using the defocus amount obtained in S404 and multiple defocus amounts, which are time-series data on the timing of past focus detection. This processing is necessary when there is a time lag between the timing of focus detection and the timing of exposure for the captured image. AF control is performed by predicting the position of the subject in the optical axis direction at the timing of exposure for the captured image, which is a predetermined time after the timing of focus detection. The subject's image plane position is predicted by performing multivariate analysis (e.g., the least squares method) using historical data on the subject's image plane position and time to find an equation for a prediction curve. The predicted image plane position wp of the subject can be calculated by substituting the timing of exposure for the captured image into the equation for the prediction curve thus found.
[0088] Furthermore, three-dimensional positions may be predicted in addition to the optical axis direction. For example, consider an XYZ vector, where the XY direction is on the screen and the Z direction is the optical axis direction. In this case, the position of the subject at the timing of exposure of the captured image may be predicted from the XY position of the subject obtained in the subject detection and tracking process of S402 and time-series data of the Z direction position based on the defocus amount obtained in S405. Furthermore, the position of the subject may be predicted from time-series data of the joint positions of the person who is the subject.
[0089] The above predictions make it possible to estimate the position of each object even if the ball or person is hidden during the shot, or if some of the person's joints become invisible. Predictions are made not only for the main subject, but also for multiple detected subjects. By performing predictive AF processing on multiple subjects, when the main subject is switched, there is no need to re-accumulate the defocus amount history for the new main subject, and predictive AF can be continued without any time loss.
[0090] In S405, the predicted AF processing result is used to calculate the amount of focus lens drive, and the focus actuator 114 is driven in accordance with a focus drive command from the camera CPU 121, thereby performing focus adjustment processing by moving the third lens group 105 in the optical axis direction.
[0091] When the process of S405 ends, the camera CPU 121 ends the subroutine of the subject tracking AF process and proceeds to the process of S15 in FIG.
[0092] Next, the subject detection and tracking process subroutine executed by the camera CPU 121 in S402 of FIG. 11 will be described with reference to the flowchart shown in FIG.
[0093] In S421, the camera CPU 121 sets dictionary data according to the type of subject to be detected based on the data detected from the image data acquired in S12 of FIG. 9. Based on the preset priority of the subject and the settings of the imaging device, dictionary data to be used in this process is selected from multiple dictionary data stored in the dictionary data storage unit. For example, multiple dictionary data are stored, each classified by subject, such as "people," "vehicles," and "animals." In this embodiment, one or more dictionary data may be selected. When one dictionary data is selected, subjects that can be detected using one dictionary data can be detected repeatedly at a high frequency. On the other hand, when multiple dictionary data are selected, the dictionary data can be set sequentially according to the priority of the subject to be detected, thereby allowing subjects to be detected sequentially.
[0094] In S422, subject detection unit 140 performs subject detection using the dictionary data set in step S421 and the image data read out in S12 of Fig. 9 as an input image. At this time, subject detection unit 140 outputs information such as the position, size, and reliability of the detected subject. At this time, camera CPU 121 may cause display unit 131 to display the above information output by subject detection unit 140.
[0095] In S422, multiple regions of the subject are detected hierarchically from the image data. For example, if "person" or "animal" is set as dictionary data, multiple organs such as the "whole body" region, the "face" region, and the "eye" region are detected. Local regions such as a person's eyes or face are regions where it is desired to adjust the focus and exposure as the subject, but they may not be detectable due to surrounding obstacles or the direction of the face. Even in such cases, the whole body is detected to continue robustly detecting the subject and detect the subject hierarchically. Similarly, if a "vehicle" such as a motorcycle is set as dictionary data, the entire body including the driver and vehicle body and the helmet (head) as a local region are detected hierarchically.
[0096] In S423, the camera CPU 121 performs a known template matching process using the subject detection area obtained in S422 as a template. Using the multiple images obtained in S12, the subject detection area obtained in a past image is used as a template to search for a similar area in the most recently obtained image. As is well known, any information may be used for template matching, such as brightness information, color histogram information, or feature point information such as corners and edges. Various matching methods and template update methods are possible, and any of these methods may be used. The tracking process performed in S423 is performed to achieve stable subject detection and tracking when a subject is not detected in S422 by detecting an area similar to past subject detection data from the most recently obtained image data.
[0097] When the processing of S423 ends, the camera CPU 121 ends the subject detection and tracking processing subroutine and proceeds to S403 in FIG.
[0098] (Virtual space photography processing) The flowchart in FIG. 13 illustrates the operation of the virtual space imaging process shown in S1000 in FIG. 8. The virtual space imaging process is a process of generating an image by extracting information from a virtual space, which changes over time, at a specific moment. Imaging in a virtual space does not require a physical imaging optical system or imaging element, but for ease of explanation, the same terms as those used for imaging in real space are used. For example, image generation in a virtual space is expressed as imaging or capturing an image. More specifically, FIG. 13 illustrates the process from the pre-imaging operation of displaying a video image in the virtual space as a live view image on the display 131 of the camera 100 to capturing a still image. The camera CPU 121 and CPU 1001, which are computers, execute this process according to a computer program. In the following description, unless otherwise specified, the camera CPU 121 and CPU 1001 are the subjects of the operations.
[0099] In S1001, settings related to virtual space photography, such as the virtual photography space where photography will be performed and the equipment to be used, are made. In setting the virtual photography space, as described in FIG. 7, foreground objects are placed in appropriate positions relative to background objects. Information regarding the position and shape of the foreground objects that change over time is also obtained. There may be one or more foreground objects placed.
[0100] Furthermore, when setting up virtual space photography, the camera CPU 121 outputs model information of the equipment being operated to the external computing device 1000 as the settings for the camera and lens used in virtual space photography. When using equipment for virtual space photography that is different from the model being actually operated, the photographer sets a unique symbol for that model, and the camera CPU 121 outputs the set information to the external computing device 1000. This allows the operator (photographer) to experience photography using a camera or lens that they do not actually own. For example, while operating a camera equipped with a so-called wide-angle lens with a short focal length, they can experience photography in the virtual space with a telephoto lens with a long focal length. The operation of a model different from the model being actually operated may not be limited to the lens, but may also be the camera, or both. This allows for a more flexible photography experience that is not dependent on the weight or size of the equipment. Similarly, by using a camera that they do not actually own, they can experience the camera's new functions, improved performance due to newly installed algorithms, and the like.
[0101] In addition, when setting up virtual space photography, initial values for the viewpoint position and direction (the camera's position and direction in the virtual space) when starting virtual space photography are set. A position at an appropriate distance based on information such as the type of foreground object described above can be set as the initial value. Alternatively, a photography position previously set within a background object can be set as the initial value.
[0102] In S1002, the CPU 1001 instructs the camera 100 to initialize the camera driving unit. Details will be described later.
[0103] In S1003, the CPU 1001 acquires camera information, lens information, and camera and lens operation information from the camera 100 and the camera / lens information storage device 2000.
[0104] (Explanation of communication between camera / lens and external computing device) An example of information communicated between the camera / lens and the external computing device 1000 will be described with reference to the table in FIG.
[0105] The following describes the information in the camera / lens information storage device 2000, the camera / lens information, and the external computing device 1000. Camera information and lens information are recorded in the camera / lens information storage device 2000. The camera / lens information storage device 2000 also acquires and stores information from the camera / lens.
[0106] The camera information includes display resolution, recorded image resolution, image sensor size, focus frame mode, autofocus (AF) mode (e.g., one-shot, servo), and continuous shooting settings previously acquired from the camera 100. It also includes camera settings such as the shooting difficulty setting set by the photographer, camera algorithm information such as the AF algorithm, autoexposure (AE) and continuous shooting drive sequence, and camera detection information such as temperature. It also includes image sensor characteristic information such as S / N information for each ISO sensitivity, and shading correction values that represent image sensor signal characteristic correction and light intensity unevenness. It also includes a defocus conversion coefficient that converts image shift amount into defocus amount, focus-related correction information, information on best focus position correction that corrects the difference between the focus detection result and the best image plane position, and focus-related correction information that is defocus error information. It also includes general information such as the camera / lens model name and firmware versions of various algorithms.
[0107] Lens information includes focal length range, current value, and resolution, F-number range, increments, and current value, focus lens drive range and current focus information, and focus control information related to focus drive control characteristics. It also includes sensitivity for converting focus lens drive into image plane movement amount, image stabilization information related to image stabilization range, current value, and correction resolution, and image stabilization control information related to image stabilization control characteristics. It also includes aperture control information related to aperture drive control characteristics, frame information (position, diameter) related to vignetting, peripheral light falloff information, distance information related to focus lens position and distance, and information related to the point spread function.
[0108] The camera / lens generates operational information when the photographer operates the camera / lens body. This operational information includes framing, zooming, focusing, shutter release, and other button operations. This operational information is sent to the external computing device 1000 and reflected in the generation of a virtual image.
[0109] The external computing device 1000 acquires camera information, lens information, and operation information, and generates a display image, a recording image, subject information that is shooting difficulty information, a virtual defocus amount, and various shooting-related information.
[0110] The acquired lens information includes focal length, F-number, information about the settable range and current position of the focus lens, mechanical controllability of the lens, and the amount of movement (sensitivity) of the imaging plane associated with movement of the focus lens. It also includes frame information (position, diameter) related to vignetting, information about peripheral light falloff, and shooting distance (distance to the subject at which the focus is achieved).
[0111] The acquired camera information also includes general information such as the model name, firmware version, resolution of EVF images and still images, and image sensor size. It also includes camera setting information such as the AF frame setting for setting the AF range, AF mode settings such as one-shot and servo AF, and continuous shooting mode settings such as continuous shooting speed. The camera setting information also includes shooting difficulty information (shooting difficulty setting) set by the photographer.
[0112] The correction values for the signals used for autofocus focus detection include a correction value for signal characteristics that depend on the characteristics of the image sensor 107, a shading correction value that represents unevenness in the amount of light, and a defocus conversion coefficient that converts the phase difference between a pair of signals into a defocus amount.The correction values also include a best focus correction value that corrects the deviation between the focus detection result and the best image plane position.
[0113] The camera information also includes, as characteristic information of the image sensor 107, S / N information of the signal for each ISO sensitivity, various algorithm information such as continuous shooting sequence and photometry when taking pictures with the camera, and autofocus-related algorithm information such as AF frame selection and predictive AF. Some of this information changes as the camera is operated, so information that may change is periodically acquired from S1003 onwards.
[0114] The camera and lens operation information also includes information about the amount and speed of operations for panning, zooming, and focusing the camera held by the photographer, and information about button press operations such as release operations, which are shooting instructions.
[0115] 13, in S2000, an image of the virtual space is generated based on the settings made up to this point, and output to the camera 100. Details of this process will be described later.
[0116] In S1005, the camera CPU 121 acquires the image output in S2000 and displays it on the display device 131. The displayed image is thereafter updated at, for example, 60 fps. By combining this with the camera operation information and lens operation information, an image in which the range of the displayed virtual space varies in accordance with the panning and zooming operations of the camera is updated on the display device 131.
[0117] In S1006, it is determined whether or not the mode is one in which the viewpoint in the virtual space being observed through the display 131 is moved (viewpoint movement mode). If the mode is one in which viewpoint movement processing is performed, the answer is Yes in S1006 and the process proceeds to S3000. In S3000, viewpoint movement processing is performed to determine the position and shooting direction of the camera in the virtual space. Details will be described later. When S3000 is completed, the process returns to S2000.
[0118] 9, the camera CPU 121 determines whether or not the switch Sw1, which instructs the start of an image capture preparation operation, has been turned on by half-pressing the release switch included in the operation switch group 132. If Sw1 is not turned on, the camera CPU 121 returns to S2000 and repeats the determination to monitor the timing at which Sw1 is turned on. On the other hand, if Sw1 is turned on, the camera CPU 121 advances the process to S4000 and performs virtual subject tracking processing.
[0119] In S4000, various corrections are made to the generated image in accordance with the photographer's operation, the movement of the subject, etc., making it possible to photograph a subject that is at least a part of a foreground object. Details of this process will be described later.
[0120] In S1008, similar to S15 in Fig. 9, it is determined whether or not the switch Sw2, which instructs the start of imaging operation, has been turned on by fully pressing the release switch. If Sw2 is not turned on, the camera CPU 121 returns to S2000. On the other hand, if Sw2 is turned on, the camera CPU 121 proceeds to S5000, where the virtual space photography subroutine is executed. Details of the virtual space photography subroutine will be described later. When the virtual space photography subroutine ends, this process ends.
[0121] (Virtual space image generation and output subroutine) Next, the subroutine for generating and outputting an image of a virtual space executed by the external computing device 1000 in S2000 of FIG. 13 will be described with reference to the flowchart shown in FIG.
[0122] In S2001, the CPU 1001 acquires a foreground object. The photographer first selects the type of subject (for example, a person, an animal, a vehicle, etc.) that he or she wishes to photograph using the virtual space reproduction device 1100. Next, the photographer selects the shape and color of the subject, and also selects the type of movement (speed, direction of movement, etc.) that the subject will have. The user interface for selection may be configured to display information stored in the foreground object storage unit 1101 of the virtual space reproduction device 1100 on the camera's display 131, and allow the photographer to operate and select. As described above, a foreground object, which is a three-dimensional model of the subject, is acquired by multiple methods.
[0123] In step S2002, the CPU 1001 acquires a background object, which is a three-dimensional model of an object other than the subject, by using a plurality of methods as described above.
[0124] In S2003, the CPU 1001 composites the objects. Object composite is a process for composite- ing the foreground object and background object described above. Object composite is performed by determining how to position the background object in three-dimensional space and where in three-dimensional space to place the foreground object relative to the background object. First, the background object is positioned in three-dimensional space, and the photographer selects where in three-dimensional space to place the foreground object. The foreground object can be placed only in positions that are possible relative to the background object (for example, outside the interior of the background object) based on the three-dimensional model of the background object and the coordinates at which it is positioned in three-dimensional space. The photographer selects the position at which to place the foreground object from the three-dimensional space of the background object. Object composite is performed in this manner.
[0125] In S2004, the CPU 1001 acquires camera / lens information. The camera / lens information acquired here is information for generating and outputting an image in a virtual space, which will be described later. Specifically, the camera information includes the display resolution of the display 131 used for display, the size and number of pixels of the camera's image sensor, etc. Furthermore, the lens information includes the focal length range and current value, the aperture range and current value, the focus lens range and current value, information on peripheral light falloff and point spread function, etc.
[0126] In S2005, the CPU 1001 acquires viewpoint position information. In order to generate a virtual image (described later), virtual viewpoint information in a three-dimensional space is acquired. The virtual viewpoint information may be a predetermined value as an initial value, or may be a virtual viewpoint changed by viewpoint movement processing (described later) in S3000.
[0127] In step S2006, the CPU 1001 acquires camera / lens operation information, which includes information on framing, zooming, focusing, release operation, and other button operations.
[0128] In S2007, the CPU 1001 acquires the image correction amount. The image correction amount is the amount of correction related to framing, zooming, and focus. Details will be explained in the subflow of the virtual subject tracking process in S4000, which will be described later. Here, a predetermined initial value of the image correction amount is acquired.
[0129] In S2008, the CPU 1001 generates a display image in a virtual space. It renders an image based on the foreground and background objects arranged in three-dimensional space and the viewpoint position information. The range of the display image is determined based on the focal length information (as lens information) and the image sensor size, display resolution, camera settings, and framing and zooming information (operation information). The range is then determined by modifying the range based on the image correction amount. Furthermore, a display image with a modified aperture value and defocus amount is generated based on the lens information, such as aperture value information, peripheral light falloff information, point spread function information, focus lens position information, and focusing-related image correction amount information. The display image differs from the recorded image described below and is not recorded. Therefore, after correctly determining the display image range, the display image may be generated simply with less information than the recorded image, omitting some of the information, such as focus lens position information and peripheral light falloff information.
[0130] In S2009, the CPU 1001 outputs the display image generated in S2008. The output image is transmitted from the external computing device 1000 to the camera 100 and displayed on the display device 131.
[0131] In S2010, the CPU 1001 saves the video-related information. The video-related information includes subject information, shooting-related information, virtual defocus amount, and AF log information. The video-related information is temporarily saved in the RAM 1003 of the external computing device 1000 and recorded as video-related information in a virtual space shooting subroutine, which will be described later in detail.
[0132] This completes the virtual space image generation and output process of S2000 in Fig. 13. In this embodiment, the virtual space image generation and output are performed by the external calculation device 1000, but the virtual space image generation and output process may also be performed within the camera 100.
[0133] (Subroutine for virtual subject tracking processing) The virtual subject tracking process in S4000 of FIG. 13 will be described with reference to the flowchart of FIG.
[0134] Of the various types of corrections described below, framing corrections are made when the photographer's framing is off and the subject they want to photograph is off-screen or cut off, so that the subject fits neatly within the camera's display angle of view.
[0135] Zooming correction is performed to ensure that the subject is displayed at an appropriate size on the screen when the photographer's zooming (lens focal length) is off and the subject they want to photograph goes outside the frame or is too small. In addition to single-timing zoom correction, correction is also performed when shooting an approaching subject while keeping it within the frame at a constant size (zooming photography). If the photographer's zooming results in jerky successive images, zooming correction is performed that takes into account the timing before and after the change in focal length to achieve a smooth change in focal length.
[0136] In focusing corrections, for example, during autofocus, the camera's tracking algorithm (tracking limit performance) is used to correct blurred focus when the subject is moving at high speed or when there are large changes in speed. This allows for images to be obtained that are in focus and have reduced blur. In manual focus, the camera also corrects for out-of-focus images caused by the photographer's focusing operation. It also corrects for other phenomena, such as when the photographer's framing is off and the focus lens is moved to the background, causing the focus to shift away from the subject.
[0137] First, in S4001, the CPU 1001 acquires camera / lens information from the camera lens information acquisition unit 1202. The camera lens information acquired here is information for determining whether correction is on or off, acquiring subject difficulty information, and calculating the amount of correction, which will be described later. Specifically, the camera information includes camera settings related to correction, such as a shooting difficulty setting for shooting set by the photographer. The lens information includes information related to the focal length, focus lens position, and the on / off setting of the image stabilization switch.
[0138] In S4002, the CPU 1001 acquires setting information related to correction using the camera lens information acquired in S4001. The setting information related to correction includes, for example, setting information such as ON / OFF of correction settings in the camera, mode settings such as difficulty settings, and ON / OFF setting of the image stabilization switch in the lens.
[0139] In S4003, the CPU 1001 detects framing and acquires information such as whether the camera is being swung (panned), in which direction, and at what speed.
[0140] In S4004, the CPU 1001 detects zooming and acquires information such as whether the zoom lens is being operated, in which direction (Tele / Wide) and at what speed.
[0141] In S4005, the CPU 1001 detects focusing and acquires information such as whether the focus ring is being operated, and in which direction (close focus / infinity) the focus ring is being operated and at what speed.
[0142] Furthermore, the detection in S4003 to S4005 includes not only manual operations by the photographer but also auto operations (auto-framing / auto-zoom / auto-focus, etc.) performed on the camera side.
[0143] In S4006, the CPU 1001 sets the subject area. Here, it determines which foreground object in the image generated by the object composition unit 1103 will be the main subject, and simultaneously sets the area to be subjected to AF. Furthermore, by determining the main subject, it is possible to obtain information relating to the subject, such as the velocity, acceleration, angular velocity, angular acceleration, size of the subject, contrast value of the subject, and distance between the subject and the photographer, from the foreground object storage unit 1101.
[0144] There are various methods for setting the subject area, and in this embodiment, it is possible to set it in three-dimensional space. On the other hand, in the camera when shooting in real space, the subject area is set based on the framing of the photographer and the detection results of the subject detection unit in the video (two-dimensional) space. In shooting in virtual space in this embodiment, the subject area can be set in accordance with the method described above, such as information about foreground objects, objects that are closer, or objects closer to the center of the shooting range. Furthermore, as with shooting in real space, the main subject may be detected from the obtained image. This makes it possible to more closely reproduce the performance of the camera when shooting in real space.
[0145] In S4007, the CPU 1001 determines whether correction is ON or OFF in the image correction amount calculation unit 1206 based on the various information acquired in S4002 to S4006. For example, in S4002, if the information related to in-camera correction is ON, correction is turned ON; if it is OFF, correction is turned OFF. In addition, the photographer's intention extraction unit 1262 extracts and determines the photographer's intention, such as which subject the photographer is aiming at, whether the photographer is tracking the subject with framing, or whether the photographer is attempting to switch framing to another subject, from the framing information detected in S4003. In the former case, turning correction ON allows the photographer's framing error to be covered by correction, and in the latter case, turning correction OFF allows the photographer to frame the other subject (fit it into the angle of view) as intended.
[0146] Furthermore, from the zooming and focusing information detected in S4004 and S4005, it is determined that the photographer's intention is strong during manual operations such as manual zoom and manual focus, and compensation is turned off.It is also possible to determine that the photographer's intention is weak during operation of auto functions such as auto zoom and auto focus, and compensation is turned on.
[0147] As described above, by turning off the correction when it is determined that the photographer's intention is strong, it is possible to provide shooting results and a shooting experience that are close to the photographer's operational sense. On the other hand, even if the intention is weak, turning on the correction will allow for a good captured image to be obtained through correction without impairing the photographer's shooting experience. In addition, the camera itself has a defined correction capability value, and there is also a judgment method in which correction is turned on only if the camera's correction capability value exceeds the shooting difficulty, as described below, by comparing it with the shooting difficulty.
[0148] In S4100, the CPU 1001 acquires photographing difficulty information using the photographing difficulty calculation unit 1261. The photographing difficulty information is used to calculate various correction amounts, which will be described later. For example, by reducing the correction amount as the photographing difficulty increases, the more difficult the subject, the more difficult it becomes to keep the subject on the screen and maintain focus. On the other hand, by setting a large correction amount for a subject with a low level of difficulty, good photographing results can be obtained through correction even if a major mistake is made, thereby increasing the success rate of photographing a subject with a low level of difficulty.
[0149] With a subject that is easy to photograph, it is often not possible to concentrate on or immerse oneself in the experience of photographing itself, and therefore even if the amount of correction is increased, it is possible to reduce failed photographs without impairing the photographic experience. On the other hand, with a subject that is difficult to photograph, increasing the amount of correction may impair the fulfillment of the photographic experience, so in this embodiment, the correction is reduced. This is also true for camera photography in real space, so adjusting the amount of correction according to the difficulty of the subject leads to the provision of a more realistic photographic experience.
[0150] The acquisition of the photographing difficulty level information will be described with reference to FIG.
[0151] In S4101, the CPU 1001 acquires the subject's velocity and acceleration information from the foreground object storage unit 1101, and in S4102, the CPU 1001 acquires the subject's angular velocity and angular acceleration information from the foreground object storage unit 1101. This information may be information for each timing, or may be information that is fixedly defined such as maximum velocity and maximum acceleration, and the larger these values are, the higher the degree of difficulty of photographing calculated in S4106.
[0152] In S4103 , the CPU 1001 acquires size information of the subject from the foreground object storage unit 1101 .
[0153] In S4104, the CPU 1001 acquires the contrast value of the subject from the foreground object storage unit 1101. The lower the contrast value, the higher the difficulty of photographing.
[0154] In S4105, the CPU 1001 acquires the distance between the subject and the photographer from information from the foreground object storage unit 1101 and information from the viewpoint information acquisition unit 1201. The subject size on the imaging surface is determined by combining this with zooming (focal length) information acquired in S4001 and S4004 and the subject size information acquired in S4003. The smaller this value, the higher the difficulty of capturing the image. In addition, differences in the part of the subject (whether it is a person's eyes or face) also affect the difficulty of capturing the image.
[0155] In S4106, the CPU 1001 calculates photographing difficulty information for the subject from the information acquired in S4101 to S4105. The photographing difficulty information defined here may be defined as one piece of information that includes all elements. Alternatively, it may be defined as multiple types of information corresponding to framing correction / zooming correction / focusing correction, which will be described later, respectively (framing difficulty / zooming difficulty / focusing difficulty, etc.). The photographing difficulty information may be calculated from various pieces of information in this way, or the difficulty level itself may be stored in the foreground object storage unit 1101. Furthermore, in this embodiment, the photographing difficulty level is calculated (the photographing difficulty level changes) each time the speed or distance of the subject changes, but it may also be defined as being always fixed.
[0156] 16, in S4009, the CPU 1001 calculates a virtual defocus amount for the subject region set in S4006. The calculated virtual defocus amount may be a defocus map calculated from multiple regions as described with reference to Fig. 5, or may be one output for a single part of the subject, such as the face. In the former case, there is a process of selecting one region from multiple regions, but a detailed description thereof will be omitted in this embodiment.
[0157] In step S4200, the CPU 1001 performs processing of the virtual defocus amount, the details of which will be described later with reference to FIG.
[0158] In S4011, the CPU 1001 calculates the focus drive amount. The focus drive amount may be a value converted into a drive amount for the focus lens based on the virtual defocus amount calculated in S4009. Alternatively, the future subject position may be predicted from the subject position in multiple past frames, and the focus drive amount may be set for that predicted position. Various prediction methods are possible, but as this is not the main focus of this embodiment, a description thereof will be omitted.
[0159] In shooting a virtual space, since there is no need to drive a physical focus lens, it is possible to instantly switch to a desired focus state without taking time for focus drive. However, in this embodiment, the object is to provide the photographer with an experience similar to that of shooting in real space by shooting a virtual space using a camera capable of shooting in real space. Therefore, a process of changing the focus state during shooting (for example, focusing on an out-of-focus subject) is performed over a predetermined time. The time spent on focus drive may be set to a time that matches the functions / performance of the camera and lens actually being used using camera / lens information, or may be set based on the virtual camera and lens.
[0160] In S4012 to S4014, CPU 1001 causes image correction amount calculation unit 1206 to calculate various correction amounts.
[0161] In step S4012, the CPU 1001 calculates the amount of correction for framing. The correction for framing will be described with reference to FIG.
[0162] If there are subjects A and B, and subject A is defined as being more difficult to photograph, then the maximum correction amount for framing will be smaller for subject A in accordance with the degree of difficulty (maximum framing correction amount A<maximum framing correction amount B). Here, the dashed rectangular area in FIG. 18 is the area actually framed by the photographer, and the solid rectangular area is the framing area after framing correction has been applied. In this case, in FIG. 18(a), the photographer's framing is misaligned with respect to subject A, but by performing correction that falls within the maximum framing correction amount A, the corrected framing area is able to fit subject A within the frame. On the other hand, in FIG. 18(b), the photographer's framing is even more misaligned with respect to subject A, and even if the maximum framing correction amount A was applied, subject A could not be fitted within the frame.
[0163] Next, taking subject B as an example, in FIG. 18(c), the photographer's framing deviation is small, so similar to FIG. 18(a), the corrected framing area is able to fit subject B within the screen. In FIG. 18(d), the amount of deviation is large, and in the case of subject A, the maximum framing correction amount A was smaller than the amount of deviation, so correction was insufficient. Even in such a case, in FIG. 18(d), the amount of deviation is within the maximum framing correction amount B, so the corrected framing area is able to fit subject B within the screen. In this way, by changing the amount of framing correction depending on the difficulty of shooting, it is possible to provide a realistic shooting experience in which the more difficult a subject is to photograph, the more difficult it is to frame it.
[0164] 16, in step S4013, the CPU 1001 calculates the amount of correction for zooming. Correction for zooming will be described with reference to FIG.
[0165] If two subjects C and D are present, and subject C is defined as being more difficult to photograph, then the maximum correction amount for zooming will be smaller for subject C in accordance with the degree of difficulty (maximum zoom correction amount C<maximum zoom correction amount D). Here, the dashed rectangular area in FIG. 19 represents the angle of view actually adjusted by the photographer through zooming, and the solid rectangular area represents the angle of view after zoom correction has been applied. In this case, in FIG. 19(a), the photographer's zooming is off-center relative to subject C, but by applying a correction that falls within the maximum zoom correction amount C, the corrected zoom angle of view is able to fit subject C within the frame. On the other hand, in FIG. 19(b), the photographer's zooming is even further off-center relative to subject C, and even applying the maximum zoom correction amount C would not allow subject C to fit within the frame.
[0166] Next, taking subject D as an example, in FIG. 19(c), the zooming error caused by the photographer is small, so similar to FIG. 19(a), the corrected zoom angle of view is able to fit subject D within the frame. In FIG. 19(d), the amount of error is large, and in the case of subject C, the maximum zoom correction amount C was smaller than the amount of error, so correction was insufficient. Even in this case, the amount of error is within the maximum zoom correction amount D, so the corrected zoom angle of view is able to fit subject D within the frame. In this way, by changing the zoom correction amount depending on the difficulty of shooting, it is possible to provide a realistic shooting experience in which the more difficult a subject is to photograph, the more difficult it is to keep the subject within the angle of view through zooming.
[0167] 16, in step S4014, the CPU 1001 calculates the amount of correction related to focusing. Correction related to focusing will be described with reference to FIG.
[0168] If there are subjects E and F, and subject E is defined as being more difficult to photograph, then the maximum amount of focusing correction for subject E will be smaller in accordance with the difficulty of photographing (maximum focusing correction amount E < maximum focusing correction amount F). Here, FIG. 20(a) shows subject E, and FIG. 20(b) shows subject F, approaching the photographer from a distance as time passes. The solid lines show the trajectory of the position of each subject, the dotted lines show the trajectory of the focus position when the focus is moved to bring the subject into focus (trajectory of the actual focus position), and the dashed lines show the trajectory after focus correction has been made.
[0169] This indicates that if the solid line, which indicates the subject position, matches the dashed or dotted line, an image with the subject in focus can be captured. When using manual focus, the cause of the deviation between the subject position and the actual focus position is largely due to the photographer's focusing operation, but when using autofocus, it is largely due to the results of the camera's tracking algorithm (tracking limit performance). In this case, one example of the cause is the subject's speed being fast or the speed change being large. For this reason, in Figure 20, the autofocus tracking limit is reached over time, causing the subject position to deviate from the actual focus position.
[0170] In Figure 20(a), the actual focus position is shifted with respect to subject E, but the corrected focus position matches the subject position up to the range within the maximum focusing correction amount E. However, as time passes and the subject gets closer, the amount of shift in the actual focus position increases, and ultimately the maximum focusing correction amount E is not enough to correct the subject, and it becomes impossible to achieve focus.
[0171] On the other hand, in FIG. 20(b), even if the actual focus position shifts significantly with respect to subject F, it is still within the range of the maximum focusing correction amount F, so it is possible to capture an image in which subject F is in focus until the end. As mentioned above, focusing corrections include out-of-focus caused by manual focus operation and out-of-focus caused by framing shifts, which result in the background area being in focus. The corrections for these may be the same as or different from the correction amount for out-of-focus caused by tracking performance limit. Considering that these types of out-of-focus blur can occur simultaneously, a combination of correction amounts for each cause may be applied. In this way, by changing the focusing correction amount depending on the difficulty of shooting, it is possible to provide a realistic shooting experience, in which the more difficult a subject is to photograph, the more difficult it is to focus on it.
[0172] In calculating the various correction amounts in S4012 to S4014 above, as long as detection of Sw1 continues in the virtual space shooting process of Fig. 13, calculations may be performed using past recorded image information and past image correction amounts stored in the storage unit 1004. This allows for continuity in the correction results between images, reducing the sense of incongruity felt as a series of recorded images.
[0173] 16, in S4015, the CPU 1001 performs focus driving, which reflects the focus driving amount calculated in S4011 and the focus correction amount calculated in S4014.
[0174] As described above, by changing the amount of correction effect applied to the recorded image based on the photographer's operation information, subject information, and camera information, it is possible to take photos in a virtual space without compromising the shooting experience itself.
[0175] In this embodiment, an example has been described in which the amount of image correction decreases as the shooting difficulty level increases, but conversely, it is also possible to increase the amount of image correction as the shooting difficulty level increases. By doing so, it is possible to take successful photos (photos in which the subject is within the angle of view or in which the subject is in focus) at a certain level regardless of the shooting difficulty level. Furthermore, it may be possible to switch between decreasing the amount of image correction as the shooting difficulty level increases and increasing the amount of image correction as the shooting difficulty level increases in the camera settings.
[0176] As described above, the processes such as defocus amount calculation and focus driving performed in the virtual subject tracking process of S4000 can be performed using an AF algorithm stored in the ROM of the camera 100. It is also possible to perform the processes using an AF algorithm of another camera.
[0177] (Subroutine for virtual defocus amount processing) Next, the subroutine for the virtual defocus amount processing in S4200 of FIG. 16 will be described using the flowchart shown in FIG. 21. In this subroutine, a process is performed to add an error amount to the virtual defocus amount calculated in S4009 according to the settings at the time of shooting and characteristic information of the image sensor. In virtual space shooting, the defocus amount is calculated from a known distance to the subject, so calculation errors beyond those caused by the number of significant digits of each numerical value are not generated. On the other hand, in real space shooting, errors occur constantly due to the characteristics of the image sensor and errors that vary with each shooting. In virtual space shooting, to reproduce focus state behavior similar to that in real space shooting, it is necessary to generate errors that occur in real space shooting in virtual space shooting as well. In this embodiment, based on the above, a defocus processing process is performed to add errors to the virtual defocus amount calculated in S4009, which does not include errors, in order to simulate real space shooting.
[0178] In S4201, the CPU 1001 acquires, from the camera / lens information storage device 2000, camera information for the virtual image generated by the virtual image generation device 1200, including the resolution of the recorded image, the image sensor size, the AF frame mode, the AF algorithm, and S / N information for each ISO sensitivity. The CPU 1001 also acquires focus-related correction information, such as information related to correction of the defocus conversion coefficient and the focus position, and defocus error information. The CPU 1001 also acquires, as lens information, the focal length and resolution of the lens, the F-number, focus lens information, focus drive control information, and the sensitivity for converting focus lens drive into the amount of movement of the image plane. The CPU 1001 also acquires image stabilization control information, aperture control information, lens frame information, peripheral light falloff information, and distance information related to the focus lens position and distance. The CPU 1001 then stores these as information associated with the virtual image in the external computing device 1000.
[0179] In S4202, the CPU 1001 acquires the defocus error information stored in the RAM 1003. The defocus error information will be described later.
[0180] In S4203, the CPU 1001 assigns the defocus error amount acquired in S4202 to the virtual defocus amount calculated in S4009 of Fig. 16. If virtual defocus amounts have been calculated for multiple focus detection areas, the defocus error amount is assigned to the multiple focus detection areas.
[0181] The defocus error information acquired in S4202 will be described using Fig. 22. Fig. 22 is a graph showing the relationship between the contrast value of the subject acquired in S4006 of Fig. 16 and the amount of error that is expected to occur in the virtual defocus amount. Generally, in shooting in real space, if there are few patterns on the subject and the contrast is low, the amount of error included in the detected defocus amount will be large.
[0182] In FIG. 22, the horizontal axis represents the contrast of the subject and the vertical axis represents the amount of error. Line 24101 indicates that the greater the contrast of the subject (toward the right on the horizontal axis), the smaller the amount of error. This relationship between contrast and error varies depending on the signal-to-noise ratio (SN ratio) of the pixel section and readout circuit of the image sensor, the number of pixels used for the focus detection signal, the gain applied to the signal, which is set by the ISO sensitivity, and other factors. Therefore, in this embodiment, the relationship shown in FIG. 22 is stored for each mode in which the SN ratio of the image sensor changes, the focus detection signal specifications, and the ISO sensitivity. The storage format may be discrete values stored as a table, or a graph may be represented as a function and stored in the form of function coefficients. The above-mentioned error amount also varies because the defocus conversion coefficient, which is camera information, changes depending on the F-number and lens frame information, which are part of the lens information. In this embodiment, the error amount calculated using the above-mentioned lens information and camera information, is multiplied by a predetermined coefficient. The predetermined coefficient may be a ratio value relative to a reference value stored as a table. For example, a table storing defocus conversion coefficients and coefficients by which the above-mentioned error amounts are multiplied is stored using F-numbers and lens frame information as indexes, and the error amounts are calculated according to the conditions at the time of shooting.
[0183] 23(a) and 23(b) show examples of virtual defocus when the virtual defocus error generated in S4203 of FIG. 21 is not applied and when it is applied, superimposed as a map on the main subject.
[0184] 23(a) shows a case where no virtual defocus error is added. In the virtual defocus map 25102, for the head focus position of a person 25106 within an AF frame 25101 in the virtual space image, the entire area of the AF frame 25101 is shown as an in-focus area 25103 (hatched in a horizontal and vertical grid pattern).
[0185] 23(b) shows the case where a virtual defocus error is added. The area of the head of a person 25106 within an AF frame 25101 is shown as an in-focus area 25103, a front focus position 25104 (hatched with a diagonal grid), and a rear focus position 25105 (hatched with black dots).
[0186] In Figure 23(a), all of the AF frames are in-focus areas, whereas in Figure 23(b), the addition of errors has resulted in AF frames that are in front focus and rear focus, in addition to the in-focus areas.
[0187] In this way, when capturing a virtual image, by adding a defocus error according to changes in camera / lens information and applying an algorithm for selecting an AF frame, it is possible to obtain defocus detection results similar to those obtained when capturing an image in real space. While this embodiment illustrates an example in which a defocus error is added to a defocus map for ease of explanation, a defocus error may also be added to a single AF frame. Because this error affects predictive AF, a similar effect can be expected. This makes it possible to reproduce the same focus adjustment behavior when capturing an image in virtual space as when capturing an image in real space, allowing users to evaluate product performance and check new features before purchasing a camera or lens.
[0188] In this embodiment, defocus variation is added to make the image captured in virtual space closer to the image captured in real space, but adding an error is not necessarily required. If it is not necessary to make the image captured in virtual space closer to the image captured in real space, adding an error may be omitted, or switching may be performed depending on the situation.
[0189] (Virtual space photography subroutine) Next, the virtual space photographing subroutine executed by the external processing device 1000 in S5000 of FIG. 13 will be described with reference to the flowchart shown in FIG.
[0190] In S5001, the CPU 1001 outputs the set F-number and the time when Sw2 was detected. How this information is used will be explained in the section on real camera operation linked to virtual space shooting operation in Fig. 28, which will be described later.
[0191] In S5002, the CPU 1001 acquires camera / lens information. Specifically, the camera information includes the resolution for recording, the size and number of pixels of the camera's image sensor, and the lens information includes the range and current value of the focal length, the range and current value of the F-number, the range and current value of the focus lens position, and information on peripheral light falloff and point spread function.
[0192] In S5003, the CPU 1001 acquires the image correction value described above.
[0193] In S5004, the CPU 1001 generates a recorded image of the virtual space. It renders an image based on the foreground and background objects arranged in three-dimensional space and the viewpoint position information. The range of the image to be displayed is determined based on the focal length in the lens information, the image sensor size, resolution, and camera settings in the camera information. Furthermore, the recorded image is generated based on the F-number information, vignetting information, point spread function information, and focus lens position information in the lens information. Unlike the displayed image, the recorded image is an image to be recorded, so the recorded image range is correctly determined and generated using various optical information, such as focus lens position information, vignetting information, and point spread function information. Unlike the displayed image, the recorded image does not need to be displayed to the photographer in real time, so the generation of the recorded image can be delayed compared to the generation of the displayed image. Therefore, the recorded image can be generated using more detailed data from the camera information and lens information used to generate the displayed image.
[0194] In S5005, the CPU 1001 records the recorded video of the virtual space in the recording unit 1004 of the external computing device 1000. Alternatively, the recorded video generated in S5004 described above may be transferred to the camera 100 and recorded in the flash memory 133 of the camera 100.
[0195] In S5006, the CPU 1001 records various types of video-related information. Video-related information refers to information including subject information (photographing difficulty information), photography-related information, and virtual defocus amount. The video-related information is recorded in the recording unit 1004 of the external computing device 1000. Alternatively, the video-related information may be transferred to the camera 100 and recorded in the flash memory 133 of the camera 100. This completes the virtual space photography subroutine.
[0196] (Virtual space photography reflecting operation information) Virtual space photography that reflects operation information will be described with reference to Fig. 25. Fig. 25(a) shows an example of zooming, and Fig. 25(b) shows an example of framing.
[0197] 25(a), a still virtual space display image 18003 generated by virtual image generation device 1200 is displayed on display 131 of camera 100. When the photographer rotates the zoom ring of the lens to change the focal length to the telephoto side, virtual image generation device 1200 acquires the change in focal length due to the operation of the zoom ring as operation information. Then, by changing the display range when generating the virtual space display image, virtual space display image 18004 is generated that reflects the change in focal length due to the zooming operation by the photographer, and is displayed on display 131.
[0198] In this embodiment, the display image of the virtual space is generated using camera / lens information, so it is possible to generate the display image of the virtual space with a focal length range that cannot be controlled by the lens operated by the photographer. For example, even if the photographer is using a lens with a short focal length, by generating the display image of the virtual space using lens information of a telephoto lens with a long focal length, it is possible to experience shooting with a telephoto lens with a long focal length. Generally, lenses with long focal lengths used to shoot in real space are large, heavy, and expensive, but the shooting of virtual space in this embodiment can provide a shooting experience without such constraints.
[0199] In Figure 25(b), a display image 18006 of a still virtual space generated by virtual image generation device 1200 is displayed on display 131 of camera 100. Figure 25(b) shows a case where the photographer performs a framing operation on the camera / lens, moving the camera / lens horizontally in the direction of arrow 18005. Framing information, which is the position of the camera / lens as a result of the framing operation, is acquired by virtual image generation device 1200 as operation information. Then, by changing the display range when generating a display image of the virtual space, a display image 18007 of the virtual space that reflects the change in framing due to the photographer's framing operation is generated and displayed on display 131.
[0200] As a result, a virtual space display image that reflects the photographer's operational information can be realized by capturing a still image in the virtual space.
[0201] (Viewpoint movement subroutine) Next, the subroutine of the viewpoint movement process of S3000 in Fig. 13 will be described with reference to Fig. 26. Specifically, this shows the process of changing the viewpoint position (the position of the camera in the virtual space) when taking a picture in the virtual space. This process is performed by the CPU 1001.
[0202] First, in S3001, the CPU 1001 adjusts the focal depth and angle of view of the image to be displayed when moving the viewpoint. When moving the viewpoint, it is desirable to make it easier to view a wider range of subjects and to confirm the in-focus distance. This is because, as will be described later, the destination of the viewpoint is set based on the angle of view displayed on the display 131 and the object at the in-focus distance. In this embodiment, in S3001, the angle of view is widened to a preset angle of view, and the range of in-focus distance is adjusted to a preset depth of field shallower than that of the shooting mode. This process may be omitted because it facilitates the operation of moving the viewpoint.
[0203] In S3002, the CPU 1001 displays an image of the virtual space based on the settings made in S3001, i.e., an image from the viewpoint position set in advance for shooting in the virtual space or the viewpoint position set as the initial position.
[0204] In S3003, the CPU 1001 adjusts the focus and the direction of the index. First, in adjusting the focus, as described in S3001, a state in which the focus is achieved within a predetermined distance range within the shooting range is displayed, and the photographer performs the same focus adjustment operation as when shooting. Specifically, with one index (I(n,m)) for focus detection described in FIG. 5 displayed for an object located at a position within the shooting range where the viewpoint is to be moved, the focus is adjusted by half-pressing the release switch included in the operation switch group 132. This makes it possible to show the photographer the in-focus distance in the image displayed on the display 131.
[0205] The distance at which the focus is achieved is set as the viewpoint movement distance. Although the method for setting the viewpoint movement distance has been described using an operation similar to automatic focus adjustment (autofocus), it may also be set in the same manner as manual focus operation, in which the focus lens (third lens group 105) of the imaging optical system is manually operated. By rotating a focus ring (not shown) provided on the imaging optical system, the focused distance is adjusted toward infinity or a close distance, and the distance intended by the photographer is adjusted as the viewpoint movement distance.
[0206] The direction of the index is adjusted to set the direction in which the viewpoint will move. The photographer changes the position of the index (I(n,m)) that performs focus detection within the screen of the display 131, or moves the camera 100 by panning or the like, and aligns the index with the direction in which the viewpoint will move. This allows the photographer to set the direction in which the viewpoint will move while checking the image of the virtual space displayed on the display 131.
[0207] In S3004, the CPU 1001 determines whether or not there is an instruction to move the viewpoint. If the viewpoint position change button included in the operation switch group 132 is pressed, the process proceeds to S3005. If the viewpoint position change button is not pressed, the process returns to S3003 and continues adjusting the focus and index direction.
[0208] In S3005, CPU 1001 determines whether or not the viewpoint can be moved. The photographer instructs the viewpoint to be moved by pressing a viewpoint position change button, and the direction and distance to move the viewpoint are determined. If the viewpoint after movement is below the ground (underground) of the background object, inside the foreground object, or if the camera after the viewpoint movement interferes with another object, it is determined that the viewpoint cannot be moved. Furthermore, if the distance at which the focus was achieved when the viewpoint position change button was pressed is infinity, it is determined that the viewpoint cannot be moved to infinity. If the viewpoint movement distance is farther than a predetermined distance, the viewpoint movement distance may be reset using a predetermined distance set in advance as the maximum value.
[0209] In S3006, if the result of the viewpoint movement determination in S3005 indicates that viewpoint movement is possible, the CPU 1001 proceeds to S3007. On the other hand, if it is determined that viewpoint movement is not possible, the CPU 1001 proceeds to S3008. In S3007, the viewpoint is moved by the viewpoint movement distance and in the direction set in S3003.
[0210] In S3008, the fact that the viewpoint cannot be moved is notified as a warning to the photographer via the display 131. The photographer may be notified only that the viewpoint cannot be moved, or may also be notified of the reason, such as interference with an object or the set movement distance being too far.
[0211] After the viewpoint has been moved in S3007 or the notification that the viewpoint cannot be moved in S3008 has been completed, the process proceeds to S3009.
[0212] In S3009, if the CPU 1001 receives an instruction to end the viewpoint movement mode, it ends this subroutine. If there is no instruction to end the viewpoint movement mode, the process returns to S3003.
[0213] Next, a specific example of the viewpoint movement subroutine explained in Fig. 26 will be explained with reference to Fig. 27. Fig. 27 shows an example of a display on the display 131 in the viewpoint movement mode.
[0214] 27(a) shows a display example in which viewpoint movement is set in viewpoint movement mode in a virtual space in which a person and a dog are placed as foreground objects. Foreground objects 27003 of the person and the dog are displayed on a display screen 27001 of the display device 131, and the state before the viewpoint movement is such that the camera is positioned to view from the upper left of the person.
[0215] Reference numeral 27002 is one of the indices (I(n,m)) described in FIG. 5, and indicates the direction in which the viewpoint moves on the display screen. Reference numeral 27005 indicates the distance the viewpoint can be moved, along with the range in which the viewpoint can be moved. FIG. 27(a) shows that the viewpoint can be moved from 0.45 m to 10 m, and the current adjustment distance is 1 m. In the example of FIG. 27(a), an object that is 1 m away from the current viewpoint position (camera position) is in focus, but the in-focus and out-of-focus distances are not shown in the figure, and the focus is shown from close up to infinity.
[0216] The index 27002 is movable within the display screen 27001. Furthermore, by operating the camera by panning or the like, the index 27002 can be superimposed on an object such as a dog, and the focus can be adjusted to that distance, i.e., the viewpoint movement distance can be set. The sub-display screen 27004 displays a preview image of the foreground object 27003 observed from the currently set viewpoint after the viewpoint has been moved. The direction of the image to be previewed from the viewpoint after the viewpoint has been moved may be set automatically from position information of the foreground object, or may be set by the photographer operating an operation switch. Furthermore, a rectangular frame may be displayed within the sub-display screen 27004 to indicate the shooting range corresponding to the focal length of the lens being used.
[0217] In this way, by setting the distance and direction of viewpoint movement using the screen displayed on the display screen 27001 and the indicator 27002, the photographer can intuitively and easily move the viewpoint using the same operations as when taking a photo.
[0218] A modified example of viewpoint movement will be explained using Fig. 27(b). This method allows the photographer to move the viewpoint more easily by positioning and displaying a viewpoint movement target as a target position within the display screen.
[0219] FIG. 27(b) shows a view point movement target (target position, marker) displayed on the display screen in the view point movement mode. View point movement targets 27006 are displayed in a grid pattern on the ground, which is part of the background object of the display screen 27001. Among the grid intersections, the intersection in the second row and first column is indicated as view point movement target 27006(2,1), and the intersection in the fourth row and fourth column is indicated as view point movement target 27006(4,4). The photographer can set the view point movement distance by moving the indicator 27002 close to the view point movement target 27006 closest to the distance at which the photographer wishes to move the view point and selecting it. The view point movement target may be displayed superimposed on the foreground object, or may be displayed together with the numerical value of the view point movement target distance. Displaying the view point movement target in this manner allows the photographer to more easily set the view point movement distance.
[0220] (Feedback on the operation feel of shooting virtual space by driving the camera drive unit) Next, the operation of the camera during virtual space shooting will be described using FIG. 28 . When shooting in real space, the photographer receives tactile feedback such as vibrations and sounds as the shutter and lens are driven in response to shooting operations, leading to an improved quality of the shooting experience. On the other hand, as described above, in virtual space shooting, virtual subject tracking processing and virtual space shooting are achieved by operating the release switch of the operation switch group 132 of the camera 100. In this case, there is no drive unit required for image generation, and there is no need to drive the shutter or lens. This state may result in a degradation of the quality of the shooting experience. In this embodiment, the drive unit of the camera 100 is driven in synchronization with the virtual space shooting operation, providing the photographer with tactile feedback such as vibrations and sounds, thereby achieving a more realistic shooting experience.
[0221] 28 is a flowchart illustrating the camera operation when capturing an image in a virtual space. Each process is executed by the camera CPU 121.
[0222] In S1201, the camera CPU 121 receives an instruction to initialize the camera driving unit from the CPU 1001 in S1002 in FIG. 13 and executes initialization of the camera driving unit. The camera CPU 121 drives the shutter 106 to match the open / closed state with the camera settings in the virtual space. For example, if shooting is about to begin, the shutter is driven to an open state. The camera CPU 121 also drives the zoom actuator 111, aperture actuator 112, and focus actuator 114 to the focal length, aperture value, and focus lens position corresponding to the angle of view, depth of field, and in-focus distance at the start of shooting in the virtual space. In this embodiment, the initial positions of the driving units of the camera 100 are set in response to an instruction from the CPU 1001. However, position information of each driving unit of the camera 100 may be output to the external computing device 1000, and the external computing device 1000 may determine the initial state.
[0223] In S1202, the camera CPU 121 outputs camera information, lens information, and operation information in a manner corresponding to S1003 in Fig. 13. The contents of the information are as described in S1003.
[0224] In S1203, the camera CPU 121 acquires the image generated in S2000 of FIG.
[0225] In S1204, the camera CPU 121 monitors whether the focus drive instruction performed in S4015 in Fig. 16 is input to the camera 100. If the focus drive instruction has not been acquired, the process proceeds to S1206, and if the focus drive instruction has been acquired, the process proceeds to S1205. In S1205, the camera CPU 121 causes the focus actuator 114 to drive the focus lens (third lens group 105) in accordance with the focus drive instruction.
[0226] In S1206, the camera CPU 121 monitors whether the aperture value input to the camera 100 in S5001 in Fig. 24 is different from the current setting. If there is no change in the aperture value, the process proceeds to S1208, and if there is a change in the aperture value, the process proceeds to S1207.
[0227] In S1207, the camera CPU 121 drives the aperture 102 using the aperture actuator 112 in accordance with the change in the aperture value.
[0228] In S1208, it is monitored whether or not the time when Sw2 was detected is input to the camera 100 in S5001 in Fig. 24. If the time when Sw2 was detected is not input, the process proceeds to S1210, and if the time when Sw2 was detected is input, the process proceeds to S1209.
[0229] In S1209, the camera CPU 121 drives the shutter 106 in the same way as for capturing an image in real space after a predetermined time has elapsed since the input detection time of Sw2.
[0230] As described above, by driving the focus lens, aperture, and shutter when taking photos in a virtual space, the photographer can receive feedback such as vibrations and sounds caused by the operation of the camera they are operating, providing a more realistic shooting experience.
[0231] In this embodiment, the driving unit of camera 100 is configured to be driven in response to an operation instruction related to shooting in virtual space. However, there are cases where camera 100 cannot drive in response to the issued operation instruction. For example, this may occur when an operation instruction is issued to drive a continuous shooting speed faster than camera 100 can drive, or when an operation instruction is issued to drive a focus lens with a distance longer than the lens currently attached.
[0232] In such a case, the driving of the driving unit of the camera 100 may be prohibited based on the generated operation instruction, or the operation instruction may be edited and changed to an instruction that can be driven by the driving unit of the camera 100, and then the driving may be performed. For example, if the operation instruction is faster than the continuous shooting speed of the camera 100, the operation instruction may be thinned out at regular intervals to drive the shutter. Furthermore, with regard to driving the aperture and focus, the specifications of the lens used in the virtual space may be compared with the specifications of the lens attached in the real space, and standardization may be performed to align the driving ranges, and when an operation instruction is received, the driving amount may also be standardized in the same way.
[0233] In this embodiment, feedback of the operational feeling to the photographer has been described, but the sounds and vibrations that occur during this process may also be recorded in other recording means. For example, a blur corresponding to the amount of camera vibration may be added to a still image, or the generated sounds may be recorded in a video. This allows for shooting in a virtual space that is closer to shooting in real space.
[0234] (Method of taking and evaluating the photographic results) The flowchart in Figure 29 shows a method for reproducing and evaluating images after shooting using the camera 100 of this embodiment. Specifically, the method involves reproducing images captured in real space, reproducing images captured in virtual space, and displaying a defocus map of images and calculating the degree of focus of a series of consecutive images under conditions different from those used during actual shooting. This makes it possible to display the causes of poor focus and the best settings for improving the focus. In the following explanation, "S" represents a step.
[0235] First, in S1101, camera CPU 121 selects an image to be played back from flash memory 133. The photographer operates operation switches 132 to instruct playback, thereby displaying the image captured immediately before or the image played back the previous time. Thereafter, the photographer operates operation switches 132 to play back the image they want to play back.
[0236] In S1102, the camera CPU 121 determines whether the image to be played back is an image taken in virtual space or an image taken in real space. If the camera CPU 121 determines that the image to be played back is an image taken in virtual space, the process proceeds to S1103, where the virtual image stored in the storage unit 1004 of the external computing device 1000 is played back on the display 131. On the other hand, if the camera CPU 121 determines that the image was not taken in virtual space but in real space, the process proceeds to S1104, where the image stored in the flash memory 133 is played back on the display 131.
[0237] In S1105, the camera CPU 121 determines whether to perform shooting evaluation of the virtual space captured image (S1103) or the real space captured image (S1104) being played back. If shooting evaluation is to be performed, the process proceeds to S1106, and if shooting evaluation is not to be performed, the process ends.
[0238] In S1106, the camera CPU 121 acquires shooting-related information when the image to be played back was captured. The shooting-related information refers to various information about the camera settings and lens used when shooting. The shooting-related information includes the lens and camera settings used when shooting, such as focal length, F-number, continuous shooting mode, AF mode, subject detection AF tracking setting, AF frame setting, and shutter method. The shooting-related information is information used when evaluating the focus state of the image, which will be described later, and may be any information that affects the focus state. The shooting-related information may be stored in the flash memory 133 or the storage unit 1004, or may be attached to the image being played back as meta information.
[0239] In S1107, the camera CPU 121 acquires AF log information attached to the playback image as meta information. The AF log information includes defocus information used when the playback image was captured, AF frame setting information, and tracking information (the subject detection AF function focuses on a detected object, such as a person, animal, or vehicle, automatically using a preset algorithm). It also includes servo AF characteristics (which assign various servo AF parameters to set the focus priority), and action recognition information (subject posture information, and information on how to prioritize subject recognition when the subject performs a specific action). It also includes shutter method information (which allows you to select a shutter mode, such as a mechanical shutter mode that drives a mechanical shutter or an electronic shutter mode that determines the exposure time solely using the image sensor without a mechanical shutter, and confirm the frame rate setting for continuous shooting, such as 30, 20, or 10 frames per second for the electronic shutter).
[0240] In S1108, the camera CPU 121 sets one or more images, including the image being played back, as an evaluation image group. The evaluation image group may be set by setting images taken at a time close to the shooting time of the image being played back, or by setting a group of images taken in a single continuous shot. Furthermore, the image group may be set by setting the first and last images in the evaluation image group.
[0241] In S1109, the camera CPU 121 sets the evaluation sequence. The camera CPU 121 (or the CPU 1001 of the external computing device in the case of shooting a virtual space) determines the equipment and various algorithms to be used and determines the type of evaluation to be performed. Details will be described later.
[0242] In S1110, the camera CPU 121 performs evaluation under each setting condition. Evaluation is performed under the conditions of the various information and various setting contents acquired in S1106 to S1109 described above. As a result, for the image being played back, the defocus amount of the captured image is calculated from the focus control result that differs from that at the time of shooting, and the degree of focus is calculated by evaluating the amount of focus deviation based on a threshold determined from the defocus amount.
[0243] When calculating the degree of focus, image analysis is also performed to simultaneously analyze the causes of the good or bad degree of focus. In addition, the amount of defocus is calculated from the results of focus control that differ from the time of shooting, and a defocus map is created and displayed superimposed on the captured image.
[0244] This makes it possible to evaluate the difference in focus state between an image taken under the conditions set at the time of image acquisition and an image taken under conditions set different from those set.
[0245] In S1111, the camera CPU 121 displays the evaluation results performed in S1110 on the display 131 or an external display such as a PC. The method for displaying the evaluation results is not limited to one format. The display method will be described later. By displaying the evaluation results, the photographer can confirm the cause of the poor focus, which can lead to corrections to the shooting method or changes to the shooting settings, thereby improving the photography technique.
[0246] In S1112, the camera CPU 121 presents the best settings. Based on the evaluation results of S1111, the best settings are presented from the evaluation results regarding the degree of focus during shooting. The contents of the presentation and examples of display will be described later. In the best settings presented here, the photographer confirms and changes the best settings displayed on the display 131, but before evaluating the settings of the playback image in S1101, a menu can also be provided to determine whether to automatically change the settings using the evaluation results. By selecting automatic change, the best camera settings can be automatically changed based on the evaluation results.
[0247] In addition, similar processing can be performed in parallel during shooting, and settings can be changed without the photographer's confirmation, for example, by automatically switching to the optimal settings for the fourth shot based on the results of evaluating three shots during continuous shooting and continuing continuous shooting.
[0248] (Defocus map display) Next, with reference to Figures 30(a) to (c), a description will be given of the display in which a defocus map is superimposed on a captured image in S1111 based on the evaluation result of S1110. Figure 30(a) shows an example of a captured scene of a person skiing.
[0249] 30(b) shows a superimposed display of a defocus map 30001 based on the calculation of the focus control results from the shooting-related information in the meta information attached to the captured image. This defocus map displays on the captured image whether the focus position is in the plus direction (front focus) or minus direction (back focus) relative to 0 for each block displayed in a 10x8 grid.
[0250] A diamond-shaped frame 30002 displayed in each block indicates that the focus position is near 0 and the subject is in focus.
[0251] A diagonal line frame 30003 displayed in each block indicates a state in which the defocus amount is positive, and illustrates a tendency toward front focus.
[0252] A dotted frame 30004 displayed in each block indicates a state in which the defocus amount is negative, illustrating a tendency toward back focus.
[0253] The block that overlaps with the skier is displayed as a roughly diamond-shaped frame 30002, indicating that the subject is in focus.
[0254] 30(b) is an example, and the defocus map does not need to be a 10 x 8 grid, but may be displayed in more detail. Although the defocus amount for an area roughly including the main subject is displayed, the defocus amount is not limited to this and may be displayed for the entire captured image.
[0255] Figure 30(c) shows a defocus map 30001 superimposed on the image data based on the calculation of the focus control results from the shooting-related information in the meta information attached to the image captured using the combination of camera (product name CA) and lens (product name LA) used by the photographer.
[0256] AF frame 30000 is the result of shooting under conditions where only one point from the camera's AF frame setting information was used as the focus detection area. The evaluation results show that the AF frame covers half of the subject's face, and is affected by the contrast of the background subject, resulting in the focus being set further away than the main subject, reducing the degree of focus. Looking at the defocus amount for the focus position of each block, the right side of the subject is displayed as a diagonal line, indicating a tendency for forefocus.
[0257] FIG. 30(d) shows a defocus map taken with the combination of camera (product name CA) and lens (product name LA) used by the photographer described above.
[0258] This shows the defocus result when the AF frame 30000 is changed from the above-mentioned one-point AF to a wide-range AF frame in the evaluation sequence setting performed in S1109.
[0259] In step S1110, the defocus amount is calculated based on the focus control information for the single-point AF frame in the AF frame setting information actually captured and the focus control information when the AF frame is changed to a wide-range AF frame, and the result is displayed as a defocus map 30001 in Figures 30(c) and 30(d).
[0260] By showing the changes in the defocus map 30001 resulting from changes in the AF frame setting, it is possible to compare the amount of defocus during single-point AF and area expansion AF. In Figure 30(d), many diamond-shaped frames 30002 are superimposed on the subject, indicating that an image in focus on the subject can be obtained without being affected by the background. In the example shown in Figure 30, it can be seen that, depending on the photographer's framing technique, better defocus results can be obtained by shooting with an AF frame setting for wide-area AF rather than single-point AF.
[0261] In addition to changing the AF frame, the system also obtains the information necessary for AF from the virtual camera information and existing lens information of a different camera (product name CB) than the camera (product name CA). Then, by overwriting the focus-related information of a captured image with the focus control information of the camera (product name CB), it is possible to compare the difference in AF performance between the cameras (product name CA) and (product name CB). This makes it possible to evaluate the degree of performance improvement for each shooting scene, especially when a new product is expected to offer improved performance. This can be used when considering purchasing a new product.
[0262] Similarly, when switching from a lens (product name LA) to a different lens (product name LB), lens information for a virtual lens different from the lens used in the playback image is obtained. Then, a virtual defocus amount can be obtained for the playback image by combining different focal lengths, F-numbers, etc. As a result, the defocus information for the virtual lens (product name LB) can be compared with the defocus information for an image already captured using the lens (product name LB), allowing you to display and check the performance when changing lenses.
[0263] These displays may be displayed on a display device of an external PC, rather than on display 131 of camera 100. The method for comparing the changed defocus map images may be to line up the images before and after the change, or it may be possible to change only the defocus amount of the defocus map.
[0264] The defocus amount in FIG. 30 is displayed by dividing the focus position into three stages: near focus, front focus, and back focus, but it may be divided into more detailed stages, or the defocus amount may be displayed in units of mm, for example.
[0265] Photographers can check the performance difference by changing the focus control results based on information from a camera or lens different from the one they used for shooting.In addition, since the performance of the camera or lens they want can be checked before purchasing, they can choose the camera or lens that best suits their needs.
[0266] 30(a) to (c) explain the shooting in real space, but the present invention may be applied not only to the real space but also to images obtained by shooting in a virtual space using the external computing device 1000.
[0267] (Focus level) FIG. 31 shows an example of calculating the degree of focus for a series of continuously captured images based on the results of the photography evaluation flow shown in FIG.
[0268] FIG. 31 shows how a photographer 31003 captures multiple images of a series of scenes of a subject skiing using continuous shooting to obtain image 31001. The images handled here may be captured in real space or virtual space. The defocus amount is calculated from the evaluation results S1110 of the captured series of multiple images, and if the calculated result is within a predetermined threshold range centered around a defocus amount of 0, the focus degree result is determined to be O, and any other focus degree is determined to be X. Each image is judged to be O or X, and the proportion of O's among all images captured in the series of continuous shooting is displayed as focus degree display 31002.
[0269] In Figure 31, an example is shown in which the proportion of focus degrees that were rated as ◯ was 70%, and it was determined that there was room for improvement and was marked as △. Although a series of focus degree results were displayed as 70% △, focus degrees of 80% or more could also be displayed as ◯ to make the focus degree assessment results easier for the photographer to understand.
[0270] Furthermore, the display of the sign may be freely set, such as displaying an X if the result of the degree of focus is 60% or less, or only the degree of focus may be displayed without any display.
[0271] In this embodiment, the calculation result of the focus degree is shown in two stages, O or ×, but this is only an example and the method of determining the focus degree can be freely determined. Dispersion, etc., may be displayed using the unit of mm used in defocus calculation. The display method may also be freely displayed without detailed settings. A series of focus degree results are displayed on the display 131 of the camera, but can also be confirmed on other display devices such as a PC.
[0272] (List of setting changes) 32 is a table showing examples of changeable setting items and evaluation conditions for the setting change proposals evaluated in S1110. Examples of changeable setting items are shown horizontally.
[0273] The items written in the white frame are examples of camera information settings, such as AF frame setting, tracking that enables subject detection AF, AF mode that switches between one-shot AF and servo AF, servo AF characteristics that change various servo AF parameters, and shutter method. The gray frame is an example of lens information settings, showing the lens and focal length.
[0274] In FIG. 32, the initial settings shown are the shooting settings at the time of image acquisition. Recommended settings 1 and recommended settings 2 are examples of evaluation sequences set in S1109 in FIG. 29. The number of evaluation sequences is not limited to two. By changing all combinations of configurable parameters and evaluating them in S1110, it is possible to find the best settings (evaluation sequence) that maximize the degree of focus. However, since this increases the computational load, it is also possible to reduce the computational load by changing only parameters that are effective in improving the degree of focus from the initial settings.
[0275] Using the recommended settings in FIG. 32, settings that are effective in improving the degree of focus will be described.
[0276] The following situations can be considered as reasons why the AF frame may move away from the subject when using the initial settings set by the photographer. When single-point AF is combined with tracking off, the AF frame visible in the viewfinder is fixed. This means that the photographer needs to keep the AF frame aligned with the subject, making framing more difficult as it can handle unexpected subject movement.
[0277] With recommended setting 1, even with single-point AF, subject detection is performed by turning on and using tracking AF, allowing the AF frame to automatically capture and continue tracking the subject. This reduces the difficulty of framing, as the photographer only needs to align the subject with the AF frame and begin tracking when starting shooting, allowing them to concentrate on getting the subject into the field of view. Settings related to the AF frame and subject detection AF tracking can be changed and evaluated for images taken in both real space and virtual space. A new AF frame can be selected using the image and defocus map information at the time the image was acquired, using the AF frame setting algorithm to be applied after the change. A new subject detection area can also be set for the image using the tracking algorithm to be applied after the change.
[0278] Similarly, Recommended Setting 2 selects the AF area expansion setting to expand the AF range for single-point AF. This setting makes framing easier without using tracking.
[0279] With electronic front-curtain shutter systems, the shutter curtain is activated with each release, causing blackouts. Blackouts cause you to lose sight of the subject for a moment, and the display update rate also slows down, making framing difficult. Since electronic shutters do not use a shutter curtain, the display update rate does not slow down during continuous shooting, and blackouts, where the entire screen becomes black, do not occur. Therefore, when continuously shooting moving subjects in particular, the electronic shutter makes framing easier without losing sight of the subject.
[0280] Regarding shutter method settings, for images captured in real space, it is possible to change the frame rate to slower (thin out images), but changing the frame rate to faster is difficult due to a lack of information. For images captured in virtual space, it is possible to generate images anew in the shooting environment, so it is also possible to change the frame rate to faster. Therefore, when evaluating images captured in real space, if the frame rate is slow, such as when an electronic front-curtain shutter is used, it may be possible to recommend to the photographer that they use an electronic shutter using other information, such as the speed of the subject being photographed.
[0281] Regarding the focal length of the lens, the photographer uses a 70mm-200mm zoom lens, and sets the focal length to 200mm when shooting, which narrows the angle of view of the subject, making it easier for the subject to go out of frame due to unexpected movement or fast sliding movements. This makes framing difficult. By widening the angle of view, there is more room in the angle of view for the subject's unexpected movement or fast sliding movements, reducing the risk of the subject going out of frame. Therefore, by setting the focal length to the wide end of 70mm, the difficulty of framing can be reduced.
[0282] Regarding the setting of the focal length, it is possible to narrow the angle of view for images taken in real space, but it is difficult to widen the angle of view due to the lack of information (images). For images taken in virtual space, it is possible to regenerate the image in the shooting environment, so it is also possible to widen the angle of view. Therefore, when evaluating images taken in real space, if the image was taken with a long focal length, it may be possible to recommend to the photographer to use a lens with a wider focal length using other information such as the speed of the subject when taking the image.
[0283] As described above, by changing the settings that have factors that reduce the degree of focus, it is possible to find recommended settings that can improve the degree of focus more efficiently, thereby reducing the calculation load.
[0284] The setting method in Figure 32 is merely an example, and settings may be made based on various considerations. Analysis results of the shooting environment, such as whether the subject type is a human or an animal, or whether the scene is one in which multiple subjects intersect, can also be utilized. In addition, the recommended settings may be narrowed down by determining the photographer's skill based on the photography history entered in advance, the movement of the subject during photography, and the photographer's framing.
[0285] (Explanation of information display for shooting settings) Next, the display of information on the captured image relating to the photographer's framing technique based on the evaluation results of the focus degree described above will be described with reference to FIG.
[0286] Figure 33(a) shows an out-of-focus image in the evaluation result S1111 of a series of continuously captured images, where the defocus amount was large and the focus level was determined to be NG. This image assumes that the framing was unable to keep up with the subject's high-speed skiing, resulting in poor focus because the AF frame 32001 was off the subject's face. It is assumed that the camera settings set by the photographer, such as the shutter method, the AF frame selection, and the angle of view due to the lens focal length, make framing difficult. In such a situation, the camera's display 131 or a display device such as a PC displays a suggested optimal setting based on image evaluation during image playback. The suggested optimal setting and its display will be explained in Figure 33(b) below.
[0287] FIG. 33(b) shows an example in which suggestions to the photographer regarding the best settings are displayed as a result of evaluating various shooting sequences.
[0288] In the evaluation of each setting condition in S1110, in addition to the evaluation, many types of conditions are calculated with various combinations from each piece of information and setting conditions in S1106 to S1109 to find the setting condition that will result in the highest degree of focus.
[0289] For example, in Figure 33(a), AF frame 32001 is off-center from the subject, but in the evaluation of S1110, the degree of focus is calculated when the AF frame is widened based on the AF frame setting in the AF log information of S1107. The AF mode is not changed, and the degree of focus is confirmed when subject detection AF tracking information is added. Various other conditions are also changed to calculate the degree of focus, and the setting combination that provides the highest degree of focus is selected.
[0290] In the example of FIG. 33, it is assumed that the evaluation result of S1110 for the captured image shows that the following conditions increase the focus degree based on the results of the focus degree under various conditions. (1) Leave the AF frame at one point and turn on subject detection AF tracking. (2) The shutter system has been changed to an electronic shutter. (3) Change the angle of view to the wide-angle side Based on the evaluation results of S1110 described above, the best information can be proposed to the photographer. Display 32003 displays the best information for the above factors.
[0291] Fig. 33(c) shows a display 32004 that allows the photographer to select whether or not to change the tracking for subject detection in the shooting-related information for the best setting plan shown in Fig. 33(b). The photographer can check the display and make settings that can further improve the degree of focus.
[0292] Next, although not shown, a similar display is made to prompt the photographer to select whether or not to change the shutter method. When photographing in real space, the lens angle of view cannot be selected, so a display is made to advise the photographer to change the focal length. When photographing in virtual space, the focal length can be virtually changed, so a display regarding the change is made, as in Figure 33(c), to prompt the photographer to make a selection.
[0293] The display method and display order of these setting change suggestions may be such that various settings are displayed together and selected, or such that only the focus level is displayed and the focus level selected by the photographer is changed to the calculated setting all at once.
[0294] Although the display prompting the user to make settings has been described above, the camera may automatically change the settings.
[0295] FIG. 33(d) shows the evaluation results of the photography in which the settings were changed in FIG. 33(c).
[0296] AF frame 32001 has been changed to a dotted line, which is the tracking AF frame, and as a result of subject detection, the subject is constantly tracked, allowing you to concentrate on framing.
[0297] The shutter method has also been changed to an electronic shutter, which eliminates blackouts and keeps the subject visible in the viewfinder at all times, reducing the chance of losing sight of the subject. The lens has been changed to the wide-angle 70mm, which gives more room for the subject and angle of view, reducing the chance of the subject going out of frame. After changing to the best setting, the focus rate was 85%, a dramatic improvement over the 70% rate before the change.
[0298] By evaluating a series of images taken and quantifying the degree of focus, photographers can understand the capabilities of their framing techniques. By analyzing the causes of poor focus and displaying the best settings, photographers can improve their framing techniques.
[0299] 33(a) to (c) are explained using real space photography. However, as in Fig. 31, by evaluating the focus level of the result of photography not only in real space but also in a virtual space environment using the external computing device 1000, the best settings may be displayed and the photography settings may be changed or automatically changed.
[0300] (Variation) In this embodiment, a configuration has been described in which focus area detection is achieved by area detection based on machine learning. However, the focus area detection method is not limited to this. For example, the focus area can be set using the aspect ratio of the subject detection area, the size of the subject detection area, depth information of the subject using a defocus map, etc.
[0301] <Second embodiment> Next, a second embodiment will be described. In this embodiment, in the generation and output processing of an image in a virtual space, an image of the virtual space is generated and output using a captured image of the real space as well. The configuration of the imaging system 10 of this embodiment is the same as that of the first embodiment, but the generation and output processing of an image in a virtual space is partially different. Here, the description will focus on the differences from the first embodiment in the generation and output processing of an image in a virtual space.
[0302] The virtual space image generation and output processing of the second embodiment will be described with reference to the virtual space image generation and output processing sub-flowchart of FIG.
[0303] In S3501, the CPU 1001 acquires a captured image. The captured image may be an image captured in the real space imaging process described above, or may be an image that has been captured in advance.
[0304] In S3502, the CPU 1001 acquires and composites a foreground object. A 3D model of the subject may be generated from a captured image in real space using a trained model that estimates a 3D model for an image, and then composited with the 3D model of the subject in the virtual space described above. For example, a face may be acquired as a foreground object from a captured image in real space, and other body parts, such as the torso, may be acquired from the foreground object storage unit 1104 of the virtual space reproduction device 1100. Alternatively, the 3D model of the subject from the captured image and the 3D model of the subject in the virtual space described above may be acquired alternately in time series, and displayed at different times as a display image in the virtual space, as described below. Alternatively, a foreground object in the virtual space may be acquired and composited with a captured image in real space during the generation of a display image in the virtual space in S3503 described below.
[0305] In S3503, the CPU 1001 acquires and composites a background object. As with the acquisition of the foreground object in S3502, a 3D model may be generated from the captured image and used as the background object, or the background object may be composited with a separate area of the background object in the virtual space acquired by the background object acquisition unit 1105. As another method, the background object from the captured image and the background object in the virtual space may be separated in time series. Alternatively, the background object in the virtual space may be acquired and composited with the captured image in the generation of a display image in the virtual space in S3505, which will be described later.
[0306] In S3505, the CPU 1001 generates a display image for the virtual space. When a captured image is composited with a foreground object and a background object, the CPU 1001 performs processing similar to the virtual space display image generation processing of S2008 in FIG. 15 of the first embodiment. When generating a display image by compositing a foreground object, a background object, and a captured image in the virtual space, the CPU 1001 aligns and composites the captured image with the display image generated by the foreground object and the background object in the virtual space to generate the display image. For example, only the face portion of the subject in the captured image is cut out and composited with the display image in the virtual space.
[0307] Furthermore, if the image range or viewpoint position differs due to differences in the angle of view from the captured image, a pre-trained model that estimates a 3D model for the image is used to generate a 3D model from the captured image, and the image range or viewpoint position is changed to generate a display image.The display image of the virtual space is then divided into regions and synthesized to generate a display image of the virtual space that also includes the captured image.
[0308] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more of the functions.
[0309] The disclosure of this specification includes the following image generating apparatus, method, program, and storage medium.
[0310] (Item 1) A generating means for generating a virtual space image; a first acquisition means for acquiring operation information of a photographer with respect to a camera; a second acquiring means for acquiring information about the camera; a third acquisition means for acquiring information about the subject; a change means for changing a correction to the virtual space image captured by the camera based on at least one of the pieces of information obtained by the first to third acquisition means; An image generating device comprising:
[0311] (Item 2) 2. The image generating device according to item 1, wherein the change unit changes the amount of correction based on at least one of the pieces of information obtained by the first to third acquisition units.
[0312] (Item 3) 2. The image generating device according to item 1, wherein the change unit changes whether or not to perform the correction based on at least one of the pieces of information obtained by the first to third acquisition units.
[0313] (Item 4) 4. The image generating device according to any one of items 1 to 3, wherein the operation information for the camera includes operation information for a lens attached to the camera.
[0314] (Item 5) 5. The image generating device according to any one of items 1 to 4, wherein the information about the camera includes information about a lens attached to the camera.
[0315] (Item 6) 6. The image generating device according to any one of items 1 to 5, wherein the change means changes the correction based on information about an image captured before the timing of the intended capture.
[0316] (Item 7) The image generating device according to any one of items 1 to 6, characterized in that the change means determines the photographer's intention from the photographer's operation information and changes the correction based on the photographer's intention.
[0317] (Item 8) 8. The image generating device according to any one of items 1 to 7, wherein the operation information is information on at least one of a framing operation, a zooming operation, and a focusing operation.
[0318] (Item 9) 9. The image generating device according to any one of items 1 to 8, wherein the change means changes the correction based on the difficulty of photographing.
[0319] (Item 10) The image generating device described in item 9, characterized in that the change means determines the difficulty of shooting based on at least one of information on the subject's speed, acceleration, angular velocity, angular acceleration, subject's size, subject's contrast value, and distance between the subject and the photographer.
[0320] (Item 11) 11. The image generating device according to item 9 or 10, wherein the change means reduces the amount of correction as the degree of difficulty of photographing increases.
[0321] (Item 12) 11. The image generating device according to item 9 or 10, wherein the change means increases the amount of correction as the degree of difficulty of photographing increases.
[0322] (Item 13) a generation step of generating a virtual space image; a first acquisition step of acquiring operation information of the photographer with respect to the camera; a second acquiring step of acquiring information about the camera; a third acquisition step of acquiring information about the subject; a modification step of modifying a correction to the virtual space image captured by the camera based on at least one of the pieces of information obtained in the first to third acquisition steps; An image generating method comprising:
[0323] (Item 14) Item 14. A program for causing a computer to execute each step of the image generating method according to Item 13.
[0324] (Item 15) Item 14. A computer-readable storage medium storing a program for causing a computer to execute each step of the image generating method according to Item 13.
[0325] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0326] Camera: 100, 107: imaging element, 111: zoom actuator, 112: aperture actuator, 114: focus actuator, 121: camera CPU, 131: display, 1000: external computing device, 1001: CPU, 1100: virtual space reproduction device, 1200: virtual image generation device, 2000: camera / lens information storage device
Claims
1. A generating means for generating a virtual space image; a first acquisition means for acquiring operation information of a photographer with respect to a camera; a second acquiring means for acquiring information about the camera; a third acquisition means for acquiring information about the subject; a change unit that changes a correction to the virtual space image captured by the camera based on at least one of the pieces of information obtained by the first to third acquisition units; An image generating device comprising:
2. 2. The image generating apparatus according to claim 1, wherein the change unit changes the amount of correction based on at least one of the pieces of information obtained by the first to third acquisition units.
3. 2. The image generating apparatus according to claim 1, wherein the change unit changes whether or not the correction is to be performed based on at least one of the pieces of information obtained by the first to third acquisition units.
4. The image generating device according to claim 1 , wherein the operation information for the camera includes operation information for a lens attached to the camera.
5. 2. The image generating device according to claim 1, wherein the information about the camera includes information about a lens attached to the camera.
6. 2. The image generating apparatus according to claim 1, wherein the change means changes the correction based on information about an image captured before the timing of the image capture.
7. 2. The image generating apparatus according to claim 1, wherein said change means determines the photographer's intention from operation information of the photographer, and changes said correction based on the photographer's intention.
8. 2. The image generating device according to claim 1, wherein the operation information is information on at least one of a framing operation, a zooming operation, and a focusing operation.
9. 2. The image generating device according to claim 1, wherein the change means changes the correction based on the difficulty of photography.
10. The image generating device according to claim 9, wherein the change means determines the difficulty of photographing based on at least one of information on the subject's speed, acceleration, angular velocity, angular acceleration, size of the subject, contrast value of the subject, and distance between the subject and the photographer.
11. 10. The image generating device according to claim 9, wherein the change means reduces the amount of correction as the degree of difficulty of photographing increases.
12. 10. The image generating device according to claim 9, wherein the change means increases the amount of correction as the degree of difficulty of photography increases.
13. a generation step of generating a virtual space image; a first acquisition step of acquiring operation information of a photographer with respect to a camera; a second acquiring step of acquiring information about the camera; a third acquisition step of acquiring information about the subject; a modification step of modifying a correction to the virtual space image captured by the camera based on at least one of the pieces of information obtained in the first to third acquisition steps; An image generating method comprising:
14. A program for causing a computer to execute each step of the image generating method according to claim 13.
15. A computer-readable storage medium storing a program for causing a computer to execute each step of the image generating method according to claim 13.
Citation Information
Patent Citations
Imaging apparatus, imaging method and program
JP2008078908A