Image processing device and image processing method

The image processing device and method allow users to simulate camera operations in a virtual space, reflecting real camera actions in virtual images, enhancing the photography experience.

JP2025169797APending Publication Date: 2025-11-14CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024074920
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-02
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing image capturing technologies do not allow users to experience taking pictures in a virtual space that reflects the operation of the camera they are using.

Method used

An image processing device and method that generate a virtual image by capturing a virtual space with a virtual camera, reflecting the operations on a real camera, using an acquisition means to gather camera operation information and a generation means to output the virtual image for display.

Benefits of technology

Enables users to experience taking pictures in a virtual space that reflects the operation of their real camera, providing a simulated photography experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169797000001_ABST
    Figure 2025169797000001_ABST
Patent Text Reader

Abstract

To provide an image processing device and an image processing method that enable a user to experience taking pictures of a virtual space that reflects the operation of the camera they are using.SOLUTION: An image processing device generates a virtual image by capturing an image of a virtual space represented by a three-dimensional model with a virtual camera. The image processing device acquires information on operations performed on the real camera, and generates a virtual image in which the operations performed on the real camera are reflected in the virtual camera on the basis of the information on the operations. The image processing device outputs the virtual image for display on the real camera.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device and an image processing method. [Background technology]

[0002] Patent Document 1 discloses a camera that can generate a composite image of a simulated image of a scene that can be photographed at another location, generated from pre-stored three-dimensional data, and an actual image obtained by an imaging element. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-78908 Summary of the Invention [Problem to be solved by the invention]

[0004] By using the camera disclosed in Patent Document 1, a user can simulate the experience of taking pictures at a remote location without actually going to that location. However, the technology disclosed in Patent Document 1 does not allow the user to experience taking pictures in a virtual space that reflects the operation of the camera they are using.

[0005] The present invention has been made in consideration of the above-mentioned problems of the conventional technology. In one embodiment, the present invention provides an image processing device and an image processing method that enable a user to experience capturing images of a virtual space that reflect the operation of the camera they are using. [Means for solving the problem]

[0006] In one aspect, the present invention provides an image processing device that generates a virtual image by capturing an image of a virtual space represented by a three-dimensional model with a virtual camera, the image processing device comprising: an acquisition means for acquiring information about operations on the real camera; a generation means for generating a virtual image in which the operations on the real camera are reflected in the virtual camera based on the information about the operations; and an output means for outputting the virtual image for display on the real camera. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide an image processing device and an image processing method that enable a user to experience taking pictures of a virtual space that reflect the operation of the camera that the user is using. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of an imaging system according to an embodiment. [Figure 2] FIG. 1 is a block diagram showing an example of the functional configuration of a camera according to an embodiment; [Figure 3] A diagram showing the pixel arrangement of an imaging element [Figure 4] FIG. 1 shows an example of a pixel configuration. [Figure 5] A diagram showing an example of a focus detection area and its indicator [Figure 6] FIG. 1 is a block diagram showing an example of the hardware configuration of an external computing device according to an embodiment; [Figure 7] Block diagram showing an example of the functional configuration of an external computing device [Figure 8] 1 is a flowchart illustrating the operation of a camera according to an embodiment of the present invention; [Figure 9] 1 is a flowchart illustrating a real space imaging process according to an embodiment. [Figure 10] Flowchart for photographing processing in an embodiment [Figure 11] Flowchart for subject tracking AF processing in an embodiment [Figure 12] Flowchart for subject detection and tracking processing in an embodiment [Figure 13] Flowchart for virtual space photography processing in an embodiment [Figure 14] Flowchart for generating and outputting images of a virtual space in the first embodiment [Figure 15] Flowchart of virtual subject tracking processing in an embodiment [Figure 16] Flowchart for virtual photography in an embodiment [Figure 17] FIG. 1 is a diagram showing information on a camera lens storage device, a camera / lens, and an external computing device according to an embodiment. [Figure 18] Flowchart for virtual photography in an embodiment [Figure 19] Flowchart for acquiring photographing difficulty information in an embodiment [Figure 20] 10 is a flowchart illustrating a defocus amount processing process according to an embodiment. [Figure 21] 1 is a diagram showing corrections related to framing in an embodiment; [Figure 22] Zooming correction diagram according to an embodiment [Figure 23] 10A and 10B are diagrams relating to correction of focusing in an embodiment; [Figure 24] FIG. 10 is a diagram illustrating calculation of a virtual defocus amount according to an embodiment. [Figure 25] FIG. 1 is a diagram illustrating an example of a virtual defocus map according to an embodiment. [Figure 26] 10 is a flowchart of a subroutine of a viewpoint movement process in an embodiment. [Figure 27] An example of viewpoint movement in an embodiment [Figure 28] 1 is a flowchart illustrating the operation of a camera during virtual photography in an embodiment. [Figure 29] Flowchart for playback of photographed results and evaluation of photographed images in an embodiment [Figure 30] FIG. 10 is an explanatory diagram of a defocus map display in the embodiment. [Figure 31] FIG. 10 is an explanatory diagram of a focus degree display of a series of captured images in an embodiment. [Figure 32] FIG. 10 is an explanatory diagram of the best setting display after the evaluation result in the embodiment. [Figure 33] FIG. 10 is an explanatory diagram of a setting change list in an embodiment. [Figure 34] Flowchart for generating and outputting images in a virtual space in the second embodiment DETAILED DESCRIPTION OF THE INVENTION

[0009] The present invention will be described in detail below based on exemplary embodiments with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the claimed invention. Furthermore, although multiple features are described in the embodiments, not all of them are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] 1 is a block diagram showing an example of the configuration of an imaging system 10 using an image processing device according to an embodiment. A camera 100 is communicably connected to an external computing device 1000 as an image processing device according to an embodiment. Note that the functions of the external computing device 1000 may be performed by the camera 100. In this specification, the physical camera 100 may be referred to as a real camera to distinguish it from a virtual camera.

[0011] The external computing device 1000 has a virtual space reproduction unit 1100 and a virtual image generation unit 1200. The virtual space reproduction unit 1100 and the virtual image generation unit 1200 are schematic representations of the functions realized by the external computing device 1000. The virtual space reproduction unit 1100 manages three-dimensional data of the virtual space.

[0012] The virtual image generation unit 1200 acquires information about the configuration and state of the camera 100 from the camera 100. The information about the configuration and state of the camera 100 may include information about the current imaging optical system and imaging conditions, information about the state of the operation members, position information of the movable members, movement information of the camera 100, and the like.

[0013] The virtual image generator 1200 obtains information about a camera and / or lens different from the camera 100 from the camera / lens information storage device 2000. The camera / lens information storage device 2000 may be a separate device capable of communicating with the external computing device 1000, or may be a component of the external computing device 1000.

[0014] The virtual image generation unit 1200 generates an image (virtual image) of the virtual space captured by the virtual camera by rendering the three-dimensional data managed by the virtual space reproduction unit 1100 using information acquired from the camera 100 and the camera / lens information storage device 2000. The virtual image may be a two-dimensional image or a three-dimensional image (e.g., a stereo image). The virtual image generated by the virtual image generation unit 1200 is an image obtained by simulating the operation of the camera and / or lens (virtual camera) that acquired information from the camera / lens information storage device 2000, based on information acquired from the camera 100. The virtual image for display is a moving image for use in live view display on the display unit 131 of the camera 100. Furthermore, the virtual image for recording is expected to be a still image, but may also be a moving image.

[0015] 2 is a block diagram showing an example of the functional configuration of camera 100. Camera 100 has an imaging optical system that includes a first lens group 101, an aperture 102, a second lens group 103, and a third lens group 105, and forms an optical image of a subject on the imaging surface of an image sensor 107.

[0016] The first lens group 101 is disposed at the front (closest to the subject) of the multiple lens groups included in the photographic optical system and is movable along the optical axis OA. The position of the first lens group 101 is controlled by a zoom actuator 111. The zoom actuator 111 moves the first lens group 101 and the second lens group 103 in conjunction with each other along the optical axis by, for example, driving a cam barrel (not shown).

[0017] The aperture 102 has an opening that can be adjusted by an aperture actuator 112 .

[0018] The second lens group 103 moves along the optical axis OA integrally with the aperture 102 and in conjunction with the first lens group 101. The angle of view (focal length) of the photographic optical system is determined by the positions of the first lens group 101 and the second lens group 103.

[0019] The third lens group 105 is movable along the optical axis OA. The position of the third lens group 105 is controlled by a focus actuator 114. The position of the third lens group determines the focal length of the photographic optical system. The third lens group 105 is also called a focus lens.

[0020] The shutter unit 106 includes a focal plane shutter and its drive circuit. Based on instructions from the CPU 121, the shutter unit 106 controls the operation of the focal plane shutter via the drive circuit so that light is incident on the image sensor 107 only during the exposure period. Note that if the diaphragm 102 is also used as a mechanical shutter, the shutter unit 106 is not necessary.

[0021] The optical low-pass filter 106 is provided to reduce false colors and moire that occur in captured images.

[0022] The image sensor 107 is a CMOS image sensor or a CCD image sensor having a rectangular pixel array (also called a pixel area) in which m pixels are arranged horizontally and n pixels are arranged vertically in a two-dimensional manner. Each pixel is provided with a color filter in a primary color Bayer array, for example, and an on-chip microlens. The image sensor 107 may also be a three-chip color image sensor.

[0023] In this embodiment, the lens unit is configured as an interchangeable lens that can be attached to and detached from the body of the camera 100 via a mount part M. The lens unit has a photographic optical system, a zoom actuator 111, an aperture actuator 112, a focus actuator 114, a focus drive circuit 126, an aperture drive circuit 128, and a zoom drive circuit 129. The lens unit may be configured as an integral part of the body of the camera 100.

[0024] The flash 115 is a light source that illuminates the subject. The flash 115 is equipped with a flashlight device using a xenon tube or a continuous-light-emitting LED (light-emitting diode). The AF (autofocus) auxiliary light source 116 projects a predetermined pattern image through a projection lens. This improves focus detection capability for low-brightness or low-contrast subjects.

[0025] The CPU 121 controls the overall operation of the camera 100. The CPU 121 loads a program stored in the ROM 135 into the RAM 136 and executes it to control each unit of the camera 100 and realize the functions of the camera 100, such as autofocus detection (AF), image capture, image processing, and recording. Some of the functions realized by the CPU 121 executing the program may be implemented by a hardware circuit separate from the CPU 121. Some of the circuits may also use a reconfigurable circuit such as an FPGA.

[0026] A flash control circuit 122 controls the lighting of the flash 115 in synchronization with the imaging operation. An auxiliary light source drive circuit 123 controls the lighting of the AF auxiliary light source 116 in synchronization with the focus detection process. An image sensor drive circuit 124 controls the imaging operation of the image sensor 107, and also A / D converts the signal obtained by the imaging operation and transmits it to the CPU 121.

[0027] Image processing circuit 125 can apply various image processing to image data, such as gamma conversion, white balance adjustment, color interpolation, encoding, decoding, evaluation value generation, and feature region detection. Image processing circuit 125 determines a main subject region from the subject regions detected by subject detection unit 140 using one or more of the position and size of the subject region and the positions of objects present in the shooting scene. Information on the determined main subject region may be used for image processing (e.g., evaluation value generation, white balance adjustment, etc.).

[0028] The focus driving circuit 126 drives the focus actuator 114 based on a command including the driving amount and driving direction of the focus lens given from the CPU 121. This causes the third lens group 105 to move along the optical axis OA, changing the focal length of the photographic optical system.

[0029] The aperture drive circuit 128 drives the aperture actuator 112 to control the aperture diameter and opening / closing of the aperture 102. The zoom drive circuit 129 drives the zoom actuator 111 in response to, for example, a user's instruction, and changes the focal length (angle of view) of the photographic optical system by moving the first lens group 101 and the second lens group 103 along the optical axis OA.

[0030] The display unit 131 has, for example, an LCD (liquid crystal display device). The display unit 131 displays information about the imaging mode of the camera 100, a preview image before imaging, a confirmation image after imaging, or an in-focus state display image during focus detection.

[0031] The operation unit 132 is a collective term for input devices (buttons, switches, dials, etc.) provided for the user to input various instructions to the camera 100. The input devices constituting the operation unit 117 are named according to the functions assigned to them. For example, the operation unit 117 includes a release switch, a video recording switch, a shooting mode selection dial for selecting a shooting mode, a menu button, directional keys, an enter key, etc. The release switch is a switch for recording still images, and the control unit 101 recognizes the half-pressed state of the release switch (SW1 is on) as an instruction to prepare for shooting and the full-pressed state (SW2 is on) as an instruction to start shooting. Furthermore, the control unit 101 recognizes pressing the video recording switch in the shooting standby state as an instruction to start video recording, and pressing the switch during video recording as an instruction to stop recording. Note that if the lens unit has input devices such as a zoom ring or focus ring, these also constitute the operation unit 132.

[0032] The recording medium 133 is, for example, a semiconductor memory card that is detachable from the camera 100. Still image data and video data obtained by capturing images are recorded on the recording medium 133 by the CPU 121.

[0033] Note that if the display unit 131 is a touch display, a touch panel or a combination of a touch panel and a GUI displayed on the display unit 131 may be used as the operation unit 132. For example, when a tap operation on the touch panel is detected during live view display, the image area corresponding to the tap position can be configured to perform focus detection as a focus detection area.

[0034] It is also possible to calculate contrast information of captured image data using the image processing circuit 125, and have the CPU 121 perform contrast AF. In contrast AF, contrast information is calculated sequentially while moving the focus lens group 105 to change the in-focus distance of the photographing optical system, and the focus lens position at which the contrast information reaches its peak is set as the in-focus position.

[0035] In this way, camera 100 can perform both image-plane phase-difference AF and contrast AF, and can selectively use one or a combination of both depending on the situation.

[0036] The subject detection unit 140 can be configured using, for example, a convolutional neural network (CNN). By configuring the CNN using parameters (dictionary data) generated by machine learning for each type of subject, the area of ​​a specific subject present in the image represented by the image data is detected. The subject detection unit 140 may be implemented using dedicated hardware configured to enable high-speed execution of CNN-based processing operations, such as a graphics processing unit (GPU) or a neural processing unit (NPU).

[0037] Machine learning for generating dictionary data can be performed by any known method, such as supervised learning. Specifically, a CNN can be trained using a data set in which, for each type of subject, input images are associated with whether or not a target subject is captured. The trained CNN or its parameters can be stored in ROM 135 as dictionary data. Note that CNN training may be performed on a device different from camera 100. When a trained CNN is used for subject detection processing of a captured image, an image of the same size as the input image used for CNN training is cut out from the captured image and input to the CNN. By inputting the image to the CNN while sequentially changing the cutout position, the area in which the target subject is captured can be estimated.

[0038] Note that the object region may be detected using other methods, such as detecting an object region in an image and then determining the type of object in the object region using feature values ​​for each type of object. The configuration and learning method of the neural network can be changed depending on the detection method used.

[0039] The subject detection unit 140 can be implemented using any known method as long as it can output the number, position, size, and reliability of areas in the input image that are estimated to contain a predetermined type of subject.

[0040] The subject detection unit 140 can apply subject detection processing for multiple types of subjects to one frame of image data by repeatedly switching between dictionary data. The CPU 121 can determine which dictionary data to use for the subject detection processing from among the multiple dictionary data stored in the ROM 135 based on the priority of the subject types set in advance, the setting values ​​of the camera 100, etc.

[0041] The types of subjects may be, for example, but are not limited to, human bodies, human organs (faces, eyes, torsos, etc.), non-human subjects (animals, non-living objects (tools, vehicles, buildings, etc.)), etc. Separate dictionary data is prepared for subjects with different characteristics.

[0042] The dictionary data for detecting a human body may be prepared separately as dictionary data for detecting a human body (contour) and dictionary data for detecting organs of the human body. The dictionary data for detecting organs of the human body may be prepared separately for each type of organ.

[0043] The video input unit 141 and the information output unit 142 are communication interfaces with external devices. The video input unit 141 and the information output unit 142 support, for example, one or more well-known wired or wireless communication standards. If the video input unit 141 and the information output unit 142 support the same standard, the video input unit 141 and the information output unit 142 may be implemented as a single input / output interface. There are no particular restrictions on the wired or wireless communication standards that the video input unit 141 and the information output unit 142 may support, but examples include USB, Thunderbolt, HDMI (registered trademark), DVI, SDI, WiFi, and Ethernet (registered trademark).

[0044] The video input unit 141 receives a virtual space image from the virtual image generation unit 1200 of the external computing device 1000. The CPU 121 temporarily stores the received virtual space image in the RAM 136. The virtual space image can then be used for display on the display unit 131, for synthesis with a real-life image, for recording on the recording medium 133, and so on.

[0045] The information output unit 142 is used to output various types of information about the camera 100 to the external computing device 1000. For example, information about the operation, settings, control, etc. of the camera 100 can be output. The information about the operation may include information about the state of the input devices included in the operation unit 132. Specifically, it may include information about the state of the release switches (SW1 and SW2), zooming operations, focus operations, etc. The information about the settings may include information about the shooting mode (single shooting or continuous shooting mode for still images, video mode, etc.), AF area, AF mode, metering mode, exposure conditions (aperture value, shutter speed, sensitivity), frame rate, image processing, lens control, etc. The information about the control may include information about correction values ​​and thresholds used in shooting and image processing, information about the movement of the camera 100, etc. The information supplied from the camera 100 to the external computing device 1000 is not limited to these.

[0046] ●(image sensor) The pixel array and pixel structure of the image sensor 107 will be described with reference to Figures 3 and 4. The left-right direction in Figure 3 is the x-direction (horizontal direction), the up-down direction is the y-direction (vertical direction), and the direction perpendicular to the x-direction and y-direction (direction perpendicular to the paper surface) is the z-direction (optical axis direction). In the example shown in Figure 3, the pixel (unit pixel) array of the image sensor 107 is shown as an area of ​​4 columns x 4 rows, and the sub-pixel array is shown as an area of ​​8 columns x 4 rows.

[0047] In the pixel group 200 of 2 columns x 2 rows, for example, a pixel 200R having a spectral sensitivity of a first color R (red) is arranged at the upper left position, a pixel 200G having a spectral sensitivity of a second color G (green) is arranged at the upper right and lower left position, and a pixel 200B having a spectral sensitivity of a third color B (blue) is arranged at the lower right position. Furthermore, each pixel (unit pixel) is divided into two in the x direction (Nx divisions) and one in the y direction (Ny divisions), resulting in a division number of 2 (division number N LF =Nx×Ny) first subpixel 201 and second subpixel 202 (the first to Nth subpixels LF It is composed of multiple sub-pixels (sub-pixels).

[0048] In the example shown in FIG. 3, each pixel of the image sensor 107 is divided into two sub-pixels arranged in the horizontal direction, and the number of divisions N is calculated from the image signal obtained by one image capture. LF It is possible to generate a number of viewpoint images equal to the number of sub-pixels in the image sensor 107, and a captured image obtained by combining all the viewpoint images. Note that the pixels may be divided in two directions, and there is no limit to the number of divisions in each direction. Therefore, it can be said that the viewpoint images are images generated from signals of some of the sub-pixels among the plurality of sub-pixels, and the captured image is an image generated from signals of all the sub-pixels. In this embodiment, as an example, the pixel period P in the horizontal and vertical directions of the image sensor 107 is set to 4 μm, and the number of horizontal pixels N H =5575, number of vertical pixels N V = 3725. Therefore, the total number of pixels N = N H ×N V = approximately 20.75 million. In addition, the horizontal period of the sub-pixels P S If the total number of sub-pixels is N S =N H ×(P / P S )×N V = approximately 41.5 million.

[0049] FIG. 4(a) shows a plan view of one pixel 200G of the image sensor 107 shown in FIG. 3, as viewed from the light receiving surface side (+z side) of the image sensor 107. The z-axis is set perpendicular to the plane of FIG. 4(a), with the positive direction of the z-axis defined as the front side. The y-axis is set perpendicular to the z-axis, with the positive direction of the y-axis being the up-down direction, and the x-axis is set perpendicular to the z-axis and y-axis, with the positive direction of the x-axis being the right side. FIG. 4(b) shows a cross-sectional view taken along the aa section line in FIG. 4(a), as viewed from the -y side.

[0050] As shown in Figures 4(a) and 4(b), in the pixel 200G, a microlens 305 is formed on the light receiving surface side (+z direction) of each pixel, and incident light is condensed by this microlens 305. Furthermore, a plurality of photoelectric conversion units, namely, first photoelectric conversion units 301 and second photoelectric conversion units 302, are formed, which are divided into two in the x (horizontal) direction and one in the y (vertical) direction, with a division number of 2. The first photoelectric conversion unit 301 and the second photoelectric conversion unit 302 correspond to the first subpixel 201 and the second subpixel 202 in Figure 2, respectively. More generally, the photoelectric conversion unit of each pixel is divided into Nx divisions in the x direction and Ny divisions in the y direction, and the division number of the photoelectric conversion unit is N LF = Nx × Ny, the first to Nth LF The photoelectric conversion units are 1st to Nth LF It corresponds to a sub-pixel.

[0051] The first photoelectric conversion unit 301 and the second photoelectric conversion unit 302 are two independent pn junction photodiodes, each consisting of a p-type well layer 300 and two divided n-type layers 301 and 302. If necessary, an intrinsic layer may be sandwiched between them to form a pin structure photodiode. In each pixel, a color filter 306 is formed between the microlens 305 and the first photoelectric conversion unit 301 or second photoelectric conversion unit 302. If necessary, the spectral transmittance of the color filter 306 may be changed for each pixel or each photoelectric conversion unit, or the color filter may be omitted.

[0052] Light incident on pixel 200G is collected by microlens 305 and further dispersed by color filter 306, after which it is received by first photoelectric conversion unit 301 and second photoelectric conversion unit 302, respectively. In first photoelectric conversion unit 301 and second photoelectric conversion unit 302, pairs of electrons and holes (positive holes) are generated according to the amount of received light, and after they are separated by a depletion layer, the electrons are accumulated. Meanwhile, the holes are discharged to the outside of image sensor 107 through a p-type well layer connected to a constant voltage source (not shown). The electrons accumulated in first photoelectric conversion unit 301 and second photoelectric conversion unit 302 are transferred to a capacitance unit (FD) via a transfer gate and converted into a voltage signal.

[0053] In this embodiment, the microlenses 305 correspond to the optical system in the image sensor 107. The optical system in the image sensor 107 may be configured to use microlenses as in this embodiment, or may be configured to use materials with different refractive indices, such as waveguides. The image sensor 107 may be a back-illuminated image sensor having circuits and the like on the surface opposite to the surface having the microlenses 305, or may be a stacked image sensor having some circuits, such as the image sensor drive circuit 124 and the image processing circuit 125. A material other than silicon may be used for the semiconductor substrate, and for example, an organic material may be used as the photoelectric conversion material.

[0054] Light incident from the photographing optical system is collected by microlens 305, dispersed by color filter 306, and then reaches photoelectric conversion units 301 and 302, where it is photoelectrically converted. Camera 100 having image sensor 107 is capable of autofocus detection based on the phase difference between a pair of signal sequences obtained by splitting a light beam passing through the photographing optical system, using known technology such as that described in JP 2023-95509 A, for example.

[0055] The following describes the pixel areas of the image sensor 107 that are used to generate a pair of signal strings (first and second focus detection signals) used for autofocus detection (focus detection areas). FIG. 5 shows an example of focus detection areas set in the effective pixel area 300 of the image sensor 107, superimposed with focus detection area indicators displayed on the display unit 131 during focus detection. In this embodiment, a total of nine focus detection areas are set, three in the row direction and three in the column direction, but this is merely an example, and more or fewer focus detection areas may be set. The size, position, and spacing of the focus detection areas may also vary.

[0056] Furthermore, when every pixel in the effective pixel area 300 has a first sub-pixel 201 and a second sub-pixel 202, as in the image sensor 107, the position and size of the focus detection area may be set dynamically. For example, a predetermined range may be set as the focus detection area, centered on a position specified by the user. In this embodiment, the focus detection area is set so that focus detection results can be obtained with higher resolution when acquiring a defocus map, which will be described later. For example, the effective pixel area 1000 is divided into 120 areas horizontally and 80 areas vertically, for a total of 9600 areas, each of which is set as a focus detection area.

[0057] 5, the nth focus detection area in the row direction and the mth focus detection area in the column direction is represented as A(n,m), and the rectangular frame-shaped index representing the focus detection area of ​​A(n,m) is represented as I(n,m). Images A and B used to detect the defocus amount in that focus detection area are generated from signals obtained from the first sub-pixel 201 and the second sub-pixel 202 in that focus detection area. Furthermore, the index I(n,m) is usually displayed superimposed on a live view image.

[0058] 6 is a block diagram showing an example of the hardware configuration of the external computing device 1000. The external computing device 1000 may be any computer device capable of executing a program, such as a personal computer or a tablet device. The external computing device 1000 may include components not shown. The functions of the external computing device 1000, which will be described later, can be realized, for example, by executing a specific application program on the computer device.

[0059] The CPU 1001 loads programs stored in the ROM 1002 and the storage unit 1004 into the RAM 1003 and executes them to control the components of the external processing device 1000 and realize the functions of the external processing device 1000. The CPU 1001 may be one or more processors.

[0060] The RAM 1003 is used as a main memory and a work memory for the CPU 1001. A part of the RAM 1003 is also used as a video memory for the display unit 1008.

[0061] The ROM 1002 stores programs executed by the CPU 1001 and associated data, various setting values ​​of the external processing device 1000, etc. The ROM 1002 may be a rewritable nonvolatile memory.

[0062] The storage unit 1004 stores programs executed by the CPU 1001, data used when executing the programs (GUI data, setting values, etc.), user data, etc. The storage unit 1004 can be a known storage device or storage medium such as an SSD, HDD, or optical disk drive.

[0063] The display unit 1008 displays a screen, GUI, etc. provided by a program executed by the CPU 1001. The display unit 1008 may be an external display device connected to the external computing device 1000. The display unit 1008 may be a liquid crystal display, an organic EL display, an HMD, etc.

[0064] The operation unit 1009 is a general term for input devices that can be used by a user to input instructions to the external computing device 1000. The operation unit 1009 can include, for example, one or more of a keyboard, a mouse, a joystick, a touchpad, etc. When the CPU 1001 detects an operation on the operation unit 1009, it executes processing according to the detected operation.

[0065] The input interface (IF) 1005 is a communication interface through which the external processing device 1000 receives signals or data from external devices, and the output interface (IF) 1006 is a communication interface through which the external processing device 1000 transmits signals or data to external devices.

[0066] The input IF 1005 and the output IF 1006 support, for example, one or more well-known wired or wireless communication standards. If the input IF 1005 and the output IF 1006 support the same standard, the input IF 1005 and the output IF 1006 may be implemented as a single input / output interface. There are no particular limitations on the wired or wireless communication standards that the input IF 1005 and the output IF 1006 may support, but these may include USB, Thunderbolt, HDMI (registered trademark), DVI, SDI, WiFi, Ethernet (registered trademark), etc.

[0067] The video input unit 141 of the camera 100 is connected to the output IF 1006, and the information output unit 142 of the camera 100 is connected to the input IF 1005. The external calculation device 1000 can transmit, via the output IF 1006, to the camera 100, image data to be displayed on the display unit 131 and image data to be recorded on the storage medium 133. In addition, the external calculation device 1000 can receive, via the input IF 1005, various types of information related to the camera 100 (main body and lens unit), captured image data, and the like from the camera 100.

[0068] The above-mentioned components of the external computing device 1000 are connected to each other via a system bus 1007 so that they can communicate with each other.

[0069] The camera / lens information storage device 2000 is a storage device that can be accessed directly or indirectly by the external computing device 1000. The camera / lens information storage device 2000 may be a known storage device such as an SSD or HDD. When the camera / lens information storage device 2000 is an external device of the external computing device 1000, the external computing device 1000 accesses the camera / lens information storage device 2000 via the input IF 1005 and the output IF 1006. When the camera / lens information storage device 2000 is a component of the external computing device 1000, the camera / lens information storage device 2000 is connected to the system bus 1007.

[0070] The camera / lens information storage device 2000 stores information necessary for the external computing device 1000 to simulate the operation of a virtual camera. The camera / lens information storage device 2000 stores information necessary for simulating the operation of various models of real cameras and lenses. Therefore, the external computing device 1000 can simulate the operation of various models of camera and lens combinations as the operation of a virtual camera based on the information acquired from the camera / lens information storage device 2000. The camera / lens information storage device 2000 can store information about the camera, such as its configuration and performance, configurable parameters, and operating parameters for each setting. The camera / lens information storage device 2000 can also store information about the lens unit, such as its angle of view range, configurable parameters, and operating parameters for each setting. The camera / lens information storage device 2000 can also store any information necessary for simulating the operation of the camera and lens unit.

[0071] Next, the virtual image generation process performed by the external computing device 1000 will be described with reference to Fig. 7. Fig. 7 is a block diagram showing an example of the functional configuration of the virtual space reproduction unit 1100 and the virtual space generation unit 1200 of the external computing device 1000. The operation of each functional block (excluding the storage unit) shown in Fig. 7 is realized by the CPU 1001 of the external computing device 1000 reading a program stored in the ROM 1002 and the storage unit 1004 into the RAM 1003 and executing it. Note that the operation may also include an operation realized by hardware of the external computing device 1000 that is different from the CPU 1001.

[0072] First, the virtual space reproduction unit 1100 will be described. The foreground object storage unit 1101 (for example, the storage unit 1004) stores a 3D model of a foreground subject. A foreground subject is a subject that exists in front of the background, such as a human subject. A 3D model is 3D shape data that describes information indicating shape and color. A 3D model is composed of a textured mesh model or a 3D point cloud where each point is colored. Note that the 3D model does not have to be colored. The 3D model may also have information such as velocity, acceleration, angular velocity, angular acceleration, size, and contrast as subject information.

[0073] The subjects whose three-dimensional models are stored in the foreground object storage unit 1101 may be of various types, such as human subjects of different races, genders, and ages, various animals, and moving objects (cars, airplanes, trains, ships, etc.).

[0074] The 3D model can be generated by any known generation method. Alternatively, an image of a specific subject may be input into a trained machine learning model to generate a 3D model of the specific subject. In this case, information on the size and contrast of the 3D model can also be obtained during generation. Furthermore, information on the speed, acceleration, angular velocity, and angular acceleration of the subject for which the 3D model is to be generated can be obtained based on information on the movement of the subject between multiple time-series images.

[0075] As another alternative, a 3D model may be generated and stored based on multiple viewpoint images and a parameter set received from a camera, for example, using a method described in Japanese Patent Application Laid-Open No. 2017-211827. Multiple cameras capture images of an imaging area from multiple directions. This imaging area may be, for example, an indoor imaging studio or a theatrical stage. The multiple cameras are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple cameras do not need to be installed around the entire periphery of the imaging area; depending on installation location restrictions, they may be installed in only certain directions of the imaging area. The number of cameras can be set using various methods. For example, if the imaging area is a soccer stadium, approximately 30 cameras may be installed around the stadium. Cameras with different functions, such as telephoto and wide-angle cameras, may also be installed.

[0076] A parameter set may be described for each camera, including parameters representing the three-dimensional position of each of multiple cameras, parameters representing the camera's direction in the pan, tilt, and roll directions, and the size of the camera's field of view (angle of view) and resolution. The information included in the parameter set is calculated in advance using a known camera calibration procedure and stored in an appropriate storage device. That is, points in multiple images captured by multiple cameras are associated with each other and calculated using geometric calculations. Note that the content of the information included in the parameter set is not limited to the above. For example, the parameter set may have multiple parameter sets corresponding to multiple frames constituting a video captured by the camera, and the information may indicate the position and direction of the camera at each of multiple consecutive points in time.

[0077] The three-dimensional model may be generated by the external computing device 1000. For example, the external computing device 1000 receives images from the above-described multiple cameras via the input IF 1005. Then, the foreground object acquisition unit 1102 generates a three-dimensional model using the images and the parameter set stored in the foreground object storage unit 1101, and stores the generated three-dimensional model in the foreground object storage unit 1101.

[0078] The background object storage unit 1104 (e.g., storage unit 1004) stores a three-dimensional model (background object) representing an environment or space in which the foreground object is placed. The background object may represent a large-scale environment such as a concert hall, a soccer stadium, or a baseball field, or a small-scale environment such as a room in a house.

[0079] Background objects can be generated using design data such as CAD, 3D data measured using a laser scanner, or 3D data obtained from a group of images from multiple viewpoints using computer vision techniques such as Structure from Motion. These are merely examples, and background objects can be generated based on any known method. The present invention does not depend on the method for generating foreground and background objects.

[0080] The foreground object acquisition unit 1102 acquires one or more foreground objects stored in the foreground object storage unit 1101. The foreground object acquisition unit 1102 outputs the acquired foreground objects to the object composition unit 1103.

[0081] Furthermore, the background object acquisition unit 1105 acquires one or more background objects stored in the background object storage unit 1104. The background object acquisition unit 1102 outputs the acquired background object to the object composition unit 1103.

[0082] The object composition unit 1103 places the foreground object in the space represented by the background object. The information about the foreground object that the object composition unit 1103 receives from the foreground object acquisition unit 1102 may include a three-dimensional model of the foreground object associated with a plurality of different times.

[0083] The object composition unit 1103 generates a composite object by placing the foreground object on the background object so that it does not interfere with other objects and so that the positional relationship with the background object does not look unnatural. For example, if the foreground object should be in contact with the ground, the object composition unit 1103 places the foreground object so that it appropriately contacts the ground of the background object. The object composition unit 1103 may automatically place the foreground object according to object placement information (position, orientation) that the background object has, or may place the foreground object according to a user instruction via the operation unit 1109. The object composition unit 1103 outputs the generated composite object to the virtual image generation unit 1200.

[0084] Next, the virtual image generator 1200 will be described. The viewpoint information acquisition unit 1201 acquires information about a virtual viewpoint (virtual viewpoint information) including the position of the viewpoint (virtual viewpoint) in the virtual space and the line of sight direction (pan (horizontal angle), tilt (vertical angle), roll (rotation angle)). The virtual viewpoint information may be stored in the ROM 1002 or the storage unit 1004 as an initial value, a registered value, a previous history position, etc., or may be set by the user.

[0085] The camera lens information acquisition unit 1202 acquires information about the camera and lens used in virtual photography (camera lens information) from the camera / lens information storage device 2000 or the camera 100. Details of the acquired information will be described later.

[0086] The camera lens information update unit 1203 acquires and updates the camera lens information, which is updated over time, each time.

[0087] The operation information acquisition unit 1205 acquires information (operation information) relating to operations performed on the camera 100 and the lens unit from the camera 100. Details of the acquired information will be described later.

[0088] The virtual viewpoint information, camera lens information, and operation information are input to an image correction amount calculation unit 1206. The image correction amount calculation unit 1206 calculates the image correction amount using a shooting difficulty calculation unit 1261 and a user intention extraction unit 1262. Details of the processing will be described later.

[0089] The display image generation unit 1204 renders the composite object acquired from the object composition unit 1103 using the virtual viewpoint information and camera lens information, and generates an image (virtual image) of the virtual space captured by the virtual camera. The display image generation unit 1204 transmits the generated virtual image to the camera 100 via the output IF 1006. In the camera 100, the virtual image is displayed on the display unit 131, for example, or recorded in the storage medium 133. The display image generation unit 1204 may also store the generated virtual image in the storage unit 104.

[0090] ●(Photography processing) Next, the photographing operation of camera 100 will be described using the flowchart shown in Fig. 8. Here, it is assumed that still image photographing in real space or virtual space can be selectively performed from the photographing standby state until the main switch is turned off.

[0091] In S1, the CPU 121 starts live view display on the display unit 131 as processing in a shooting standby state. The operation of generating a moving image (live view image) used for the live view display will be described later. Furthermore, the CPU 121 displays a menu setting screen on the display unit 131 to allow the user to select whether to shoot in real space or in virtual space at a predetermined timing such as at startup or in response to detection of a specific user operation. If the setting has already been made, it is not necessary to display the menu setting screen.

[0092] In S2, the CPU 121 determines whether or not virtual space photography is set, and if it is determined that it is set, executes S1000, and if not, executes S10. The CPU 121 executes real space photography processing in S10 and virtual space photography processing in S1000. Details of these processes will be described later.

[0093] In S3, the CPU 121 determines whether the main switch (power switch) included in the operation unit 132 has been turned off, and if it is determined that it has been turned off, ends the processing of the flowchart in Figure 8, and if it is not determined that it has been turned off, executes S1 again.

[0094] ●(Real-space photography processing) The details of the real space shooting process executed in S10 of FIG. 8 will be described with reference to the flowchart shown in FIG. In S11, the CPU 121 starts driving the image sensor 107 via the image sensor drive circuit 124 in order to capture a moving image to be displayed on the display unit 131. Thereafter, the image sensor 107 outputs an analog image signal at a predetermined frame rate.

[0095] When the CPU 121 acquires one frame's worth of analog image signals from the image sensor 107, it applies correlated double sampling, A / D conversion, and the like to generate a digital image signal. The CPU 121 outputs the digital image signal to the signal processing circuit 125. The signal processing circuit 125 applies demosaic processing and the like to the digital image signal to generate image data for display. The signal processing circuit 125 writes the image data for display to, for example, a video memory area of ​​the RAM 136. The signal processing circuit 125 also generates an evaluation value to be used in AE processing from the digital image signal and outputs it to the CPU 121. Furthermore, the signal processing circuit 125 generates first and second focus detection signals for each of a plurality of focus detection areas based on signals read from pixels included in the focus detection area, and outputs the evaluation values ​​to the CPU 121.

[0096] When the first and second subpixels 201 and 202 are configured as separate pixels (do not share the same microlens), the pixel coordinates at which a signal is obtained from the first subpixel 201 are different from the pixel coordinates at which a signal is obtained from the second subpixel 202. Therefore, the signal processing circuit 125 generates first and second focus detection signals by interpolating the signals so that a signal pair of the first and second subpixels 201 and 202 exists at the same pixel position.

[0097] In S12, the CPU 121 supplies the display image data stored in the video memory area of ​​the RAM 135 to the display unit 131, and displays it as one frame of a live view image. The user can adjust the imaging range, exposure conditions, etc. while viewing the live view image displayed on the display unit 131. The CPU 121 determines the exposure conditions based on the evaluation value obtained from the signal processing circuit 125, and displays an image indicating the determined exposure conditions (shutter speed, aperture value, imaging ISO sensitivity) on the display unit 131, superimposed on the live view image.

[0098] Thereafter, the CPU 121 executes the operation of S2 each time the imaging of one frame is completed, thereby causing the display unit 131 to function as an EVF.

[0099] In S13, CPU 121 determines whether or not a half-press operation (SW1 ON) of the release switch included in operation unit 132 has been detected. If CPU 121 determines that SW1 ON has not been detected, CPU 121 repeatedly executes S3. On the other hand, if CPU 121 determines that SW1 ON has been detected, CPU 121 executes S400.

[0100] In S400, the CPU 121 executes subject tracking autofocus (AF) processing. In S400, the CPU 121 applies subject detection processing to image data for display and determines a focus detection area. The CPU 121 also executes predictive AF processing to prevent a decrease in AF accuracy due to the time lag between the execution of AF processing and the detection of a full press of the release switch (SW2 ON). Details of the operation in S400 will be described later.

[0101] In S15, the CPU 121 determines whether or not ON of SW2 is detected. If the CPU 121 determines that ON of SW2 is not detected, it executes S13. On the other hand, if it determines that ON of SW2 is detected, the CPU 121 executes the photographing and recording process of S300. Details of the operation of S300 will be described later.

[0102] Although the subject detection process and AF process have been described as being executed in response to the detection of SW1 being turned on, they may be executed at other times. If the subject tracking AF process of the S300 is executed before the detection of SW1 being turned on, it becomes possible to start capturing images immediately by pressing the shutter button all the way down, omitting the halfway press operation.

[0103] ●(Photography and recording processing) Next, the photographing and recording process executed by the CPU 121 in S300 of FIG. 8 will be described with reference to the flowchart shown in FIG. In S301, the CPU 121 determines exposure conditions (shutter speed, aperture value, imaging ISO sensitivity, etc.) through AE processing based on the evaluation value generated by the signal processing circuit 125. Then, the CPU 121 controls the operation of each unit so that a still image is captured according to the determined exposure conditions.

[0104] That is, the CPU 121 transmits the aperture value to the aperture drive circuit 128 and the shutter speed to the shutter 106 to drive the aperture 102 and the shutter 106. The CPU 121 also controls the charge accumulation operation of the image sensor 107 via the image sensor drive circuit 124.

[0105] In S302, the CPU 121 reads out one frame's worth of analog image signals from the image sensor 107 via the image sensor drive circuit 124. Note that for at least pixels within the focus detection area, the CPU 121 also reads out the signals of one of the first and second sub-pixels 201 and 202.

[0106] In S303, the CPU 121 A / D converts the signal read out in S402 to a digital image signal. The CPU 121 also applies defective pixel correction processing to the digital image signal using the image processing circuit 125. The defective pixel correction processing is a process in which a signal read out from a pixel (defective pixel) from which a normal signal cannot be read out is complemented by a signal read out from a surrounding normal pixel.

[0107] In S304, the CPU 121 causes the image processing circuit 125 to generate a still image data file for recording and first and second focus detection signals. The image processing circuit 125 applies image processing, encoding processing, etc. to the digital image signal after the defective pixel correction processing to generate still image data for recording. The image processing may be, for example, demosaic (color interpolation) processing, white balance adjustment processing, gamma correction (tone correction) processing, color conversion processing, edge enhancement processing, etc. The image processing circuit 125 also applies encoding processing to the still image data in a format that corresponds to the format of the data file in which the still image data is stored.

[0108] In S305, the CPU 121 records, on the recording medium 133, an image data file that stores the still image data generated in S304 and the sub-pixel signals read from the focus detection area in S302.

[0109] In S306, the CPU 121 records device characteristic information as characteristic information of the camera 100 in the recording medium 133 in association with the image data file recorded in S305.

[0110] The device characteristic information includes, for example, the following information: Imaging conditions (aperture value, shutter speed, sensitivity, etc.) Information about the image processing applied to the digital image signal by the image processing circuit 125 Information about the light-receiving sensitivity distribution of the imaging pixels and sub-pixels of the imaging element 107 Information about vignetting of the imaging light beam within the camera 100 Information on the distance from the mounting surface (mount part M) of the photographing optical system (lens unit) of the camera 100 to the image sensor 107 Information about the manufacturing tolerances of the Camera 100

[0111] Information on the light sensitivity distribution of the imaging pixels and sub-pixels (hereinafter simply referred to as light sensitivity distribution information) is information on the light sensitivity of the imaging element 107 according to the distance from the intersection of the imaging element 107 and the optical axis. Since the light sensitivity depends on the microlens 305 and photoelectric conversion units 301 and 302 of the pixel, the information may also be information on these. The light sensitivity distribution information may also be information on the change in sensitivity with respect to the angle of incidence of light.

[0112] In S307, the CPU 121 records the lens characteristic information as characteristic information of the imaging optical system in the recording medium 133 in association with the still image data file recorded in S305.

[0113] The lens characteristic information includes, for example, the following information: Exit pupil information Information about the frame of the lens barrel, etc. that blocks the light beam -Information about focal length and F-number at the time of shooting Information about aberrations in the imaging optical system -Information about manufacturing errors in the imaging optical system Position of the focus lens 105 during imaging (subject distance)

[0114] Next, in S308, the CPU 121 associates the image-related information as information related to the still image data with the still image data file recorded in S305 and records it on the recording medium 133. The image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.

[0115] In steps S306 to S308, the CPU 121 may also store the device characteristic information, lens characteristic information, and image-related information in the ROM 136 in association with the image data file recorded in step S305.

[0116] In S409, the CPU 121 causes the image processing circuit 125 to scale the still image data to generate image data for display, and causes the image data to be displayed on the display unit 131. This allows the user to check the captured image. When a predetermined display time has elapsed, the CPU 121 ends the image capturing and recording process.

[0117] ●(Subject tracking AF processing) Next, the subject tracking AF process in S400 of FIG. 9 will be described with reference to the flowchart shown in FIG.

[0118] In S401, the CPU 121 calculates the amount of image shift (phase difference) between the first and second focus detection signals generated for each of the multiple focus detection areas in S12. The amount of image shift between the signals can be determined as the relative position at which the correlation between the signals is maximized. The CPU 121 calculates the amount of defocus as the degree of focus for each focus detection area from the calculated amount of image shift.

[0119] As described above, in this embodiment, the effective pixel area 1000 is divided into 120 areas horizontally and 80 areas vertically, resulting in a total of 9,600 areas, each of which is set as a focus detection area. The CPU 121 generates data (a defocus map) that associates the defocus amount calculated for each area with the position of the area. The CPU 121 stores the generated defocus map in, for example, the RAM 136.

[0120] In S402, CPU 121 executes subject detection processing using subject detection unit 140. Subject detection unit 140 detects the areas of one or more types of subjects, and outputs to CPU 121 the detection results including the type of subject, the position and size of the area, the reliability of detection, etc. for each detected area.

[0121] Furthermore, CPU 121 performs a process (tracking process) to detect the position of the subject in the current frame based on the result of the subject detection process in the current frame and the result of the subject detection process in a past frame. If the subject cannot be detected by the subject detection process using the trained CNN of subject detection unit 140, CPU 121 can estimate the position of the subject in the current frame by a tracking process using another method such as template matching. Details will be described later.

[0122] In S403, the CPU 121 sets a focus detection area based on the result of the subject detection process obtained in S402. For example, of the 9600 focus detection areas that can be set, the CPU 121 sets one or more focus detection areas that are included in the subject area and whose detected defocus amount satisfies a condition. The condition may be, for example, that a value indicating the reliability of the defocus amount is equal to or greater than a threshold and that a defocus amount indicating the closest subject distance has been obtained.

[0123] In S404, the CPU 121 acquires the defocus amount for the focus detection area set in S403. The defocus amount acquired here may be the one calculated in S401, or may be the defocus amount newly calculated for a new frame.

[0124] In S405, CPU 121 executes predictive AF processing for each of the subject regions detected by subject detection unit 140 in S402. Predictive AF processing is processing for predicting the defocus amount of a subject region at the time of capturing the next frame. CPU 121 generates time-series data of the defocus amount for each subject region based on, for example, the defocus map generated in S401 for one or more past frames and the current frame. Then, CPU 121 obtains an equation for a prediction curve using multivariate analysis (for example, the least squares method) based on the time-series data of the defocus amount. CPU 121 predicts the defocus amount corresponding to the subject distance when capturing the next frame by substituting the capture time of the next frame into the obtained equation for the prediction curve. Note that time-series data of the position of the subject region may be generated to predict the three-dimensional position of the subject when capturing the next frame.

[0125] For example, the three-dimensional position (X, Y, Z) of the subject is expressed in an XYZ Cartesian coordinate system with the intersection of the imaging plane and the optical axis as the origin and the optical axis as the Z axis. The three-dimensional position of the subject when capturing the next frame can be predicted from the image coordinates (X, Y) of the subject area and the time-series data of the defocus amount Z.

[0126] For human subjects, the defocus amount corresponding to the subject distance when capturing the next frame may be predicted from time-series data of joint positions. By using time-series data, it is possible to estimate the position even when the joint position cannot be detected because it is hidden by another subject. Whether a part of the subject is hidden or the subject has fallen out of frame can be determined from the number and positions of undetectable joint positions. By performing predictive AF processing on multiple subjects, when the main subject is switched, there is no need to re-accumulate the defocus amount history for the new main subject, allowing predictive AF to continue quickly.

[0127] When the process of S405 ends, the CPU 121 ends the subject tracking AF process and executes S15 of S9.

[0128] ●(Subject detection and tracking processing) Next, the subject detection and tracking process in S402 of FIG. 11 will be described in detail with reference to the flowchart shown in FIG.

[0129] In S421, CPU 121 sets dictionary data to be used by subject detection unit 140 by determining the type of subject to be detected by subject detection unit 140. The type of subject to be detected can be determined based on predetermined priorities, settings of camera 100 (e.g., shooting mode), etc. For example, assume that dictionary data for "people," "vehicles," and "animals" is stored in ROM 135. Note that subject types may be categorized more finely. For example, dictionary data such as "dogs," "cats," "birds," and "cows" may be stored instead of "animals," and dictionary data such as "four-wheeled vehicles," "two-wheeled vehicles," "trains," and "airplanes" may be stored instead of "vehicles."

[0130] When a shooting mode for shooting a specific type of subject is set in camera 100, CPU 121 sets dictionary data for that type of subject. For example, when portrait mode or sports mode is set, dictionary data for "people" is set. When "panning mode" is set, dictionary data for "vehicles" is set.

[0131] If a shooting mode for shooting a specific type of subject is not set in camera 100, CPU 121 sets dictionary data for the subject according to a predetermined priority. For example, dictionary data for "people" and "animals" can be set.

[0132] The type of dictionary data and the method of determining the dictionary data to be set are not limited to the methods described here. One or more dictionary data may be set. When one dictionary data is set, subjects that can be detected by one dictionary data can be detected with a high frequency. When multiple dictionary data are set, multiple types of subjects can be detected by switching dictionaries for each frame. Note that multiple types of subjects may be detected for the same frame if the processing time is sufficient. When detecting one type of subject per frame, the detection frequency of a type of subject with a first priority may be set higher than the detection frequency of a type of subject with a lower second priority.

[0133] In S422, CPU 121 uses subject detection unit 140 to apply subject detection processing to the image of the current frame. Here, it is assumed that the type of subject is "person." Subject detection unit 140 applies subject detection processing to the image of the current frame using "person" dictionary data stored in dictionary data storage unit 141. Subject detection unit 140 outputs the detection result to CPU 121. At this time, CPU 121 may display the subject detection result on display unit 131. Furthermore, CPU 121 saves the detected subject area in RAM 136.

[0134] When dictionary data for "people" is set, the subject detection unit 140 hierarchically detects multiple types of regions related to people with different granularities, such as the "whole body" region, the "face" region, and the "eye" region. It is desirable to detect local regions such as a person's eyes and face for use in focus detection and exposure control, but there is a possibility that they cannot be detected if the face is not facing forward or is hidden by other subjects. On the other hand, it is unlikely that the whole body will be completely undetectable. Therefore, multiple types of regions with different granularities are detected, increasing the possibility of detecting some region of a "people." Dictionary data can also be configured to hierarchically detect multiple types of regions with different granularities for subjects other than people, such as "animals."

[0135] In S423, the CPU 121 performs object tracking processing by applying template matching processing to the current frame using the object region most recently detected in S422 as a template. The image of the object region itself may be used as the template, or information obtained from the object region, such as brightness information, color histogram information, or feature point information such as corners and edges, may be used as the template. Any known method may be used for the matching method and template update method. The result of the tracking processing may be the position and size of the region in the current frame that is most similar to the template.

[0136] The tracking process of S423 may be executed only if a subject is not detected in S422. By detecting an area in the current frame that is similar to a subject area detected previously, stable subject detection and tracking process can be achieved. Upon completion of the tracking process, the CPU 121 ends the subject detection and tracking process and executes S403.

[0137] ●(Virtual space photography processing) Next, details of the virtual space shooting process executed by CPU 121 in S1000 of Fig. 8 will be described using the flowchart of Fig. 13. The virtual space shooting process is a process for generating a virtual image by shooting a virtual space that changes over time from a viewpoint within the virtual space. Although a physical shooting optical system or image sensor is not required to shoot a virtual space, for ease of explanation and understanding, the description will be given assuming that the virtual space is shot using a virtual camera that has the same functions as a physical camera.

[0138] 13 is, more specifically, processing from a shooting standby state in which a virtual image is displayed as a live view image on display unit 131 of camera 100 to the execution of shooting processing of the virtual space in response to a shooting operation on camera 100. The operations described below are realized by CPU 121 of camera 100 and CPU 1001 of external computing device 1000 each executing a program.

[0139] In S1001, the CPU 1001 (virtual space reproduction unit 1100) sets up the virtual space and the virtual camera. Specifically, the CPU 1001 generates a composite object representing the virtual space, determines the initial values ​​of the virtual viewpoint position and shooting direction, and determines the configuration of the virtual camera. The initial value of the virtual viewpoint position may be a position away from the foreground object by a distance depending on the type of the foreground object, etc. Alternatively, a position set in advance within the background object may be used as the initial value of the virtual viewpoint. The initial value of the shooting direction can be determined, for example, so that the foreground object is located in the center of the screen. Furthermore, if the foreground object is to move over time, the CPU 1001 also sets information such as the position and shape of the foreground object at each time. There may be one or more foreground objects.

[0140] Furthermore, when determining the configuration of the virtual camera, the CPU 1001 may use a camera and / or lens model specified by the camera 100. Alternatively, the CPU 1001 may automatically select an appropriate combination from among cameras and lenses whose information is stored in the camera / lens information recording device 2000. Here, an appropriate combination may be a combination that can actually be used. For example, combinations of cameras and lenses with different mounts or combinations of cameras and lenses whose image circle and image sensor size do not match can be excluded.

[0141] By allowing the user of camera 100 to specify the configuration of a virtual camera, the user of camera 100 can simulate the experience of taking pictures with a desired camera and / or lens that is different from the physical camera and lens unit that they have in their hands.

[0142] For example, while operating camera 100 equipped with a lens having a short focal length (wide-angle), a user can simulate the experience of taking pictures using a lens having a long focal length (telephoto). In this case, the body of the virtual camera may be the same model as the body of camera 100, or a different model.

[0143] This allows the user of camera 100 to simulate photography using a desired model without being restricted by the weight or size of actual equipment. Also, by simulating photography using a camera or lens that the user does not own, the user of camera 100 can confirm whether the intended photography is possible, experience functions that the camera or lens that the user owns does not have, and check differences in performance.

[0144] In S1002, the CPU 1001 instructs the camera 100 to initialize the camera driving unit via the output IF 1006. Details will be described later.

[0145] In S1003, the CPU 1001 acquires camera information, lens information, and operation information from the camera 100 and the camera / lens information storage device 2000.

[0146] ●(Information used by external computing devices) Fig. 17 is a diagram showing an example of information used by the external processing device 1000. Of the various types of information shown in Fig. 17, some of the camera information and lens information is acquired from the camera 100, while others are stored in the camera / lens information recording device 2000. Information about the camera 100 and the lens unit attached to the camera 100 can be acquired when the camera 100 is connected to the external processing device 1000 and communication is established. Of the camera information and lens information, information about current values ​​and operation information are acquired periodically from the camera 100 via the input IF 1005. Some of the information is stored in the ROM 136 by the external processing device 1000.

[0147] The camera information may include the following information: Display resolution (resolution of the display unit 131 or display resolution set in the camera 100) Recording resolution (resolution of the image recorded on the recording medium 133) Image sensor size Camera Settings Number and arrangement of focus detection areas Autofocus (AF) mode (one-shot, servo, etc.) Rapid fire settings, -Shooting difficulty settings, etc. AF algorithm type Camera algorithm information (auto exposure (AE) method type, continuous shooting driving sequence, etc.) Camera detection information (image sensor and environmental temperature, etc.) Image sensor characteristic information (S / N information for each shooting sensitivity, etc.) Shading correction value (correction value for the signal characteristics of the image sensor and uneven light intensity on the screen) Defocus conversion coefficient (coefficient for converting the image shift amount to the defocus amount) Focus-related correction information (information for correcting the deviation between the focus detection result and the best image plane position, and defocus error information) General information (camera body, lens unit model name, firmware version, etc.) These are merely examples, and some of them may not be included as shown in FIG. 17, or other information may be included.

[0148] The lens information may include the following information: Focal length (range, current value, and resolution) Aperture value (F-stop) (range, current value, and increments) Focus information (focus lens driving range and current position) Focus control information (information about the control characteristics of the focus drive) Sensitivity (parameter for converting focus lens drive amount into image plane movement amount) Image stabilization information (drive range, current drive amount, and correction resolution of the image stabilization mechanism (shift lens and / or image sensor)) Image stabilization control information (information about image stabilization control characteristics) Aperture control information (information about the control characteristics of the aperture drive) Lens frame information (position and diameter) regarding vignetting - Vignetting information Distance information (information about the focus lens position and subject or focal distance) Information about the point spread function These are merely examples, and some may not be included, or other information may be included.

[0149] The operation information is information indicating the content of a user operation on the camera 100 or the lens unit. The user operation is not limited to an operation on the operation unit 132 of the camera 100 (including the components of the lens unit), but also includes an operation to move the shooting direction of the camera 100 for framing, etc. Therefore, the output of the motion detection unit 143 is also supplied to the external computing device 1000 as operation information.

[0150] Camera information and lens information can be stored for multiple models. The external computing device 1000 may also be configured to allow the user of the camera 100 to select the camera information and lens information used when generating a virtual image. For example, the CPU 1001 transmits to the camera 100 menu screen data displaying a selectable list of camera and lens model names stored in the camera / lens information recording device 2000. The CPU 1001 then acquires from the camera 100 the model name of the camera and / or lens selected through operation of the operation unit 132 of the camera 100. The CPU 1001 uses the camera information and / or lens information corresponding to the acquired model name to generate a virtual image. This allows the user of the camera 100 to view the virtual image received from the external computing device 1000 while operating the camera 100, thereby simulating the experience of taking pictures using the specified model.

[0151] The external computing device 1000 generates a virtual image for display and / or recording using the camera information, lens information, and operation information. The external computing device 1000 also generates subject information, which is shooting difficulty information, a virtual defocus amount, and various shooting-related information.

[0152] From step S1003 onwards, the CPU 1001 periodically acquires information included in the camera information and lens information that may change in accordance with the operation of the camera or lens.

[0153] Next, in S2000, the CPU 1001 (virtual image generation unit 1200) generates a virtual image based on the above-described settings and outputs it to the camera 100. Here, it is assumed that the virtual image generation unit 1200 generates a moving image to be displayed on the camera 100 as a virtual image. Details of this process will be described later.

[0154] In S1005, CPU 121 of camera 100 receives the virtual image output by external computing device 1000 in S2000 through video input unit 141. CPU 121 displays the received virtual image on display unit 131. Thereafter, generation, output, and display of the virtual image are repeatedly executed in accordance with the frame rate of the virtual image (e.g., 60 fps). Virtual image generation unit 1200 generates a virtual image in which the operation on camera 100 indicated in the operation information is reflected in the virtual camera, and therefore a virtual image in which the shooting range has changed in accordance with the panning and zooming operations of camera 100 is displayed on display unit 131.

[0155] In S1006, CPU 121 determines whether or not a mode (viewpoint movement mode) for moving the viewpoint of the virtual space being observed through display unit 131 is set in camera 100. If it is determined that the viewpoint movement mode is set, CPU 121 executes S3000, and if not, executes S1007.

[0156] In S3000, CPU 121 and CPU 1001 execute viewpoint movement processing to change the viewpoint position and shooting direction in the virtual space. Details of this processing will be described later with reference to Fig. 26. When the viewpoint movement processing ends, CPU 1001 executes S2000.

[0157] 8, in S1007, CPU 121 determines whether or not a half-press operation (SW1 ON) of the release switch included in operation unit 132 has been detected. If CPU 121 determines that SW1 ON has not been detected, CPU 1001 executes S2000. On the other hand, if CPU 121 determines that SW1 ON has been detected, CPU 121 executes S4000.

[0158] In S4000, CPU 121 and CPU 1001 execute processing for tracking and photographing a foreground object in response to the operation of camera 100, the movement of a subject (foreground object), etc. Details of the processing will be described later with reference to FIG.

[0159] In S1008, CPU 121 determines whether or not ON of SW2 is detected, similarly to S15. If CPU 121 determines that ON of SW2 is not detected, CPU 1001 executes S2000. On the other hand, if CPU 121 determines that ON of SW2 is detected, CPU 1001 executes shooting processing in the virtual space in S5000. Details of the operation in S5000 will be described later. When the shooting processing is completed, the virtual space shooting processing is terminated.

[0160] ●(Virtual image generation and output processing) Next, the virtual image generation and output process executed by the external calculation device 1000 (the virtual space reproduction unit 1100 and the virtual image generation unit 1200) in S2000 of FIG. 13 will be described with reference to the flowchart shown in FIG.

[0161] In S2001, the virtual space reproduction unit 1100 (foreground object acquisition unit 1102) acquires a foreground object from the foreground object storage unit 1101. The foreground object to be acquired may be predetermined, or may be selectable by the user.

[0162] For example, the CPU 121 displays a menu screen supplied from the external computing device 1000 on the display unit 131. Then, the CPU 121 transmits information about a foreground object (main subject in the virtual space) that the user has specified from the menu screen via the operation unit 132 to the external computing device 1000. For example, the menu screen may enable the user to select the type of foreground object (e.g., person, animal, vehicle, etc.), shape, color, and movement (speed, direction of movement, etc.). Data and values ​​required for generating the menu screen can be stored in the foreground object storage unit 1101 of the virtual space reproduction unit 1100.

[0163] Instead of or in addition to acquiring a foreground object from the foreground object storage unit 1101, the foreground object acquisition unit 1102 may generate a foreground object at this stage.

[0164] In S2002, the virtual space reproduction unit 1100 (background object acquisition unit 1105) acquires a background object from the background object storage unit 1104. The background object to be acquired may be predetermined or may be selectable by the user. As with the foreground object, the background object may also be selectable by the user of the camera 100.

[0165] In S2003, the virtual space reproduction unit 1100 (object synthesis unit 1103) places the foreground object on the background object to generate a synthesized object. The foreground object may be placed automatically by the object synthesis unit 1103, or the user of the camera 100 may be able to select the position. For example, the initial placement of the foreground object may be automatically performed by the object synthesis unit 1103, and the user of the camera 100 may change the position of the foreground object from the initial position to a desired position.

[0166] For example, the virtual image generation unit 1200 generates a virtual image by rendering a composite object in which a foreground object is placed at an initial position using the initial state of the virtual camera. The generated virtual image is then supplied to the camera 100 in a selectable manner, together with information on a position where the foreground object can be placed (for example, a position where the background object and the foreground object do not interfere with each other). The CPU 121 selectably displays, on the display unit 131, an area of ​​the virtual image corresponding to a position where the foreground object can be placed. The user selects a desired position via the operation unit 132. The CPU 121 transmits information on the selected position to the external computing device 1000. The CPU 1001 (object composition unit 1103) updates the position of the foreground object in the composite object based on the received information.

[0167] After the processes from S2001 to S2003 have been executed once, they may be executed only when explicitly requested by the user. After the composite object is generated in S2003, the processes from S2001 to S2003 may be skipped, and the processes from S2004 onward may be executed to continuously generate virtual images based on operation information.

[0168] In S2004, the CPU 1001 (virtual image generation unit 1200) acquires camera information and lens information. The camera information and lens information can be acquired from the camera 100 and the camera / lens information storage device 2000. The camera information and lens information acquired here is information necessary to generate a virtual image to be sent to the camera 100.

[0169] Specifically, the camera information includes the display resolution of the display unit 131, and the size and number of pixels of the image sensor of the camera 100. Furthermore, the lens information includes the range and current value of the focal length, the range and current value of the aperture, the focus lens range and current value, information on peripheral light falloff, and the point spread function.

[0170] In S2005, the virtual image generation unit 1200 determines the viewpoint and shooting direction of the virtual camera to be used when rendering the composite object. The initial values ​​can be determined as described in S1001. If the viewpoint has been changed by the viewpoint movement processing in S3000, the viewpoint position after the change is used. Also, if the shooting direction has been changed by operation information, the shooting direction after the change is used.

[0171] In S2006, the virtual image generator 1200 acquires operation information from the camera 100. As described above, the operation information includes movement information of the camera 100 in addition to the operation on the operation unit 132 of the camera 100 (including the lens unit).

[0172] In S2007, the virtual image generation unit 1200 (image correction amount calculation unit 1206) acquires the image correction amount. The image correction amount is the amount of correction related to framing, zooming, and focus. Details will be described later with reference to FIG. 15. The initial value is assumed to be stored in advance in the ROM 136, for example.

[0173] In S2008, the virtual image generation unit 1200 (display image generation unit 1204) generates a virtual image obtained by capturing an image of the composite object with a virtual camera. The display image generation unit 1204 determines the shooting range based on the camera information, lens information, and operation information. The display image generation unit 1204 then corrects the determined shooting range using an image correction amount, and generates a virtual image based on the corrected shooting range. The display image generation unit 1204 applies image processing to the generated virtual image to reflect the current aperture value of the camera 100, information about peripheral light falloff of the lens, information about the point spread function, and the like, so that the virtual image simulates the results of shooting using specific equipment with the settings of the camera 100.

[0174] The virtual image for display is not intended to be recorded. Therefore, parameters that affect the display, such as the shooting range, are determined with precision, while some image quality-related processing, such as reflecting peripheral light reduction, may be omitted. Such omissions are not performed on the virtual image for recording, or are omitted to a lesser extent than on the virtual image for display.

[0175] In S2009, the virtual image generation unit 1204 outputs the virtual image (frame) for display generated in S2008 to the camera 100. The virtual image is displayed as a live view on the display unit 131 of the camera 100 as described above.

[0176] In S2010, the virtual image generation unit 1204 saves the video-related information. The video-related information is information including subject information, shooting-related information, virtual defocus amount, AF log information, etc. At this point, the video-related information is temporarily saved in, for example, the RAM 1003, and is recorded as video-related information in the storage unit 1004 during the shooting process in S5000. Details will be described later. The above is the virtual image generation and output process performed in S2000. Note that at least a part of the process described as being performed by the external computing device 1000 may be performed by the camera 100.

[0177] ●(Virtual subject tracking processing) The virtual subject tracking process performed in S4000 of FIG. 13 will be described with reference to the flowchart shown in FIG.

[0178] The framing correction, which will be described later, corrects the shooting range so that the subject fits into the virtual image when at least a part of the subject (foreground object) is out of frame.

[0179] Furthermore, when the focal length of the lens is inappropriate and the subject area in the virtual image becomes too large or too small, the zooming correction corrects the focal length of the lens so that the subject area in the virtual image becomes an appropriate size. The zooming correction also includes correction to obtain a virtual image with a smoothly changing angle of view even when the user's zooming operation is not smooth.

[0180] The focusing correction is a correction that improves the degree of focus of the subject when the subject cannot be brought into focus due to performance limitations of the tracking AF algorithm of the camera 100, for example, when the subject is moving at a high speed or the change in moving speed is large.

[0181] In manual focus mode, the focus correction function corrects out-of-focus situations caused by the user's focusing operation. Focusing correction also includes corrections to improve the degree of focus on the subject when the background area is in focus due to framing.

[0182] In S4001, the virtual image generation unit 1200 (camera lens information acquisition unit 1202) acquires camera information and lens information. The information acquired here is information for determining whether correction is ON / OFF, acquiring shooting difficulty information, and calculating the amount of correction, which will be described later. Specifically, the camera information includes camera settings related to correction, such as shooting difficulty settings related to shooting, which the user sets in the camera 100. Furthermore, the lens information includes information related to the focal length, focus lens position, and ON / OFF setting of the image stabilization switch. These are merely examples, and other information may be used.

[0183] In S4002, the camera lens information acquisition unit 1202 acquires information about correction settings from the information acquired in S4001. The correction settings may include, for example, ON / OFF settings for correction in the camera 100, mode settings related to the strength of correction, and ON / OFF settings for the image stabilization switch of the lens unit.

[0184] In S4003, the operation information acquisition unit 1205 detects the presence or absence of movement (panning) related to the framing of the camera 100 from the movement information obtained from the camera 100. Then, if the operation information acquisition unit 1205 detects that the camera 100 is being panned, it also detects the direction and speed of the panning.

[0185] In S4004, the operation information acquisition unit 1205 detects whether a zooming operation has been performed based on the operation information. If the operation information acquisition unit 1205 detects that a zooming operation has been performed on the camera 100, it also detects the direction (Tele side or Wide side) and speed of the zooming.

[0186] In S4005, the operation information acquisition unit 1205 detects whether or not a focusing operation has been performed based on the operation information. If the operation information acquisition unit 1205 detects that a focusing operation has been performed on the camera 100, it also detects the direction (either to the close side or the infinity side) and speed of the focusing operation.

[0187] In steps S4003 to S4005, the operation information acquisition unit 1205 detects not only operations of the camera 100 by the user but also automatic operations of the camera 100 (auto-framing / auto-zoom / auto-focus, etc.).

[0188] In step S4006, the CPU 1001 sets a main subject area within the virtual image. Here, the main subject area is set from the area of ​​a foreground object included in the virtual image. The CPU 1001 also sets a focus detection area (AF area) to include the set main subject area.

[0189] By identifying the foreground object corresponding to the main subject region, information about the subject can be acquired from the foreground object storage unit 1101. Specifically, this information includes the velocity (or acceleration), angular velocity (or angular acceleration), size, contrast value, distance from the virtual viewpoint position, and the like.

[0190] There are various methods for setting the main subject region, and in this embodiment, it can be set in three-dimensional space. On the other hand, in the camera when shooting in real space, it is set based on the user's framing or the detection results of the subject detection unit in the video (two-dimensional) space. In this embodiment, the foreground object that is to be the main subject region when shooting in virtual space can be determined based on information about the foreground object. Alternatively, it may be the foreground object that is closest within a certain range of distance from the virtual viewpoint position, or the foreground object that is closest to the center of the shooting range. Furthermore, as is commonly used in shooting in real space, it may be determined using, for example, the detection results of the subject region for the virtual image. This makes it possible to more closely simulate the operation of the camera when shooting in real space.

[0191] In S4007, image correction amount calculation unit 1206 determines whether correction is ON or OFF from the various information acquired or detected in S4002 to S4006. For example, if the information related to in-camera correction acquired in S4002 is ON, image correction amount calculation unit 1206 determines that correction is ON, and if it is OFF, it determines that correction is OFF.

[0192] Alternatively, if framing is detected in S4003, the user intention extraction unit 1262 is used to determine whether the user is framing to track a specific subject or to switch subjects. If it is determined that the user is framing to track a specific subject, the image correction amount calculation unit 1206 turns on correction to cover up framing errors. On the other hand, if it is determined that the user is framing to switch subjects, the image correction amount calculation unit 1206 turns off correction and prioritizes the user's framing.

[0193] Furthermore, if a manual zoom operation and / or a manual focus operation is detected in S4004 and / or S4005, image correction amount calculation unit 1206 may determine that the user's intention regarding the detected operation is strong and turn correction OFF. On the other hand, if an auto zoom operation and / or auto focus operation is detected, image correction amount calculation unit 1206 may determine that the user's intention regarding those operations is weak and turn correction ON.

[0194] By turning off correction for operations that are determined to be strongly user-intentional, it is possible to provide a virtual image that simulates shooting in line with the user's intention. Furthermore, even if correction is turned on for operations that are determined to be weakly user-intentional, it is unlikely that the image will go against the user's intention. In fact, by providing a virtual image in which the subject remains within the frame through correction, it is possible to improve usability, especially for users who are unfamiliar with the operation.

[0195] If a correction capability value is defined for the camera 100, the image correction amount calculation unit 1206 may compare the correction capability value of the camera 100 with the shooting difficulty described below and turn on correction only if the correction capability value of the camera 100 exceeds the shooting difficulty.

[0196] In S4100, the image correction amount calculation unit 1206 acquires shooting difficulty information using the shooting difficulty calculation unit 1261. The image correction amount calculation unit 1206 calculates various correction amounts, which will be described later, using the shooting difficulty information. For example, the higher the shooting difficulty, the more difficult it is to keep a subject within the frame or to keep the subject in focus. In other words, the shooting difficulty indicates the difficulty of appropriate framing and focusing. For example, the higher the shooting difficulty, the smaller the correction amount the image correction amount calculation unit 1206 reduces. This is because if the shooting difficulty decreases by increasing the correction amount, the realism of the simulated shooting experience provided by the virtual image will be impaired.

[0197] On the other hand, when the shooting difficulty level is low, it is thought that in many cases the user is simply looking for good shooting results rather than enjoying shooting itself. Therefore, the image correction amount calculation unit 1206 increases the amount of correction. This increases the likelihood that the correction will cover up any operational errors made by the user, making it possible to provide good shooting results even if the user makes some operational mistakes.

[0198] The details of the process of acquiring the photographing difficulty level information in S4100 will be described with reference to the flowchart shown in FIG.

[0199] In S4101, the photographing difficulty calculation unit 1261 acquires the velocity and acceleration information of the subject from the foreground object storage unit 1101 and stores it in, for example, the RAM 1004. In S4102, the photographing difficulty calculation unit 1261 acquires the angular velocity and angular acceleration information of the subject from the foreground object storage unit 1101 and stores it in, for example, the RAM 1004. This information may be information at the time of acquisition, or may be a fixed value such as a maximum value. In any case, the photographing difficulty calculated in S4106 increases as the (angular) velocity and (angular) acceleration of the subject increase.

[0200] In S4103, the photographing difficulty calculation unit 1261 acquires size information of the subject from the foreground object storage unit 1101. The smaller the size, the higher the photographing difficulty. In S4104, the photographing difficulty calculation unit 1261 acquires the contrast value of the subject from the foreground object storage unit 1101. The lower the contrast value, the higher the photographing difficulty.

[0201] In S4105, the photographing difficulty calculation unit 1261 calculates the distance between the main subject and the user in the virtual space from the position of the foreground object in the composite object generated by the object composition unit 1103 and the virtual viewpoint position obtained from the viewpoint information acquisition unit 1201.

[0202] The photographing difficulty calculation unit 1261 can identify the size of the subject area on the imaging surface from the zooming (focal length) information acquired in S4001 and S4004 and the subject size information acquired in S4103. The smaller the subject area on the imaging surface, the higher the photographing difficulty. The difference in size of the part to be focused on within the subject area (for example, whether it is a person's eyes or face) also affects the photographing difficulty. Between the face and the eyes, the eyes are more difficult to photograph.

[0203] In S4106, the photographing difficulty calculation unit 1261 calculates photographing difficulty information for the subject from the information acquired in S4101 to S4105. The photographing difficulty calculation unit 1261 may calculate one piece of photographing difficulty information common to all correction items, or may calculate photographing difficulty information for each type of correction. In the latter case, the photographing difficulty information for framing correction / zooming correction / focusing correction will be framing difficulty, zooming difficulty, and focusing difficulty, respectively. One piece of photographing difficulty information can be calculated based on the photographing difficulty information calculated for each type of correction. For example, it may be an average value, a median value, a maximum value, etc.

[0204] Furthermore, instead of calculating the photographing difficulty information, photographing difficulty calculation unit 1261 may acquire a photographing difficulty level associated with a foreground object if such a level is stored in foreground object storage unit 1101. Furthermore, in this embodiment, the photographing difficulty level is calculated each time the moving (angular) velocity of the main subject or the distance from the virtual viewpoint position changes (the photographing difficulty level may change over time), but each foreground object may have a static photographing difficulty level.

[0205] 15, in S4009, the display image generation unit 1204 calculates a virtual defocus amount for the subject region set in S4006. The calculated virtual defocus amount may be a defocus map as described using Fig. 11, or may be a single defocus amount for a part of the subject region, such as the face region of the subject. In the former case, one region is selected from multiple regions using an arbitrary method, and the defocus amount for that region is used.

[0206] In S4200, the display image generation unit 1204 performs processing on the virtual defocus amount calculated in S4009. Details of the processing will be described later with reference to FIG.

[0207] In S4011, the display image generation unit 1204 calculates the focus lens drive amount of the virtual camera. The display image generation unit 1204 may directly convert the virtual defocus amount obtained after processing in S4200 into a focus lens drive amount. Alternatively, the display image generation unit 1204 may predict the future position of the subject area from the position of the subject area in multiple past frames and calculate the focus drive amount for the predicted position. The method of predicting the future position of the subject area described here is merely an example, and any other known method may be used. The virtual camera capturing the virtual space does not have a physical focus lens. Therefore, it is possible to instantly bring the focus detection area into focus without considering the focus lens drive time. However, in this embodiment, to provide a user operating a physical camera 100 with a realistic simulated shooting experience, a virtual image is generated so that the focus detection area is brought into focus over a time period corresponding to the defocus amount, just like a physical camera. The time required to reach the focus state may be the same as that of the actual device being simulated, based on camera information and lens information, or it may be based on the focus lens drive speed predetermined as a parameter of the virtual camera.

[0208] S4012 to S4014 are processes in which image correction amount calculation unit 1206 calculates the amount of correction for each type of operation. In S4012, image correction amount calculation unit 1206 calculates the amount of correction related to framing. Framing correction will be described using FIG. 21. Assume that there are subjects A and B, and subject A is more difficult to photograph. In this case, as described above, the maximum framing correction amount A for subject A, which is more difficult to photograph, will be smaller than the maximum framing correction amount B for subject B. In other words, maximum framing correction amount A<maximum framing correction amount B. In the figure, the length of the arrow indicates the magnitude of the maximum framing correction amount.

[0209] In Fig. 21, the rectangular area indicated by the dashed line is the actual framing area (shooting range) of camera 100, and the rectangular area indicated by the solid line is the corrected framing area. Fig. 21(a) shows a case where subject A is not included in part of the actual framing area, but subject A can be entirely included in the corrected framing area by correction equal to or less than the maximum framing correction amount A.

[0210] On the other hand, Figure 21(b) shows a case where the area of ​​subject A that does not fit within the actual framing area is large, and even if the maximum framing correction amount A is applied, the entire subject A cannot be fit within the frame.

[0211] Regarding subject B, in Figure 21(c), as in Figure 21(a), there is a small area of ​​subject B that is not included in the actual framing, so the entire subject B can be included in the framing area within the range of the maximum framing correction amount B.

[0212] On the other hand, in Figure 21(d), the area of ​​subject B that is not actually contained within the framing is larger than in Figure 21(c), but the entire subject B can be contained within the framing area within the range of the maximum framing correction amount B.

[0213] In this way, by varying the maximum framing correction amount depending on the degree of difficulty of photography, it is possible to generate a virtual image that provides a realistic photography experience in which the framing is more difficult for a subject that is more difficult to photograph.

[0214] Returning to FIG. 15, in S4013, image correction amount calculation unit 1206 calculates the amount of correction related to zooming. Correction related to zooming will be explained using FIG. 22. Assume that moving subjects C and D are present, and that moving subject C has a faster moving (angular) velocity and is more difficult to photograph. In this case, as described above, the maximum zoom correction amount C for moving subject C, which is more difficult to photograph, is smaller than the maximum zoom correction amount D for moving subject D. In other words, maximum zoom correction amount C<maximum zoom correction amount D. In the diagram, the length of the arrow indicates the magnitude of the maximum zoom correction amount.

[0215] 22, the rectangular area indicated by the dashed line is the actual angle of view (shooting range) of camera 100, and the rectangular area indicated by the solid line is the corrected angle of view. Fig. 22(a) shows a case where part of subject C is not included in the actual angle of view, but the entire subject C can be included within the corrected angle of view by correction equal to or less than the maximum zoom correction amount C.

[0216] On the other hand, Figure 22(b) shows a case where a large area of ​​subject C does not fit within the actual angle of view, and even if the maximum framing correction amount C is applied, the entire subject C cannot be fit within the angle of view.

[0217] With respect to subject D, in FIG. 22(c), as in FIG. 22(a), there is a small area of ​​subject D that is not included in the actual angle of view, so that the entire subject D can be included in the angle of view within the range of the maximum zoom correction amount D.

[0218] On the other hand, in Figure 22(d), the area of ​​subject D that does not fit within the actual angle of view is larger than in the state of Figure 22(c), but the entire subject D can be included within the angle of view within the range of the maximum framing correction amount D.

[0219] In this way, by varying the maximum zoom correction amount depending on the difficulty of shooting, it is possible to generate a virtual image that provides a realistic shooting experience, in which the more difficult a moving subject is to shoot, the more difficult it is to continuously shoot it at an appropriate angle of view.

[0220] Returning to FIG. 15, in S4014, image correction amount calculation unit 1206 calculates the amount of correction related to focusing. Correction related to focusing will be explained using FIG. 23. Assume that moving subjects E and F are present, and that moving subject E has a faster moving (angular) velocity and is more difficult to photograph. In this case, as described above, the maximum focusing correction amount E for moving subject E, which is more difficult to photograph, is smaller than the maximum focusing correction amount F for moving subject F. In other words, maximum focusing correction amount E<maximum focusing correction amount F. In the diagram, the length of the arrow indicates the magnitude of the maximum focusing correction amount.

[0221] Figure 23(a) shows the change in distance as subject E approaches, and Figure 23(b) shows the change in distance as subject F approaches. Subject E moves faster than subject F. The origin of the vertical axis corresponds to the virtual viewpoint position. The solid lines show the change in distance to each subject over time, the dotted lines show the change in focal distance (actual focal distance) over time due to camera focusing, and the dashed lines show the change in focal distance over time after focusing correction. The smaller the deviation between the solid line and the dotted or dashed line, the higher the degree of focus of the subject.

[0222] The magnitude of the difference between the focus distance and the subject distance depends largely on the user's photography technique when using manual focus. On the other hand, when using autofocus, it depends largely on the camera's tracking ability and algorithm (tracking limit performance). In general, the difference between the focus distance and the subject distance tends to be large for subjects that move quickly or whose moving speed changes significantly.

[0223] Figures 23(a) and (b) show typical examples of deviations between the focus distance and the subject distance due to the autofocus tracking limit. In Figure 23(a), focus correction works effectively within the range of the maximum focusing correction amount E, and the deviation between the focus distance and the subject distance after correction is kept within a very small range. However, as the subject gets closer, once the difference between the actual focus distance and the subject distance exceeds the maximum focusing correction amount E, the deviation between the focus distance and the subject distance after correction increases.

[0224] On the other hand, Figure 23(b) shows a case where a larger maximum focusing correction amount F can be applied to subject F, which moves slower than subject E. Even when subject F approaches and the difference between the actual focus distance and the subject distance increases, the degree of focus of subject F is maintained high to the end because it is within the range of maximum focusing correction amount F. Even in the case of subject E, which moves faster, the degree of improvement in focus as it approaches becomes greater.

[0225] As mentioned above, focusing correction can be applied to a variety of subjects, including deviations between the actual focus distance and the subject distance due to manual focus operation and deviations between the actual focus distance and the subject distance due to AF on an unintended subject caused by framing. Corrections for these deviations may use the same amount of correction as the amount of correction for deviations between the actual focus distance and the subject distance due to tracking limit performance, or a different amount of correction. Furthermore, corrections for multiple factors that may occur simultaneously may be applied in combination. In this way, by varying the amount of focusing correction depending on the difficulty of the subject, it is possible to provide a more realistic shooting experience, in which the more difficult the subject is to photograph, the more difficult it is to focus on it.

[0226] In calculating the various correction amounts in S4012 to S4014, while SW1 is continuously detected as being on in S1007 in the virtual space shooting process of Fig. 13, past recorded image information and video correction amounts stored in storage unit 1004 may be used. This allows for continuity in the correction results between frames, reducing the possibility that the viewer will feel uncomfortable.

[0227] In S4015, the display image generation unit 1204 performs virtual focus drive. The display image generation unit 1204 changes the position of the virtual focus lens (the focal length of the virtual lens) by using a drive amount that reflects the focus correction amount calculated in S4014, relative to the focus drive amount calculated in S4011. The drive amount and drive direction used for the virtual focus drive may be output to the camera 100.

[0228] As described above, by adjusting the correction amounts for framing, zooming, and focusing based on operation information, subject information, camera information, and lens information, it is possible to adjust the reality of the shooting experience with camera 100. Therefore, it is possible to provide an appropriate and flexible shooting experience in a virtual space according to the requests and skills of the user of camera 100.

[0229] In this embodiment, an example has been described in which realism is emphasized and the amount of image correction is reduced as the difficulty of shooting increases. However, if the success rate of shooting is emphasized over realism, the amount of image correction may be increased as the difficulty of shooting increases. Successful shooting may be achieved by satisfying one or more of the following: the subject is framed at an appropriate size, the subject is highly focused, etc. The user may be able to switch in the settings of camera 100 whether the amount of correction is increased or decreased as the difficulty of shooting increases.

[0230] Note that the processes such as defocus amount calculation and focus driving executed in the virtual subject tracking process described above may be performed using an AF algorithm stored in the ROM 135 of the camera 100 or the camera / lens information storage device 2000.

[0231] ●(Defocus amount processing) Next, the defocus amount processing in S4200 will be described in detail using the flowchart shown in Fig. 20. The defocus amount processing adds an error amount to the virtual defocus amount calculated in S4009 according to the settings at the time of shooting and characteristic information of the image sensor.

[0232] In virtual space photography, the subject distance is known, so the only error included in the defocus amount is a calculation error caused by the number of significant digits of each numerical value used in the calculation. On the other hand, in real space photography, the defocus amount includes errors that occur steadily due to the characteristics of the image sensor and errors that vary with each image capture. In virtual space photography, to reproduce the same focus state behavior as in real space photography, it is necessary to reflect the defocus amount error that occurs in real space photography in the virtual defocus amount. Therefore, to improve the realism of the photography experience in virtual space, an error is added to the virtual defocus amount calculated in S4009.

[0233] In S4201, the display image generation unit 1204 acquires the camera information and lens information of the virtual image generated by the virtual image generation unit 1200 in S2000 from the camera / lens information storage device 2000, and stores it in the RAM 1003 as information to be associated with the virtual image.

[0234] The camera information may be, for example, recording resolution, image sensor size, AF frame mode, AF algorithm, S / N information for each ISO sensitivity, information regarding correction of defocus conversion coefficients and focus position, and focus-related correction information including defocus error information.

[0235] The lens information may be, for example, focal length and resolution, F-number, focus lens information, focus drive control information, sensitivity, image stabilization control information, aperture control information, lens frame information, peripheral light falloff information, and distance information related to the focus lens position and distance. Sensitivity is a parameter for converting the focus lens drive amount into the image plane movement amount.

[0236] In S4202, the display image generation unit 1204 acquires the defocus error information stored in the RAM 1003. The defocus error information will be described later.

[0237] In S4203, the display image generation unit 1204 adds (adds) an error amount to the virtual defocus amount calculated in S4009 according to the defocus error information acquired in S4202. If a virtual defocus amount has been calculated for each of multiple focus detection areas, the display image generation unit 1204 adds (adds) a defocus error amount to each virtual defocus amount. This completes the defocus amount processing process.

[0238] An example of the defocus error information acquired in S4202 will be described using FIG. 24. FIG. 24 shows an example of the relationship between the contrast value of the subject acquired in S4006 and the amount of error that is expected to occur in the virtual defocus amount. Generally, when capturing an image in real space, the amount of error contained in the defocus amount detected for a low-contrast subject becomes large. In FIG. 24, the horizontal axis represents the contrast of the subject, and the vertical axis represents the amount of error contained in the defocus amount, and line 24101 indicates that the greater the contrast of the subject (to the right on the horizontal axis), the smaller the amount of error that occurs.

[0239] The relationship between contrast and error amount depends on the S / N ratio of the pixel section and readout circuit section of the image sensor, the number of pixels used for focus detection signals, and the signal gain amount according to ISO sensitivity. Therefore, in this embodiment, the relationship between the contrast value of the subject and the error amount included in the defocus amount is stored as defocus error information for each mode in which the S / N ratio of the image sensor changes, the focus detection signal specifications, and the ISO sensitivity.

[0240] The defocus error information can be stored as a table storing discrete values ​​or as a function representing a relationship. The error amount also depends on the F-number and lens frame information, which are part of the lens information. This is because the defocus conversion coefficient, which is camera information, depends on this lens information. In this embodiment, the lens information and camera information are used to multiply the calculated error amount by a predetermined coefficient. The predetermined coefficient can be stored in a table as a ratio to a reference value. For example, a table storing the defocus conversion coefficient and the coefficient by which the error amount is multiplied can be stored in, for example, ROM 1002, using the F-number and lens frame information as an index, and the error amount can be calculated according to the shooting situation.

[0241] 25(a) and (b) are schematic diagrams showing the virtual defocus amount without an error and the virtual defocus amount with an error, in association with an image. Here, it is assumed that an AF frame 25101 is set.

[0242] 25(a) shows a defocus map 25102 of virtual defocus amounts to which no error is added. The focused position of the head of a person 25106 within an AF frame 25101 is an in-focus area 25103.

[0243] 25(b) shows a defocus map 25102 of the virtual defocus amount with an error added. In addition to a focused region 25103, a region 25104 where the focus distance is shifted forward and a region 25105 where the focus distance is shifted backward are shown.

[0244] In Fig. 25(a), where no error is added, the entire area within AF frame 25101 is an in-focus area. On the other hand, in Fig. 25(b), where an error is added, AF frame 25101 includes not only in-focus area 25103, but also area 25104 where the in-focus distance is shifted forward and area 25105 where the in-focus distance is shifted backward.

[0245] When capturing a virtual image, it is possible to obtain defocus detection results similar to those obtained when capturing images in real space by incorporating an error amount corresponding to changes in camera / lens information into the defocus amount and applying an AF frame selection algorithm. While this embodiment illustrates an example in which a defocus error is added to a defocus map for ease of explanation, a defocus error may also be added to a single AF frame. The effect of adding an error amount can also be achieved in predictive AF operation, etc. This allows the same focus adjustment behavior as in real space to be reproduced when capturing images in virtual space, enabling users to evaluate product performance and check new features before purchasing a camera or lens.

[0246] In this embodiment, an error amount is added to the virtual defocus amount in order to make the shooting results in the virtual section closer to the shooting results in real space. However, adding an error amount is not essential. Whether or not to add an error amount may be switched depending on the degree of reality required.

[0247] ●(Shooting process in virtual space) Next, the photographing process executed by the external processing device 1000 in S5000 of FIG. 13 will be described with reference to the flowchart shown in FIG.

[0248] In S5001, the CPU 1001 (operation information acquisition unit 1206) acquires the F-number set in the camera 100 and the time when it was detected that SW2 was turned on from the operation information. How this information is used will be explained in the section on real camera operation linked to virtual shooting operation in Fig. 28, which will be described later.

[0249] In S5002, the CPU 1001 (camera lens information acquisition unit 1202) acquires camera / lens information. Specifically, the acquired camera information includes the recording resolution, the size and number of pixels of the image sensor of the camera 100, etc. The acquired lens information includes the range and current value of the focal length of the camera 100, the range and current value of the F-number, the focus lens driving range and current position, and information related to peripheral light falloff and point spread function.

[0250] In S5003, the CPU 1001 (display image generation unit 1204) acquires the correction amount calculated by the image correction amount calculation unit 1206 described above.

[0251] In S5004, the CPU 1001 (display image generation unit 1204) generates a virtual image for recording. The display image generation unit 1204 renders the above-mentioned composite object based on the virtual viewpoint position and shooting direction to generate a virtual image. The range of the composite object to be rendered can be determined by the focal length from the above-mentioned lens information, the image sensor size and resolution from the camera information, the recording resolution settings, etc.

[0252] Furthermore, the rendering reflects the F-number, peripheral light falloff information, point spread function information, and focus lens position information obtained from lens information. Virtual images for recording are processed with higher precision than virtual images for display. For example, the precision of the rendering range can be increased, and various optical information such as focus lens position information, peripheral light falloff information, and point spread function information can be reflected in the rendering. Because real-time processing is not required for virtual images for recording, processing is performed with a higher emphasis on quality (reality) than for virtual images for display.

[0253] In S5005, the CPU 1001 (display image generation unit 1204) records the generated virtual image for recording in the storage unit 1004. Alternatively, the CPU 1001 may transfer the virtual image for recording generated in S5004 to the camera 100 and record it in the storage medium 133 of the camera 100.

[0254] In S5006, the CPU 1001 records related information for the virtual image in association with the virtual image recorded in S5005. The related information may be subject information (photography difficulty information), photography-related information, virtual defocus amount, etc. The recording destination may be the storage unit 1004 of the external computing device 1000 or the storage medium 133 of the camera 100, as long as it is the same as the virtual image. This completes the photography process in the virtual space.

[0255] ● (A virtual space photography experience that reflects the operating information of a real camera) The photographing experience in a virtual space that reflects the operation information of a real camera (camera 100) realized in this embodiment will be described with reference to Fig. 18. Fig. 18(a) shows an example of zooming operation, and Fig. 18(b) shows an example of framing operation.

[0256] FIG. 18(a) shows a schematic diagram of a virtual image 18003 for display generated by the virtual image generating unit 1204 being displayed as a live view on the display unit 131 (eye-view type EVF) of the camera 100.

[0257] Assume that the user of camera 100 has changed the focal length of the lens unit of camera 100 to the telephoto side by performing a zooming operation. The zooming operation may be, for example, operation of a zoom ring provided on the lens unit. Information related to this zooming operation (for example, the direction of operation or the current value of the focal length) is transmitted from camera 100 to external computing device 1000 as operation information.

[0258] As described above, the virtual image generation unit 1200 generates a virtual image that reflects the operation information from the camera 100. Therefore, the angle of view of the virtual image 18004 that the camera 100 receives from the external computing device 1000 is changed to reflect the zooming operation. Because the virtual image is generated by rendering a virtual space using camera / lens information, it is also possible to generate a virtual image with a focal length that cannot be realized by the lens unit of the camera 100.

[0259] For example, even if the lens unit of the camera 100 is a wide-angle lens, by generating a virtual image using information from a telephoto lens, it is possible to provide a shooting experience with an angle of view that would not be possible with the lens being operated. Generally, lenses with long focal lengths are large, heavy, and expensive, but according to this embodiment, a shooting experience can be provided without such constraints.

[0260] FIG. 18(b) shows a schematic diagram of a virtual image 18006 for display generated by the virtual image generating unit 1204 being displayed on the display unit 131 (eye-viewing EVF) of the camera 100.

[0261] It is assumed that the user of the camera 100 has performed a framing operation to move the shooting direction of the camera 100 horizontally to the right. Information about the movement of the camera 100 due to the framing operation is transmitted from the camera 100 to the external computing device 1000 as operation information.

[0262] When the virtual image generation unit 1200 detects the movement of the camera 100 as operation information from the camera 100, it changes the viewpoint position and shooting direction used for rendering the virtual image according to the movement. Therefore, the virtual image 18007 that the camera 100 receives from the external computing device 1000 is changed to a shooting range that reflects the framing operation.

[0263] As described above, the external calculation device 1000 according to this embodiment provides a virtual image that reflects the user's operation on the camera 100. Therefore, the user can simulate various shooting experiences according to the operation through physical camera operation.

[0264] ●(Viewpoint movement processing) Next, the viewpoint movement processing performed in S3000 of Fig. 13 will be described in detail with reference to Fig. 26. The viewpoint movement processing is processing for changing the viewpoint position of the virtual camera when generating a virtual image.

[0265] In S3001, CPU 1001 adjusts the depth and angle of field of an image to be generated for viewpoint movement. It is desirable that the image for viewpoint movement has a wide shooting range and allows the focal distance to be easily visible. This is because, as will be described later, the destination of the viewpoint is set based on the angle of field displayed on display unit 131 and an object at the focal distance. CPU 1001 adjusts the angle of field and depth of field for generating a virtual image to values ​​set in advance for viewpoint movement. Note that the processing of S3001 may be omitted because it facilitates the operation of viewpoint movement.

[0266] In S3002, the CPU 1001 (virtual image generation unit 1200) generates a virtual image for viewpoint movement based on the settings made in S3001. The viewpoint position can be the viewpoint position used immediately before or a preset initial viewpoint position. The generated virtual image is sent to the camera 100 and displayed on the display unit 131.

[0267] In S3003, the CPU 121 performs a process of adjusting the focus and the index direction. First, in the focus adjustment, as described in S3001, a state in which the focus is set within a predetermined distance range within the shooting range is displayed, and the user of the camera 100 performs a focus adjustment operation similar to that performed when shooting.

[0268] Specifically, with one focus detection area index (I(n,m)) described above displayed, the user adjusts the focus on an object located at a position within the shooting range where the viewpoint is to be moved by half-pressing the release switch included in the operation unit 132. The external calculation device 1000 generates a virtual image for viewpoint movement that reflects this focus adjustment operation and feeds it back to the camera 100, thereby making it possible to show the user the focused distance in the image displayed on the display unit 131. The CPU 1001 also sets the distance (focus distance) at which the focus was achieved by the user's operation as the distance of the viewpoint position after the movement.

[0269] Although the method of setting the distance of the viewpoint position after movement by autofocus operation has been described, it may also be set by manually operating the focus lens (third lens group 105), that is, by manual focus operation. The external calculation device 1000 generates a virtual image for viewpoint movement that reflects the rotation operation of a focus ring provided on the lens unit, and feeds it back to the camera 100. The CPU 1001 also sets the distance at which the image is in focus by user operation (focus distance) as the distance of the viewpoint position after movement.

[0270] The direction of movement can be determined by changing the position of the index (I(n, m)) in the focus detection area on the screen of the display unit 131, or by moving the camera 100 to change the shooting direction to the direction in which the viewpoint position is to be moved. This allows the user to set the direction of viewpoint movement while checking the image of the virtual space displayed on the display unit 131.

[0271] In S3004, CPU 1001 determines whether an instruction operation to move the viewpoint has been detected from the operation information. If, for example, operation of the enter button on operation unit 132 is detected, CPU 1001 recognizes this as an instruction operation to change the viewpoint position. If it is determined from the operation information that an instruction operation to move the viewpoint has been detected, CPU 1001 executes S3005, and if not, continues to execute S3003.

[0272] In S3005, the CPU 1001 determines whether the viewpoint can be moved to a position specified based on the currently set distance and direction of the viewpoint after movement. The CPU 1001 determines that movement is not possible if the specified viewpoint position after movement satisfies a predetermined condition, such as coordinates inside an object or below the ground (underground) of a background object, or if the distance to the object is too close. The CPU 1001 also determines that movement is not possible if the distance after movement is infinite. If the distance after movement is farther than a threshold, the distance after movement may be changed to a predetermined upper limit distance.

[0273] In S3006, the CPU 1001 determines whether it was determined in S3005 that the viewpoint can be moved. If it was determined that the viewpoint can be moved, the CPU 1001 executes S3007, and if not, the CPU 1001 executes S3008.

[0274] In S3007, CPU 1001 moves the viewpoint position to the set position. CPU 1001 may generate an image indicating that the setting was successful and transmit it to camera 100. Camera 100 displays the received image on display unit 131.

[0275] In S3008, CPU 1001 generates an image indicating that the viewpoint cannot be moved to the set position, and transmits the image to camera 100. Camera 100 displays the received image on display unit 131. The image may simply notify that the viewpoint cannot be moved, or may also notify of a specific reason, such as interference with an object or the set movement distance being too far.

[0276] In S3009, CPU 1001 determines from the operation information whether an operation to end the viewpoint shift mode has been detected. For example, if CPU 1001 detects an operation of the back or cancel button on operation unit 132, it recognizes this as an operation to end the viewpoint shift mode. If CPU 1001 determines from the operation information that an operation to end the viewpoint shift mode has been detected, it ends the viewpoint shift processing, and if not, it continues to execute S3003.

[0277] Next, a specific example of the viewpoint movement processing described in FIG. 26 will be described with reference to FIG. 27. FIG. 27 shows a display example of the display unit 131 in viewpoint movement mode. FIG. 27(a) shows a display example in which viewpoint movement is set in viewpoint movement mode in a virtual space in which a person and a dog are placed as foreground objects. A virtual image 27001 including foreground objects 27003 of a person and a dog is displayed on the display unit 131, and the viewpoint position before movement is assumed to be above and to the left of the person.

[0278] 27002 is one of the indices (I(n,m)) described in FIG. 5. The direction of movement of the viewpoint is set according to the position of the index 27002. 27005 indicates the range of distance from the current viewpoint position to which the viewpoint position can be changed. FIG. 27(a) shows that the viewpoint distance can be set in the range from 0.45 m to 10 m, and that the current setting is to move 1 m away. In reality, in virtual image 27001 in FIG. 27(a), objects that are 1 m away from the current viewpoint position (camera position) are in focus, and parts at other distances are blurred.

[0279] The user of the camera 100 can move the index 27002 within the display screen 27001 using the operation unit 132. The user can also move the index 27002 to the position of the foreground object by panning the camera 100. By performing a focus adjustment operation to bring the image at the position of the index 27002 into focus, the user can set the distance of the viewpoint position after movement.

[0280] The sub-display screen 27004 displays a preview of a virtual image obtained from the currently set viewpoint position after movement. If the currently set viewpoint position after movement cannot be set, the image generated in S3008 may be displayed on the sub-display screen. The shooting direction when generating the virtual image to be previewed may be set automatically as the direction from the viewpoint position to the foreground object, or may be set by the user via the operation unit 132. A rectangular frame may be displayed within the sub-display screen 27004 to indicate the shooting range corresponding to the focal length of the lens of the camera 100.

[0281] In this way, in this embodiment, the distance and direction of viewpoint movement can be set using the EVF display and the AF pointer 27002. This allows the user of camera 100 to intuitively change the viewpoint position using the same operations as when taking a picture.

[0282] A modified example of viewpoint movement will be described using FIG. 27(b). A virtual image 27001 is generated as a virtual image for viewpoint movement, in which a viewpoint movement target (target position) 27006 is placed on a background object. The viewpoint movement target 27006 is in a grid pattern, with the intersection of the second row and first column of the grid indicated as viewpoint movement target 27006(2,1), and the intersection of the fourth row and fourth column of the grid indicated as viewpoint movement target 27006(4,4). The user can set the distance corresponding to the target by moving the AF indicator 27002 closer to the viewpoint movement target 27006 that is closest to the distance to which the user wishes to move the viewpoint. The viewpoint movement target may be superimposed on a foreground object, or a numerical value indicating the corresponding distance may be displayed nearby. In this way, by including viewpoint movement targets according to distance in the virtual image for viewpoint movement, the user can more easily set the movement distance of the viewpoint position.

[0283] ● (Feedback on the operation feel while shooting in virtual space by driving the camera drive unit) Next, the operation of the camera 100 while capturing an image in a virtual space will be described with reference to Fig. 28. When capturing an image in a real space, the user receives feedback such as vibrations and sounds due to the operation of the shutter and lens accompanying the operation of the camera.

[0284] In this embodiment, an experience of taking pictures of a virtual space is provided by operating the camera 100. Although it is not necessary to drive the shutter or focus lens of the camera 100 in the experience of taking pictures of a virtual space, the quality of the experience can be improved by providing feedback similar to that of real-life photography. The operation of the camera 100 for providing such feedback will be described below.

[0285] In S1201, the CPU 121 executes initialization of the driving unit of the camera 100 in response to the driving unit initialization instruction output by the external computing device 1000 in S1002 of Fig. 13. The driving unit refers to movable members such as the shutter, aperture, and focus lens, as well as driving members such as motors and actuators that drive them.

[0286] During this initialization, the CPU 121 sets the drive units to operate in accordance with the operation of the virtual camera. For example, when shooting is about to begin, the CPU 121 drives the shutter to an open state. The CPU 121 also controls the zoom actuator 111, aperture actuator 112, and focus actuator 114 so that the focal length, aperture value, and focus lens position correspond to the angle of view, depth of field, and in-focus distance at the start of shooting.

[0287] Here, it has been described that the initial positions of the driving units of the camera 100 are set in response to instructions from the external computing device 1000. However, position information of each driving unit of the camera 100 may be output to the external computing device 1000, and the external computing device 1000 may adjust the initial state of the virtual camera to the state of the camera 100.

[0288] In S1202, the CPU 121 outputs the camera information, lens information, and operation information to the external calculation device 1000 in a format corresponding to S1003 in Fig. 13. The contents of each piece of information are as described above.

[0289] In S1203, the CPU 121 acquires the image generated in S2000 of FIG.

[0290] In S1204, the CPU 121 determines whether or not it has received the drive amount and drive direction used for the virtual focus drive in S4015 of Fig. 15. If it is determined that they have been received, the CPU 121 executes S1205, and if not, it executes S1206.

[0291] In S1205, the CPU 121 controls the focus actuator 114 in accordance with the received drive amount and drive direction to drive the focus lens (third lens group 105).

[0292] In S1206, the CPU 121 determines whether the aperture value supplied to the external calculation device 1000 in S5001 of Fig. 16 is different from the aperture value of the camera 100. If it is determined that the aperture values ​​are different, the CPU 121 executes S1207, and if not, it executes S1208.

[0293] In S1207, the CPU 121 controls the aperture actuator 112 to drive the aperture value of the aperture 102 so that it matches the aperture value of the virtual camera.

[0294] In S1208, the CPU 121 determines whether the time when ON of SW2 is detected is supplied to the external calculation device 1000 in S5001 of Fig. 16. If it is determined that the time when ON of SW2 is detected is supplied to the external calculation device 1000, the CPU 121 executes S1209, and if not, executes S1210.

[0295] In S1209, the CPU 121 drives the shutter 106 in the same manner as in photographing in real space after a predetermined time has passed since the time when ON of SW2 was detected.

[0296] In S1210, the CPU 121 determines whether or not the main switch (power switch) of the operation unit 132 has been turned off, and if it is determined that it has been turned off, ends the shooting operation of the virtual space, and if it is not determined that it has been turned off, executes S1204.

[0297] As described above, in a virtual space photography experience, by controlling the operation of the actual camera being operated in accordance with the operation of the aperture and shutter of the virtual camera, feedback such as vibration and sound can be obtained regarding photography in the virtual space, thereby providing the user with a more realistic photography experience.

[0298] Here, a case has been described in which the driving unit of camera 100 can be driven in accordance with the operation of the virtual camera. However, if the capability of the virtual camera is higher than the capability of camera 100, the operation of camera 100 cannot be adjusted to match the operation of the virtual camera. Specifically, this occurs when the continuous shooting speed of the virtual camera exceeds the continuous shooting capability of camera 100, or when the focus lens needs to be driven beyond the amount of drive that can be achieved by camera 100.

[0299] In such a case, the driving of camera 100 in accordance with the virtual camera may be prohibited or may be changed to a driving range that can be realized by camera 100. For example, in the case of a continuous shooting speed, the shutter of camera 100 is driven only once for every two shutter operations of the virtual camera. In addition, with regard to driving of the aperture and focus lens, the driving amount may be converted so that the driving range of the virtual camera matches the driving range of camera 100.

[0300] We have explained the sound and vibration feedback associated with operations, but the sound and vibration can also be recorded along with the image. For example, a still image can be blurred according to the amount of camera vibration, or the generated sound can be recorded in a video. This allows for shooting in a virtual space that is closer to shooting in real space.

[0301] ● (Playback and evaluation of shooting results) Next, the playback and evaluation process of images (still images) captured by camera 100 will be described using the flowchart shown in Fig. 29. Specifically, the process plays back images obtained by capturing images in real space and virtual space, displays a defocus map of images captured with settings different from those used during actual capture, and calculates the degree of focus of a series of consecutive images. This makes it possible to show the user the cause of poor focus and to suggest settings that are expected to improve the degree of focus.

[0302] In S1101, the CPU 121 selects an image to be played back from the storage medium 133. In accordance with a user instruction via the operation unit 132, the CPU 121 selects an image that was captured immediately before or an image selected by the user as the image to be played back.

[0303] In S1102, the CPU 121 determines whether the image to be played back is an image in virtual space or an image in real space. If the image to be played back is determined to be an image in virtual space, the CPU 121 executes S1103, and if the image is determined to be an image in real space, the CPU 121 executes S1104.

[0304] In S1103, the CPU 121 acquires the image to be played back from the external computing device 1000 and plays it back on the display unit 131. In S1104, the CPU 121 acquires the image to be played back from the storage medium 133 and displays it on the display unit 131. Note that the recording destination of the image does not necessarily have to depend on the type of space in which the image was captured.

[0305] In S1105, the CPU 121 determines whether or not to perform shooting evaluation of the image being played back, and if it is determined that shooting evaluation is to be performed, executes S1106, and if not, ends the processing shown in FIG.

[0306] In S1106, the CPU 121 acquires information related to the shooting of the image to be played back (shooting-related information). The shooting-related information is various information about the settings of the camera used at the time of shooting and the lens. The shooting-related information includes the lens and camera settings at the time of shooting, such as focal length, F-number, continuous shooting mode, AF mode, subject detection AF tracking setting, AF frame setting, and shutter method. These are examples, and the shooting-related information may include any information that affects the focus state of the image. The shooting-related information may be stored in the storage medium 133 or the storage unit 1004, or may be included in accompanying information stored in the header of the image being played back, for example.

[0307] In step S1107, the CPU 121 acquires AF log information included in the accompanying information of the image being played back. The AF log information is a group of information related to AF when the image was captured, and may include, for example, one or more of the following: Defocus information AF frame setting information Tracking information (information about the AF function that tracks detected subjects (such as people, animals, vehicles, etc.)) Servo AF characteristics (various servo AF parameters) Action recognition information (posture information of the subject, information on whether the subject performed a specific action) Shutter information (type of shutter used (mechanical and / or electronic), burst speed setting)

[0308] In S1108, CPU 121 sets one or more images, including the image being played back, as a group of images to be evaluated (group of evaluation images). The group of evaluation images may be set by images taken at a time close to the time of the image being played back, or by a group of images taken in a single continuous shot. Alternatively, the user may be allowed to select images to be included in the group of evaluation images from a list of images.

[0309] In step S1109, the CPU 121 sets an evaluation sequence, the details of which will be described later.

[0310] In S1110, the CPU 121 performs evaluation on each image included in the evaluation image group in accordance with the information acquired in S1106 to S1109 and the set evaluation sequence.

[0311] The CPU 121 generates a defocus map based on the focus control result included in the shooting-related information included in the accompanying information of the image. The CPU 121 also calculates the defocus amount when a focus control different from that used during shooting is applied to the image, and generates a defocus map.

[0312] The CPU 121 displays each of the generated defocus maps superimposed on the image to be evaluated, and then calculates the degree of focus based on the deviation of the defocus amount from a predetermined threshold value.

[0313] The CPU 121 simultaneously calculates the degree of focus and performs image analysis to analyze the causes of the good or bad degree of focus, thereby making it possible to evaluate the difference between the degree of focus achieved with the settings actually used when taking a photograph and the degree of focus that would be achieved if a photograph were taken with different settings.

[0314] In S1111, CPU 121 displays the results of the evaluation performed in S1110 on display unit 131 or an external display device connected to camera 100. The user of camera 100 can check the cause of the poor focus from the evaluation results. Knowing the cause makes it possible to take appropriate measures, allowing the user to efficiently learn and master photography techniques.

[0315] In S1112, the CPU 121 presents the best settings to the user. Based on the evaluation results in S1110, the settings that are considered to be best for improving the focus of the image are presented to the user. Along with presenting the best settings, the user may be allowed to select whether or not to change to the best settings. Alternatively, it may be possible to set in advance whether or not to automatically change the settings of the camera 100 to the best settings based on the evaluation results. Furthermore, the evaluation process may be executed during continuous shooting, and the settings of the camera 100 may be automatically changed to the best settings based on the evaluation results while continuous shooting is being performed.

[0316] ●(Defocus map display in evaluation process) An example of defocus map display performed in S1111 based on the evaluation result of S1110 will be described using FIGS. 30(a) to 30(c). 30(a) shows a schematic diagram of the state of the image to be evaluated when it was taken, in which a person skiing is taken.

[0317] Fig. 30(b) shows an example in which a defocus map 30001 corresponding to focus control during shooting is superimposed on the image obtained by shooting shown in Fig. 30(a). Note that the use of a specific camera or lens is not taken into consideration here. The defocus map is generated by the CPU 121 based on shooting-related information included in the accompanying information of the image.

[0318] Here, an area that roughly encompasses the subject area in the image is divided into a 10x8 grid, and for each area, a defocus amount of 0 is defined as the in-focus state, and a defocus map is generated that indicates whether the defocus amount is in the positive direction (front focus) or negative direction (back focus). The method of dividing the areas is merely an example. The defocus map may also be displayed superimposed on the entire image.

[0319] In the figure, a diamond-shaped frame 30002 indicates that the defocus amount is 0 or nearly 0, indicating that the subject is in focus. A diagonally-lined frame 30003 indicates that the defocus amount is positive, indicating front focus. A dotted frame 30004 indicates that the defocus amount is negative, indicating back focus.

[0320] In the defocus map, most of the area of ​​the human subject is framed by a diamond pattern 30002, which indicates that almost the entire subject is in focus.

[0321] Figure 30(c) shows a display similar to Figure 30(b), except that here a defocus map 30001 generated by the CPU 121 based on shooting-related information included in the accompanying information of an image captured using a combination of a camera (product name CA) and a lens (product name LA) is superimposed and displayed.

[0322] AF frame 30000 indicates the position of the AF frame used when the photo was taken. Because AF frame 30000 includes the subject's face area and the background, the accuracy of the defocus amount is reduced by the influence of the low-contrast background. Specifically, the background has caused the subject to be in back focus. Looking at the defocus amount for each block, we can see that there is a large area to the right of the subject that is in front focus.

[0323] 30(d) shows a state in which a defocus map 30001 is superimposed and displayed when the AF frame setting is changed using a combination of a camera (product name CA) and a lens (product name LA). In this case, in the evaluation sequence setting in S1109, the AF frame 30000 is changed to a wide-range AF frame that is different from the single-point AF frame that was actually used. Then, in S1110, the CPU 121 generates a defocus map based on the defocus amount calculated based on the focus control information when the wide-range AF frame is used.

[0324] In this way, by displaying a defocus map based on the settings actually used during shooting (Figure 30(c)) and a defocus map based on different settings (Figure 30(d)), it is possible to evaluate the effect of changing the settings on the degree of focus.

[0325] Compared to Fig. 30(c), the defocus map 30001 in Fig. 30(d) has an increased number of diamond-shaped frames 30002, which indicates that the influence of the background on the defocus amount is suppressed. Therefore, in the example of Fig. 30, it can be seen that an image with a better degree of focus can be obtained by changing the AF frame setting from single-point AF to wide-range AF.

[0326] While the case where a defocus map is displayed when the AF frame setting is changed has been described above, the camera may also be changed. In this case, CPU 121 acquires information related to AF control from virtual camera information of a camera (product name CB) different from the camera (product name CA) used for shooting and existing lens information. CPU 121 then rewrites the focus-related information of the image to be evaluated with the focus control information of the camera (product name CB) to generate a defocus map. This makes it possible to compare the AF performance differences between individual cameras. For example, this is useful when considering purchasing a product, as it allows easy performance comparison of a product currently owned with a new or higher-end product for a desired shooting scene.

[0327] Furthermore, regarding the lens, it is possible to obtain lens information for a lens (product name LB) different from the lens (product name LA) used for shooting, and generate a defocus map for the image to be evaluated for different combinations of focal length, F-number, etc. Therefore, by comparing the defocus information for the lens (product name LB) with that of an image already shot using the lens (product name LA), it is possible to confirm the change in performance when changing lenses without actually using the lens.

[0328] The display as shown in FIG. 30 may be performed on an external display device instead of on the display unit 131 of the camera 100. The display for comparing defocus maps may be a method of displaying two identical images side by side with different defocus maps superimposed on each, or a method of displaying defocus maps on a single image while switching between them. In the example shown in FIG. 30, a defocus map is generated in which the amount of defocus is classified into three levels: near focus, front focus, and back focus, but the number of classifications may be increased, or the amount of defocus may be displayed numerically (in units of mm, for example) within a region. The image to be evaluated may also be a virtual image.

[0329] ●(Focus level) An example of calculating the degree of focus for a series of images captured in time series by continuous shooting or the like will be described with reference to FIG. FIG. 31 shows an image group 31001 obtained by continuously shooting a person skiing as a subject. The image group may be virtual images. CPU 121 calculates the defocus amount in the evaluation of S1110, where a series of images is used as an evaluation image group. Then, for each image, CPU 121 determines, based on the calculation result of the defocus amount, that the degree of focus is good (◯) if the defocus amount is within a predetermined threshold range centered around 0, and determines that the degree of focus is not good (×) if the defocus amount is not within the range. CPU 121 then displays the proportion of images determined to have a good degree of focus among the multiple images included in the evaluation image group as the degree of focus for the series of images.

[0330] FIG. 31 shows an example in which the proportion of images determined to have a good degree of focus among a series of images obtained by continuous shooting is 70%. CPU 121 displays the degree of focus on the image (representative image of the series of images) 31002 currently being displayed. In addition, since the degree of focus is 70%, a mark △ indicating that there is room for improvement is displayed in association with the numerical value of the degree of focus. The determination result of the degree of focus may be shown to the user in an easy-to-understand manner by displaying a mark ◯ if the degree of focus is over 80%, or a mark × if it is below 60%. Displaying the mark is not essential, and only the numerical value of the degree of focus may be displayed.

[0331] In this embodiment, the focus level of each image is classified into two levels, good and bad, but the number of classifications may be increased. The focus levels for multiple images may also be determined using other methods, such as the variance or median of the defocus amount (unit: mm). Similar to the defocus map, the focus level may be displayed on an external display device instead of on the camera's display unit 131.

[0332] ●(List of setting changes) 33 is a diagram showing examples of changeable settings and evaluation conditions for the evaluation in S1110. Here, five items of camera information and two items of lens information can be set.

[0333] Camera information includes AF frame setting, tracking (subject tracking AF) on / off, AF mode (one-shot AF or servo AF), and servo AF characteristics, which are the parameter set types used for servo AF. Shutter method is the type of shutter used. Lens information includes the lens focal length (for prime lenses) or its range (for zoom lenses) and the focal length used during shooting.

[0334] FIG. 33 shows three combinations of setting values ​​for a group of items: initial setting, recommended setting 1, and recommended setting 2. The initial setting is the combination of setting values ​​used when capturing an image. Recommended setting 1 and recommended setting 2 are predetermined combinations and can be selected as an evaluation sequence in S1109 of FIG. 29. There may be three or more evaluation sequences (recommended settings). By preparing recommended settings for all combinations of values ​​that can be set for each item and sequentially setting them in S1110 to evaluate the image, it is possible to identify the best settings (evaluation sequence) that maximize the degree of focus. However, evaluating all combinations would require a large computational load. Therefore, the computational load may be reduced by performing evaluation only for recommended settings that change the values ​​of only items that are effective in improving the degree of focus from the initial setting.

[0335] We will now explain settings that are effective in improving the degree of focus. When shooting using the initial settings shown in Figure 33, the following situations can be considered as factors that cause the AF frame to move away from the subject. With the initial setting of one-point AF and tracking off, the position of the AF frame 32001 observed in the EVF is fixed. This means that the user needs to keep the AF frame aligned with the subject, making it difficult to respond to unexpected movement of the subject.

[0336] On the other hand, with recommended setting 1, there is only one AF frame, but by turning on tracking, the camera can track the subject. Therefore, if the user starts shooting by aligning the subject with the AF frame and starting AF, they can then achieve subject tracking by framing the image so that the subject does not leave the shooting range, making framing easier. Settings related to the AF frame and subject detection AF tracking can be changed and evaluated for both images taken in real space and images taken in virtual space. Using the captured image and defocus map information at the time of shooting, and applying an algorithm corresponding to the changed AF frame setting, it is possible to evaluate the new AF frame setting. The same method can be used to evaluate whether tracking is turned on or off.

[0337] Recommended setting 2 is the same as the default setting in that tracking is turned off, but the AF frame setting is changed to use a larger AF frame, which makes framing easier even when tracking is off.

[0338] The default shutter setting is electronic first curtain (mechanical second curtain). In this case, the shutter curtain is activated each time the shutter is released, causing a blackout in the EVF display. This can cause you to lose sight of the subject for a moment, and the display update rate also slows down, making framing more difficult. On the other hand, by changing the setting to use only the electronic shutter (not the mechanical shutter), you can prevent the display update rate from slowing down during continuous shooting and prevent blackouts in the EVF.

[0339] Therefore, by using an electronic shutter, it is possible to reduce the difficulty of framing without losing sight of the subject, especially when taking continuous shots of a moving subject. Regarding the shutter method settings, for images taken in real space, it is possible to change the direction of decreasing the continuous shooting speed (thinning out images), but changing the direction of increasing the continuous shooting speed is difficult due to the lack of information. On the other hand, for images taken in virtual space, it is possible to redo the shooting operation, so it is also possible to change the continuous shooting speed to a faster speed. Therefore, when evaluating images taken in real space, if the continuous shooting speed is low, such as when an electronic front-curtain shutter is used, the use of an electronic shutter may be recommended to the user based on information such as the subject speed at the time of shooting.

[0340] The default focal length for the lens is a 70mm to 200mm zoom lens, with a telephoto end focal length of 200mm, resulting in a narrow angle of view that makes it difficult to keep a moving subject in the frame. Therefore, the recommended setting is a wide-angle end focal length of 70mm, reducing the risk of a moving subject being out of frame. While it is possible to narrow the angle of view for images captured in real space, it is difficult to widen the angle of view because this creates areas without information (images). On the other hand, for images captured in virtual space, the shooting action can be redone, making it possible to widen the angle of view. Therefore, when evaluating images captured in real space, if the image was captured with a long focal length, the user may be recommended to use a wider focal length based on information such as the subject's speed during shooting.

[0341] As described above, by changing the settings based on factors that reduce the focus level, it is possible to find recommended settings that can more efficiently improve the focus level, thereby reducing the computational load. The setting method in FIG. 33 is merely an example, and settings may be set based on various considerations. Analysis results of the shooting environment, such as whether the subject type is a human or an animal, or whether the scene is one in which multiple subjects intersect, can also be utilized. In addition, the user's skill may be determined based on the shooting history entered in advance, the movement of the subject during shooting, and the user's framing, and the recommended settings may be narrowed down.

[0342] ●(Explanation of information display for shooting settings) Next, the operation of displaying information about the user's framing technique based on the evaluation result of the focus degree described above will be described with reference to FIG.

[0343] FIG. 32(a) shows an image in which a series of continuous shots were evaluated and determined to have a large defocus amount and poor focus (×). This image depicts a case in which the framing of a fast-moving human subject was inappropriate, and the AF frame 32001 was off the subject, resulting in poor focus on the subject. It is assumed that one or more of the user settings during shooting—the shutter method, the size of the AF frame, and the lens focal length—makes it difficult to achieve proper framing. In such a situation, the camera 100's display unit 131 or an external display device displays a suggested optimal setting based on image evaluation during image playback. The suggested optimal setting and its display will be described in FIG. 32(b).

[0344] FIG. 32(b) shows an example of a display for suggesting to the user the best settings identified based on image evaluation for various shooting sequences. The CPU 121 sets various evaluation sequences in S1109, and compares the evaluation results for each evaluation sequence in S1110 to identify the evaluation sequence that will achieve the highest degree of focus on the subject. For example, the CPU 121 performs evaluation on an evaluation sequence that uses an AF frame larger than the AF frame used during shooting, or an evaluation sequence that turns tracking on if tracking was off during shooting. The CPU 121 also performs evaluation on evaluation sequences in which the setting values ​​for other items have been changed to reduce the difficulty of framing.

[0345] Here, it is assumed that the evaluation result shows that the degree of focus of the subject is improved compared to when the photograph was taken by changing the settings as follows. (1) Leave the AF frame at one point and turn on subject detection AF tracking. (2) Shutter method changed to electronic shutter (3) Change the focal length of the lens to the wide-angle side Based on the results of such evaluation, the CPU 121 suggests the best settings to the user through a display 32003 .

[0346] Fig. 32(c) shows an example of display 32004 (GUI) for allowing the user to select whether or not to change the settings of the camera 100, specifically the tracking settings, to ON, in accordance with the best setting proposal proposed in display 32003 in Fig. 32(b). The user can change the settings directly from display 32004 without operating a menu screen or the like.

[0347] When a user instruction as to whether or not to change the tracking settings is detected, the CPU 121 sequentially displays the same display as display 32004 for other items with the best settings. When photographing in real space, the CPU 121 cannot change the focal length of the lens, so a display is displayed to prompt the user to change the focal length. When photographing in virtual space, the focal length of the lens in the virtual camera can be changed, so the CPU 121 displays the same display as display 32004 to prompt the user to make a selection.

[0348] Regarding the change to the best settings, the user may be allowed to select whether or not to change all items at once. Alternatively, the top focus degrees obtained in the evaluation may be displayed and the user may be allowed to select. Then, the CPU 121 automatically sets the recommended settings that have resulted in the selected focus degree to the camera 100. Alternatively, the best settings may be automatically changed without the user's confirmation.

[0349] FIG. 32(d) shows an example of the evaluation results of an image obtained by shooting after changing the settings in FIG. 32(c). Tracking has been turned on, and AF frame 32001 has changed to a dotted line, indicating that tracking is on. The shutter method has also been changed to an electronic shutter, eliminating EVF blackout. Furthermore, the focal length of the lens has been changed from 200mm to 70mm, reducing the proportion of the subject in the shooting range. These factors have made framing easier, and as a result, the focus rate has increased to 85%.

[0350] By evaluating a series of captured images and quantifying the degree of focus, users can understand the capabilities of their own framing skills. In addition, by suggesting changes to settings that may be causing poor focus, users can take photos with improved focus using their current photography skills.

[0351] 32 has been explained assuming that the image is captured in a real space. However, the same evaluation, proposal of the best settings, and setting changes can be performed for the image capture in a virtual space.

[0352] ●(Variation) In this embodiment, the AF area is detected using area detection based on machine learning. However, the AF area may be detected using other methods. For example, the AF area can be set using the aspect ratio of the subject detection area, the size of the subject detection area, depth information of the subject using a defocus map, etc.

[0353] ●(Second embodiment) Next, a second embodiment of the present invention will be described. In this embodiment, a captured image in real space is used in the virtual image generation and output process described in the first embodiment with reference to FIG. 14. This embodiment can be implemented by the camera 100 and external computing device 1000 described in the first embodiment. Therefore, only the virtual image generation and output process in this embodiment will be described.

[0354] FIG. 34 is a flowchart of the virtual image generation and output process in this embodiment, and the same reference numerals as in FIG. 14 are used for steps that perform the same operations as in the first embodiment, and the description thereof will be omitted.

[0355] In S3501, the CPU 1001 acquires a captured image. The captured image may be an image captured in the real space imaging process described above, or may be an image prepared in advance in the storage unit 1004. The CPU 1001 stores the acquired image in the RAM 1003.

[0356] In S3502, the CPU 1001 acquires and composites a foreground object. The CPU 1001 may generate a 3D model of the subject from the captured image acquired in S3501 using, for example, a trained model that estimates a 3D model of an object in an image, stored in the memory unit 1004, and composite the generated model with the foreground object. For example, a 3D model of the head of a human subject may be generated from the captured image, and the 3D model of the person other than the head may be a 3D model of the person acquired from the foreground object memory unit 1104 of the virtual space reproduction unit 1100.

[0357] As another method, generation of a 3D model of the subject from the captured image and acquisition of the 3D model from the foreground object storage unit 1104 may be performed alternately. Alternatively, a foreground object in the virtual space may be acquired and combined with the subject area in the captured image in the real space in the process of generating a virtual image for display in S3503, which will be described later.

[0358] In S3503, the CPU 1001 acquires and composites a background object. As with the acquisition of the foreground object in S3502, a 3D model generated from a captured image may be used as the background object. Alternatively, the background object acquired by the background object acquisition unit 1105 may be composited with a 3D model generated from a captured image. Alternatively, generation of a 3D model of the background object from a captured image and acquisition of a 3D model from the background object storage unit 1105 may be performed alternately. Alternatively, a background object in a virtual space may be acquired and composited with a background area in a captured image in real space in the process of generating a virtual image for display in S3503, which will be described later.

[0359] Next, steps S2003 to S2007 are the same as those in the first embodiment, and therefore the explanation will be omitted.

[0360] After acquiring the correction amount in S2007, the display image generation unit 1204 generates a virtual image for display in S3505. If a composite object has already been generated by combining a 3D object generated from a captured image with a foreground or background object, the display image generation unit 1204 generates a virtual image in the same manner as in S2008 in the first embodiment.

[0361] When generating a display image by combining a captured image with a foreground object and a background object in a virtual space, the captured image is aligned with the display image generated by the foreground object and the background object in the virtual space and the display image is generated. For example, only the face portion of the subject in the captured image is cut out and combined with the display image in the virtual space.

[0362] Furthermore, when the angle of view or viewpoint position of the captured image and the virtual image differ, the CPU 1001 generates a 3D model from the captured image using a trained model that estimates a 3D model of the subject from the image.The CPU 1001 then uses the 3D model to generate a display image with a changed range or viewpoint position, and synthesizes the image with the virtual image for display, thereby generating a virtual image for display that also includes the captured image.

[0363] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0364] The disclosure of the present embodiment includes the following image processing device, image processing method, imaging device and control method thereof, and program. (Item 1) An image processing device that generates a virtual image by capturing an image of a virtual space represented by a three-dimensional model with a virtual camera, an acquisition means for acquiring information on operations performed on the real camera; a generating means for generating a virtual image in which the operation on the real camera is reflected in the virtual camera based on the information on the operation; an output means for outputting the virtual image for display by the real camera; 1. An image processing device comprising: (Item 2) 2. The image processing device according to item 1, wherein the virtual camera simulates the operation of a model different from that of the real camera. (Item 3) 3. The image processing device according to item 2, wherein the user can select the model of the virtual camera to be simulated. (Item 4) 4. The image processing device according to any one of items 1 to 3, characterized in that the operation information includes information on operations performed on the operation unit and / or lens of the real camera and information on the movement of the real camera. (Item 5) 5. The image processing device according to any one of items 1 to 4, wherein the virtual space has a background object and a foreground object, and the foreground object moves over time. (Item 6) 6. The image processing device according to item 5, further comprising a correction amount calculation means for calculating a correction amount for an operation on the real camera based on information about the operation and information about the position and / or size of the foreground object in the virtual image. (Item 7) 7. The image processing device according to item 6, wherein the correction amount calculation means calculates the correction amount for one or more of a framing operation, a zooming operation, and a focusing operation. (Item 8) 8. The image processing device according to item 6 or 7, further comprising a difficulty calculation means for calculating the difficulty of photographing based on one or more of the speed, acceleration, size, contrast, and distance of the foreground object. (Item 9) 9. The image processing device according to item 8, wherein the correction amount calculation means decreases or increases the maximum value of the correction amount as the degree of difficulty increases. (Item 10) 10. The image processing device according to any one of items 1 to 9, characterized in that when an operation involving movement of a focus lens is performed in the real camera based on the operation information, driving of a virtual focus lens is also performed in the virtual camera. (Item 11) 11. The image processing device according to any one of items 1 to 10, characterized in that the viewpoint position of the virtual camera is changed in accordance with an operation that changes the focal length of the real camera and the position or shooting direction of the focus detection area. (Item 12) 12. The image processing device according to any one of items 1 to 11, which is a part of the real camera. (Item 13) An image processing method executed by an image processing device for generating a virtual image by capturing an image of a virtual space represented by a three-dimensional model with a virtual camera, the method comprising: Acquiring information about operations on a real camera; generating a virtual image in which the operation on the real camera is reflected on the virtual camera based on the information on the operation; outputting the virtual image for display by the real camera; An image processing method comprising: (Item 14) 12. A program for causing a computer to function as each of the means possessed by the image processing device according to any one of items 1 to 11. (Item 15) An imaging element; an imaging device having a display device, a real space imaging mode for imaging using the imaging element and a virtual space imaging mode for imaging an image of a virtual space, an imaging device characterized in that, in the virtual space shooting mode, an image of the virtual space is displayed in live view on the display device, and a still image of the virtual space is recorded by operating a release switch of the imaging device. (Item 16) Item 16. The imaging device according to item 15, wherein the driving of the movable member is performed in the same manner as in the imaging mode of the real space, even in the imaging mode of the virtual space. (Item 17) Item 17. The imaging device according to item 16, wherein the movable member includes one or more of an aperture, a mechanical shutter, and a focus lens. (Item 18) An imaging element; A control method for an imaging device having a display device, the method comprising: the imaging device has a real space imaging mode for imaging using the imaging element and a virtual space imaging mode for imaging an image of a virtual space, In the virtual space photography mode, displaying a live view of an image of the virtual space on the display device; and recording a still image of the virtual space in response to operation of a release switch of the imaging device. (Item 19) A program for causing a computer included in an imaging device having an imaging element and a display device to implement the imaging device control method described in item 18. (Item 20) a generating means for generating a first defocus map according to settings at the time of photographing the image and a second defocus map according to settings different from the settings at the time of photographing the image, based on accompanying information recorded together with the image; a display means for displaying the first and second defocus maps on a display device; 1. An image processing device comprising: (Item 21) An image processing method implemented by an image processing device, generating a first defocus map according to settings at the time of capturing the image based on accompanying information recorded together with the image; generating a second defocus map according to settings different from the settings at the time of shooting; displaying the first and second defocus maps on a display device; An image processing method comprising: (Item 22) a generating means for generating a first defocus map for each of a series of images captured in time series based on accompanying information recorded together with the image in accordance with settings at the time the image was captured; a calculation means for determining a degree of focus for each of the series of images based on the first defocus map, and calculating a ratio of the images determined to have a good degree of focus as a degree of focus for the series of images; 1. An image processing device comprising: (Item 23) the generating means generates a plurality of second defocus maps for each of the series of images, each of which is based on a setting different from the setting at the time of shooting; the calculation means calculates a degree of focus for the series of images based on the first defocus map and a degree of focus for the series of images based on each of the plurality of second defocus maps; 23. The image processing device according to item 22, further comprising a specifying means for specifying a setting that maximizes the degree of focus for the series of images based on the first defocus map and the plurality of second defocus maps. (Item 24) 24. The image processing device according to item 23, wherein the image processing device is an imaging device, and the setting specified by the specifying means is reflected in the imaging device. (Item 25) An image processing method implemented by an image processing device, generating a defocus map for each of a series of images captured in time series based on accompanying information recorded together with the image and in accordance with settings at the time of capturing the image; determining a degree of focus for each of the series of images based on the defocus map; calculating a ratio of the images determined to have a good degree of focus as a degree of focus for the series of images; An image processing method comprising: (Item 26) A program for causing a computer to function as each of the means possessed by the image processing device according to any one of items 20, 22 to 24.

[0365] The present invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Therefore, the following claims are appended to clarify the scope of the invention. [Explanation of symbols]

[0366] 100... camera, 107... image sensor, 121... CPU, 131... display unit, 1000... external computing device, 1001... CPU, 1002... ROM, 1003... RAM, 1004... storage unit, 2000... camera / lens information storage device

Claims

1. An image processing device that generates a virtual image by capturing an image of a virtual space represented by a three-dimensional model with a virtual camera, an acquisition means for acquiring information on operations performed on the real camera; a generating means for generating a virtual image in which the operation on the real camera is reflected in the virtual camera based on the information on the operation; an output means for outputting the virtual image for display by the real camera; 1. An image processing device comprising:

2. 2. The image processing device according to claim 1, wherein the virtual camera simulates the operation of a model different from that of the real camera.

3. 3. The image processing device according to claim 2, wherein a user can select a model of the virtual camera to be simulated.

4. 2. The image processing device according to claim 1, wherein the operation information includes information on operations performed on an operation unit and / or a lens of the real camera, and information on the movement of the real camera.

5. 2. The image processing apparatus according to claim 1, wherein the virtual space has a background object and a foreground object, and the foreground object moves with time.

6. 6. The image processing device according to claim 5, further comprising a correction amount calculation means for calculating a correction amount for an operation on the real camera based on information about the operation and information about the position and / or size of the foreground object in the virtual image.

7. 7. The image processing apparatus according to claim 6, wherein the correction amount calculation means calculates the correction amount for one or more of a framing operation, a zooming operation, and a focusing operation.

8. 7. The image processing apparatus according to claim 6, further comprising difficulty calculation means for calculating the difficulty of photographing based on one or more of the speed, acceleration, size, contrast, and distance of the foreground object.

9. 9. The image processing apparatus according to claim 8, wherein the correction amount calculation means increases or decreases the maximum value of the correction amount as the degree of difficulty increases.

10. 2. The image processing device according to claim 1, wherein when an operation involving movement of a focus lens is performed in the real camera based on the operation information, a virtual focus lens is also driven in the virtual camera.

11. 2. The image processing device according to claim 1, wherein the viewpoint position of the virtual camera is changed in accordance with an operation for changing the focal length of the real camera, and the position or photographing direction of the focus detection area.

12. 2. The image processing device according to claim 1, wherein the image processing device is a part of the real camera.

13. An image processing method executed by an image processing device for generating a virtual image by capturing an image of a virtual space represented by a three-dimensional model with a virtual camera, the method comprising: Acquiring information about operations on a real camera; generating a virtual image in which the operation on the real camera is reflected on the virtual camera based on the information on the operation; outputting the virtual image for display by the real camera; An image processing method comprising:

14. A program for causing a computer to function as each of the means included in the image processing device according to any one of claims 1 to 11.

15. An imaging element; an imaging device having a display device, a real space imaging mode for imaging using the imaging element and a virtual space imaging mode for imaging an image of a virtual space, an imaging device characterized in that, in the virtual space shooting mode, an image of the virtual space is displayed in live view on the display device, and a still image of the virtual space is recorded by operating a release switch of the imaging device.

16. 16. The imaging device according to claim 15, wherein the movable member is driven in the virtual space imaging mode in the same manner as in the real space imaging mode.

17. 17. The imaging device according to claim 16, wherein the movable member includes one or more of an aperture, a mechanical shutter, and a focus lens.

18. An imaging element; A control method for an imaging device having a display device, the method comprising: the imaging device has a real space imaging mode for imaging using the imaging element and a virtual space imaging mode for imaging an image of a virtual space, In the virtual space photography mode, displaying a live view of an image of the virtual space on the display device; and recording a still image of the virtual space in response to operation of a release switch of the imaging device.

19. 20. A program for causing a computer included in an imaging device having an imaging element and a display device to execute the imaging device control method according to claim 18.

20. a generating means for generating a first defocus map according to settings at the time of photographing the image and a second defocus map according to settings different from the settings at the time of photographing the image, based on accompanying information recorded together with the image; a display unit for displaying the first and second defocus maps on a display device; 1. An image processing device comprising:

21. An image processing method implemented by an image processing device, generating a first defocus map according to settings at the time of capturing the image based on accompanying information recorded together with the image; generating a second defocus map according to settings different from the settings at the time of shooting; displaying the first and second defocus maps on a display device; An image processing method comprising:

22. a generating means for generating a first defocus map for each of a series of images captured in time series based on accompanying information recorded together with the image in accordance with settings at the time the image was captured; a calculation means for determining a degree of focus for each of the series of images based on the first defocus map, and calculating a proportion of the images determined to have a good degree of focus as a degree of focus for the series of images; 1. An image processing device comprising:

23. the generating means generates a plurality of second defocus maps for each of the series of images, each of which is based on a setting different from the setting at the time of shooting; the calculation means calculates a degree of focus for the series of images based on the first defocus map and a degree of focus for the series of images based on each of the plurality of second defocus maps; 23. The image processing device according to claim 22, further comprising: a specifying unit that specifies a setting that maximizes the degree of focus for the series of images based on the first defocus map and the plurality of second defocus maps.

24. 24. The image processing apparatus according to claim 23, wherein the image processing apparatus is an image capturing apparatus, and the setting specified by the specifying means is reflected in the image capturing apparatus.

25. An image processing method implemented by an image processing device, generating a defocus map for each of a series of images captured in time series based on accompanying information recorded together with the image and in accordance with settings at the time of capturing the image; determining a degree of focus for each of the series of images based on the defocus map; calculating a ratio of the images determined to have a good degree of focus as a degree of focus for the series of images; An image processing method comprising:

26. A program for causing a computer to function as each of the means included in the image processing device according to any one of claims 20 and 22 to 24.

Citation Information

Patent Citations

  • Imaging apparatus, imaging method and program

    JP2008078908A