Imaging device and its control method, program, and storage medium
The imaging device synchronizes virtual space photography with real-space controls, providing a seamless shooting experience by replicating the sense of control in virtual environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-17
- Publication Date
- 2026-03-30
AI Technical Summary
Existing imaging technologies using VR goggles fail to replicate the operation feeling of shooting in real space due to differences in user interface and equipment, lacking a seamless experience between virtual and real-space photography.
An imaging device with operating, driving, and control means that synchronizes the drive unit with the photographing sequence in virtual space, replicating the sense of control akin to real-space shooting.
Enables photographers to experience the same sense of control in virtual space as in real space, bridging the gap in operation feeling.
Smart Images

Figure 2026054958000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology for reproducing the operation feeling in shooting in a virtual space.
Background Art
[0002] Patent Document 1 discloses a method of synthesizing a video of a virtual space by a 3D model and a video of a real space shot by a camera, and further creating a video reflecting the operation information of the photographer. In Patent Document 1, shooting can be performed while the photographer views a video in which a subject existing in the real space and a background generated in the virtual space are synthesized.
[0003] In recent years, a usage method of capturing an image in a virtual space using a VR goggle or the like has been proposed. At that time, the user is fed back that an image has been captured using a controller held by hand or sound or the like.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, Patent Document 1 does not mention shooting in an environment where both the subject and the background are generated only in the virtual space. Further, in the case of capturing an image using a VR goggle, there is a problem that the operation feeling is different from that in shooting in the real space because the user interface and the equipment to be held are different from those of the equipment (so-called camera) used in shooting in the real space.
[0006] The present invention has been made in view of the above-described problems, and an object thereof is to provide an imaging device that enables a photographer to experience the operation feeling obtained in shooting in a real space in shooting in a virtual space. [Means for solving the problem]
[0007] The imaging device according to the present invention comprises an operating means for performing operations for taking photographs, a driving means for driving a drive unit for taking photographs in real space, and a control means for controlling the driving means, wherein the control means controls the driving means so that when the operating means instructs taking photographs in a virtual space, the drive unit is driven in synchronization with the photographing sequence in the virtual space. [Effects of the Invention]
[0008] According to the present invention, when shooting in a virtual space, the photographer can experience the same sense of control as when shooting in a real space. [Brief explanation of the drawing]
[0009] [Figure 1] A block diagram showing the configuration of an imaging system according to a first embodiment of the present invention. [Figure 2] A block diagram showing the configuration of the camera of the present invention. [Figure 3] A diagram showing the pixel arrangement in the camera of the first embodiment. [Figure 4] Plan view and cross-sectional view of a pixel in the first embodiment. [Figure 5] A diagram showing the focus detection region in the first embodiment. [Figure 6] A block diagram showing the hardware configuration of the external computing device in the first embodiment. [Figure 7] A block diagram showing the functional configuration of the external computing device in the first embodiment. [Figure 8] A flowchart illustrating the processing of real-world space photography and virtual space photography in the first embodiment. [Figure 9] A flowchart illustrating the real-world space imaging process in the first embodiment. [Figure 10] A flowchart illustrating the imaging process in the first embodiment. [Figure 11] Flowchart for explaining subject tracking AF processing in the first embodiment. [Figure 12] Flowchart for explaining subject detection / tracking processing in the first embodiment. [Figure 13] Flowchart for explaining virtual space shooting processing in the first embodiment. [Figure 14] Diagram for explaining information of the camera lens information storage device, camera / lens, and external arithmetic device in the first embodiment. [Figure 15] Flowchart for explaining video generation / output in the virtual space in the first embodiment. [Figure 16] Flowchart for explaining virtual subject tracking processing in the first embodiment. [Figure 17] Flowchart for obtaining shooting difficulty information in the first embodiment. [Figure 18] Diagram for explaining correction related to framing in the first embodiment [Figure 19] Diagram for explaining correction related to zooming in the first embodiment. [Figure 20] Diagram for explaining correction related to focusing in the first embodiment. [Figure 21] Flowchart for explaining defocus amount processing in the first embodiment. [Figure 22] Diagram showing a graph of virtual defocus amount calculation in the first embodiment. [Figure 23] Diagram showing an example of a virtual defocus map in the first embodiment. [Figure 24] Flowchart for explaining a subroutine of virtual space shooting in the first embodiment. [Figure 25] Diagram for explaining virtual space shooting reflecting operation information in the first embodiment. [Figure 26] Flowchart for explaining a subroutine of viewpoint movement processing in the first embodiment. [Figure 27] Diagram showing an example of viewpoint movement in the first embodiment. [Figure 28] Flowchart illustrating the camera operation during virtual space photography in the first embodiment. [Figure 29] A flowchart illustrating the playback of the shooting results and evaluation of the captured images in the first embodiment. [Figure 30] An explanatory diagram of the defocus map display in the first embodiment. [Figure 31] An explanatory diagram showing the degree of focus display of a series of captured images in the first embodiment. [Figure 32] An explanatory diagram of the setting change list table in the first embodiment. [Figure 33] An explanatory diagram of the best setting display after evaluation in the first embodiment. [Figure 34] A flowchart illustrating the generation and output of images in a virtual space in the second embodiment. [Modes for carrying out the invention]
[0010] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.
[0011] <First Embodiment> Figure 1 shows the configuration of an imaging system 10, which includes an imaging device, an external computing device (information processing device), and a camera / lens information storage device, according to a first embodiment of the present invention.
[0012] In Figure 1, the imaging device (camera) 100 has the function of capturing images of subjects existing in the real world, the function of instructing the imaging of subjects existing in a virtual space, and the function of displaying the captured images. The camera 100 also functions as a tactile sensation reproduction device.
[0013] The external computing unit 1000 is connected to the camera 100 by wired or wireless means to exchange information, and is configured to include a virtual space reproduction device 1100 and a virtual image generation device 1200. The virtual space reproduction device 1100 places subjects as virtual space objects whose position and shape change moment by moment within a set virtual space (background space). The virtual image generation device 1200 acquires camera and lens setting information, control information, operation information of operating components, and position information including shooting direction from the camera 100. It also uses the information obtained from the camera 100 to acquire related information from the camera / lens information storage device 2000.
[0014] The camera / lens information storage device 2000 may be a server on a network such as the cloud, or it may be located within the external computing device 1000.
[0015] The virtual image generation device 1200 uses the information acquired as described above to generate (capture) an image from the virtual space constructed by the virtual space reproduction device 1100. The image generated here may be a two-dimensional image or a three-dimensional image containing information that allows for stereoscopic display. In Figure 1, the virtual space reproduction and image generation are performed by the external processing unit 1000, but these functions may also be implemented within the camera 100.
[0016] Figure 2 shows the configuration of a camera 100 as an imaging device in the first embodiment of the present invention. In Figure 2, the first lens group 101 is positioned on the subject side (front side) of the imaging optical system, which is an imaging optical system, and is held so as to be movable in the optical axis direction. The aperture 102 adjusts the amount of light by adjusting its aperture diameter. The second lens group 103 moves in the optical axis direction together with the aperture 102 and performs magnification (zoom) together with the first lens group 101 which moves in the optical axis direction.
[0017] The third lens group (focusing lens) 105 moves along the optical axis to adjust the focus. The optical low-pass filter 108 is an optical element that reduces false colors and moiré in the captured image. The imaging optical system is composed of the first lens group 101, the aperture 102, the second lens group 103, the third lens group 105, and the optical low-pass filter 108.
[0018] The zoom actuator 111 rotates a cam cylinder (not shown) around the optical axis, causing a cam on the cam cylinder to move the first lens group 101 and the second lens group 103 in the optical axis direction, thereby changing the magnification. The aperture actuator 112 drives a plurality of light-shielding vanes (not shown) in the opening and closing direction to adjust the light intensity of the aperture 102. The focus actuator 114 moves the third lens group 105 in the optical axis direction to adjust the focus.
[0019] The focus drive circuit 126 drives the focus actuator 114 in response to a focus drive command from the camera CPU 121, moving the third lens group 105 in the optical axis direction. The aperture drive circuit 128 drives the aperture actuator 112 in response to an aperture drive command from the camera CPU 121. The zoom drive circuit 129 drives the zoom actuator 111 in response to the user's zoom operation.
[0020] In this embodiment, the interchangeable lens, which includes the imaging optical system, actuators 111, 112, 114, and drive circuits 126, 128, 129, is configured to be detachable from the camera body using a mount portion M that enables electrical and mechanical connections. However, the imaging optical system, actuators 111, 112, 114, and drive circuits 126, 128, 129 may be provided integrally with the camera body, which includes the image sensor 107.
[0021] The electronic flash 115 has a light-emitting element such as a xenon tube or LED, and emits light to illuminate the subject. The AF assist light-emitting unit 116 has a light-emitting element such as an LED, and projects an image of a mask having a predetermined aperture pattern onto the subject via a projection lens, thereby improving focus detection performance for dark or low-contrast subjects. The electronic flash control circuit 122 controls the electronic flash 115 to light up in synchronization with the imaging operation. The assist light drive circuit 123 controls the AF assist light-emitting unit 116 to light up in synchronization with the focus detection operation.
[0022] The camera CPU 121 controls various functions in the camera 100. The camera CPU 121 includes an arithmetic unit, ROM, RAM, A / D converter, D / A converter, and communication interface circuitry. The camera CPU 121 drives various circuits within the camera 100 according to computer programs stored in the ROM, and controls a series of operations such as autofocus, imaging, image processing, and recording. The camera CPU 121 also functions as an image processing device.
[0023] The image sensor 107 consists of a two-dimensional CMOS photosensor containing multiple pixels and its peripheral circuits, and is positioned on the imaging plane of the imaging optical system. The image sensor 107 converts the subject image formed by the imaging optical system into a digital signal. The image sensor drive circuit 124 controls the operation of the image sensor 107 and performs A / D conversion on the analog signal generated by the photoelectric conversion to transmit the digital signal to the camera CPU 121.
[0024] The shutter 106 has a focal-plane shutter configuration and is driven by a shutter drive circuit built into the shutter 106 based on instructions from the camera CPU 121. During signal readout from the image sensor 107, the shutter blocks light from the image sensor 107. When exposure is taking place, the focal-plane shutter opens, and the photographic light beam is directed to the image sensor 107.
[0025] The image processing circuit 125 performs predetermined image processing on the image data stored in the RAM of the camera CPU 121. The image processing performed by the image processing circuit 125 includes, but is not limited to, so-called development processing such as white balance adjustment, color interpolation (demosaic) processing, and gamma correction processing, as well as signal format conversion processing and scaling processing. Furthermore, the image processing circuit 125 determines the main subject based on the posture information of the subject and the position information of objects specific to the scene (hereinafter referred to as specific objects). The result of the determination processing may be used for other image processing (for example, white balance adjustment processing). The image processing circuit 125 stores the processed image data, joint position information of each subject, position and size information of specific objects, center of gravity position information of the subject determined to be the main subject, and face and pupil position information in the RAM of the camera CPU 121.
[0026] The display unit (display means) 131 is equipped with a display element such as an LCD and displays information related to the camera 100's imaging mode, a preview image before imaging, a confirmation image after imaging, an indicator of the focus detection area, and a focused image. The operation switch group 132 includes a main (power) switch, a release (shooting trigger) switch, a zoom operation switch, a shooting mode selection switch, etc., and is operated by the user. The flash memory 133 records the captured images. The flash memory 133 is detachable from the camera 100.
[0027] The subject detection unit 140, acting as a subject detection means, performs subject detection based on dictionary data generated by machine learning. In this embodiment, the subject detection unit 140 uses dictionary data for each subject in order to detect multiple types of subjects. Each dictionary data is, for example, data in which the characteristics of the corresponding subject are registered. The subject detection unit 140 performs subject detection by sequentially switching between the dictionary data for each subject. The dictionary data for each subject is stored in the dictionary data storage unit (ROM in the camera CPU 121). Therefore, multiple dictionary data are stored in the dictionary data storage unit. The camera CPU 121 determines which dictionary data to use for subject detection from among the multiple dictionary data based on the pre-set subject priority and imaging device settings.
[0028] The video input unit 141 receives the generated image when shooting (image generation) in the virtual space, and the camera CPU 121 processes the input image by displaying it on the display unit 131 or storing it in the flash memory 133. The information output unit 142 outputs various information to the external processing unit 1000 when shooting in the virtual space. The camera operation information output includes release operations for issuing shooting commands, lens zooming, and focus operations. The camera setting information output includes setting information related to the mode for continuous shooting, autofocus, metering, exposure condition settings, image generation, and lens control. The camera control information output includes information related to correction values and thresholds used in various algorithms used for shooting and image generation. Information indicating the camera position and shooting direction is also output. Details will be described later.
[0029] Examples of dictionary data for subject detection include dictionary data for detecting "people," dictionary data for detecting "animals," and dictionary data for detecting "vehicles." Furthermore, dictionary data for detecting "the whole person" and dictionary data for detecting "the person's face" may be stored separately in the dictionary data storage unit.
[0030] In this embodiment, the subject detection unit 140 is composed of a machine learning-trained CNN (Convolutional Neural Network) and estimates the position of subjects included in the image data. The subject detection unit 140 may be implemented using a GPU (Graphics Processing Unit) or a circuit specialized for estimation processing using a CNN.
[0031] Machine learning of a CNN can be performed using any method. For example, a designated computer, such as a server, may perform machine learning of the CNN, and the camera 100 may acquire the trained CNN from the designated computer. For example, the designated computer may perform supervised learning using training image data as input and the positions of subjects corresponding to the training image data as training data, thereby training the CNN of the subject detection unit 140. As a result, a trained CNN is generated. The CNN training may be performed by the camera 100 or the image processing device described above.
[0032] Next, the image array of the image sensor 107 will be explained using Figure 3. Figure 3 shows the pixel array of the image sensor 107 in a range of 4 pixel rows × 4 pixel rows, viewed from the optical axis direction (z direction).
[0033] Each pixel unit 200 contains four imaging pixels arranged in a 2x2 grid. By arranging a large number of pixel units 200 on the image sensor 107, photoelectric conversion of a two-dimensional subject image can be performed. Within each pixel unit 200, an imaging pixel 200R with spectral sensitivity for red (R) is located in the upper left, and imaging pixels 200G with spectral sensitivity for green (G) are located in the upper right and lower left. Furthermore, an imaging pixel 200B with spectral sensitivity for blue (B) is located in the lower right. Each imaging pixel also contains a first focus detection pixel 201 and a second focus detection pixel 202, which are divided in the horizontal direction (x direction).
[0034] In the image sensor 107 of this embodiment, the pixel pitch P of the imaging pixels is 4 μm, and the number of imaging pixels N is approximately 20.75 million pixels (5575 horizontal columns × 3725 vertical rows). The pixel pitch PAF of the focus detection pixels is 2 μm, and the number of focus detection pixels NAF is approximately 41.5 million pixels (11150 horizontal columns × 3725 vertical rows).
[0035] In this embodiment, the case in which each imaging pixel is divided into two horizontally is described, but it may also be divided vertically. Furthermore, although the image sensor 107 in this embodiment has a plurality of imaging pixels, each including a first and a second focus detection pixel, the imaging pixels and the first and second focus detection pixels may be provided as separate pixels. For example, the first and second focus detection pixels may be discretely arranged among the plurality of imaging pixels.
[0036] Figure 4(a) shows one imaging pixel (200R, 200G, 200B) as viewed from the light-receiving side (+z direction) of the image sensor 107. Figure 4(b) shows the cross-section aa of the imaging pixel in Figure 4(a) as viewed from the -y direction. As shown in Figure 4(b), one imaging pixel is provided with one microlens 305 for focusing incident light.
[0037] Furthermore, the imaging pixel is provided with photoelectric conversion units 301 and 302, which are divided into N sections (2 sections in this embodiment) in the x direction. The photoelectric conversion units 301 and 302 correspond to the first focus detection pixel 201 and the second focus detection pixel 202, respectively. The centers of gravity of the photoelectric conversion units 301 and 302 are eccentric to the -x side and the +x side, respectively, with respect to the optical axis of the microlens 305.
[0038] A color filter 306 of R, G, or B is provided between the microlens 305 and the photoelectric conversion units 301 and 302 in each imaging pixel. The spectral transmittance of the color filter may be changed for each photoelectric conversion unit, or the color filter may be omitted.
[0039] Light incident on the imaging pixel from the imaging optical system is focused by the microlens 305, spectrally separated by the color filter 306, and then received by the photoelectric conversion units 301 and 302, where it is photoelectrically converted. The camera 100 having the image sensor 107 shown in Figures 3 and 4 can perform so-called phase-difference focus detection, which detects the phase difference from a pair of signal sequences obtained by dividing the light beam passing through the imaging optical system, using known techniques (for example, Japanese Patent Application Publication No. 2023-95509). Phase-difference focus detection makes it possible to detect the amount of defocus in a predetermined area within the imaging range, including its direction. A detailed explanation is omitted.
[0040] Next, the focus detection region of the image sensor 107, which is the region from which a pair of signal sequences for detecting phase difference is acquired, will be explained using Figure 5. In Figure 5, A(n,m) represents the nth focus detection region in the x-direction and the mth focus detection region in the y-direction, out of a total of nine focus detection regions (three each in the x-direction and y-direction) set in the effective pixel region 300 of the image sensor 107. A pair of signal sequences is generated from multiple pixels contained in the focus detection region A(n,m). I(n,m) represents an index that displays the position of the focus detection region A(n,m) on the display unit 131.
[0041] Note that the nine focus detection regions shown in Figure 5 are merely examples, and the number, position, and size of the focus detection regions are not limited. For example, one or more regions may be set as focus detection regions within a predetermined range centered on a position specified by the user or the position of the subject detected by the subject detector. In this embodiment, the focus detection regions are arranged to obtain a higher resolution focus detection result when acquiring the defocus map described later. For example, a total of 9600 focus detection regions are arranged on the image sensor with 120 horizontal divisions and 80 vertical divisions.
[0042] Figure 6 is a block diagram showing an example of the hardware configuration of the external computing device 1000. The external computing device 1000 consists of a CPU 1001, RAM 1003, ROM 1002, storage unit 1004, input interface 1005, output interface 1006, and system bus 1007. The input interface 1005 is connected to the camera 100 and the camera / lens information storage device 2000. The output interface 1006 is connected to the camera 100.
[0043] The CPU 1001 is a processor that comprehensively controls each component of the external arithmetic unit 1000. The RAM 1003 is memory that functions as the main memory and work area of the CPU 1001. The ROM 1002 is memory that stores programs and other data used for processing within the external arithmetic unit 1000. The CPU 1001 uses the RAM 1003 as a work area and executes programs stored in the ROM 1002 to perform various processes described later.
[0044] The storage unit 1004 is a storage device that stores image data used for processing by the external computing device 1000, as well as parameters (i.e., setting values) for said processing. The storage unit 1004 can be an HDD, optical disc drive, flash memory, or the like.
[0045] The input interface 1005 is, for example, a serial bus interface such as USB or IEEE1394. The external computing device 1000 can acquire the various types of information described above from the camera 100 via the input interface 1005. The output interface 1006 is, for example, a video output terminal such as DVI or HDMI (registered trademark). The external computing device 1000 can output image data processed by the external computing device 1000 to the display 131 of the camera 100 via the output interface 1006. It can also output images for recording to the flash memory 133 of the camera 100. The external computing device 1000 may include components other than those described above, but since they are not the main focus of the present invention, a detailed explanation of them is omitted.
[0046] Next, the virtual image generation process performed by the external computing device 1000 using the hardware configuration of Figure 6 will be explained using Figure 7. Figure 7 is a block diagram showing the functional configuration of the external computing device 1000. In this embodiment, each block shown in Figure 7 is realized by the CPU 1001 executing the program stored in the ROM 1002. However, the CPU 1001 does not need to perform all functions, and processing circuits that perform each function may be provided in each part of the external computing device 1000.
[0047] First, let's describe the virtual space reproduction device 1100. The foreground object acquisition unit 1102 acquires 3D objects of foreground subjects, such as performers on stage, stored in the foreground object storage unit 1101. A 3D object is 3D shape data that describes information indicating shape and color, and is composed of textured mesh models or 3D point clouds with colored points. Note that 3D objects do not necessarily have color. Objects stored in the foreground object storage unit 1101 can be various 3D objects such as people of different races, genders, and ages, various animals, and moving objects such as cars. The foreground object acquisition unit 1102 is not limited to acquiring one foreground object, but may acquire multiple objects. The 3D object also has subject information such as velocity, acceleration, angular velocity, angular acceleration, size, and contrast. In addition, a pre-trained model that estimates a 3D model for an image may be used to generate a 3D object from an image of a subject that the photographer wants to photograph in virtual space photography, and size and contrast information may also be stored. Furthermore, by using multiple time-series images, it may be possible to generate and store information on the velocity, acceleration, angular velocity, and angular acceleration of the 3D object.
[0048] Alternatively, a 3D object may be generated and stored from images captured using multiple imaging devices with different viewpoints. Multiple imaging devices capture the imaging area from multiple directions. This imaging area could be, for example, an indoor imaging studio or a stage where a play is performed. The multiple imaging devices are installed at different positions surrounding this imaging area and perform imaging synchronously. Note that the multiple imaging devices do not need to be installed around the entire circumference of the imaging area; depending on the limitations of the installation location, they may be installed only in a part of the imaging area. The number of imaging devices can be set in various ways; for example, if the imaging area is a soccer field, about 30 imaging devices may be installed around the field. Also, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0049] For each of the multiple imaging devices, a parameter set may be described for each imaging device, including parameters representing the three-dimensional position, parameters representing the orientation of the imaging device in the pan, tilt, and roll directions, and the size of the imaging device's field of view (angle of view) and resolution. The information included in the parameter set is calculated in advance using a known camera calibration procedure and stored in a suitable storage device (e.g., foreground object storage unit 1101). That is, points in multiple images based on imaging by the multiple imaging devices are associated and calculated by geometric calculation. Note that the content of the information included in the parameter set is not limited to the above. For example, there may be multiple parameter sets corresponding to multiple frames that constitute a video of the imaging device, and the information may indicate the position and orientation of the imaging device at each of multiple consecutive time points.
[0050] The foreground object acquisition unit 1102 generates a 3D object of a person, such as a stage performer, which is a foreground subject, based on images from multiple viewpoints and a parameter set received from the imaging device, for example, according to the method described in Japanese Patent Application Publication No. 2017-211827.
[0051] Similarly, the background object acquisition unit 1105 acquires a 3D object stored in the background object storage unit 1104, which serves as a space for placing foreground objects, such as a stage or a stadium. The background objects stored in the background object storage unit 1104 could include 3D objects of various spaces, such as a large concert hall, a soccer stadium, or a small indoor room. The background object may use design data from CAD or other software, or shape and color data scanned with a laser scanner. Alternatively, it may be generated from a set of images from multiple viewpoints using computer vision techniques such as Structure from Motion.
[0052] The object compositing unit 1103 places the foreground object within the space of the acquired background object. The information about the foreground object acquired by the object compositing unit 1103 may include three-dimensional models of the subject at multiple times, corresponding to the shape and color of the subject at multiple times. During placement, the unit ensures that the foreground object does not appear to float relative to the ground contained within the background object, excluding interference between objects and actions such as jumping. The foreground object may be placed according to the object placement information (position, orientation) of the background object, or it may be placed based on external instructions such as those from the user.
[0053] Next, the virtual image generation device 1200 will be described.
[0054] The viewpoint information acquisition unit 1201 acquires virtual viewpoint parameters, including the position and direction (pan, tilt, roll) of the virtual viewpoint within the virtual space. The virtual viewpoint parameters may be set as initial values, registered values, previous history positions, etc., within the virtual space, or they may be set according to user instructions.
[0055] The camera lens information acquisition unit 1202 acquires information about the camera and lens used for virtual space shooting from the camera / lens information storage device 2000 or the camera 100. Details of the information will be described later. The camera lens information update unit 1203 acquires and updates the camera lens information, which is updated over time, as needed.
[0056] The operation information acquisition unit 1205 acquires operation information for the camera and lens from the camera 100. Details of the information will be described later.
[0057] Viewpoint information, camera lens information, and operation information are input to the image correction amount calculation unit 1206, which calculates the image correction amount. The image correction amount calculation unit 1206 calculates the image correction amount using information obtained from the shooting difficulty calculation unit 1261 and the photographer's intention extraction unit 1262. Details of the process will be described later.
[0058] The display image generation unit 1204 uses foreground and background object information obtained from the object composition unit 1103, along with virtual viewpoint information and camera lens information, to perform rendering and generate a virtual image. The generated virtual image is output to the camera 100 and displayed on the camera 100's display unit 131. It is also recorded in the camera 100's flash memory 133 and the external processing unit 1000's storage unit 1004.
[0059] (Photo processing) The flowchart in Figure 8 shows the process of causing the camera 100 of this embodiment to perform real-space and virtual-space imaging. Specifically, it shows the process from the pre-image operation of displaying an image on the camera 100's display unit 131 to the actual still image capture. The camera CPU 121, which is a computer, executes this process according to the computer program. In the following description, S means step.
[0060] First, in S1, the camera CPU 121 starts displaying settings menus and live view images of the real or virtual space on the display unit 131. The generation of the live view images to be played back will be described later. Upon initial startup or user operation, the display unit 131 displays a menu setting screen in which the user can select whether to shoot in the real space or the virtual space. The displayed content may be determined based on the history of the previous startup. If shooting in the real space or virtual space has already been set, the previously started live view display will continue.
[0061] In S2, the camera CPU 121 determines whether or not to perform virtual space photography based on user instructions or previous history. If the answer in S2 is Yes, the process proceeds to S1000 and virtual space photography is performed. On the other hand, if the answer in S2 is No, the process proceeds to S10 and real space photography is performed. After completing the processes in S10 or S1000, the process proceeds to S3.
[0062] In S3, the camera CPU 121 determines whether the main switch included in the operation switch group 132 is turned off or not. If the main switch is turned off, the camera CPU 121 terminates this process; otherwise, it returns to S1.
[0063] (Real-world image processing) The flowchart in Figure 9 illustrates the real-world image capture process shown in S10 of Figure 8. Specifically, it shows the process from the pre-image capture operation of displaying a live view image on the camera 100's display unit 131 to the capture of a still image. The camera CPU 121, which is a computer, executes this process according to the computer program. In the following description, S represents a step.
[0064] First, in S11, the camera CPU 121 drives the image sensor 107 using the image sensor drive circuit 124 and acquires imaging data from the image sensor 107. Then, from the acquired imaging data, the camera CPU 121 acquires pairs of focus detection signals from the pairs of focus detection pixels included in each of the focus detection regions shown in Figure 5. The camera CPU 121 also generates an imaging signal by adding the pairs of focus detection signals from all the effective pixels of the image sensor 107, and has the image processing circuit 125 perform image processing on the imaging signal (imaging data) to acquire image data. If imaging pixels and focus detection pixels are provided separately, the camera CPU 121 performs interpolation processing on the focus detection pixels to acquire image data.
[0065] In S12, the camera CPU 121 instructs the image processing circuit 125 to generate a live view image from the image data obtained in S11, and displays this on the display unit 131. The live view image is a scaled-down image matched to the resolution of the display unit 131, allowing the user to adjust the imaging composition, exposure conditions, etc., while viewing it. Therefore, the camera CPU 121 adjusts the exposure based on the metering values obtained from the image data and displays it on the display unit 131. Exposure adjustment is achieved by appropriately adjusting the exposure time, opening and closing the aperture of the shooting lens, and adjusting the gain relative to the image sensor output.
[0066] Next, in S13, the camera CPU 121 determines whether the switch Sw1, which instructs the start of the image preparation operation, has been turned on by half-pressing the release switch included in the operation switch group 132. If Sw1 is not turned on, the camera CPU 121 repeats the determination in S13 to monitor the timing when Sw1 will be turned on. On the other hand, if Sw1 is turned on, the camera CPU 121 proceeds to S400 and performs subject-tracking autofocus (AF) processing. Here, from the obtained image signal, it performs subject area detection from the focus detection signal, setting the focus detection area, and predictive AF processing to suppress the effect of the time lag between the focus detection processing and the image recording processing. Details will be described later.
[0067] In S15, the camera CPU 121 determines whether the switch Sw2, which instructs the start of the imaging operation, has been turned on by fully pressing the release switch. If Sw2 is not turned on, the camera CPU 121 returns to S13. On the other hand, if Sw2 is turned on, the process proceeds to S300 and executes the imaging subroutine. Details of the imaging subroutine will be described later. Once the imaging subroutine finishes, this process ends.
[0068] In this embodiment, subject detection processing and AF processing are performed after Sw1 is detected as ON in S3, but the timing of these processes is not limited to this. By performing the subject tracking AF processing in S400 before Sw1 is turned ON, the photographer's preparatory actions before shooting can be eliminated.
[0069] Next, using the flowchart shown in Figure 10, we will explain the imaging subroutine executed by the camera CPU 121 at S300 in Figure 9.
[0070] In step S301, the camera CPU 121 performs exposure control processing and determines the imaging conditions (shutter speed, aperture value, ISO sensitivity, etc.). This exposure control processing can be performed using brightness information obtained from the image data of the live view image.
[0071] The camera CPU 121 then transmits the determined aperture value to the aperture drive circuit 128 to drive the aperture 102. The camera CPU 121 also transmits the determined shutter speed to the shutter 106 to open the focal-plane shutter. Furthermore, the camera CPU 121 causes the image sensor 107 to accumulate charge during the exposure period via the image sensor drive circuit 124.
[0072] In S302, the camera CPU 121, which has performed exposure control processing, instructs the image sensor drive circuit 124 to read out all pixels of the imaging signal from the image sensor 107 due to still image capture. The camera CPU 121 also instructs the image sensor drive circuit 124 to read out one of the pair of focus detection signals from the focus detection area (focus target area) within the image sensor 107. The focus detection signal read out at this time is used to detect the focus state of the image during image playback, which will be described later. The other focus detection signal can be obtained by subtracting one of the pair of focus detection signals from the imaging signal.
[0073] In S303, the camera CPU 121 instructs the image processing circuit 125 to perform defective pixel correction processing on the image data read out in S302 and converted by A / D.
[0074] In S304, the camera CPU 121 instructs the image processing circuit 125 to perform image processing and encoding processes such as demosaicing (color interpolation), white balance processing, gamma correction (tone correction), color conversion, and edge enhancement on the captured data after the defective pixel correction process.
[0075] In S305, the camera CPU 121 records the still image data obtained as image data through the image processing and encoding process in S304, and the focus detection signal read out in S302, as an image data file in the memory 133.
[0076] In S306, the camera CPU 121 records camera characteristic information, which is characteristic information of the camera 100, in memory 133 and in the memory of the camera CPU 121, corresponding to the still image data recorded in S305. The camera characteristic information includes, for example, the following information. • Imaging conditions (aperture value, shutter speed, ISO sensitivity, etc.) • Information regarding image processing performed by image processing circuit 125 • Information regarding the light-receiving sensitivity distribution of the imaging pixels and focus detection pixels of the image sensor 107. • Information regarding vignetting of the imaging light beam within Camera 100 • Information on the distance from the mounting surface of the imaging optical system in camera 100 to the image sensor 107. • Information regarding manufacturing tolerances for Camera 100 Information regarding the light sensitivity distribution of the imaging pixels and focus detection pixels (hereinafter simply referred to as light sensitivity distribution information) is information regarding the sensitivity of the image sensor 107 according to its distance (position) from the optical axis. Since this light sensitivity distribution information depends on the microlenses 305 and the photoelectric conversion units 301 and 302, it may also be information regarding these. Furthermore, the light sensitivity distribution information may also be information regarding the change in sensitivity with respect to the angle of incidence of light.
[0077] In S307, the camera CPU 121 records lens characteristic information as characteristic information of the imaging optical system in memory 133 and the memory within the camera CPU 121, corresponding to the still image data recorded in S305. The lens characteristic information includes, for example, information about the exit pupil, information about the frame such as the lens barrel that emits the light beam, information about the focal length and F number at the time of imaging, information about aberrations of the imaging optical system, information about manufacturing errors of the imaging optical system, and information about the position (subject distance) of the focus lens 105 at the time of imaging.
[0078] In S308, the camera CPU 121 records image-related information, which is information related to still image data, in memory 133 and in memory within the camera CPU 121. Image-related information includes, for example, information related to the focus detection operation before image capture, information related to the movement of the subject, and information related to the focus detection accuracy.
[0079] In S309, the camera CPU 121 displays a preview of the captured image on the display unit 131. This allows the user to easily check the captured image.
[0080] Once processing S309 is complete, the camera CPU 121 terminates this imaging subroutine.
[0081] Next, using the flowchart shown in Figure 11, we will explain the subject tracking AF processing subroutine executed by the camera CPU 121 in S400 in Figure 9.
[0082] In S401, the camera CPU 121 calculates the amount of image shift between pairs of focus detection signals obtained in each of the multiple focus detection regions acquired in S11, calculates the amount of defocus for each focus detection region from the amount of image shift, and obtains a defocus map. As described above, in this embodiment, the group of focus detection results obtained from focus detection regions arranged on the image sensor in a total of 9600 points (120 horizontal divisions and 80 vertical divisions) is called a defocus map.
[0083] In S402, the camera CPU 121 performs subject detection and tracking. Subject detection is performed by the subject detection unit 140 described above. Subject detection may not be possible depending on the state of the obtained image; in such cases, tracking is performed using other means such as template matching to estimate the position of the subject. Details will be described later.
[0084] In S403, the camera CPU 121, acting as a local area selection means, sets the focus detection area using the subject detection area information obtained in S402. The camera CPU 121 acquires information such as the subject's position, size, and reliability as the subject detection area information obtained as the output of the subject detection and tracking process performed in S402. For setting the focus detection area, it is sufficient to select a focus detection result that indicates a subject that is highly reliable and relatively close in distance from the results of the focus detection area within the area set as the subject detection area. Alternatively, for setting the focus detection area, a focus detection area may be placed again within the area set as the obtained subject detection area, image data and focus detection signals may be acquired again, and the focus detection result may be selected in the same way.
[0085] In S404, the camera CPU 121 acquires the focus detection result for the set focus detection area. The focus detection result acquired here may be selected from the focus detection results calculated in S401 to be close to the desired area, or the amount of defocus may be calculated using a newly set focus detection signal corresponding to the set focus detection area. Furthermore, the focus detection area for calculating the amount of defocus is not limited to one; multiple areas may be placed around the camera to calculate the amount of defocus.
[0086] In S405, the camera CPU 121 performs predictive AF processing using the defocus amount obtained in S404 and multiple defocus amounts, which are time-series data of past focus detection timings. This processing is necessary when there is a time lag between the timing of focus detection and the timing of image exposure. AF control is performed by predicting the position of the subject in the optical axis direction at the time of image exposure, which is a predetermined time after the timing of focus detection. The prediction of the subject's image plane position is obtained by performing multivariate analysis (e.g., least squares method) using historical data of past subject image plane positions and times to obtain the equation of the prediction curve. By substituting the time of image exposure into the obtained prediction curve equation, the predicted image plane position wp of the subject can be calculated.
[0087] Furthermore, the position may be predicted not only in the optical axis direction but also in three dimensions. For example, consider a vector in the XYZ direction, where the screen is the XY direction and the optical axis direction is the Z direction. In this case, the position of the subject at the time of exposure of the captured image may be predicted from the time-series data of the XY position of the subject obtained in the subject detection and tracking process of S402 and the position in the Z direction based on the defocus amount obtained in S405. Furthermore, the position of the subject may be predicted from the time-series data of the joint positions of the person who is the subject.
[0088] Based on the above prediction, even if the ball or person is hidden or part of a person's joint position becomes invisible, the system can still estimate the position of each object. Prediction is performed not only on the main subject but also on multiple detected subjects. By performing predictive AF processing on multiple subjects, when the main subject changes, there is no need to re-accumulate the history of the defocus amount of the new main subject, allowing predictive AF to continue without any time loss.
[0089] In S405, the amount of drive for the focus lens is calculated using the predictive AF processing result, and the focus actuator 114 is driven in response to the focus drive command from the camera CPU 121, thereby performing focus adjustment processing by moving the third lens group 105 in the optical axis direction.
[0090] Once processing in S405 is complete, the camera CPU 121 terminates the subject-tracking AF processing subroutine and proceeds to processing in S15 as shown in Figure 9.
[0091] Next, using the flowchart shown in Figure 12, we will explain the subject detection and tracking subroutine executed by the camera CPU 121 in S402 of Figure 11.
[0092] In S421, the camera CPU 121 sets dictionary data according to the type of subject to be detected, based on the data detected from the image data acquired in S12 in Figure 9. Based on the pre-set subject priority and imaging device settings, it selects the dictionary data to be used in this process from multiple dictionary data stored in the dictionary data storage unit. For example, multiple dictionary data may be stored, such as dictionary data classified by subject, such as "person," "vehicle," and "animal." In this embodiment, one or more dictionary data may be selected. If one is selected, it becomes possible to repeatedly detect subjects that can be detected by one dictionary data at a high frequency. On the other hand, if multiple dictionary data are selected, subjects can be detected sequentially by setting the dictionary data sequentially according to the priority of the subject to be detected.
[0093] In step S422, the subject detection unit 140 uses the image data read in step S12 of Figure 9 as the input image and performs subject detection using the dictionary data set in step S421. At this time, the subject detection unit 140 outputs information such as the position, size, and confidence level of the detected subject. At this time, the camera CPU 121 may display the above information output by the subject detection unit 140 on the display unit 131.
[0094] In S422, multiple regions of a subject are detected hierarchically from image data. For example, if "person" or "animal" is set as the dictionary data, multiple organs such as the "whole body," "face," and "eyes" will be detected. Local regions such as the eyes and face of a person are areas where the focus and exposure should be adjusted, but they may not be detectable due to surrounding obstacles or the direction of the face. Even in such cases, the subject is detected robustly by performing whole-body detection, and the subject is detected hierarchically. Similarly, if "vehicle" such as a motorcycle is set as the dictionary data, the driver, the entire vehicle including the body, and the helmet (head) as a local region will be detected hierarchically.
[0095] In S423, the camera CPU 121 performs known template matching processing using the subject detection area obtained in S422 as a template. Using multiple images obtained in S12, it searches for similar areas in the most recently obtained image, using the subject detection area obtained in past images as a template. As is well known, any of the following information can be used for template matching: brightness information, color histogram information, feature point information such as corners and edges, etc. Various matching methods and template update methods can be considered, and any of them can be used. The tracking processing performed in S423 is done to achieve stable subject detection and tracking processing by detecting areas similar to past subject detection data from the most recently obtained image data when no subject was detected in S422.
[0096] Once the processing in S423 is complete, the camera CPU 121 terminates the subject detection and tracking subroutine and proceeds to S403 in Figure 11.
[0097] (Virtual space image processing) The flowchart in Figure 13 illustrates the operation of the virtual space imaging process shown in S1000 of Figure 8. Virtual space imaging is the process of extracting information from a virtual space that changes over time and generating an image for a specific moment. Although imaging in a virtual space does not require a physical imaging optical system or image sensor, the same terminology as for imaging in the real space is used for ease of explanation. For example, image generation in a virtual space is expressed as imaging or capture. More specifically, Figure 13 shows the process from the pre-image operation of displaying the image in the virtual space as a live view image on the display unit 131 of the camera 100 to the capture of a still image. The camera CPU 121 and CPU 1001, which are computers, execute this process according to the computer program. In the following, unless otherwise specified, the main operators are the camera CPU 121 and CPU 1001.
[0098] In S1001, settings related to virtual space shooting, such as the virtual shooting space where the shooting will take place and the equipment to be used, are configured. In setting up the virtual shooting space, as explained in Figure 7, foreground objects are placed in appropriate positions relative to background objects. Information regarding the position and shape of foreground objects that change over time is also acquired. There may be one or more foreground objects to be placed.
[0099] Furthermore, in the virtual space shooting settings, the camera CPU 121 outputs the model information of the equipment being operated to the external processing unit 1000 as settings for the camera and lens to be used in the virtual space shooting. If a different model of equipment is used for the virtual space shooting than the one actually being operated, the photographer sets a unique symbol for that model, and the camera CPU 121 outputs the set information to the external processing unit 1000. This allows the operator (photographer) to experience shooting with cameras and lenses they do not actually own. For example, while operating a camera equipped with a short focal length, a so-called wide-angle lens, the virtual space can be used to experience shooting with a long focal length telephoto lens. Operating equipment different from the actual equipment is not limited to lenses; it can also be a camera, or both. This enables a more flexible shooting experience that is not limited by the weight or size of the equipment. Similarly, by using a camera that is not actually owned, it is possible to experience the new functions of that camera, the improved performance due to newly installed algorithms, etc.
[0100] Furthermore, the virtual space shooting settings allow you to set the initial values for the viewpoint position and direction (the camera's position and direction in the virtual space) when virtual space shooting begins. You should set the initial value to a position at an appropriate distance from the foreground object type and other information mentioned above. Alternatively, you can set a pre-configured shooting position within the background object as the initial value.
[0101] In S1002, the CPU 1001 issues an initialization command to the camera 100 for the camera drive unit. Details will be described later.
[0102] In S1003, the CPU 1001 acquires camera information, lens information, and camera and lens operation information from the camera 100 and the camera / lens information storage device 2000.
[0103] (Explanation of communication between the camera / lens and the external processing unit) An example of information communicated between the camera / lens and the external computing unit 1000 will be explained using the table in Figure 14.
[0104] This section describes the information stored in the camera / lens information storage device 2000, the camera / lens information, and the external processing unit 1000. The camera / lens information storage device 2000 stores camera information and lens information. The camera / lens information storage device 2000 also acquires and stores information from the camera / lens.
[0105] Camera information includes the display resolution, recorded video resolution, image sensor size, metering frame mode, autofocus (AF) modes such as one-shot and servo, and continuous shooting settings, all previously acquired from camera 100. It also includes camera settings such as shooting difficulty settings set by the photographer, AF algorithms, camera algorithm information such as automatic exposure (AE) and continuous shooting drive sequences, and camera detection information such as temperature. Furthermore, it includes image sensor characteristic information such as S / N information for each ISO sensitivity, and shading correction values that represent signal characteristic correction of the image sensor and unevenness of light intensity. It also includes a defocus conversion coefficient that converts the amount of image shift into the amount of defocus, focus-related correction information, information on best focus position correction that corrects the difference between the focus detection result and the best image plane position, and focus-related correction information which is defocus error information. In addition, it includes general information such as the camera / lens model name and firmware versions of various algorithms.
[0106] Lens information includes the range, current value, and resolution of the focal length; the range, increment, and current value of the f-number; the drive range and current focus information of the focusing lens; and focus control information regarding the control characteristics of the focus drive. It also includes sensitivity for converting the focus lens drive into image plane movement, image stabilization information regarding the range, current value, and correction resolution of the image stabilization, and image stabilization control information regarding the control characteristics of the image stabilization. Furthermore, it includes aperture control information regarding the control characteristics of the aperture drive, frame information (position, diameter) regarding vignetting, peripheral light falloff information, distance information regarding the position and distance of the focusing lens, and information regarding the point image distribution function.
[0107] Cameras and lenses contain operational information generated by the photographer's actions. This operational information includes framing, zooming, focusing, shutter release, and other button operations. This operational information is transmitted to the external processing unit 1000 and reflected in the virtual image generation.
[0108] The external processing unit 1000 acquires camera information, lens information, and operation information, and generates display images, recorded images, subject information which is shooting difficulty information, virtual defocus amount, and various shooting-related information.
[0109] The lens information acquired includes focal length, f-number, information on the configurable range and current position of the focus lens, the mechanical controllability of the lens, and the amount of image plane movement (sensitivity) associated with the movement of the focus lens. Furthermore, it includes information on vignetting (position, diameter), peripheral light falloff, and shooting distance (distance to the subject in focus).
[0110] The camera information acquired includes general information such as the model name, firmware version, resolution of EVF and still images, and image sensor size. Furthermore, it includes camera settings such as the AF frame setting (which determines the AF range), AF mode settings (such as One Shot and Servo AF), and continuous shooting mode settings (such as continuous shooting speed). The camera settings also include difficulty settings (shooting difficulty settings) set by the photographer.
[0111] Furthermore, the signal correction values used for autofocus focus detection include signal characteristic correction values that depend on the characteristics of the image sensor 107, shading correction values that represent unevenness in light intensity, and a defocus conversion coefficient that converts the phase difference of a pair of signals into a defocus amount. It also includes a best focus correction value that corrects the discrepancy between the focus detection result and the best image plane position.
[0112] Furthermore, camera information includes characteristic information of the image sensor 107, such as signal-to-noise ratio information for each ISO sensitivity, various algorithm information such as continuous shooting sequences and metering when taking pictures with the camera, and autofocus-related algorithm information such as AF frame selection and predictive AF. Some of this information changes with camera operation, so information that may change is acquired periodically from S1003 onwards.
[0113] Furthermore, camera and lens operation information includes information regarding the amount and speed of panning, zooming, and focusing operations performed by the photographer, as well as information regarding button presses such as shutter release operations, which are used to instruct shooting.
[0114] Returning to the explanation of Figure 13, S2000 generates video of the virtual space based on the settings made so far and outputs it to camera 100. The details of this process will be described later.
[0115] In S1005, the camera CPU 121 acquires the image output in S2000 and displays it on the display unit 131. The displayed image is updated thereafter, for example, at 60fps. By combining this with camera operation information and lens operation information, the display unit 131 updates with images that show different ranges of the virtual space according to the camera's panning and zooming operations.
[0116] In S1006, it is determined whether the mode is one in which the viewpoint in the virtual space being observed through the display unit 131 is moved (viewpoint movement mode). If the mode is one in which viewpoint movement processing is performed, S1006 is answered with Yes, and the process proceeds to S3000. In S3000, viewpoint movement processing is performed to determine the camera's position and shooting direction in the virtual space. Details will be described later. After S3000 is completed, the process returns to S2000.
[0117] If the answer in S1006 is No, the process proceeds to S1007. In S1007, similar to S13 in Figure 9, the camera CPU 121 determines whether the switch Sw1, which instructs the start of the image preparation operation, has been turned on by a half-press operation of the release switch included in the operation switch group 132. If Sw1 is not turned on, the camera CPU 121 returns to S2000 and repeats the determination in order to monitor the timing when Sw1 will be turned on. On the other hand, if Sw1 is turned on, the camera CPU 121 proceeds to S4000 and performs virtual subject tracking processing.
[0118] The S4000 applies various corrections to the generated image in response to the photographer's actions and the movement of the subject, enabling the shooting of the subject, which is at least part of the foreground object. Details of this process will be described later.
[0119] In S1008, similar to S15 in Figure 9, the camera CPU 121 determines whether the switch Sw2, which instructs the start of the imaging operation, has been turned on by fully pressing the release switch. If Sw2 is not turned on, the camera CPU 121 returns to S2000. On the other hand, if Sw2 is turned on, the process proceeds to S5000 and the virtual space shooting subroutine is executed. Details of the virtual space shooting subroutine will be described later. Once the virtual space shooting subroutine finishes, this process ends.
[0120] (Subroutine for generating and outputting images in a virtual space) Next, using the flowchart shown in Figure 15, we will explain the subroutines for generating and outputting virtual space images that the external computing device 1000 executes in S2000 in Figure 13.
[0121] In S2001, the CPU 1001 acquires the foreground object. The photographer first uses the virtual space reproduction device 1100 to select the type of subject they want to photograph (e.g., person, animal, vehicle, etc.). Next, they select the shape and color of the subject, and also select what kind of movement (speed, direction of movement, etc.) the subject should have. The user interface for selection may be configured to display the information stored in the foreground object storage unit 1101 of the virtual space reproduction device 1100 on the camera's display 131, allowing the photographer to operate and select. As described above, the foreground object, which is a 3D model of the subject, is acquired by multiple methods.
[0122] In S2002, CPU1001 acquires background objects. As described above, background objects, which are 3D models other than the subject, are acquired using multiple methods.
[0123] In S2003, CPU1001 performs object compositing. Object compositing is the process of combining the foreground object and background object as described above. Object compositing is performed by deciding how to position the background object in 3D space and where to position the foreground object in 3D space relative to the background object. First, the background object is placed in 3D space, and the photographer selects where to place the foreground object in 3D space. The foreground object is positioned so that it can only be placed in positions that are possible relative to the background object (for example, outside the interior of the background object) based on the 3D model of the background object and the coordinates placed in 3D space. The photographer selects the position to place the foreground object from the 3D space of the background object. Object compositing is performed in this way.
[0124] In S2004, CPU 1001 acquires camera / lens information. This camera / lens information is for generating and outputting images in the virtual space, which will be described later. Specifically, camera information includes the display resolution of the display unit 131, the size of the camera's image sensor, and the number of pixels. Lens information includes the range and current value of the focal length, the range and current value of the aperture, the focus lens range and current value, and information regarding peripheral light falloff and point image distribution function.
[0125] In S2005, CPU1001 acquires viewpoint position information. It acquires virtual viewpoint information in 3D space in order to generate the virtual image described later. The virtual viewpoint information may be a predetermined value as an initial value, or it may be a virtual viewpoint that has been changed by the viewpoint movement process in S3000 described later.
[0126] In the S2006, CPU1001 acquires camera / lens operation information. This operation information includes framing, zooming, focusing, shutter release, and other button operations.
[0127] In S2007, CPU1001 acquires the image correction amount. The image correction amount refers to the correction amount related to framing, zooming, and focus. Details will be explained later in the subflow of the virtual subject tracking process in S4000. Here, the initial values of the predetermined image correction amount are acquired.
[0128] In S2008, CPU1001 generates display images in a virtual space. It renders an image based on the aforementioned foreground and background objects placed in 3D space and viewpoint position information. The range to be included in the display image is determined by the aforementioned lens information (focal length), camera information (image sensor size, display resolution, camera settings), and framing and zooming information based on operation information. Furthermore, the range to be included in the display image is determined by modifying the range using image correction amount. In addition, display images with changed aperture value and defocus amount are generated from lens information (aperture value, peripheral light falloff, point image distribution function, focus lens position information, and image correction amount information related to focusing). The display image is different from the recorded image described later; it is not recorded. Therefore, after correctly determining the display image range, the display image may be generated simply with less information than that used for generating the recorded image, by omitting some information such as other information such as focus lens position information and peripheral light falloff information.
[0129] In S2009, CPU 1001 outputs the display image generated in S2008. The output image is transmitted from the external processing unit 1000 to the camera 100 and displayed on the display unit 131.
[0130] In the S2010, CPU 1001 is responsible for saving video-related information. Video-related information includes subject information, shooting-related information, virtual defocus amount, and AF log information. Video-related information is temporarily stored in RAM 1003 of the external processing unit 1000 and recorded as video-related information in the virtual space shooting subroutine described later. Details will be described later.
[0131] With the above steps, the video generation and output processing of the virtual space S2000 in Figure 13 is completed. In this embodiment, the video generation and output of the virtual space was performed by the external computing device 1000, but the video generation and output processing of the virtual space may also be performed within the camera 100.
[0132] (Subroutine for virtual subject tracking) The virtual subject tracking process of the S4000 shown in Figure 13 will be explained using the flowchart in Figure 16.
[0133] Among the various corrections described later, the framing correction is performed when the photographer's framing is off and the subject they wanted to photograph is outside the frame or cut off, so that the subject is properly contained within the camera's displayed field of view.
[0134] Zooming correction addresses situations where the photographer's zoom (lens focal length) is inaccurate, causing the subject to extend beyond the frame or appear too small. The correction ensures the subject is displayed at an appropriate size within the frame. In addition to single-timing zooming correction, the system also corrects zooming for situations where, for example, an approaching subject is being captured while maintaining a consistent size within the frame (zooming). Furthermore, to achieve smooth focal length changes, the system considers the timing of preceding and succeeding shots, preventing jerky frames in consecutive images caused by the photographer's zooming.
[0135] Focusing corrections, for example, in autofocus mode, correct for blurred focus caused by fast-moving subjects or large changes in speed, based on the camera's tracking algorithm (tracking limit performance). This allows for images that are in focus or have reduced blur. In manual focus mode, it corrects for focus blur caused by the photographer's focusing operation. It also corrects for phenomena such as the focus lens being driven towards the background due to the photographer's framing being off, resulting in the subject being out of focus.
[0136] First, in S4001, the CPU 1001 acquires camera / lens information from the camera lens information acquisition unit 1202. The camera lens information acquired here is used for determining whether correction is ON / OFF, acquiring subject difficulty information, and calculating the correction amount, which will be described later. Specifically, the camera information includes camera settings related to correction, such as the shooting difficulty setting set by the photographer. The lens information includes information such as focal length, focus lens position, and ON / OFF setting of the image stabilization switch.
[0137] In S4002, CPU1001 uses the camera lens information acquired by S4001 to obtain setting information related to correction. This setting information related to correction includes, for example, the ON / OFF setting of correction settings within the camera, mode settings such as difficulty settings, and setting information such as the ON / OFF setting of the image stabilization switch on the lens.
[0138] In the S4003, the CPU1001 detects framing and obtains information such as whether the camera is panning, in which direction and at what speed.
[0139] In the S4004, the CPU1001 detects zooming and obtains information such as whether the zoom lens is being operated, and in which direction (Tele / Wide) and at what speed.
[0140] In the S4005, the CPU1001 detects focusing and acquires information such as whether the focus ring is being operated, and at what speed in the direction of near focus or infinity focus.
[0141] Furthermore, detection in S4003-S4005 includes not only manual operation by the photographer but also automatic operation performed by the camera (such as auto framing, auto zoom, and autofocus).
[0142] In the S4006, the CPU 1001 sets the subject area. Here, it determines which foreground object in the image generated by the object synthesis unit 1103 will be the main subject, and simultaneously sets the area to be AF'd. Once the main subject is determined, information regarding the subject's velocity and acceleration, angular velocity and angular acceleration, size of the subject, contrast value of the subject, and distance between the subject and the photographer can be obtained from the foreground object storage unit 1101.
[0143] There are various methods for setting the subject area, and in this embodiment, it is possible to set it in three-dimensional space. On the other hand, in the case of a camera used for real-world shooting, the subject area is set by the photographer's framing and the detection results of the subject detection unit in the (two-dimensional) image space. In the virtual space shooting of this embodiment, the subject area can be set by combining information such as the information contained in foreground objects, objects that are closer, and objects that are closer to the center of the shooting range, as described above. Alternatively, the main subject may be detected from the obtained image, similar to when shooting in real space. This makes it possible to reproduce the performance of the camera used for real-world shooting more accurately.
[0144] In S4007, the CPU 1001, in the image correction amount calculation unit 1206, determines whether to turn correction ON or OFF based on various information acquired in S4002 to S4006. For example, in S4002, if the information regarding correction within the camera is ON, correction is turned ON; otherwise, correction is turned OFF. In addition, the photographer's intent extraction unit 1262 extracts and determines the photographer's intent from the framing information detected in S4003, such as which subject is being targeted, whether the subject is being tracked in the framing, or whether the photographer is trying to switch the framing to a different subject. In the former case, turning correction ON allows the photographer's framing mistakes to be covered by the correction, and in the latter case, turning correction OFF allows the photographer to frame (fit into the field of view) a different subject as intended.
[0145] Furthermore, based on the zooming and focusing information detected by S4004 and S4005, the system determines that the photographer's intent is strong during manual operations such as manual zoom and manual focus, and turns off the correction. Conversely, when using automatic functions such as auto zoom and autofocus, it determines that the photographer's intent is weak, and turns on the correction.
[0146] As described above, by turning off correction when it is determined that the photographer's intention is strong, it is possible to provide shooting results and a shooting experience that is close to the photographer's intuitive feel. Also, when the intention is weak, even if correction is turned on, good images can be obtained through correction without impairing the photographer's shooting experience. Furthermore, the camera itself has a defined correction capability value, and there is a method of determining whether to turn on correction only if the camera's correction capability value exceeds the shooting difficulty, as described later.
[0147] In the S4100, the CPU 1001 acquires shooting difficulty information using the shooting difficulty calculation unit 1261. Using this shooting difficulty information, various correction amounts described later are calculated. For example, by setting a smaller correction amount for higher shooting difficulty, it becomes more difficult to keep the subject in the frame or maintain focus for more difficult subjects. On the other hand, by setting a larger correction amount for lower difficulty subjects, even if a major failure occurs, good shooting results can be obtained through correction, thus increasing the success rate of shooting lower difficulty subjects.
[0148] With subjects that are not particularly difficult, the photographer is often not in a situation where they can concentrate on or become engrossed in the shooting experience itself, so even if the amount of correction is increased, the number of failed photos can be reduced without compromising the shooting experience. On the other hand, with subjects that are difficult, increasing the amount of correction may detract from the richness of the shooting experience, so in this embodiment, the correction is kept small. This is also true for camera shooting in real space, so adjusting the amount of correction according to the difficulty of the subject leads to providing a more realistic shooting experience.
[0149] The acquisition of shooting difficulty information will be explained using Figure 17.
[0150] In S4101, the CPU 1001 obtains the subject's velocity and acceleration information from the foreground object storage unit 1101, and in S4102, it obtains the subject's angular velocity and angular acceleration information from the foreground object storage unit 1101. This information may be time-specific information or fixed, defined information such as maximum velocity and maximum acceleration. The larger these values are, the higher the shooting difficulty calculated in S4106.
[0151] In S4103, the CPU 1001 obtains the size information of the subject from the foreground object storage unit 1101.
[0152] In S4104, the CPU 1001 obtains the contrast value of the subject from the foreground object storage unit 1101. The lower the contrast value, the more difficult the shooting becomes.
[0153] In S4105, the CPU 1001 obtains the distance between the subject and the photographer from information from the foreground object storage unit 1101 and the viewpoint information acquisition unit 1201. By combining this with the zooming (focal length) information obtained in S4001 and S4004, and the subject size information obtained in S4003, the subject size on the image sensor is determined. The smaller this value, the more difficult the shooting becomes. Differences in the part of the subject (e.g., eyes or face) also affect the difficulty of shooting.
[0154] In S4106, the CPU 1001 calculates the difficulty of photographing the subject from the information acquired in S4101 to S4105. The difficulty of photographing information defined here may be defined as a single piece of information encompassing all elements. Alternatively, it may be defined as multiple types of information corresponding to the framing correction, zooming correction, and focusing correction described later (e.g., framing difficulty, zooming difficulty, focusing difficulty). Furthermore, the difficulty of photographing information may be calculated from these various pieces of information, or the difficulty itself may be stored in the foreground object storage unit 1101. In addition, although this embodiment describes the difficulty of photographing as being calculated each time the subject's speed or distance changes (the difficulty of photographing changes), it may also be defined as always being fixed.
[0155] Returning to Figure 16, in S4009, the CPU 1001 calculates the virtual defocus amount in the subject area set in S4006. The calculated virtual defocus amount may be a defocus map calculated from multiple areas, as explained using Figure 5, or it may be a single output for a single part of the subject, such as the face. In the former case, there is a process to select one area from multiple areas, but a detailed explanation is omitted in this embodiment.
[0156] In the S4200, CPU1001 performs processing on the virtual defocus amount. Details will be described later using Figure 21.
[0157] In S4011, the CPU 1001 calculates the focus drive amount. The focus drive amount may be the value obtained by converting the virtual defocus amount calculated in S4009 into a focus lens drive amount. Alternatively, the future subject position may be predicted from the subject position in multiple past frames, and the focus drive amount may be set for that predicted position. Various methods for prediction are possible, but since this is not the main focus of this embodiment, a detailed explanation will be omitted.
[0158] In virtual space photography, since there is no need to drive a physical focus lens, it is possible to switch to the desired focus state instantaneously without taking time to drive the focus. However, in this embodiment, the aim is to provide the photographer with an experience similar to real-world photography by using a camera capable of real-world photography to photograph the virtual space. Therefore, the process of changing the focus state during shooting (for example, focusing on a subject that is out of focus) is performed over a predetermined period of time. The time spent on focus drive may be set to match the functions / performance of the actual camera and lens being used, using camera / lens information, or it may be set assuming a virtual camera and lens.
[0159] In S4012~S4014, the CPU 1001 calculates various correction amounts in the image correction amount calculation unit 1206.
[0160] In S4012, CPU1001 calculates the correction amount related to framing. The correction related to framing will be explained using Figure 18.
[0161] Assuming there are two subjects, A and B, and subject A is defined as being more difficult to photograph, the maximum framing correction amount will be smaller for subject A according to the difficulty of photography (maximum framing correction amount A < maximum framing correction amount B). Here, the dashed rectangular area in Figure 18 represents the area actually framed by the photographer, and the solid rectangular area represents the framing area after applying framing correction. In this case, in Figure 18(a), the photographer's framing is off relative to subject A, but by applying a correction that fits within the maximum framing correction amount A, the corrected framing area is able to include subject A within the frame. On the other hand, in Figure 18(b), the photographer's framing is even further off relative to subject A, and even with the application of the maximum framing correction amount A, subject A could not be included within the frame.
[0162] Next, taking subject B as an example, in Figure 18(c), the photographer's framing deviation is small, so, similar to Figure 18(a), the corrected framing area is able to fit subject B within the frame. In Figure 18(d), the amount of deviation is larger, and in the case of subject A, the maximum framing correction amount A was smaller than that deviation, so the correction was insufficient. Even in such a case, in Figure 18(d), the amount of deviation falls within the maximum framing correction amount B, so the corrected framing area is able to fit subject B within the frame. In this way, by changing the amount of framing correction according to the difficulty of shooting, it is possible to provide a realistic shooting experience where framing is more difficult for subjects that are more difficult to shoot.
[0163] Returning to the explanation of Figure 16, in S4013, CPU 1001 calculates the correction amount related to zooming. Figure 19 will be used to explain the correction related to zooming.
[0164] If we define that there are two subjects, C and D, and that subject C is more difficult to photograph, then the maximum zoom correction amount will be smaller for subject C according to the difficulty of photography (maximum zoom correction amount C < maximum zoom correction amount D). Here, the dashed rectangular area in Figure 19 represents the field of view that the photographer is actually adjusting by zooming, and the solid rectangular area represents the field of view after applying zoom correction. In this case, in Figure 19(a), the photographer's zoom is off relative to subject C, but by applying a correction that falls within the maximum zoom correction amount C, the corrected zoom field of view is able to keep subject C within the frame. On the other hand, in Figure 19(b), the photographer's zoom is even further off relative to subject C, and even if the maximum zoom correction amount C is applied, subject C cannot be kept within the frame.
[0165] Next, taking subject D as an example, in Figure 19(c), the photographer's zooming error is small, so, similar to Figure 19(a), the corrected zoom angle of view is able to keep subject D within the frame. In Figure 19(d), the amount of error is larger, and in the case of subject C, the maximum zooming correction amount C was smaller than the amount of error, so the correction was insufficient. Even in such a case, the amount of error falls within the maximum zooming correction amount D, so the corrected zoom angle of view is able to keep subject D within the frame. In this way, by changing the amount of zooming correction according to the difficulty of shooting, it is possible to provide a realistic shooting experience where it is more difficult to keep subjects that are difficult to shoot within the frame while zooming.
[0166] Returning to the explanation of Figure 16, in S4014, CPU 1001 calculates the correction amount related to focusing. Figure 20 will be used to explain the correction related to focusing.
[0167] If we define that there are two subjects, E and F, and that subject E is more difficult to photograph, then the maximum focusing correction amount will be smaller for subject E according to the difficulty of photography (maximum focusing correction amount E < maximum focusing correction amount F). Here, Figure 20(a) shows subject E approaching the photographer from a distance, and Figure 20(b) shows subject F approaching the photographer from a distance as time progresses. The solid lines are the trajectories of the positions of each subject, the dotted lines are the trajectories of moving the focus to focus on the subject (trajectories of the actual focus position), and the dashed lines are the trajectories after correction to the focus.
[0168] The solid line representing the subject's position coincides with the dashed or dotted line, indicating that an image in focus on the subject can be captured. The reason for the discrepancy between the subject's position and the actual focus position is largely due to the photographer's focusing operation in manual focus mode, but largely due to the camera's tracking algorithm (tracking limit performance) in autofocus mode. In this case, one example of the cause is that the subject is moving fast or its speed is changing rapidly. Therefore, in Figure 20, as time passes, the autofocus reaches its tracking limit, and the subject's position and the actual focus position become misaligned.
[0169] In Figure 20(a), the actual focus position is shifted relative to the subject E, but the corrected focus position matches the subject position up to the range within the maximum focusing correction amount E. However, as time passes and the subject gets closer, the amount of shift in the actual focus position increases, and eventually the maximum focusing correction amount E is insufficient, making it impossible to focus.
[0170] On the other hand, in Figure 20(b), even if the actual focus position shifts significantly relative to the subject F, it remains within the range of the maximum focusing correction amount F, making it possible to capture an image with the subject F in focus until the very end. Furthermore, as mentioned above, focusing corrections also include blurring caused by manual focus operation and blurring caused by framing errors resulting in the background being in focus. The correction for these can be the same amount as the correction value for blurring due to the tracking limit performance, or a different amount. Also, considering that these blurrings occur simultaneously, a combination of correction amounts for each cause may be applied. In this way, by changing the focusing correction amount according to the difficulty of shooting, it is possible to provide a realistic shooting experience where it is more difficult to focus on subjects that are more difficult to photograph.
[0171] In calculating the various correction amounts in S4012 to S4014 above, as long as the detection of Sw1 in the virtual space shooting process shown in Figure 13 continues, the calculations may be performed using past recorded image information and past video correction amounts stored in the memory unit 1004. Doing so will ensure continuity in the correction results between images and reduce the sense of incongruity as a series of recorded images.
[0172] Returning to the explanation of Figure 16, in S4015, CPU 1001 performs focus driving. Here, the focus driving amount calculated in S4011 and the focus correction amount calculated in S4014 are reflected in the driving.
[0173] As described above, by changing the amount of correction effect applied to the recorded image based on the photographer's operation information, subject information, and camera information, it is possible to take photos in a virtual space without compromising the shooting experience itself.
[0174] In this embodiment, we have described an example where the amount of image correction decreases as the difficulty of shooting increases, but it is also possible to do the opposite and increase the amount of image correction as the difficulty of shooting increases. By doing so, it is possible to take successful photos (such as photos in which the subject is within the frame or photos in which the subject is in focus) at a certain level regardless of the difficulty of shooting. Furthermore, the camera settings may be configured to allow switching between decreasing or increasing the amount of image correction as the difficulty of shooting increases.
[0175] As described above, the processes performed in the S4000's virtual subject tracking, such as defocus amount calculation and focus drive, can also be operated using the AF algorithm stored in the Camera 100's ROM. Furthermore, it is also possible to operate using the AF algorithm of other cameras.
[0176] (Subroutine for processing virtual defocus amount) Next, using the flowchart shown in Figure 21, the subroutine for processing the virtual defocus amount in S4200 of Figure 16 will be explained. This subroutine processes an error amount to the virtual defocus amount calculated in S4009, according to the settings at the time of shooting and the characteristics of the image sensor. In reality, when shooting in a virtual space, the defocus amount is calculated from the distance of a known subject, so no calculation error occurs beyond the error caused by the number of significant digits of each value. On the other hand, when shooting in a real space, errors occur steadily due to the characteristics of the image sensor, and errors occur differently each time an image is taken. In order to reproduce the same focus state behavior in virtual space shooting as in real space shooting, it is necessary to generate errors in virtual space shooting that occur in real space shooting. In this embodiment, for the reasons stated above, a defocus processing process is performed to add an error to the virtual defocus amount calculated in S4009, which does not contain errors, in order to simulate shooting in a real space.
[0177] In S4201, the CPU 1001 acquires camera information for the virtual image generated by the virtual image generation device 1200 from the camera / lens information storage device 2000, including the resolution of the recorded video, image sensor size, AF frame mode, AF algorithm, and S / N information for each ISO sensitivity. It also acquires information related to defocus conversion coefficients, focus position correction, and focus-related correction information, which is defocus error information. Furthermore, it acquires lens information such as the focal length and resolution of the lens, F-number, focus lens information, focus drive control information, and sensitivity that converts the focus lens drive into image plane movement amount. In addition, it acquires image stabilization control information, aperture control information, lens frame information, peripheral illumination falloff information, and distance information related to the focus lens position and distance. These are then stored in the external processing unit 1000 as correspondence information with the virtual image.
[0178] In the S4202, the CPU 1001 retrieves the defocus error information stored in the RAM 1003. The defocus error information will be described later.
[0179] In S4203, CPU 1001 adds the defocus error amount obtained in S4202 to the virtual defocus amount calculated in S4009 in Figure 16. If the virtual defocus amount has been calculated for multiple focus detection regions, the defocus error amount is added to all of the focus detection regions.
[0180] The defocus error information acquired in S4202 will be explained using Figure 22. Figure 22 is a graph showing the relationship between the contrast value of the subject acquired in S4006 in Figure 16 and the amount of error expected to occur in the virtual defocus amount. In general, when photographing in real space, if there are few patterns on the subject and the contrast is low, the amount of error included in the detected defocus amount will be large.
[0181] In Figure 22, the horizontal axis shows the contrast of the subject, and the vertical axis shows the amount of error. The straight line 24101 indicates that the greater the contrast of the subject (to the right on the horizontal axis), the smaller the amount of error. This relationship between contrast and error changes depending on the signal-to-noise ratio (SNR) of the image sensor's pixel and readout circuits, the number of pixels used for the focus detection signal, and the gain applied to the signal, which is set by the ISO sensitivity. Therefore, in this embodiment, the relationship shown in Figure 22 is stored for each mode in which the SNR of the image sensor changes, the specifications of the focus detection signal, and the ISO sensitivity. The data can be stored as discrete values in a table, or as a function represented by a graph, and stored as the coefficients of the function. Furthermore, the above-mentioned amount of error changes because the defocus conversion coefficient, which is camera information, changes depending on the F-number and lens frame information, which are part of the lens information. In this embodiment, the amount of error calculated using the above-mentioned lens information and camera information is multiplied by a predetermined coefficient. The predetermined coefficient only needs to be stored as a table of ratios to a reference value. For example, a table is stored that uses the f-number and lens frame information as indices, containing the defocus conversion coefficient and a coefficient to multiply by the aforementioned error amount. The error amount is then calculated according to the shooting conditions.
[0182] Figures 23(a) and 23(b) show examples of virtual defocus with and without the virtual defocus error generated in S4203 of Figure 21, superimposed as maps on the main subject.
[0183] Figure 23(a) shows the case without applying virtual defocus error. In the virtual defocus map 25102, the focus position of the head of person 25106 within the AF frame 25101 of the virtual spatial image is shown to be in focus area 25103 (hatched in a horizontal and vertical grid pattern), with the entire area of the AF frame 25101 being the focus area 25103.
[0184] Figure 23(b) shows the case where a virtual defocus error is applied. The area of the head of the person 25106 within the AF frame 25101 is shown as the in-focus area 25103, the front focus position 25104 (hatched in a diagonal grid pattern), and the rear focus position 25105 (hatched with black dots).
[0185] While all AF frames in Figure 23(a) represent the in-focus area, Figure 23(b) shows that, due to the addition of errors, AF frames for front focus and rear focus are generated in addition to the in-focus area.
[0186] In this way, by applying a defocus error corresponding to changes in camera / lens information during virtual image capture and applying the algorithm used when selecting an AF frame, it is possible to obtain defocus detection results similar to those obtained during real-world shooting. In this embodiment, for the sake of clarity, an example of applying a defocus error on a defocus map is shown, but a defocus error may also be applied to a single AF frame. Since this error affects predictive AF, a similar effect can be expected. As a result, even when shooting in a virtual space, the same focus adjustment behavior as when shooting in a real space can be reproduced, allowing users to evaluate product performance and check new features before purchasing cameras and lenses.
[0187] In this embodiment, defocus variation was added to the virtual space image to approximate the results of real-world shooting, but adding error is not always necessary. If it is not necessary to approximate the results of real-world shooting, it may be possible to omit the addition of error, or to switch between the two depending on the situation.
[0188] (Subroutine for virtual space photography) Next, using the flowchart shown in Figure 24, we will explain the virtual space imaging subroutine executed by the external computing device 1000 at S5000 in Figure 13.
[0189] In S5001, CPU1001 outputs the set F-number and the time Sw2 was detected. The use of this information will be explained later in Figure 28, which describes the actual camera operation in conjunction with the virtual space shooting operation.
[0190] In the S5002, the CPU 1001 acquires camera / lens information. Specifically, camera information includes the resolution for recording, the size and number of pixels of the camera's image sensor, and lens information includes the range and current value of the focal length, the range and current value of the F-number, the range and current value of the focus lens position, and information regarding peripheral illumination falloff and point image distribution function.
[0191] In S5003, CPU1001 performs the acquisition of the aforementioned image correction values.
[0192] In S5004, CPU1001 generates recorded video of the virtual space. It renders an image from the aforementioned foreground and background objects placed in 3D space and viewpoint position information. The range to be included in the display image is determined from the focal length in the aforementioned lens information, and the image sensor size, resolution, and camera settings in the camera information. Furthermore, it generates recorded video from the F-number information, peripheral light falloff information, point image distribution function information, and focus lens position information in the lens information. Unlike the aforementioned display image, the recorded video is a recorded image, so the recording video range is correctly determined, and it is generated using various optical information such as focus lens position information, peripheral light falloff information, and point image distribution function information. Unlike the display image, the recorded video does not need to be displayed to the photographer in real time, so the generation of the recorded video can be delayed compared to the generation of the display image. Therefore, the recorded video can be generated using more detailed camera and lens information data than that used to generate the display image.
[0193] In S5005, the CPU 1001 records the recorded video of the virtual space to the recording unit 1004 of the external processing unit 1000. Alternatively, the recorded video generated in S5004 may be transferred to the camera 100 and recorded to the camera 100's flash memory 133.
[0194] In S5006, the CPU 1001 records various video-related information. Video-related information includes subject information (shooting difficulty information), shooting-related information, and virtual defocus amount. The video-related information is recorded in the recording unit 1004 of the external processing unit 1000. Alternatively, the video-related information may be transferred to the camera 100 and recorded in the camera 100's flash memory 133. With this, the virtual space shooting subroutine is terminated.
[0195] (Virtual space photography reflecting operation information) Using Figure 25, we will explain virtual space imaging that reflects the operation information. Figure 25(a) shows an example of zooming, and Figure 25(b) shows an example of framing.
[0196] In Figure 25(a), a static virtual space display image 18003 generated by the virtual image generation device 1200 is displayed on the camera 100's display unit 131. When the photographer rotates the lens's zoom ring to change the focal length to the telephoto side, the virtual image generation device 1200 acquires the change in focal length due to the zoom ring operation as operation information. Then, by changing the display range when generating the virtual space display image, it generates a virtual space display image 18004 that reflects the change in focal length due to the photographer's zooming operation, and displays it on the display unit 131.
[0197] In this embodiment, since the virtual space display image is generated using camera / lens information, it is possible to generate the virtual space display image within a focal length range that cannot be operated by the lens being operated by the photographer. For example, even if the photographer is operating a lens with a short focal length, by generating the virtual space display image using the lens information of a telephoto lens with a long focal length, it becomes possible to experience shooting with a telephoto lens with a long focal length. Generally, lenses with long focal lengths used for shooting in real space are large, heavy, and expensive, but in the virtual space shooting of this embodiment, it is possible to provide a shooting experience without such constraints.
[0198] In Figure 25(b), a static virtual space display image 18006 generated by the virtual image generation device 1200 is displayed on the camera 100's display unit 131. Figure 25(b) shows the case where the photographer performs a camera / lens framing operation and moves the camera / lens in the direction of the arrow 18005, which is horizontal. The virtual image generation device 1200 acquires the framing information, which is the position of the camera / lens due to the framing operation, as operation information. Then, by changing the display range when generating the virtual space display image, it generates a virtual space display image 18007 that reflects the framing change due to the photographer's framing operation and displays it on the display unit 131.
[0199] As described above, it is possible to create a virtual space display image that reflects the photographer's operation information using still image virtual space shooting.
[0200] (Viewpoint movement subroutine) Next, using Figure 26, we will explain the subroutine for viewpoint movement processing of S3000 in Figure 13. Specifically, we will show the process of changing the viewpoint position (camera position in the virtual space) when taking images in the virtual space. This process is performed by CPU 1001.
[0201] First, in S3001, the CPU 1001 adjusts the depth of field and field of view of the displayed image when the viewpoint is moved. When moving the viewpoint, it is desirable that a wider range of subjects be easily visible and that the distance at which the focus is sharp is easily confirmed. This is because, as will be described later, the destination of the viewpoint is set from the field of view and the object at the distance at which the focus is sharp, as displayed on the display 131. In this embodiment, in S3001, the field of view is widened to a pre-set field of view, and the range of distances at which the focus is sharp is adjusted to a depth of field shallower than that of the pre-set shooting mode. This process is intended to facilitate the operation of moving the viewpoint and may therefore be omitted.
[0202] In S3002, CPU1001 displays images in the virtual space based on the settings made in S3001. It displays images from the viewpoint position used for pre-existing virtual space captures or from the viewpoint position set as the initial position.
[0203] In S3003, the CPU 1001 adjusts the focus and indicator direction. First, in adjusting the focus, as explained in S3001, the state in focus is displayed within a predetermined distance range within the shooting range, and the photographer performs the same focus adjustment operation as when taking a picture. Specifically, with one focus detection indicator (I(n,m)) as explained in Figure 5 displayed for an object at the position where the viewpoint shift is desired within the shooting range, the photographer adjusts the focus by half-pressing the release switch included in the operation switch group 132. This allows the photographer to see the distance at which the image is in focus on the display unit 131.
[0204] Furthermore, the distance at which the focus is achieved is set as the viewpoint shift distance. Although the method for setting the viewpoint shift distance has been described in the same way as autofocus adjustment, it may also be done by manually operating the focus lens (third lens group 105) of the imaging optical system, in other words, in the same way as manual focus operation. By rotating the focus ring (not shown) provided on the imaging optical system, the distance at which the focus is achieved is adjusted towards infinity and near focus, and the viewpoint shift distance is adjusted to the distance intended by the photographer.
[0205] Furthermore, adjusting the direction of the indicator is done to set the direction of viewpoint movement. The photographer changes the position of the focus detection indicator (I(n,m)) on the display 131 screen, or moves the camera 100 by panning, etc., and aligns the indicator with the direction in which they want to move their viewpoint. This allows the photographer to set the direction of viewpoint movement while checking the image of the virtual space displayed on the display 131.
[0206] In S3004, the CPU 1001 determines whether or not there is an instruction to move the viewpoint. If the viewpoint position change button included in the operation switch group 132 is pressed, the process proceeds to S3005. If the viewpoint position change button is not pressed, the process returns to S3003 and continues to adjust the focus and indicator direction.
[0207] In S3005, the CPU 1001 determines whether viewpoint movement is possible. The direction and distance of viewpoint movement are determined when the photographer instructs viewpoint movement by pressing the viewpoint position change button. If the viewpoint after movement is below the ground (underground) of a background object, inside a foreground object, or if the camera after viewpoint movement interferes with other objects, it is determined that viewpoint movement is not possible. Also, if the distance at which the focus is achieved when the viewpoint position change button is pressed is infinity, it is determined that it is not possible to move the viewpoint to infinity. If the viewpoint movement distance is farther than a predetermined distance, the viewpoint movement distance may be reset to a predetermined distance set in advance as the maximum value.
[0208] In S3006, CPU 1001 determines, based on the results of the viewpoint movement determination in S3005, that viewpoint movement is possible, and proceeds to S3007. On the other hand, if it determines that viewpoint movement is not possible, it proceeds to S3008. In S3007, the viewpoint is moved according to the viewpoint movement distance and direction set in S3003.
[0209] In S3008, the camera operator is notified via the display unit 131 that viewpoint movement is not possible. The notification may simply state that viewpoint movement is not possible, or it may also include the reason, such as interference with an object or the set movement distance being too far.
[0210] After completing the viewpoint movement in S3007, or the notification that viewpoint movement is not possible in S3008, proceed to S3009.
[0211] In S3009, CPU1001 terminates this subroutine upon receiving a command to end the viewpoint movement mode. If there is no command to end the viewpoint movement mode, it returns to S3003.
[0212] Next, using Figure 27, we will explain a specific example of the viewpoint movement subroutine described in Figure 26. Figure 27 shows an example of the display on the display unit 131 in viewpoint movement mode.
[0213] Figure 27(a) shows an example of a display in viewpoint movement mode where viewpoint movement is set in a virtual space with a person and a dog placed as foreground objects. The foreground objects 27003 of the person and the dog are displayed on the display screen 27001 of the display unit 131, and the camera is positioned at a viewpoint from the upper left of the person, as is the state before viewpoint movement.
[0214] 27002 is one of the indicators (I(n,m)) explained in Figure 5, and indicates the direction in which the viewpoint moves on the display screen. 27005 shows the viewpoint movement distance along with the range in which the viewpoint can be moved. In Figure 27(a), it is possible to move from 0.45m to 10m, and the current adjustment distance is 1m. In the example in Figure 27(a), the object is in focus at a distance of 1m from the current viewpoint position (camera position), but the in-focus and out-of-focus distances are not shown in the figure, and it shows a state where everything from very close to infinity is in focus.
[0215] The indicator 27002 is movable within the display screen 27001. By panning the camera, the indicator 27002 can be superimposed onto an object such as a dog, and the focus can be adjusted to that distance, i.e., the viewpoint shift distance can be set. The sub-display screen 27004 displays a preview image of the foreground object 27003 as observed from the currently set viewpoint after viewpoint shift. The direction of the image to preview from the viewpoint after viewpoint shift may be set automatically based on the position information of the foreground object, or it may be set by the photographer operating an operation switch. In addition, a rectangular frame may be shown within the sub-display screen 27004 to indicate the shooting range corresponding to the focal length of the lens being used.
[0216] In this way, by using the screen displayed on the display screen 27001 and the indicator 27002 to set the viewpoint movement distance and direction, the photographer can intuitively and easily move the viewpoint using the same operation as when taking a picture.
[0217] A modified example of viewpoint movement will be explained using Figure 27(b). This method involves placing and displaying a viewpoint movement target as the target position within the display screen, making it easier for the photographer to move their viewpoint.
[0218] Figure 27(b) shows how viewpoint movement targets (target positions, markers) are displayed on the screen in viewpoint movement mode. The viewpoint movement targets 27006 are displayed in a grid pattern on the ground, which is part of the background object of the display screen 27001. At the intersections of the grid, the intersection of the 2nd row and 1st column is shown as viewpoint movement target 27006(2,1), and the intersection of the 4th row and 4th column is shown as viewpoint movement target 27006(4,4). The photographer can set the viewpoint movement distance by moving the indicator 27002 closer to a viewpoint movement target 27006 that is close to the distance to which they want to move their viewpoint, and selecting one of the intersections. The viewpoint movement targets may be displayed superimposed on the foreground object, or they may be displayed together with the numerical value of the distance of the viewpoint movement target. In this way, by displaying the viewpoint movement targets, the photographer can set the viewpoint movement distance more easily.
[0219] (Feedback on the operation of virtual space shooting driven by the camera drive unit) Next, Figure 28 will be used to explain the camera's operation during virtual space photography. In real-world photography, the operation of the shutter and lens provides tactile feedback to the photographer, such as vibration and sound, which improves the quality of the photography experience. On the other hand, as mentioned above, in virtual space photography, virtual subject tracking and virtual space photography are achieved by operating the release switch of the camera 100's operation switch group 132. In this case, there is no drive unit necessary for image generation, and there is no need to drive the shutter or lens. This situation may lead to a decrease in the quality of the photography experience. In this embodiment, the drive unit of the camera 100 is driven in synchronization with the operation of virtual space photography, providing tactile feedback to the photographer, such as vibration and sound, and realizing a more realistic photography experience.
[0220] Figure 28 is a flowchart illustrating the camera operation during virtual space photography. Each process is executed by the camera CPU 121.
[0221] In S1201, the camera CPU 121 receives an initialization instruction for the camera drive unit from CPU 1001 in S1002 in Figure 13 and performs the initialization of the camera drive unit. It drives the shutter 106 to open or close to match the camera settings in the virtual space. For example, if shooting is about to begin, it drives the shutter to open. It also drives the zoom actuator 111, aperture actuator 112, and focus actuator 114 to match the angle of view, depth of field, focal length and aperture value corresponding to the distance at which the camera is in focus, and focus lens position at the start of shooting in the virtual space. In this embodiment, the initial position of the camera drive unit is set by an instruction from CPU 1001, but the position information of each drive unit of the camera 100 may be output to the external computing device 1000, and the initial state may be determined by the external computing device 1000.
[0222] In S1202, the camera CPU 121 outputs camera information, lens information, and operation information, corresponding to S1003 in Figure 13. The content of the information is as explained in S1003.
[0223] In S1203, the camera CPU 121 acquires the image generated in S2000 in Figure 13 and displays it on the display unit 131.
[0224] In S1204, the camera CPU 121 monitors whether the focus drive instruction performed in S4015 in Figure 16 is input to the camera 100. If the focus drive instruction is not received, the process proceeds to S1206; if the focus drive instruction is received, the process proceeds to S1205. In S1205, the camera CPU 121, in accordance with the focus drive instruction, has the focus actuator 114 drive the focus lens (third lens group 105).
[0225] In S1206, the camera CPU 121 monitors whether the aperture value input to the camera 100 in S5001 in Figure 24 differs from the current setting. If there is no change in the aperture value, the process proceeds to S1208; if there is a change in the aperture value, the process proceeds to S1207.
[0226] In S1207, the camera CPU 121 drives the aperture 102 using the aperture actuator 112 according to the change in aperture value.
[0227] In S1208, the system monitors whether the time when Sw2 was detected is input to camera 100 in S5001 in Figure 24. If the time when Sw2 was detected is not input, the system proceeds to S1210. If the time when Sw2 was detected is input, the system proceeds to S1209.
[0228] In S1209, the camera CPU 121 drives the shutter 106 after a predetermined time has elapsed since the detection time of the input Sw2, in the same manner as when capturing images in real space.
[0229] As described above, by operating the focus lens, aperture, and shutter during shooting in a virtual space, the photographer can receive feedback such as vibrations and sounds from the operation of the camera, resulting in a more realistic shooting experience.
[0230] In this embodiment, the drive unit of the camera 100 is configured to drive in response to operation instructions related to shooting in a virtual space. However, there are cases where the camera 100 cannot drive the generated operation instructions. For example, this may occur if the operation instruction exceeds the continuous shooting speed that the camera 100 can drive, or if the operation instruction is for driving a focus lens at a longer distance than the lens that is attached.
[0231] In such cases, based on the generated operation instructions, it may be possible to prohibit the operation of the camera 100's drive unit, or to edit the operation instructions to change them to content that can be driven by the camera 100's drive unit, and then drive it. For example, if the operation instructions are faster than the camera 100's continuous shooting speed, it is conceivable to skip operation instructions at regular intervals and drive the shutter. Furthermore, regarding the operation of aperture and focus, it is conceivable to compare the specifications of the lens used in the virtual space with the specifications of the lens attached in the real space, standardize the drive range to match, and then standardize the drive amount in the same way when an operation instruction is given.
[0232] In this embodiment, we have described the feedback of the camera operator's feel during operation, but the sounds and vibrations that occur during this process may be recorded using other recording means. For example, blur corresponding to the amount of camera vibration may be added to still images, or the resulting sounds may be recorded in video. This makes it possible to achieve virtual shooting that is closer to shooting in real space.
[0233] (Method of capturing and evaluating the results) The flowchart in Figure 29 illustrates the image playback and evaluation method after capture by the camera 100 of this embodiment. Specifically, it performs playback of images captured in real space, playback of images captured in virtual space, and display of the image's defocus map and calculation of the degree of focus of a series of consecutive images under conditions different from those used during actual shooting. This allows for display processing such as showing the causes of poor focus and suggesting the best settings for improvement. In the following description, S represents a step.
[0234] First, in S1101, the camera CPU 121 selects the image to be played back from the flash memory 133. By operating the control switch group 132 that the photographer uses to instruct playback, the most recently taken image or the previously played image is displayed. After that, the photographer operates the control switch group 132 to play back the desired image.
[0235] In S1102, the camera CPU 121 determines whether the image to be played back is an image taken in a virtual space or an image taken in a real space. If the camera CPU 121 determines that the image to be played back was taken in a virtual space, it proceeds to S1103 and plays back the virtual image stored in the memory unit 1004 of the external processing unit 1000 for the display 131. On the other hand, if it determines that the image was taken in a real space and not in a virtual space, it proceeds to S1104 and plays back the image stored in the flash memory 133 for the display 131.
[0236] In S1105, the camera CPU 121 decides whether to perform an evaluation of the virtual space image being replayed (S1103) or the real space image (S1104) as described above. If an evaluation is to be performed, the process proceeds to S1106; otherwise, this flow terminates.
[0237] In S1106, the camera CPU 121 acquires shooting-related information when the image to be played back was taken. Shooting-related information refers to the camera settings and various lens information set at the time of shooting. Shooting-related information includes lens and camera settings at the time of shooting, such as focal length, f-number, continuous shooting mode, AF mode, subject detection AF tracking setting, AF frame setting, and shutter method. Shooting-related information is used when evaluating the focus state of the image, as described later, and can be any information that affects the focus state. Shooting-related information may be stored in the flash memory 133 or the memory unit 1004, or it may be attached as metadata to the image being played back.
[0238] In the S1107, the camera CPU 121 acquires AF log information attached as metadata to the playback image. The AF log information includes defocus information from when the playback image was taken, AF frame setting information, and tracking information (the subject detection AF function focuses on detected objects, such as people, animals, or vehicles, which are automatically detected using a set algorithm). It also includes servo AF characteristics (setting the focus priority by assigning various servo AF parameters) and action recognition information (information on the subject's posture, and information on prioritizing subject recognition when the subject performs a specific action). Furthermore, it includes shutter method information (you can select the shutter mode, such as the mechanical shutter mode which drives the mechanical shutter, or the electronic shutter mode which determines the exposure time using only the image sensor without using the mechanical shutter, and check the frame rate setting for continuous shooting, such as 30, 20, or 10 frames per second for the electronic shutter).
[0239] In S1108, the camera CPU 121 sets one or more images, including the image being played back, as an evaluation image group. The evaluation image group can be set by selecting images taken at a time close to the time the image being played back was taken, or by selecting a group of images taken in a single burst of shots. Alternatively, the system can be configured to allow setting the beginning and end of the evaluation image group.
[0240] In S1109, the camera CPU 121 sets the evaluation sequence. The camera CPU 121 (or the external computing device CPU 1001 in the case of virtual space imaging) determines the equipment to be used and various algorithms, and decides what kind of evaluation to perform. Details will be described later.
[0241] In S1110, the camera CPU 121 performs evaluations under various setting conditions. It performs evaluations based on the various information and setting conditions obtained in S1106 to S1109 as described above. As a result, it calculates the amount of defocus in the captured image from the focus control results, which differ from those at the time of shooting, and evaluates the amount of focus deviation based on a threshold determined from the amount of defocus to calculate the degree of focus.
[0242] In calculating the degree of focus, image analysis is performed simultaneously to analyze the causes of good or bad focus. Furthermore, the amount of defocus is calculated from the focus control results, which differ from those during shooting, and a defocus map is created and superimposed on the captured image.
[0243] This allows us to evaluate the difference in focus state between images taken with the same conditions as when the image was acquired and images taken with different conditions.
[0244] In S1111, the camera CPU 121 displays the evaluation results performed in S1110 on the display unit 131 or an external display such as a PC. The method of displaying the evaluation results is not limited to one format. The display methods will be explained later. By displaying the evaluation results, the photographer can identify the cause of poor focus, allowing them to correct their shooting method, change shooting settings, and improve their shooting skills.
[0245] In S1112, the camera CPU 121 suggests the best settings. Based on the evaluation results from S1111, it suggests the best settings based on the evaluation results regarding the degree of focus in the image. An explanation of the suggested settings and an example of the display will be described later. With the best settings suggested here, the photographer confirms and changes the best settings displayed on the display unit 131. However, a menu can also be provided to decide whether to automatically change the settings using the evaluation results before the evaluation is performed in the playback image settings in S1101. By selecting automatic change, the camera settings can be automatically changed to the best settings based on the evaluation results.
[0246] Furthermore, similar processing can be performed in parallel during shooting. For example, during continuous shooting, the system can automatically switch to the optimal setting for the fourth shot based on the results of evaluating the first three shots, and continue continuous shooting without the photographer's confirmation.
[0247] (Defocus map display) Next, using Figures 30(a) to 30(c), we will explain how the defocus map is superimposed on the captured image in S1111 based on the evaluation results of S1110. Figure 30(a) shows an example of a scene of a person skiing.
[0248] Figure 30(b) shows a defocus map 30001 superimposed on the captured image, based on calculations of the focus control result from the shooting-related metadata attached to the captured image. In this defocus map, the focus position of each 10x8 grid-like block is displayed on the captured image, indicating whether it is in the positive direction (front focus) or the negative direction (back focus) relative to 0.
[0249] The diamond-shaped frame 30002 displayed within each block indicates that the focus position is near 0, meaning the subject is in focus.
[0250] The diagonal line pattern frame 30003 displayed within each block indicates a positive defocus amount, illustrating a tendency towards front focus.
[0251] The dotted frame 30004 displayed within each block indicates a negative defocus amount, illustrating a tendency towards back focus.
[0252] Blocks overlapping with skiers are displayed with a roughly diamond-shaped frame (30002), indicating that the focus is on the subject.
[0253] Figure 30(b) is an example; the defocus map does not need to be a 10x8 grid and may be displayed in finer detail. The amount of defocus is shown for the area roughly encompassing the main subject, but it is not limited to this area and may be displayed for the entire captured image.
[0254] Figure 30(c) shows a defocus map 30001 superimposed on the image captured using a camera (product name CA) and lens (product name LA) combination used by the photographer, based on the calculation of the focus control result from the metadata related to the image capture.
[0255] AF frame 30000 represents the shooting result under conditions where only one point is used as the focus detection area based on the camera's AF frame setting information. The evaluation results, showing that the AF frame covers half of the subject's face, indicate that the contrast of the background subject has affected the focus, resulting in a focus point further away than the main subject and a reduced degree of focus. Looking at the amount of defocus in each block, the right side of the subject is indicated by a diagonal line, suggesting a tendency towards front focus.
[0256] Figure 30(d) shows the defocus map captured using the camera (product name CA) and lens (product name LA) combination used by the photographer mentioned above.
[0257] This shows the defocus results when the AF frame 30000 is changed from the single-point AF described above to a wide-area AF in the evaluation sequence settings performed on S1109.
[0258] The S1110 calculates the amount of defocus based on the focus control information for the single-point AF frame and the focus control information when the AF frame setting is changed to a wide-area AF frame. The result is then displayed as a defocus map 30001 in Figures 30(c) and 30(d).
[0259] By showing the change in the defocus map 30001 due to the change in the AF frame setting, it is possible to compare the amount of defocus with that of single-point AF and area-expanded AF. In Figure 30(d), the diamond-shaped frame 30002 is heavily superimposed on the subject, indicating that the subject is not affected by the background and an image with the subject in focus can be obtained. In the example shown in Figure 30, it can be confirmed that, depending on the photographer's framing technique, better defocus results can be obtained by shooting with a wide-area AF frame setting than with single-point AF.
[0260] In addition to changing the AF frame, the system obtains all the necessary AF information from a virtual camera and existing lens information of a different camera (product name CB) than the camera (product name CA). Then, by rewriting the focus-related information of the captured image with the focus control information of camera (product name CB), it is possible to compare the AF performance difference between cameras (product name CA) and (product name CB). This allows for evaluation of the degree of performance improvement for each shooting scene, especially when performance improvements are expected in new products. This can be used when considering purchasing new products.
[0261] Similarly, for lenses, lens information for a different virtual lens (product name LB) is acquired from a different lens (product name LA) to a different lens (product name LB) that was used in the playback image. Then, a virtual defocus amount can be acquired for the playback image, combining different focal lengths and f-numbers. As a result, the defocus information of the virtual lens (product name LB) can be compared with the defocus information of the image taken using the lens (product name LB) to display and verify the performance when the lens is changed.
[0262] These displays may be shown on the display device of an external PC instead of the display 131 of the camera 100. Regarding the method of comparing the defocus map images after the change, the images before and after the change may be arranged side by side, or only the defocus amount of the defocus map may be changed.
[0263] The defocus amount in FIG. 30 is displayed by dividing the focus position into three levels: near focus, front pin, and rear pin. However, it may be divided more finely, or the defocus amount may be displayed in units of mm or the like.
[0264] The photographer can confirm the performance difference by changing the focus control result based on the information of a camera or lens different from the camera and lens used for shooting. Also, since the performance can be confirmed before purchasing a desired camera or lens, etc., a camera or lens that meets the photographer's requirements can be selected.
[0265] Although FIGS. 30(a) to (c) explain shooting in the real space, they may also be applied to images obtained by shooting in a virtual space using the external computing device 1000, not limited to the real space.
[0266] (Degree of focus) FIG. 31 shows an example of calculating the degree of focus based on the result of the shooting evaluation flow shown in FIG. 29 for a series of continuously shot images.
[0267] FIG. 31 shows a state where the photographer 31003 shoots a series of scenes of a subject skiing and obtains a plurality of images 31001 by continuous shooting. The images dealt with here may be those shot in the real space or those shot in the virtual space. The defocus amount is calculated from the evaluation result S1110 of the captured images of a series of captured images. When it is within a predetermined threshold range centered on the defocus amount 0 from the calculated result, the degree of focus result is determined as ○, and the degree of focus other than that is determined as ×. A determination of ○ or × is made for each image, and the ratio of ○ in the entire captured images obtained by a series of continuous shootings is displayed as the degree of focus display 31002.
[0268] Figure 31 shows an example where the percentage of images with a "○" (in focus) rating was 70%, and a "△" (in focus) is indicated to suggest room for improvement. While the series of focus ratings are shown as 70%△, images with a focus rating of 80% or higher could be shown as "○" to make the focus rating easier for the photographer to understand.
[0269] Furthermore, the display of symbols can be freely set, such as displaying an "X" if the degree of focus is 60% or less, or the degree of focus can be displayed without any symbols.
[0270] In this embodiment, the degree of focus is shown in two stages, either ○ or ×, but the method for determining the degree of focus is just an example and can be freely determined. The unit of defocus calculation, mm, may be used to display variance, etc. The display method can also be freely set without detailed configuration. The series of focus results are displayed on the camera's display unit 131, but they can also be checked on other display devices such as a PC.
[0271] (List of settings changes) Figure 32 is a table showing examples of modifiable items and evaluation conditions for evaluating proposed configuration changes in S1110. Examples of modifiable items are shown horizontally.
[0272] The items written in the white box are examples of camera information settings, showing AF frame settings, tracking (which enables subject detection AF), AF mode (which switches between one-shot AF and servo AF), servo AF characteristics (which changes various servo AF parameters), and shutter type. The gray box is an example of lens information settings, showing the lens and focal length.
[0273] In Figure 32, the initial settings shown are the shooting settings for image acquisition. Recommended settings 1 and 2 show examples of evaluation sequences set in S1109 in Figure 29. The number of evaluation sequences is not limited to two. By changing all possible combinations of configurable parameters and evaluating them in S1110, the best settings (evaluation sequence) that maximize the degree of focus can be found. On the other hand, this increases the computational load, so it may be better to reduce the computational load by changing only the parameters that are effective in improving the degree of focus from the initial settings.
[0274] Using the recommended settings in Figure 32, we will explain settings that are effective in improving the degree of focus.
[0275] The following situations can cause the AF frame to move away from the subject, even with the photographer's default settings: When single-point AF is set to one-point AF and tracking is turned off, the AF frame visible in the viewfinder is fixed. Therefore, the photographer needs to keep the AF frame aligned with the subject, making it difficult to handle unexpected subject movements and framing more challenging.
[0276] In Recommended Setting 1, even with single-point AF, enabling and utilizing tracking AF allows for subject detection, enabling the AF frame to automatically capture and track the subject. Therefore, the photographer only needs to align the subject with the AF frame at the start of shooting to begin tracking, allowing them to concentrate solely on getting the subject into the frame, thus reducing the difficulty of framing. Settings related to the AF frame and subject detection AF tracking can be changed and evaluated for both images taken in real space and images taken in virtual space. By using the image and the defocus map information at the time of image acquisition, a new AF frame can be selected using the algorithm applied to the AF frame settings after the change. Additionally, a new subject detection area can be set for the image using the tracking algorithm applied after the change.
[0277] Similarly, in Recommended Setting 2, the Expanded Area AF setting is selected to widen the AF range for single-point AF. This setting reduces the difficulty of framing without using tracking.
[0278] Regarding shutter mechanisms, electronic front curtain shutters drive the shutter curtain with each release, resulting in blackouts. These blackouts cause the subject to be momentarily lost sight of, and the display update rate also decreases, making framing difficult. Electronic shutters do not use a shutter curtain, so the display update rate does not decrease during continuous shooting, and blackouts (completely black screen) do not occur. Therefore, especially when continuously shooting moving subjects, electronic shutters can reduce the difficulty of framing without losing sight of the subject.
[0279] Regarding shutter settings, for images captured in real space, it is possible to change the frame rate in the direction of slowing down (downsampling) but it is difficult to change it in the direction of speeding up due to a lack of information. For images captured in virtual space, it is possible to change the frame rate in the direction of speeding up because the image can be generated again in the shooting environment. Therefore, when evaluating images captured in real space, if the frame rate is slow, such as when using an electronic front curtain shutter, it may be possible to recommend the use of an electronic shutter to the photographer using other information such as the speed of the subject being photographed.
[0280] Regarding the lens's focal length, the photographer is using a 70mm-200mm zoom lens and has set the focal length to 200mm during shooting. This results in a narrow field of view relative to the subject, making it easy for the subject to move out of frame unexpectedly or move quickly while sliding. Therefore, framing becomes difficult. By widening the field of view, there is more room to react to unexpected movements or fast sliding movements of the subject, reducing the risk of the subject moving out of frame. Therefore, by setting the focal length to the wider end, 70mm, the difficulty of framing can be reduced.
[0281] Regarding focal length settings, while it's possible to narrow the field of view for images captured in real space, widening it is difficult due to the lack of corresponding information (images). For images captured in virtual space, it's possible to recreate the image environment, thus widening the field of view. Therefore, when evaluating images captured in real space, if a long focal length was used, other information such as the subject's speed may be used to recommend to the photographer the use of a wider-angle lens.
[0282] As mentioned above, by changing the settings for factors that reduce the degree of focus, it is possible to find recommended settings that can improve the degree of focus more efficiently, thereby reducing the computational load.
[0283] The settings shown in Figure 32 are merely one example, and settings may be configured based on various considerations. Analysis of the shooting environment, such as whether the subject is human or animal, or whether it is a scene with multiple subjects, can also be utilized. Furthermore, the recommended settings may be narrowed down based on the photographer's skill level, determined from the previously entered shooting history and the subject's movement and the photographer's framing during shooting.
[0284] (Explanation of the display of information regarding shooting settings) Next, based on the evaluation results of the degree of focus described above, we will explain the display of information regarding the photographer's framing technique in the captured image using Figure 33.
[0285] Figure 33(a) shows a blurred image in which the defocus amount is large and the focus level result is determined to be NG in the evaluation result S1111 of a series of continuous captured images. It is assumed that this image fails to frame the movement of the subject skiing at high speed, and the AF frame 32001 has moved out of the subject's face, resulting in a poor focus level. Under the camera settings conditions set by the photographer, it is speculated that factors such as the influence of the shutter method, the selection of the AF frame, and the angle of view due to the focal length of the lens make framing difficult. In such a situation, based on the image evaluation during image playback, a proposal for the best settings is displayed on the camera's display 131 or a display device such as a PC. The proposal and display of the best settings will be explained in the next Figure 33(b).
[0286] Figure 33(b) shows an example of displaying a proposal to the photographer regarding the best settings as a result of evaluating various shooting sequences.
[0287] In the evaluation under each setting condition of S1110, in addition to the evaluation, various types of conditions with different combinations are calculated from each piece of information and setting conditions from S1106 to S1109 to find setting conditions with a higher focus level.
[0288] For example, in Figure 33(a), the AF frame 32001 is out of the subject, while in the evaluation of S1110, the focus level when the AF frame is widened from the AF frame setting in the AF log information of S1107 is calculated. Without changing the AF mode, the focus level when adding the tracking information of subject detection AF is confirmed. In addition, the focus level is calculated by swapping various conditions, and the combination of settings with the highest focus level is selected from among them.
[0289] In the evaluation result of S1110 of the captured image, it is assumed in the example of Figure 33 that the following conditions have resulted in an increase in the focus level from the focus level results of various conditions. (1) Keep the AF frame unchanged as a single point and change the tracking of subject detection AF to on (2) Change the shutter method to an electronic shutter (3) Change the angle of view to the wide-angle side Based on the evaluation results of S1110 described above, the best information can be suggested to the photographer. Display 32003 shows the best information for the factors described above.
[0290] Figure 33(c) shows a display 32004 that allows the photographer to choose whether or not to change the tracking for subject detection in the shooting-related information of the best setting proposal in Figure 33(b) described above. The photographer can check this display and make settings that can further improve the degree of focus.
[0291] Next, although not shown in the diagram, a similar display is shown indicating whether or not to change the shutter method, allowing the photographer to make a selection. In the case of shooting in real space, the lens angle of view cannot be selected, so the photographer is advised to change the focal length by displaying a message. In the case of shooting in virtual space, the focal length can be changed virtually, so, similar to Figure 33(c), a message regarding the change is displayed to prompt the photographer to make a selection.
[0292] The display method and order of these suggested setting changes may be arranged so that all settings are displayed together for selection. Alternatively, only the focus level may be displayed, and the settings may be changed together to the calculated focus level selected by the photographer.
[0293] While I've explained the prompts for settings here, the camera may also change these settings automatically.
[0294] Figure 33(d) shows the evaluation results of the shooting after changing the settings in Figure 33(c).
[0295] The AF frame 32001 has changed to a dotted line, which is the frame for tracking AF. As a result of subject detection, the camera can continuously track the subject, allowing you to concentrate on framing.
[0296] By switching to an electronic shutter, there is no blackout, and the subject remains visible in the viewfinder at all times, reducing the likelihood of losing sight of the subject. The lens has been changed to a 70mm wide-angle, which provides more leeway in the subject-to-angle ratio, reducing the chance of the subject going out of frame. As a result of these optimal settings, the focus accuracy is 85%, a significant improvement from the previous 70%.
[0297] By evaluating a series of captured images and quantifying the degree of focus, photographers can understand their framing skills. Analyzing the causes of poor focus and displaying the optimal settings can help improve photographers' framing techniques.
[0298] Figures 33(a) to (c) illustrate the process using real-world imaging. However, as with Figure 31, the system can also evaluate the degree of focus obtained from imaging in a virtual environment using an external computing device 1000, rather than just the real world. The system can then display the best settings and allow for changes to or automatic changes to the imaging settings.
[0299] (modified version) In this embodiment, a configuration was described in which the detection of the focus region is achieved by region detection based on machine learning. However, the method for detecting the focus region is not limited to this. For example, the focus region can be set using the aspect ratio of the subject detection region, the size of the subject detection region, or depth information of the subject obtained using a defocus map.
[0300] <Second Embodiment> Next, a second embodiment will be described. In this embodiment, in the virtual space image generation and output processing, images captured in the real space are also used to generate and output images in the virtual space. The configuration of the imaging system 10 in this embodiment is the same as in the first embodiment, but some parts of the virtual space image generation and output processing differ. Here, we will mainly explain the differences from the first embodiment in the virtual space image generation and output processing.
[0301] The virtual space image generation and output processing of the second embodiment will be explained using the virtual space image generation and output processing subflowchart shown in Figure 34.
[0302] In S3501, CPU1001 performs image acquisition. The image to be acquired may be an image captured in the aforementioned real-world imaging process, or it may be an image that has been captured in advance.
[0303] In S3502, the CPU 1001 acquires and synthesizes foreground objects. A trained model that estimates a 3D model for an image may be used to generate a 3D model of the subject from a captured image in real space and synthesize it with the 3D model of the subject in the virtual space described above. For example, the face may be acquired as a foreground object from a captured image in real space, while other parts such as the torso may be acquired from the foreground object storage unit 1104 of the virtual space reproduction device 1100. Alternatively, the 3D model of the subject from the captured image and the 3D model of the subject in the virtual space described above may be acquired alternately in a time series and displayed at different timings as the virtual space display image described later. Alternatively, foreground objects in the virtual space may be acquired and synthesized with the captured image in real space at the stage of generating the virtual space display image in S3503 described later.
[0304] In S3503, the CPU 1001 acquires and composites background objects. Similar to the acquisition of foreground objects in S3502, a 3D model may be generated from the captured image and used as the background object, or it may be composited with the virtual space background object acquired by the background object acquisition unit 1105, with the regions separated. Alternatively, the background object from the captured image and the background object from the virtual space may be separated chronologically. Alternatively, the background object from the virtual space may be acquired and composited with the captured image in the virtual space display image generation in S3505, which will be described later.
[0305] In S3505, the CPU 1001 generates the display image in the virtual space. If the foreground and background objects have been combined with the captured image, the process is the same as the virtual space display image generation process in S2008 in Figure 15 of the first embodiment. When generating a display image by combining the foreground and background objects of the virtual space with the captured image, the captured image is aligned and combined with the display image generated by the foreground and background objects of the virtual space to generate the display image. For example, only the face portion of the subject in the captured image is cut out and combined with the display image in the virtual space.
[0306] Furthermore, if the image range or viewpoint position differs due to differences in the field of view of the captured image, a pre-trained model that estimates a 3D model for the image is used to generate a 3D model from the captured image, and the image range and viewpoint position are changed to generate the display image. Then, the display image in the virtual space is segmented and combined to generate a display image in the virtual space that also includes the captured image.
[0307] (Other embodiments) Furthermore, the present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by a process in which one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0308] The disclosures herein include the following imaging devices and their control methods, programs, and storage media.
[0309] (Item 1) An operating means for performing operations for shooting, A drive mechanism for driving a drive unit for taking photographs in real space, Control means for controlling the drive means, Equipped with, The imaging apparatus is characterized in that the control means controls the drive unit so that when the operating means instructs the operation to take a picture in the virtual space, the drive unit is driven in synchronization with the shooting sequence in the virtual space.
[0310] (Item 2) The imaging device according to item 1, characterized in that the drive unit includes a shutter, aperture, zoom lens, and focus lens.
[0311] (Item 3) The imaging device according to item 1 or 2, further comprising a switching means for switching between a mode for performing imaging in the real space and a mode for performing imaging in the virtual space.
[0312] (Item 4) The imaging apparatus according to any one of items 1 to 3, characterized in that the control means controls the drive unit so that the drive unit is driven to a predetermined position before taking images in the virtual space.
[0313] (Item 5) The imaging device according to item 4, characterized in that the predetermined position is the position in the virtual space where the drive unit is set.
[0314] (Item 6) The imaging apparatus according to any one of items 1 to 5, characterized in that the control means controls the drive means to prohibit or change the driving of the drive unit if the drive unit cannot be driven in accordance with the shooting sequence in the virtual space.
[0315] (Item 7) The imaging device according to item 6, characterized in that the control means controls the drive unit to drive the shutter, which acts as a drive unit, at a continuous shooting speed different from the continuous shooting speed corresponding to the continuous shooting sequence in the virtual space, when the shutter cannot be driven at a continuous shooting speed corresponding to the continuous shooting sequence in the virtual space.
[0316] (Item 8) The imaging device according to item 7, characterized in that the control means controls the drive unit to drive the shutter, which acts as a drive unit, at a continuous shooting speed slower than the continuous shooting speed corresponding to the continuous shooting speed corresponding to the shooting sequence in the virtual space, when the shutter cannot be driven at a continuous shooting speed corresponding to the shooting sequence in the virtual space.
[0317] (Item 9) The imaging apparatus according to item 6, characterized in that the control means normalizes the drive range of the aperture as the drive unit and the drive range of the aperture corresponding to the shooting sequence in the virtual space.
[0318] (Item 10) The imaging apparatus according to item 6, characterized in that the control means normalizes the drive range of the focus lens as the drive unit and the drive range of the focus lens corresponding to the shooting sequence in the virtual space.
[0319] (Item 11) A method for controlling an imaging device comprising an operating means for performing operations for taking photographs and a driving means for driving a drive unit for taking photographs in real space, The drive means has a control step, A control method for an imaging device, characterized in that, in the control step, when the operating means instructs the operation to take a photograph in the virtual space, the drive means is controlled so that the drive unit is driven in synchronization with the shooting sequence in the virtual space.
[0320] (Item 12) A program to cause a computer to execute the control method for the imaging device described in item 11.
[0321] (Item 13) A computer-readable storage medium containing a program that causes a computer to execute the control method for the imaging device described in item 11.
[0322] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]
[0323] Camera: 100, 107: Image sensor, 111: Zoom actuator, 112: Aperture actuator, 114: Focus actuator, 121: Camera CPU, 131: Display, 1000: External processing unit, 1001: CPU, 1100: Virtual space reproduction device, 1200: Virtual image generation device, 2000: Camera / lens information storage device
Claims
1. An operating means for performing operations for shooting, A drive mechanism for driving a drive unit for taking photographs in real space, Control means for controlling the drive means, Equipped with, The imaging apparatus is characterized in that the control means controls the drive unit so that when the operating means instructs the operation to take a picture in the virtual space, the drive unit is driven in synchronization with the shooting sequence in the virtual space.
2. The imaging apparatus according to claim 1, characterized in that the drive unit includes a shutter, aperture, zoom lens, and focus lens.
3. The imaging device according to claim 1, further comprising a switching means for switching between a mode for performing photography in the real space and a mode for performing photography in the virtual space.
4. The imaging apparatus according to claim 1, characterized in that the control means controls the drive unit so that the drive unit is driven to a predetermined position before taking images in the virtual space.
5. The imaging apparatus according to claim 4, characterized in that the predetermined position is the position in the virtual space where the drive unit is set.
6. The imaging apparatus according to claim 1, characterized in that the control means controls the drive unit to prohibit or change the driving of the drive unit if the drive unit cannot be driven in accordance with the shooting sequence in the virtual space.
7. The imaging apparatus according to claim 6, characterized in that the control means controls the drive unit to drive the shutter, which acts as a drive unit, at a continuous shooting speed different from the continuous shooting speed corresponding to the continuous shooting speed corresponding to the shooting sequence in the virtual space, when the shutter cannot be driven at a continuous shooting speed corresponding to the shooting sequence in the virtual space.
8. The imaging apparatus according to claim 7, characterized in that the control means controls the drive unit to drive the shutter, which acts as a drive unit, at a continuous shooting speed slower than the continuous shooting speed corresponding to the continuous shooting speed corresponding to the shooting sequence in the virtual space, when the shutter cannot be driven at a continuous shooting speed corresponding to the shooting sequence in the virtual space.
9. The imaging apparatus according to claim 6, characterized in that the control means normalizes the drive range of the aperture as the drive unit and the drive range of the aperture corresponding to the shooting sequence in the virtual space.
10. The imaging apparatus according to claim 6, characterized in that the control means normalizes the drive range of the focus lens as the drive unit and the drive range of the focus lens corresponding to the shooting sequence in the virtual space.
11. A method for controlling an imaging device comprising an operating means for performing operations for taking photographs and a driving means for driving a drive unit for taking photographs in real space, The drive means has a control step, A control method for an imaging device, characterized in that, in the control step, when the operating means instructs the operation to take a photograph in the virtual space, the drive means is controlled so that the drive unit is driven in synchronization with the shooting sequence in the virtual space.
12. A program for causing a computer to execute the control method of the imaging device described in claim 11.
13. A computer-readable storage medium storing a program for causing a computer to execute the control method of the imaging device described in claim 11.
Citation Information
Patent Citations
Virtual reality space video production system
JP2011035638A