Imaging device and control method thereof
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2025-01-24
- Publication Date
- 2026-08-05
AI Technical Summary
【0008】 本発明によれば、消費電力を抑制しつつ、露光期間中に電子ビューファインダに表示する画像を適切に生成可能な撮像装置およびその制御方法を提供することができる。
Smart Images

Figure 2026126785000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an imaging device and a control method thereof.
Background Art
[0002] In an imaging device having an electronic viewfinder (EVF), an image for display on the EVF cannot be acquired during the period when the imaging element is exposed to capture a recording image. Therefore, when the exposure period is long (the shutter speed is slow), a phenomenon (blackout) in which an image is not displayed on the EVF may occur, which can be a problem.
[0003] Patent Document 1 proposes a technique for displaying an image predicted using artificial intelligence (AI) from a captured image in order to reduce the time difference from shooting to display in a wearable device that continuously performs shooting and display by the mounted imaging device.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the technique proposed in Patent Document 1 has a configuration that always predicts an image during the use of the wearable device, so the power consumption is high, and it is difficult to implement in an imaging device that generally operates on a battery.
[0006] In one aspect of the present invention, there is provided an imaging device and a control method thereof that can appropriately generate an image to be displayed on an electronic viewfinder during an exposure period while suppressing power consumption.
Means for Solving the Problems
[0007] In one aspect, the present invention provides an imaging device comprising: an image sensor; a display means; an image processing means for generating a display image for display on the display means from an image captured by the image sensor; a generation means that uses a machine learning model to generate a predictive image predicting an image to be captured at a time later than the capture time of multiple frames from multiple frames captured by the image sensor; a display control means that causes the display means to function as an electronic viewfinder by displaying the display image or the predictive image on the display means; and a determination means for determining whether or not to generate a predictive image with the generation means, wherein the determination means determines to generate a predictive image with the generation means when it is determined that the exposure time during still image capture is longer than a predetermined time, and determines not to generate a predictive image with the generation means when it is determined that the exposure time is not longer than a predetermined time. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide an imaging apparatus and a control method therefor that can appropriately generate an image to be displayed in an electronic viewfinder during the exposure period while suppressing power consumption. [Brief explanation of the drawing]
[0009] [Figure 1] Cross-sectional view showing an example configuration of an imaging device according to the embodiment. [Figure 2] Block diagram showing an example of the functional configuration of the imaging device according to the first embodiment. [Figure 3] Schematic diagram illustrating the operation of the imaging device in the first embodiment. [Figure 4] Flowchart relating to the operation of the imaging device in the first embodiment [Figure 5] Block diagram showing an example of the functional configuration of the imaging device according to the second embodiment. [Figure 6] Flowchart relating to the operation of the imaging device in the second embodiment [Figure 7] Block diagram showing an example of the functional configuration of the imaging device according to the third embodiment. [Figure 8] Flowchart relating to the operation of the imaging device in the third embodiment [Figure 9] Flowchart relating to another operation of the imaging device in the third embodiment [Figure 10] Schematic diagram illustrating the operation of the imaging device in the third embodiment. [Modes for carrying out the invention]
[0010] The present invention will be described in detail below with reference to the attached drawings, based on exemplary embodiments thereof. Note that the following embodiments do not limit the invention to the claims. Furthermore, while multiple features are described in the embodiments, not all of them are essential to the invention, and the multiple features may be combined arbitrarily. In addition, in the attached drawings, the same or similar configurations are given the same reference numeral, and redundant descriptions are omitted.
[0011] In the following, we will describe an embodiment of the present invention using a digital camera as an example of an imaging device. However, the present invention can be implemented with any electronic device having an imaging function. Such electronic devices include video cameras, computer equipment (personal computers, tablet computers, media players, etc.), smartphones, game consoles, robots, drones, dashcams, and the like. These are examples, and the present invention can be implemented with other electronic devices as well.
[0012] ●(First Embodiment) <Digital Camera Configuration> Figure 1 shows a YZ cross-section including the optical axis of a digital camera 1 according to an embodiment of the present invention. The digital camera 1 is an interchangeable-lens camera, and the photographic lens 1A is a lens unit that can be attached to and detached from the camera body 1B. The photographic lens 1A has a focus lens for adjusting the focusing distance, movable members such as an aperture, and a motor, actuator, etc. that drive them. The aperture may also function as a mechanical shutter. The operation of the focus lens and aperture is controlled by the CPU 3.
[0013] The imaging device 2 is arranged to image the optical image formed by the photographing lens 1A of the digital camera 1. The imaging device 2 may be, for example, a known CCD or CMOS color image sensor having a color filter in a primary color Bayer array. The imaging device 2 has a pixel array in which a plurality of pixels are two-dimensionally arranged, and a peripheral circuit for reading signals from each pixel. Each pixel accumulates electric charges corresponding to the amount of incident light by photoelectric conversion. By reading out from each pixel a signal having a voltage corresponding to the amount of electric charges accumulated during the exposure period, a group of pixel signals (analog image signals) representing the subject image formed by the photographing lens 1A is obtained. In the present embodiment, it is assumed that the imaging device 2 has an A / D converter and outputs image data obtained by A / D converting the analog image signal, but the A / D conversion may be performed by an external circuit such as the CPU 3.
[0014] The CPU 3 is one or more processors capable of executing programs. The CPU 3 is a control unit of the digital camera 1, and realizes the functions of the digital camera 1 by, for example, reading a program stored in a non-volatile memory included in the memory unit 4 into a RAM included in the memory unit 4 and executing it.
[0015] The memory unit 4 has a rewritable non-volatile memory and a RAM. The non-volatile memory stores programs executable by the CPU 3, set values, GUI data, etc. The non-volatile memory is also used as a recording destination for the captured images. The RAM is used to read a program executed by the CPU 3, save values necessary during the execution of the program, and temporarily store the captured images.
[0016] The display element 6 and the eyepiece lens 5 constitute an electronic viewfinder of the type that looks into the camera body 1B from the outside. The display element 6 can perform two-dimensional dot matrix display. While the digital camera 1 is operating in the shooting mode, the CPU 3 realizes the function of the electronic viewfinder by continuously performing video shooting by the imaging element 2 and displaying the captured video on the display element 6. Note that the display element 6 is not limited to being disposed within the camera body 1B, and may be configured to be provided on the housing surface of the camera 1B. That is, it may not be an electronic viewfinder of the looking-in type.
[0017] The operation unit 8 is a general term for input devices (buttons, switches, dials, etc.) provided for the user to input various instructions to the digital camera 1. For the sake of convenience, FIG. 1 shows only one input device provided on the back of the main body and the release button 7 among the input devices constituting the operation unit 8.
[0018] The input devices constituting the operation unit 8 have names corresponding to the assigned functions. For example, the operation unit 8 includes a release button 7, a video recording switch, a shooting mode selection dial for selecting a shooting mode, a menu button, direction keys, a determination key, and the like. The release button 7 is a switch for still image recording, and the CPU 3 recognizes the half-pressed state of the release button 7 (turning on of switch 1 (SW1)) as an instruction for shooting preparation and the fully pressed state (turning on of switch 2 (SW2)) as an instruction for starting shooting. Also, when the video recording switch is pressed in the shooting standby state, the CPU 3 recognizes it as an instruction for starting video recording, and when it is pressed during video recording, the CPU 3 recognizes it as an instruction for stopping recording. Note that the functions assigned to the same input device may be variable. Also, the input device may be a software button or key using a touch display. Further, the operation unit 8 may include an input device corresponding to a non-contact input method such as voice input or eye gaze input.
[0019] The inertial sensor 9 includes an accelerometer and a gyroscope. The inertial sensor 9 detects acceleration information in the XYZ axis direction and angular velocity information around the XYZ axis as motion information of the digital camera 1. The CPU 3 acquires the motion information of the digital camera 1 from the inertial sensor 9 at a predetermined period and stores it in the memory unit 4. The inertial sensor 9 may also be used for image blur prevention.
[0020] Figure 2 is a block diagram illustrating the functions of CPU3, with the same reference numerals used for components identical to those in Figure 1. The functional blocks of CPU3 schematically represent the functions that CPU3 will implement by executing programs. Therefore, the operations of the functional blocks described below are actually primarily performed by CPU3. Note that functional blocks do not necessarily have to be implemented solely by CPU3 executing programs; they may also be implemented using hardware circuits other than CPU3 (e.g., GPU, NCU, ASIC, etc.) as needed.
[0021] The imaging control unit 301 generates signals to control the operation of the image sensor 2 according to the state of the release button 7 (SW1 on, SW2 on, SW1 and SW2 off) and the shooting conditions.
[0022] The shooting condition setting unit 302 determines the shooting conditions (shutter speed (exposure time), aperture value, and ISO sensitivity) based on the settings input through the operation unit 8. The shooting condition setting unit 302 outputs the determined shooting conditions to the image generation execution determination unit 303 and stores them in the RAM of the memory unit 4. In addition to the settings via the operation unit 8, or instead of settings, the shooting condition setting unit 302 may also determine the shooting conditions based on evaluation values for AE generated by the image processing unit 306 from the image captured by the image sensor 2.
[0023] The image generation execution determination unit 303 determines whether or not to generate a display image during the next still image capture, based on the shutter speed (exposure time) among the shooting conditions determined by the shooting condition setting unit 302.
[0024] If the image generation execution determination unit 303 determines that a display image should be generated during the next still image capture, the image generation unit 304 generates an image to be displayed on the display element 6 during the exposure period of the next still image capture. Details of the operation of the image generation unit 304 will be described later.
[0025] The display control unit 305 executes control to display the image generated by the image processing unit 306 or the image generated by the image generation unit 304 on the display element 6. Details of the operation of the display control unit 305 will be described later.
[0026] The image processing unit 306 applies predetermined image processing to the image data stored in the memory unit 4 to generate signals and image data according to their intended use, and to acquire and / or generate various types of information.
[0027] The image processing applied by the image processing unit 306 to the image data may include, for example, preprocessing, color interpolation, correction, detection, data processing, evaluation value calculation, and special effects processing. Preprocessing may include signal amplification, reference level adjustment, and defective pixel correction. Color interpolation is performed when a color filter is provided on the image sensor, and it is a process that interpolates the values of color components that are not included in the individual pixel data that makes up the image data. Color interpolation is also called demosaicing. Correction processing may include white balance adjustment, gradation correction, correction of image degradation caused by optical aberrations in the imaging optical system of the photographic lens 1A (image recovery), correction of the effect of vignetting in the imaging optical system, and color correction. Detection processes may include detecting feature regions (such as face regions or human body regions) and their movements, as well as recognizing people. Data processing may include processes such as region cropping, merging, scaling, encoding and decoding, and header information generation (data file generation). The generation of display image data and recording image data is also included in data processing. Display image data also includes data for the EVF image displayed on the display element 6. Display image data is generated to correspond to the characteristics of the display device that displays it (resolution, brightness dynamic range, display frame rate, etc.). The evaluation value calculation process may include processes such as generating signals and evaluation values used for autofocus detection (AF), and generating evaluation values used for automatic exposure control (AE). Special effects processing may include adding blur effects, changing color tones, and relighting. These are merely examples of processes that the image processing unit 306 can apply, and do not limit the processes that the image processing unit 306 can apply.
[0028] Although Figure 2 shows the inertial sensor 9 directly storing motion information in the memory unit 4, in reality, the CPU 3 periodically acquires motion information from the inertial sensor 9 and stores it in the memory unit 4.
[0029] <Explanation of operation examples when taking still images> Figure 3 schematically shows the time-dependent changes between the shooting scene and the display image generated by the digital camera 1 during the period from when still image capture is performed in the shooting standby state until the camera returns to the shooting standby state. Times T1 to T4 are times that have elapsed from time T0 by a predetermined unit time (e.g., 1 / 120 second). The unit time may be, for example, the reciprocal of the display frame rate of the EVF (display element 6).
[0030] In the shooting standby state, CPU3 performs the operations necessary to enable the electronic viewfinder to function. Specifically, CPU3 performs the following: • Video recording using image sensor 2 • Generation of EVF images based on video footage obtained during shooting • Display of EVF image on display element 6 The operation of the digital camera 1 is controlled to continuously perform the following actions.
[0031] When the SW2 of the release button 7 is turned on between time T0 and time T1, the image control unit 301 controls the operation of the image sensor 2 to perform still image capture according to the shooting conditions determined by the exposure time setting unit 302. The CPU 3 also controls the operation of the focus lens and aperture as needed.
[0032] Assume that the exposure time ends between time T3 and time T4. Note that in Figure 3, for convenience, SW2 is shown to be ON from time T1 to T3, but in non-burst shooting mode, the state of SW2 after it was first turned ON is not considered in the operation.
[0033] During the exposure period for still image capture, the image sensor 2 is occupied. Therefore, while a real image can be captured and an image for EVF display can be generated at time T0 before the exposure period and time T4 after the exposure period, an image for EVF display based on a real image cannot be generated during the period from time T1 to T3.
[0034] Therefore, the CPU 3 (image generation unit 304) generates predicted images based on multiple past frames, including the image taken at time T0 immediately before the exposure period, to predict the captured images at times T1 to T3 during the exposure period, after the capture time T0. Then, the display control unit 305 displays the predicted images generated by the image generation unit 304 on the display element 6 as images for EVF display at times T1, T2, and T3.
[0035] Thus, the digital camera 1 of this embodiment generates an EVF display image to be displayed during the exposure period of still image shooting based on past images and displays it on the display element 6. Therefore, even if the exposure period of still image shooting spans the EVF display timing, the display on the display element 6 can continue to display, and it is possible to display a more appropriate image that is closer to the actual image than when past images are continuously displayed.
[0036] Figure 4 is a flowchart illustrating the operation of CPU 3 regarding EVF display operation when digital camera 1 is operating in shooting mode rather than playback mode. CPU 3 continuously performs the operations shown in the flowchart of Figure 4 at a frequency corresponding to the EVF display frame rate, in parallel with other operations such as operations for taking still images.
[0037] In S101, CPU3 determines whether or not a recording image is being captured. CPU3 determines that a recording image is being captured during the period from when SW2 of the release button 7 is turned on until the end of the exposure period. If CPU3 determines that a recording image is being captured, it executes S102; otherwise, it executes S120.
[0038] In S120, CPU3 continues to record video to generate images for EVF display, and image sensor2 stores one frame of image data in RAM of memory unit 4. Then, image processing unit 306 generates image data for EVF display based on the image data read from memory unit 4.
[0039] In S121, the image processing unit 306 stores the generated display image data in the RAM of the memory unit 4. The image processing unit 306 deletes old display image data as needed so that the RAM of the memory unit 4 stores the most recent predetermined number of display image data frames. The RAM of the memory unit 4 also stores motion information of the digital camera 1 (measured values from the inertial sensor 9) at the time the predetermined number of display image data frames were captured.
[0040] In S122, the display control unit 305 displays the display image data read from the memory unit 4 on the display element 6, and completes the processing for one frame.
[0041] In S102, the image generation execution determination unit 303 obtains the shooting conditions from the exposure time setting unit 302 and executes S103. The exposure time setting unit 302 continuously determines the shooting conditions based on the AE evaluation value generated by the image processing unit 306, for example, when SW1 of the release button 7 is turned on.
[0042] In S103, the image generation execution determination unit 303 determines whether the shutter speed (exposure time) among the shooting conditions acquired in S102 is longer than a predetermined time. If the image generation execution determination unit 303 determines that the exposure time is longer than the predetermined time, it executes S104; otherwise, it executes S110. In other words, if the exposure time is longer than the predetermined time, the image generation execution determination unit 303 decides to generate a predicted image with the image generation unit 304, and if the exposure time is not longer than the predetermined time, it decides not to generate a predicted image with the image generation unit 304.
[0043] Note that for subsequent executions during the capture of recording images, the execution of S102 may be skipped. Also, in S103, the image generation execution determination unit 303 determines whether the remaining exposure time is longer than a predetermined time.
[0044] In S110, the display control unit 305 displays the EVF display image data generated in S121 during the processing of the previous frame on the display element 6, and then completes the processing for one frame.
[0045] In S104, the image generation unit 304 generates a predicted image based on the EVF display image data stored in the memory unit 4. When considering the movement of the digital camera 1, the image generation unit 304 also uses the motion data stored in the RAM of the memory unit 4 to generate the predicted image. The image generation unit 304 generates a predicted image by inputting the EVF display image data into a trained machine learning model. By generating a predicted image using a machine learning model, it becomes possible to generate a predicted image that takes into account the type and posture of the subject and the environment in which the subject exists (such as the shape of the road).
[0046] Machine learning models can be, for example, CNNs (Convolutional Neural Networks), GANs (Generative Adversarial Networks), or generative models using Transformer layers. In the following examples, we will use a CNN.
[0047] Here, we will explain the configuration and training method of the machine learning model used to generate predictive images. A separate machine learning model can be prepared for each type of shooting scene that includes moving subjects, such as sports, motorsports, and animals. When preparing a machine learning model according to the shooting scene, the image processing unit 306 detects the subject and its movement in the image data used for the EVF display in the shooting standby state and determines the shooting scene. Then, the image generation unit 304 generates a predictive image using the machine learning model corresponding to the shooting scene most recently determined by the image processing unit 306.
[0048] The training data for the machine learning model will be video footage shot at the same frame rate as the EVF's display frame rate. While the machine learning model is configured here to output a predicted image for the next frame when multiple frames are input, it may also be configured to generate predicted images for subsequent frames when multiple frames are input.
[0049] Furthermore, camera motion information can be used to train the machine learning model. By using camera motion information, the accuracy of predicted images can be improved when the camera moves during the exposure period, such as in panning shots.
[0050] Specifically, a machine learning model is trained using video shot with panning motion, similar to panning, and information about the camera's movement during the video's capture. In this case, the video is used as input data for the machine learning model (CNN), and the motion information is input to the same layer as the CNN's output, where it can be used as a feature along with the CNN's output in the connected layer.
[0051] The learning process is supervised learning, with the training data consisting of the image of the next frame (or the next multiple frames) of the multiple frame images used as training data. The CNN is then trained by updating its parameters using gradient descent or Adam (Adaptive Moment Estimation) to minimize the error between the predicted image generated for the training data and the training data. Alternatively, a GAN configuration may be used, where a discriminator is connected after the generator to determine whether an image is a predicted image or not.
[0052] The image generation unit 304 generates a predicted image by inputting the display image data (and motion data if the movement of the digital camera 1 is to be considered) stored in the RAM of the memory unit 4 into the machine learning model that has been trained as described above.
[0053] In S105, the display control unit 305 displays the image data of the predicted image generated by the image generation unit 304 in S104 on the display element 6, and completes the processing for one frame.
[0054] According to this embodiment, since images for EVF display are generated using a machine learning model during still image capture, blackout of the EVF display during still image capture can be suppressed. Furthermore, since image generation using the machine learning model is performed only when the shutter speed (exposure time) during still image capture is longer than a predetermined time, power consumption can be suppressed, making it suitable for implementation in battery-powered devices such as imaging devices. In addition, by using camera motion information, it becomes possible to generate images for EVF display with greater accuracy during panning shots.
[0055] ●(Second Embodiment) Next, a second embodiment of the present invention will be described. In this embodiment, in addition to the exposure time, the shooting mode is taken into consideration when determining whether or not to generate a predicted image.
[0056] Figure 5 is a block diagram showing an example of the functional configuration of the digital camera 1 according to this embodiment, and the same reference numerals as in Figure 2 are used for components similar to those in the first embodiment. In this embodiment, the CPU 3 has the function of the shooting mode setting unit 306. Although the shooting mode setting unit 306 is shown instead of the exposure time setting unit 302 in Figure 2, this does not mean that the function equivalent to the exposure time setting unit 302 has disappeared; it is still included in the functions executed by the CPU 3.
[0057] The shooting mode in this embodiment differs from the shooting mode in the first embodiment. The shooting mode in the first embodiment is paired with the playback mode and refers to an operating mode that enables the shooting of video or still images.
[0058] On the other hand, the shooting modes in this embodiment are modes related to setting shooting conditions when taking still images. Specifically, these include modes that set shooting conditions suitable for photographing a specific subject all at once on the digital camera 1, and modes that allow the user to change either or both the aperture and / or shutter speed. The former includes sports mode, night scene mode, portrait mode, etc., while the latter includes shutter-priority mode, aperture-priority mode, manual mode, bulb mode, etc. These are examples, and the types of shooting modes may vary depending on the camera.
[0059] The shooting mode can be set by the user, for example, through the shooting mode setting dial included in the control unit 8 or through a menu screen that can be operated by the control unit 8. Furthermore, user setting is not mandatory; for example, the CPU 3 may automatically set the mode based on the detection results of the image processing unit 306, such as the type and movement of the subject. In this embodiment, the shooting mode setting is stored in the RAM of the memory unit 4 as one of the current settings.
[0060] Figure 6 is a flowchart illustrating the operation of the CPU 3 related to the EVF display operation of the digital camera 1 in this embodiment. In Figure 6, steps that perform the same operations as in the first embodiment are denoted by the same reference numerals and their explanations are omitted.
[0061] In S202, the image generation execution determination unit 303 obtains information regarding the currently set shooting mode from the memory unit 4.
[0062] In the following S211, the image generation execution determination unit 303 determines whether the currently set shooting mode is a shooting mode in which the exposure time is determined before shooting. For example, shooting modes in which the exposure time may change depending on the scene conditions after shooting has started, or shooting modes in which the exposure time is undetermined, such as bulb mode, are shooting modes in which the exposure time is not determined before shooting. On the other hand, for example, shutter speed priority mode and manual mode are shooting modes in which the exposure time is determined before shooting.
[0063] The image generation execution determination unit 303 executes S103 if it determines that the currently set shooting mode is a shooting mode in which the exposure time is determined before shooting, and executes S104 if it does not determine that. In other words, the image generation execution determination unit 303 decides to generate a predicted image with the image generation unit 304 if the currently set shooting mode is a shooting mode in which the exposure time is not determined before shooting. Also, if the currently set shooting mode is a shooting mode in which the exposure time is determined before shooting, the image generation execution determination unit 303 decides whether or not to generate a predicted image with the image generation unit 304 according to the length of the exposure time.
[0064] Other steps are the same as in the first embodiment and will therefore not be described. In the first embodiment and in this embodiment, the length of the set exposure time was evaluated in S103. However, if there is a shooting mode (referred to as the first shooting mode) that is determined to use an exposure time longer than a predetermined time, the length of the exposure time may be evaluated indirectly by determining whether or not the first shooting mode is set. Examples of the first shooting mode include shooting modes that use a slow shutter, such as night scene mode and slow sync mode.
[0065] According to this embodiment, if it is not possible to determine at the start of shooting whether the exposure time is longer than a predetermined time, a predictive image is generated, thereby suppressing blackout of the EVF display even if the exposure time is longer than the predetermined time.
[0066] ●(Third embodiment) Next, a third embodiment of the present invention will be described. In this embodiment, a predictive image is generated by superimposing the positional information of the subject.
[0067] Figure 7 shows a block diagram illustrating an example of the functional configuration of the digital camera 1 according to this embodiment. Configurations similar to those in the first embodiment are denoted by the same reference numerals as in Figure 2, and their explanations are omitted. Characteristic functions performed by the image processing unit 306 in this embodiment are shown as individual functional blocks. Note that not all functions performed by the image processing unit 306 are described as functional blocks; the image processing unit 306 is still capable of performing the various functions described above. The operation of the functional blocks within the image processing unit 306 is actually performed primarily by the CPU 3.
[0068] The subject detection unit 311 detects the main subject from the captured image according to the shooting scene. Specifically, this subject is a car in a car sports scene, or a human body in a sports scene. The subject detection unit 311 stores information indicating the position and size of the detected subject's area in the RAM of the memory unit 4. The subject detection unit 311 also detects movement information (direction of movement and speed) for each detected subject based on the position of the subject in frame images taken at different times.
[0069] The main subject determination unit 307 determines one subject as the main subject from among the subjects detected by the subject detection unit 306. If multiple subjects of the same type exist, the main subject can be determined using known methods, such as considering size and position within the screen. In this embodiment, a predicted image is generated by superimposing the positional information of the main subject.
[0070] The estimated movement range calculation unit 308 uses the position and movement information detected by the subject detection unit 306 for the area of the main subject determined by the main subject determination unit 307 to identify the area in the predicted image generated by the image generation unit 304 where the main subject is estimated to actually exist. Based on the past changes in the direction and speed of movement of the main subject, the estimated movement range calculation unit 308 calculates the movement probability of the main subject for each coordinate within a predetermined range in the image. Then, the estimated movement range calculation unit 308 identifies the range where the calculated movement probability is greater than or equal to a predetermined value as the estimated range.
[0071] The movement trajectory of the main subject can be predicted using polynomial or exponential approximation based on the historical position of the main subject's region. For prediction, only an approximation model with small squared error may be used, or multiple approximation models may be used.
[0072] Furthermore, machine learning models capable of trajectory prediction, such as LSTM (Long Short-Term Memory), may be used. When using LSTM, training is performed using the same training dataset as the machine learning model used to generate predicted images. Specifically, the history of the position (image coordinates) of the main subject area detected by the subject detection unit 306 can be used, with positions at multiple past time points used as training data and positions at time points later than the training data used as target data. In addition, multiple movement trajectories may be predicted using multiple LSTMs with different weight parameters. When multiple movement trajectories are predicted, the estimated movement range calculation unit 308 calculates an estimated range corresponding to each individual movement trajectory.
[0073] The image generation unit 304 generates a predicted image, and the subject detection unit 306 detects the position of the main subject. The estimated movement range calculation unit 308 calculates the squared error between the position of the main subject in the predicted image and the estimated position calculated from the predicted movement trajectory. The estimated movement range calculation unit 308 calculates the standard deviation from the time-series information of the squared error, and assuming that the squared error follows a normal distribution, the range of estimated positions in which the squared error falls within a range less than or equal to a predetermined value is defined as the estimated movement range. The position of the subject region may be the centroid position, or the position of one vertex or the intersection of the diagonals of a rectangular region inscribed with the subject region.
[0074] The subject position prediction unit 309 acquires information on the position and movement of the main subject region determined by the main subject determination unit 307, which has been previously detected by the subject detection unit 306, and the position of the subject region in the predicted image generated by the image generation unit 304. The subject position prediction unit 309 then predicts the movement trajectory of the main subject within the image from a time corresponding to the predicted image generated by the image generation unit 304 to a predetermined time thereafter.
[0075] The image superimposition unit 310 superimposes at least one of the following onto the predicted image generated by the image generation unit 304: an index indicating the estimated range identified by the estimated movement range calculation unit 308, and an index indicating the movement trajectory predicted by the subject position prediction unit 309, and outputs it to the display control unit 305.
[0076] Figure 8 is a flowchart illustrating the operation of the CPU 3 in relation to the EVF display operation of the digital camera 1 according to this embodiment, and corresponds to the operation when an indicator showing the estimated range is superimposed on the predicted image. The CPU 3 continuously executes the operations shown in the flowchart of Figure 8 at a period corresponding to the display frame rate of the EVF, in parallel with other operations such as the operation for taking still images. In Figure 8, the same reference numerals as in Figure 4 are used for the steps that perform the same operations as in the first embodiment, and the explanation is omitted.
[0077] Although Figure 8 shows that the operations from S305 onwards are executed after the image generation unit 304 generates a predicted image in S104, in practice, the processes from S305 onwards may be performed in parallel with the generation of the predicted image in S104.
[0078] In S305, CPU3 determines whether a subject has been detected in the EVF display image generated immediately before still image capture begins. If it is determined that a subject has been detected, it executes S306; otherwise, it executes S320. CPU3 can perform this determination by, for example, referring to the history of subject detection results stored in the RAM of memory unit 4.
[0079] In S306, the estimated movement range calculation unit acquires the information of the main subject determined by the main subject determination unit 307 and the history of the movement information (direction of movement and speed) of the main subject detected by the subject detection unit 311.
[0080] In S307, the estimated movement range calculation unit 308 calculates the estimated range of the position in the image where the main subject actually exists, based on the position of the main subject in the predicted image generated by the image generation unit 304 and the history of the position of the main subject region detected before the start of shooting.
[0081] S308 is executed when multiple estimated ranges are calculated, for example, when multiple movement trajectories are calculated. The estimated movement range calculation unit 308 extracts the estimated range with the highest probability from among the multiple estimated ranges. If only one estimated range is calculated, S308 is skipped.
[0082] In S309, the estimated movement range calculation unit 308 determines whether the estimated range is larger than a predetermined size. The CPU 3 executes S310 if it determines that the estimated range is larger than the predetermined size, and S320 otherwise. The size may be, for example, the number of pixels (area). If the estimated range is not larger than the predetermined size, it means that the accuracy of the position of the main subject in the predicted image is high.
[0083] In S320, the image superposition unit 310 outputs the predicted image generated by the image generation unit 304 in S104 directly to the display control unit 305. As a result, the predicted image is displayed on the display element 6 via the display control unit 305, and processing for the next frame begins.
[0084] If the estimated range is not large, the accuracy of the main subject's position in the predicted image is considered high, so the estimated range indicator is not superimposed on the predicted image. By not providing users with unnecessary information, usability can be improved.
[0085] In S310, the image superimposition unit 310 superimposes an index representing the estimated range calculated by the estimated movement range calculation unit 308 onto the predicted image generated by the image generation unit 304 and outputs it to the display control unit 305. The index may be, for example, an image showing the contour of the estimated range, but other display patterns may also be used.
[0086] In S311, the display control unit 305 displays a prediction image on the display element 6, with an index indicating the estimated range superimposed on it. Then the CPU 3 starts processing the next frame.
[0087] Figure 9 is a flowchart illustrating the operation of the CPU 3 in relation to another EVF display operation of the digital camera 1 according to this embodiment, and corresponds to the operation when an indicator showing the predicted position of the main subject is superimposed on the predicted image. The CPU 3 continuously executes the operations shown in the flowchart of Figure 9 at a period corresponding to the EVF display frame rate, in parallel with other operations such as operations for taking still images. In Figure 9, the steps that perform the same operations as in the first embodiment are referred to in Figure 4, and the steps that perform the same operations as in Figure 8 are referred to in the same reference numbers as in Figure 8, and their explanations are omitted. Only the steps specific to Figure 9 will be explained.
[0088] In S407, the subject position prediction unit 309 uses the position and movement information detected by the subject detection unit 306 for the area of the main subject to predict the position of the main subject in the image at a time later than the time corresponding to the predicted image generated by the image generation unit 304. Then, the subject position prediction unit 309 uses the past position and movement information of the main subject and the predicted position to calculate the movement trajectory of the main subject in the same manner as when calculating the estimated range. However, the movement trajectory calculated here is the movement trajectory from the time corresponding to the predicted image to a predetermined time thereafter.
[0089] In S408, the image superimposition unit 310 superimposes an index representing the movement trajectory calculated by the subject position prediction unit 309 onto the predicted image generated by the image generation unit 304 and outputs it to the display control unit 305. The index may be, for example, a linear image showing the movement trajectory, but other display patterns may also be used.
[0090] In S409, the display control unit 305 displays a predicted image on the display element 6, with an index indicating the predicted movement trajectory of the main subject superimposed on it. Then the CPU 3 starts processing the next frame.
[0091] Figures 8 and 9 illustrate the case where either the estimated range or the predicted movement trajectory is superimposed on the predicted image, but both may be superimposed. In that case, the processing from S307 onwards in Figure 8 and the processing from S407 onwards in Figure 9 should be executed in parallel. In this case, the image superimposition unit 310 always superimposes the indicator for the predicted movement trajectory, and only superimposes the indicator for the estimated range if its size is less than or equal to a predetermined size.
[0092] Figure 10 is a diagram that adds examples to Figure 3, showing a predicted image with an index of the estimated range in which the main subject actually exists superimposed, and a predicted image with an index of the predicted future movement trajectory of the main subject superimposed.
[0093] The estimated range increases as the accuracy of the predicted image decreases. Therefore, it becomes larger the longer the time elapsed since the start of predictive image generation. However, since the photographer knows the possible location of the main subject, they can frame the shot while considering the estimated range.
[0094] Furthermore, by superimposing an indicator showing the predicted future movement trajectory of the main subject, it serves as a guide when panning the camera during panning shots, improving usability.
[0095] Although this explanation assumes the configuration of the first embodiment, the same indicators can be superimposed on the predicted image in the configuration of the second embodiment as well.
[0096] According to this embodiment, in addition to the effects of the first and second embodiments, it is possible to achieve the effect of supporting the photographer's camera operation during periods when images for EVF display cannot be captured.
[0097] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0098] The disclosure of this embodiment includes the following imaging device and its control method, as well as a program. (Item 1) Image sensor and Display means and Image processing means for generating a display image for display on the display means from an image captured by the image sensor, A generation means that uses a machine learning model to generate a predicted image that predicts an image to be taken at a time later than the time the images of the multiple frames captured by the image sensor are taken, A display control means that causes the display means to function as an electronic viewfinder by displaying the display image or the prediction image on the display means, The generation means includes a determination means for determining whether or not to generate the predicted image, The determination means determines that the generation means will generate the predicted image if it is determined that the exposure time during still image capture is longer than a predetermined time, and determines that the generation means will not generate the predicted image if it is determined that the exposure time is not longer than the predetermined time. An imaging device characterized by the following features. (Item 2) The imaging device according to item 1, characterized in that the display control means displays the predicted image on the display means when still image capture is in progress, and the display image on the display means when still image capture is not in progress. (Item 3) The determination means further determines whether the shooting mode set in the imaging device is a shooting mode in which the exposure time for still image shooting is determined before shooting, If the shooting mode set in the imaging device is determined to be a shooting mode in which the exposure time for still image shooting is determined before shooting, then it is determined that the generation means will generate the predicted image if the exposure time is determined to be longer than a predetermined time, and that the generation means will not generate the predicted image if the exposure time is not determined to be longer than the predetermined time. If the shooting mode set in the imaging device is not determined to be a shooting mode in which the exposure time for still image shooting is determined before shooting, then the generation means is determined to generate the predicted image. The imaging device according to item 1 or 2, characterized in that it is an imaging device. (Item 4) The imaging device according to item 1 or 2, characterized in that the determination means determines that the exposure time for still image capture is longer than the predetermined time when a shooting mode for capturing images with an exposure time longer than the predetermined time is set in the imaging device. (Item 5) The system further includes an acquisition means for acquiring information regarding the movement of the imaging device, The imaging apparatus according to any one of items 1 to 4, characterized in that the generation means generates the predicted image from a plurality of images captured by the image sensor and information regarding the movement of the imaging apparatus when the plurality of images were captured. (Item 6) A detection means for detecting a specific subject from an image, A calculation means that calculates an estimated range in which the probability of the subject being present in the predicted image is greater than or equal to a predetermined value, based on the position of the subject detected in the multiple frames of the image, A superimposing means that superimposes an index indicating the estimated range onto the predicted image and outputs it to the display control means, The imaging device according to any one of items 1 to 5, further comprising the above. (Item 7) The imaging apparatus according to item 6, characterized in that the superimposing means outputs the predicted image without superimposing an index indicating the estimated range to the display control means when the size of the estimated range is not larger than a predetermined size. (Item 8) A detection means for detecting a specific subject from an image, A prediction means predicts the movement trajectory of the subject from a later time to a later time, based on the position of the subject detected with respect to the multiple frames of images and the predicted image. A superimposing means that superimposes an index indicating the movement trajectory onto the predicted image and outputs it to the display control means, The imaging device according to any one of items 1 to 7, further comprising the above. (Item 9) Image sensor and Display means and A control method performed by an imaging device having a generation means that uses a machine learning model to generate a predicted image that predicts an image to be taken at a time later than the time the images of multiple frames were taken, from a plurality of images captured by the image sensor, To generate a display image for display on the display means from the image captured by the image sensor, By displaying the aforementioned display image or the aforementioned prediction image on the display means, the display means is made to function as an electronic viewfinder. When it is determined that the exposure time during still image capture is longer than a predetermined time, the generation means is determined to generate the predicted image, and when it is determined that the exposure time is not longer than the predetermined time, the generation means is determined not to generate the predicted image. A control method for an imaging device, characterized by having the following features. (Item 10) A program to cause the computer of the imaging device to function as each of the means of the imaging device described in any one of items 1 to 8, excluding the display means.
[0099] The present invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]
[0100] 1A: Shooting lens, 1B: Camera body, 2: Image sensor, 3: CPU, 4: Memory unit, 5: Eyepiece, 6: Display element, 7: Shutter release button, 8: Control unit
Claims
1. Image sensor and Display means and Image processing means for generating a display image for display on the display means from an image captured by the image sensor, A generation means that uses a machine learning model to generate a predicted image that predicts an image to be taken at a time later than the time the images of the multiple frames captured by the image sensor are taken, A display control means that causes the display means to function as an electronic viewfinder by displaying the display image or the prediction image on the display means, The generation means includes a determination means for determining whether or not to generate the predicted image, The determination means determines that the generation means will generate the predicted image if it is determined that the exposure time during still image capture is longer than a predetermined time, and determines that the generation means will not generate the predicted image if it is determined that the exposure time is not longer than the predetermined time. An imaging device characterized by the following features.
2. The imaging apparatus according to claim 1, characterized in that the display control means displays the predicted image on the display means when still image capture is in progress, and the display image on the display means when still image capture is not in progress.
3. The determination means further determines whether the shooting mode set in the imaging device is a shooting mode in which the exposure time for still image shooting is determined before shooting, If the shooting mode set in the imaging device is determined to be a shooting mode in which the exposure time for still image shooting is determined before shooting, then it is determined that the generation means will generate the predicted image if the exposure time is determined to be longer than a predetermined time, and that the generation means will not generate the predicted image if the exposure time is not determined to be longer than the predetermined time. If the shooting mode set in the imaging device is not determined to be a shooting mode in which the exposure time for still image shooting is determined before shooting, then the generation means is determined to generate the predicted image. The imaging apparatus according to feature 1.
4. The imaging device according to claim 1, characterized in that the determination means determines that the exposure time for still image capture is longer than the predetermined time when a shooting mode for capturing images with an exposure time longer than the predetermined time is set in the imaging device.
5. The system further includes an acquisition means for acquiring information regarding the movement of the imaging device, The imaging apparatus according to claim 1, characterized in that the generation means generates the predicted image from a plurality of images captured by the image sensor and information regarding the movement of the imaging apparatus when the plurality of images were captured.
6. A detection means for detecting a specific subject from an image, A calculation means that calculates an estimated range in which the probability of the subject being present in the predicted image is greater than or equal to a predetermined value, based on the position of the subject detected in the multiple frames of the image, A superimposing means that superimposes an index indicating the estimated range onto the predicted image and outputs it to the display control means, The imaging apparatus according to claim 1, further comprising the following:
7. The imaging apparatus according to claim 6, characterized in that the superimposing means outputs the predicted image without superimposing an index indicating the estimated range to the display control means when the size of the estimated range is not larger than a predetermined size.
8. A detection means for detecting a specific subject from an image, A prediction means predicts the movement trajectory of the subject from a later time to a later time, based on the position of the subject detected with respect to the multiple frames of images and the predicted image. A superimposing means that superimposes an index indicating the movement trajectory onto the predicted image and outputs it to the display control means, The imaging apparatus according to claim 1, further comprising the following:
9. Image sensor and Display means and A control method performed by an imaging device having a generation means that uses a machine learning model to generate a predicted image that predicts an image to be taken at a time later than the time the images of multiple frames were taken, from a plurality of images captured by the image sensor, To generate a display image for display on the display means from the image captured by the image sensor, By displaying the aforementioned display image or the aforementioned prediction image on the display means, the display means is made to function as an electronic viewfinder. When it is determined that the exposure time during still image capture is longer than a predetermined time, the generation means is determined to generate the predicted image, and when it is determined that the exposure time is not longer than the predetermined time, the generation means is determined not to generate the predicted image. A control method for an imaging device, characterized by having the following features.
10. A program for causing the computer of an imaging device to function as each of the means of the imaging device according to any one of claims 1 to 8, excluding the display means.