Imaging device, method for controlling imaging device, and program

WO2026163538A1PCT designated stage Publication Date: 2026-08-06CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2025-10-30
Publication Date
2026-08-06

Smart Images

  • Figure JP2025038079_06082026_PF_FP_ABST
    Figure JP2025038079_06082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an imaging device that imparts a blur effect to a moving image. An imaging device (100) for capturing a moving image comprises: an intention input unit (107) that receives setting of an imaging style by a user; a scene analysis unit (106) that analyzes a scene of the moving image being captured by an imaging unit; a correction target amount generation unit 108 that generates a blur to be imparted, which is a blur to be imparted to the moving image; a correction amount calculation unit 110 that calculates a blur amount on the basis of the blur to be imparted acquired from the correction target amount generation unit (108) and shaking of the imaging device (100); and a hand movement correction control unit (112) that controls a blur correction unit on the basis of the blur amount. The correction target amount generation unit (108) generates the blur to be imparted on the basis of the imaging style and an analysis result from the scene analysis unit (106).
Need to check novelty before this filing date? Find Prior Art

Description

Imaging device, control method for imaging device, and program

[0001] The present invention relates to an imaging device that performs shake correction.

[0002] Conventionally, there is an imaging device having a shake correction function. Shake correction is a technique for suppressing hand shake that occurs during shooting of an imaging device and obtaining a clear and stable video. Conventional shake correction techniques mainly use optical or electronic methods to detect the shake of the camera during shooting and correct it to achieve an "image without blur". In conventional shake correction techniques, shake correction during shooting aims to obtain an "image without blur", and it is common to set a process for reducing the amount of shake.

[0003] On the other hand, a video with hand shake is said to be able to change the impression and emotion given to viewers. These effects can effectively change the impression given to viewers depending on the combination of the shake pattern and intensity. For example, slight shake can enhance the sense of presence, and intense shake can cause a sense of tension or uneasiness. Patent Document 1 discloses an image processing device that reproduces an image that can give a sense of presence to viewers by adding the reproduction of the shake information of the imaging device during shooting as an inverse stabilization process to the image.

[0004] Japanese Patent No. 6740900

[0005] However, although the reproduction of hand shake has the effect of giving a sense of presence or tension, the visual effect is limited only by the reproduction of hand shake, and the hand shake addition by the inverse stabilization process in Patent Document 1 is insufficient as a video effect. For example, by only reproducing a part of the hand shake during shooting, it is not possible to add blur that did not occur during shooting. Therefore, there is a need for an imaging device that maximally extracts the visual effect by adding the effect of hand shake suitable for the scene during shooting.

[0006] An object of the present invention is to provide an imaging device that imparts a blur effect to a moving image.

[0007] To solve the above problems, the present invention provides an imaging device for capturing moving images, comprising: an input unit that accepts the user's setting of a shooting style; a scene analysis unit that analyzes the scene of the moving image being captured by the imaging unit; a generation unit that generates added blur, which is blur to be added to the moving image; a calculation unit that calculates the amount of blur based on the added blur obtained from the generation unit and the vibration of the imaging device; and a correction control unit that controls a blur correction unit based on the amount of blur, wherein the added blur is generated based on the shooting style and the analysis result by the scene analysis unit.

[0008] According to the present invention, it is possible to provide an imaging device that imparts a blurring effect to moving images.

[0009] This is a diagram showing the configuration of the imaging device 100 in the first embodiment. This is a diagram showing the configuration of the scene analysis unit 106. This is a diagram showing the configuration of the motion vector calculation unit 200. This is a diagram explaining the settings. This is a diagram showing an example of a user notification method. This is a diagram explaining the target amount of blur generated by the correction target amount generation unit 108. This is a diagram explaining the relationship between the output of the gyro sensor 111 and the amount of vertical blur during shooting. This is a flowchart of the blur application process. This is a diagram showing the configuration of the imaging device 900 in the second embodiment. First Embodiment

[0010] Figure 1 shows the configuration of the imaging device 100 in the first embodiment. The imaging device 100 is an imaging device intended for capturing both still images and / or moving images. In this embodiment, the case in which the imaging device 100 captures moving images will be described as an example. The imaging device 100 is, for example, a digital single-lens reflex camera in which the main body of the imaging device and the lens device are integrated. In this embodiment, the case in which the main body of the imaging device and the lens device (lens barrel) are integrated will be described as an example, but the imaging device may also be one in which the lens device is detachable from the main body of the imaging device.

[0011] The imaging device 100 includes an optical system unit 101, an image sensor 102, an A / D conversion unit 103, a capture unit 104, a digital signal processing unit 105, a control unit 114, a gyro sensor 111, an external recording device 113, a memory 115, and an operation unit 116. The optical system unit 101 includes an optical lens group including a shift lens for vibration damping, a shutter, an aperture, a lens control unit, etc. The optical system unit 101 forms an optical image of the subject on the image sensor 102. The shift lens of the optical system unit 101 is driven based on the output of the image stabilization control unit 112, which will be described later.

[0012] The image sensor 102 is a photoelectric conversion element and is an imaging unit that converts an optical image photoelectrically and outputs an output signal (analog signal) corresponding to the optical image. The image sensor 102 is, for example, a solid image sensor in which Bayer array unit pixel cells are arranged in two dimensions. Specifically, the image sensor 102 has Bayer array unit pixel cells in which red (R), green (G), and blue (B) filters are arranged in a specific pattern. This allows each pixel to receive light of a specific color. The image sensor 102 converts the image formed via the optical system unit 101 photoelectrically and reads it out. The shutter and aperture member included in the optical system unit 101 adjust the incident time and amount of light, and control it so that an appropriate amount of light reaches the image sensor 102. The image formed via the optical system unit 101 reaches the image sensor 102 and is photoelectrically converted. Each pixel cell of the image sensor 102 generates an electric charge according to the intensity of the light it receives. The generated charges are read out sequentially and output to the A / D conversion unit 103.

[0013] The A / D conversion unit 103 converts the analog electrical signal output from the image sensor 102 into a digital signal (image data, pixel signal) and outputs it to the capture unit 104. For example, the A / D conversion unit 103 performs analog signal processing on the analog electrical signal output from the image sensor 102 in an analog signal processing unit (not shown), converts it into a digital signal, and outputs it to the capture unit 104. The capture unit 104 determines the effective area and type of the digital signal (image data) output from the A / D conversion unit 103 and outputs the effective area to the digital signal processing unit 105.

[0014] The digital signal processing unit 105 performs digital signal processing on the digital signal (image data) of the effective area acquired from the capture unit 104. The digital signal processing unit 105 performs digital signal processing such as simultaneous processing, gamma processing, and noise reduction processing on R, G, and B pixels input in Bayer array. Note that the techniques of simultaneous processing, gamma processing, and noise reduction processing are not directly related to the present invention and will not be described in detail. The image data to which digital signal processing has been applied by the digital signal processing unit 105 is output to the scene analysis unit 106 and the external recording device 113. The external recording device 113 records the image data output from the digital signal processing unit 105. The external recording device 113 is, for example, a recording medium such as an SD card or a CF card. The external recording device 113 may be detachably provided from the housing of the imaging device 100.

[0015] The control unit 114 controls the entire imaging device 100. The control unit 114 is composed of a microprocessor such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit). Alternatively, the control unit 114 may be composed of a microcontroller such as an MCU (Micro Controller Unit). The memory 115 stores programs and other data necessary for the control unit 114 to control the imaging device 100. The control unit 114 loads programs from the memory 115 and executes them to perform various controls.

[0016] The control unit 114 includes a scene analysis unit 106, an intent input unit 107, a correction target amount generation unit 108, an assigned blur type notification unit 109, a correction amount calculation unit 110, and a blur correction control unit 112. The various software modules of the control unit 114 are realized when the CPU reads and executes programs etc. from the memory 115.

[0017] The scene analysis unit 106 analyzes the scene based on the instructions of the control unit 114, setting an analysis algorithm for the video image output from the digital signal processing unit 105. The scene analysis unit 106 analyzes at least one piece of information from the captured video image in real time, including the type of scene, main object information, motion information, human body posture information, and environmental information. The scene analysis unit 106 outputs the analysis information, which is the analysis result obtained based on the analyzed video image, to the correction target amount generation unit 108. Here, the detailed configuration of the scene analysis unit 106 will be explained using Figure 2.

[0018] Figure 2 shows the configuration of the scene analysis unit 106. The scene analysis unit 106 includes an object detection unit 201, a motion vector calculation unit 200, an environmental information detection unit 202, and an analysis information integration unit 203. The object detection unit 201 detects objects in the video using deep learning technology based on the video, which is a digital signal (image data) output by the digital signal processing unit 105. Specifically, it uses a pre-trained neural network model to identify the type, posture, position, and size of important objects in the image. The detection results from the object detection unit 201 are output to the motion vector calculation unit 200 and the analysis information integration unit 203.

[0019] The training process for the deep learning model used by the object detection unit 201 is carried out as follows. First, a large amount of image data is collected, and the location and type of object are labeled for each image. A neural network model is trained using this dataset. In the training process, images are given as input, and the model predicts the location and type of the object. The error between the prediction result and the correct label is calculated, and the model parameters are updated to minimize this error. By repeating this process, the accuracy of the model is improved. For example, a CNN model can be used as a deep learning model for object detection. This model can be constructed by labeling the type of object, human body posture information, location, and size based on image data captured in various scenes. Note that the detailed algorithms and training methods of the deep learning model are not directly related to the present invention, so a description is omitted. In this embodiment, an example of detecting objects in video using deep learning technology is described, but other known methods for detecting objects in video without using deep learning technology may also be used.

[0020] The motion vector calculation unit 200 calculates the motion vector of an object based on the motion image output from the digital signal processing unit 105 and the position and size of important objects output from the object detection unit 201. The motion vector calculation unit 200 then outputs the calculated motion vector of the object to the analysis information integration unit 203. Here, the detailed configuration of the motion vector calculation unit 200 will be explained with reference to Figure 3.

[0021] Figure 3 shows the configuration of the motion vector calculation unit 200. The motion vector calculation unit 200 includes a feature extraction unit 300 and an optical flow calculation unit 301. The feature extraction unit 300 processes the image data (moving image) output from the digital signal processing unit 105 to extract the feature quantities of important objects, which are the output of the object detection unit 201. Specifically, it detects features such as edges, corners, and textures in the image, identifies the location of important objects based on these feature quantities, and outputs them to the optical flow calculation unit 301.

[0022] The optical flow calculation unit 301 calculates motion vectors between frames based on the features extracted by the feature extraction unit 300, and calculates the direction and velocity of motion, which are the positional changes of important objects. As a result, the motion vector calculation unit 200 provides basic data for accurately capturing the movement of objects and generating appropriate blur effects.

[0023] Returning to the explanation of the scene analysis unit 106 in Figure 2, the environmental information detection unit 202 uses deep learning technology to detect the shooting location and environmental information based on the image data output from the digital signal processing unit 105. Location information refers to information about the shooting location, such as a live music venue, a fighting ring, a road, or a park. Environmental information refers to information about the shooting environment, such as brightness, lighting intensity, and weather. For example, a CNN model can be used as the deep learning model for detecting environmental information. A deep learning model for detecting environmental information can be constructed by labeling brightness, weather, and background type based on image data taken in various scenes. In this embodiment, an example of detecting the shooting location and environmental information using deep learning technology is described, but other known methods for detecting the shooting location and environmental information without using deep learning technology may also be used.

[0024] The analysis information integration unit 203 detects scene information using deep learning technology based on the types of important objects in the image, human posture, position, size, direction and speed of movement, location information, and environmental information. Scene information includes at least one of the following: scene type, main object information, movement information, human posture information, and environmental information. For the deep learning model that detects environmental information, models that are robust to time-series data, such as RNN or LSTM, can be used. The deep learning model that detects environmental information can be constructed by determining the scene based on the input object information, location information, and environmental information, and then labeling it.

[0025] The scene information output by the analysis information integration unit 203 includes, for example, the following information: • Scene type: Fighting scene • Main object information: Opponents • Movement information: Speed ​​at which opponents fall • Human posture information: Opponent's falling posture (back to the ground) • Environment information: Medium brightness

[0026] Let's return to the explanation of the configuration of the imaging device 100 in Figure 1. The operation unit 116 displays information to the user and accepts operations from the user. The operation unit 116 may consist of, for example, a display and buttons for operation, or it may consist of a touch panel. By associating input coordinates and display coordinates on the touch panel, a GUI can be configured that makes it appear as if the user can directly operate the screen displayed on the touch panel.

[0027] The assigned blur type notification unit 109 functions as a display unit that controls the display of notifications on the operation unit 116's display based on the output from the correction target amount generation unit 108. The intention input unit 107 receives the user's selection input on the operation unit 116 and outputs the user's selection to the correction target amount generation unit 108. In this embodiment, in order to receive the user's selection of shooting style and blur type, the operation unit 116 displays the shooting style options and the blur type options, respectively.

[0028] First, let's explain the selection of a shooting style. The shooting style includes at least one of the content genre or the impression you want to give. The blur type notification unit 109 displays the shooting style on the operation unit 116. The intention input unit 107 accepts the user's selection of the shooting style. The intention input unit 107 then outputs the shooting style entered by the user to the correction target amount generation unit 108. Here, we will explain the specific options that the intention input unit 107 accepts using Figure 4.

[0029] Figure 4 is a diagram illustrating the settings. Figures 4(A) and 4(B) are examples of shooting style options. Figure 4(A) shows the genre of video content or streamed video content that the user wants to create. The options for the genre of video content or streamed video content are displayed on the operation unit 116, and the user selects a shooting genre from the options. The intent input unit 107 accepts the genre selection by the user and outputs it as a shooting style to the correction target amount generation unit 108.

[0030] The options for video content or streamed video content genres include, for example, at least one of the following: documentary, live, martial arts, news, and sports. The options may be only a selection of these, or other genres may be included. For example, within the documentary category, subgenres such as action-oriented perspective, vehicle, and interviews may be included as options. Alternatively, the genre options may be set based on the analysis results of the scene analysis unit 106. If documentary is selected, the desired camera shake effect should prioritize realism and stable footage. Therefore, for documentaries, a slight shake, similar to handheld footage, is desirable to enhance realism. On the other hand, if live or martial arts is selected, viewers will want a presentation that conveys the excitement and energy of the event. Therefore, for live footage, a shake that matches the surrounding movements or the performers is desirable. For martial arts, a shake effect that simulates a large camera shake when a fighter falls is desirable. As the blurring effect varies depending on the content genre, the correction target amount generation unit 108 acquires information about the content genre via the operation unit 116 and the intent input unit 107.

[0031] Furthermore, since the desired image stabilization effect may vary even within the same genre of scene, it may be possible to more directly accept instructions for the impression (visual effect) to be applied to the video in order to obtain content that better reflects the user's intentions. Figure 4(B) shows an example of the impression to be applied to video content or streamed footage. The options for the impression to be applied to the video content or streamed footage are displayed on the operation unit 116 by the applied blur type notification unit 109, and the user selects the option for the impression they wish to create from among the choices. The intention input unit 107 accepts the user's selection of the visual effect to be applied to the video and outputs it to the correction target amount generation unit 108.

[0032] The visual effects that can be added to video content or streamed footage include, for example, a sense of realism, excitement / passion, tension / thrill, and calmness. The options may be only a part of these, or other visual effects may be included as options. Alternatively, the options for the impression to be conveyed may be set based on the analysis results of the scene analysis unit 106. If "sense of realism" is set as the visual effect, a visual effect that easily conveys a sense of realism can be obtained by adding camera shake, such as adding a stable, still state before the start of a fight, slight shaking when the match begins, and large shaking when the character falls down, such as during a knockdown. On the other hand, if "calm" is selected as the visual effect, a visual effect that is gentler in shaking can be obtained by adding camera shake, such as adding slight shaking like that of a handheld camera to increase realism.

[0033] The correction target amount generation unit 108 generates added blur, which is the amount of blur to be added to the moving image, based on the shooting style acquired from the intention input unit 107 and the scene information acquired from the scene analysis unit 106. When generating added blur, the correction target amount generation unit 108 generates two or more added blurs. Then, the correction target amount generation unit 108 outputs the type of blur corresponding to each of the generated added blurs to the added blur type notification unit 109.

[0034] The assigned blur type notification unit 109 presents the user with candidates for the blur to be assigned by displaying a plurality of blur types to be assigned, which are the output of the correction target amount generation unit 108, on the operation unit 116. The assigned blur type notification unit 109 superimposes the candidates for the blur type to be assigned on the live viewing screen, for example, on which the captured image is displayed in real time. An example of the types of blur that the assigned blur type notification unit 109 displays on the operation unit 116 will be explained using Figure 4(C).

[0035] Figure 4(C) shows examples of the types of blur that can be added. The types of blur that can be added include, for example, audience sway, floating sway, large sway, walking sway, stillness, small sway, and rhythmic. Note that the types of blur that can be added may be only a part of these, or other types of blur that can be added may be included in the options. In addition, the options for the types of blur that can be added may be set based on the analysis results of the scene analysis unit 106. The user selects the desired type of blur from among the multiple types of blur that correspond to the multiple types of added blur generated by the correction target amount generation unit 108. The intention input unit 107 accepts the user's selection of the type of blur and outputs the selected type of blur to the correction target amount generation unit 108. Note that the added blur and the types of blur may be associated and managed in advance, or the association between added blur and types of blur may be performed by machine learning or the like.

[0036] Here, we will explain an example of a screen displayed during shooting. Figure 5 shows an example of a shooting screen in a sports scene. On the shooting screen, the intention input options 500 and the assigned blur type notification area 501 are superimposed on the real-time captured image 502. The intention input options 500 is an area where the user selects the content genre they want to create. The intention input options 500 displays the intention input options 500, accepts the user's selection for the intention input options 500, and outputs the selection result to the correction target amount generation unit 108.

[0037] The assigned blur type notification area 501 is an area that displays the type of assigned blur generated by the correction target amount generation unit 108. The type of assigned blur is determined by the correction target amount generation unit 108 based on the content genre selected by the user and the output of the scene analysis unit 106. In the assigned blur type notification area 501, the type of blur selected by the user and currently assigned to the video is highlighted prominently at the top, as shown in Figure 5, "Audience Blur". The highlighting may be done using color or by changing the size of the display. On the other hand, types of blur that are not currently assigned to the video are displayed as selection candidates. The user checks the type of assigned blur currently assigned to the image in the assigned blur type notification area 501, and if they want to change the type of blur to be assigned, they select and specify the type of blur to be assigned from the candidate types of blur to be assigned. When the user selects a type of blur to be assigned in the assigned blur type notification area 501, the intention input unit 107 accepts the user's selection and outputs the selected type of blur to be assigned to the correction target amount generation unit 108.

[0038] Here, using a sports scene as an example, we will explain the candidate types of added blur displayed on the operation unit 116. The intention input option 500 displays a list of content types (genres) of video content or streamed video content shown in Figure 4(A). The user selects "Sports" from the intention input option 500. The intention input unit 107 accepts the user's selection and outputs to the correction target amount generation unit 108 that "Sports" has been selected as the shooting style. The correction target amount generation unit 108 generates multiple types of added blur based on the selected shooting style and the analysis results of the scene analysis unit 106. In the example shown in Figure 5, the types of blur corresponding to each of the generated multiple types of added blur are displayed in the added blur type notification area 501 via the added blur type notification unit 109: "Audience shake," "Silent," "Large shake," and "Slight shake." The user confirms the types of blur to be added in the added blur type notification area 501 and selects the desired type of blur from the candidate types of blur. The intent input unit 107 accepts the user's selection of the type of blur and outputs it to the correction target amount generation unit 108. The correction target amount generation unit 108, having received the selection of the type of blur, outputs the amount of blur to be applied corresponding to the selected type of blur to the correction amount calculation unit 110 as the target amount of blur to be applied to the moving image. For example, if the user selects "audience sway," a sway synchronized with the rhythm of the cheering in the surrounding area will be applied.

[0039] In this embodiment, the correction target amount generation unit 108 notifies the user of the type of shake corresponding to the generated assigned shake as an assigned shake candidate, and outputs the assigned shake corresponding to the type of shake selected by the user to the correction amount calculation unit 110. However, this is not the only example. For example, the correction target amount generation unit 108 may generate two or more assigned shakes and output the assigned shake corresponding to the shake type with the most past usage history to the correction amount calculation unit 110. Alternatively, the correction target amount generation unit 108 may generate two or more assigned shakes and output the assigned shake with the smallest integral value of the absolute values ​​of the assigned shake amounts to the correction amount calculation unit 110. By automatically selecting the assigned shake to output to the correction amount calculation unit 110, the correction target amount generation unit 108 can apply appropriate shake according to the scene, even in rapidly changing scenes such as sports scenes. For example, at the moment a goal is scored in soccer, the correction target amount generation unit 108 generates a blur with the "large blur" attribute based on the selected content genre and the scene analysis unit 106, which identifies it as a goal scene. In this case, "large blur" is displayed at the top of the assigned blur type notification area 501. In either case, when determining the assigned blur to be output to the correction amount calculation unit 110 based on past usage history or based on the integral value of the absolute value of the blur amount, the correction target amount generation unit 108 displays multiple types of blur in the assigned blur type notification area 501. If the user selects a type of assigned blur different from the type of assigned blur determined by the correction target amount generation unit 108, the user's selection takes precedence, and the assigned blur corresponding to the type of blur selected by the user is output to the correction amount calculation unit 110.

[0040] Here, an example of the configuration of the correction target amount generation unit 108 will be described. The correction target amount generation unit 108 generates blur information to be added, for example, by using multiple trained models obtained by machine learning. The correction target amount generation unit 108 has an inference machine based on an artificial intelligence model. The artificial intelligence model is generated by machine learning, taking the shooting style and analysis results as input and using the type of blur information and the blur to be added to the video as training data.

[0041] The correction target quantity generation unit 108 has a feature extraction layer, an integration layer, and a generation layer. The feature extraction layer analyzes the features and motion patterns of the scene using a time-series reference model to extract important features from the input data and extracts feature quantities. Examples of time-series reference models include Long Short Term Memory (LSTM) and Recurrent Neural Network (RNN). The integration layer is a network for fusing different features within the feature quantities, and uses a time-series reference model such as LSTM or RNN to integrate the dynamic changes of the scene as a high-dimensional feature vector. The generation layer generates and outputs blur information to which the high-dimensional feature vector is added using methods such as generative adversarial networks (GAN) or variational autoencoders (VAE). These models can be constructed by labeling the type of blur added by content genre for scene judgment results based on a large amount of video data containing diverse scenes and motion patterns. More detailed algorithms and learning methods for deep learning models are not directly related to the present invention and will therefore not be explained. In this embodiment, an example in which the correction target amount generation unit 108 is composed of a feature extraction layer, an integration layer, and a generation layer has been described, but it is not limited to this. Any method that can accurately generate blur information that matches the scene and the photographer's requirements can be used for generating the blur information to be added by the correction target amount generation unit 108.

[0042] Here, an example of the shake information to be imparted will be described with reference to FIG. 6. The correction target amount generation unit 108 generates a target amount of shake as the shake to be imparted. FIG. 6 is a diagram for explaining the target amount of shake generated by the correction target amount generation unit 108. In FIG. 6, the vertical axis represents the target amount of shake, and the horizontal axis represents time. FIG. 6(A) shows the amount of shake when the user holds the imaging device 100 and takes a picture. In the present embodiment, as the "minute shake" in FIG. 2(C), the shake when the user holds the imaging device 100 and takes a picture, shown in FIG. 6(A), is imparted. FIG. 6(B) shows the amount of shake when the user walks while holding the imaging device 100. In the present embodiment, as the "walking shake" in FIG. 2(C), the shake when the user walks while holding the imaging device 100, shown in FIG. 6(B), is imparted. FIG. 6(C) shows the amount of shake when the user fixes the imaging device 100 to a stabilizing device such as a gimbal and walks. In the present embodiment, as the "floating shake" in FIG. 2(C), the shake when the user fixes the imaging device 100 to a stabilizing device such as a gimbal and walks, shown in FIG. 6(C), is imparted.

[0043] Thus, the correction target amount generation unit 108 outputs the target amount of shake corresponding to the type of shake to be imparted to the correction amount calculation unit 110. Note that in a normal imaging device that does not impart shake such as hand shake as a video expression, these correction target amounts are zero, or if it is during a panning operation, the vector amount in the operation direction is notified.

[0044] The correction amount calculation unit 110 calculates the amount of shake of the shift lens that maximally expresses the intended shake effect based on the outputs of the correction target amount generation unit 108 and the gyro sensor 111, and outputs it to the shake correction control unit 112. That is, in the present embodiment, instead of calculating the image shake correction amount for correcting the image blur that occurs in the moving image, the amount of shake for imparting the desired shake to the moving image is calculated. The correction amount calculation unit 110 calculates the amount of shake for driving the shift lens, which is the shake correction unit, based on the imparted shake, which is the output of the correction target amount generation unit 108, and the shake of the imaging device 100, which is the output of the gyro sensor 111.

[0045] Figure 7 illustrates the relationship between the output of the gyro sensor 111 and the amount of vertical shake during shooting. In Figure 7, the output of the gyro sensor 111 is converted into the amount of vertical shake during shooting and shown. The correction amount is obtained by subtracting the correction target amount corresponding to the output of the gyro sensor 111 from the applied shake information output by the correction target amount generation unit 108. Therefore, the correction amount calculation unit 110 calculates the amount of shift lens shake that expresses the intended shake effect to the maximum extent by subtracting the correction target amount corresponding to the output of the gyro sensor 111 from the applied shake output by the correction target amount generation unit 108.

[0046] The gyro sensor 111 is a measuring device for measuring the angular velocity of the imaging device 100. The output of the gyro sensor 111 indicates the shake of the imaging device 100. The shake correction control unit 112 applies the intended shake by controlling the shake correction unit used for image shake correction based on the shake amount, which is the output of the correction amount calculation unit 110. The shake correction unit is a combination of one or more correction units, including an electronic correction unit, an optical correction unit that drives a shift lens, and an optical correction unit that drives the image sensor 102. Electronic correction is, for example, correction of shake by cropping the image. Optical correction is correction by driving the shift lens or the image sensor 102 toward a plane perpendicular to the optical axis of the optical system unit 101. In this embodiment, the case in which the shift lens of the optical system unit 101 is controlled as the shake correction unit will be explained as an example. The image stabilization control unit 112 drives and controls the shift lens, which is the image stabilization unit, not to correct image blur occurring in the moving image, but to impart a desired blur to the moving image. The image stabilization unit may be a pan-tilt mechanism having a rotation axis in the pan direction and a rotation axis in the tilt direction, allowing rotation in the pan and tilt directions, or a gimbal mechanism capable of rotation in three directions: pan, tilt, and roll.

[0047] Next, the operation of the imaging device 100 will be explained using the flowchart in Figure 8. Figure 8 is a flowchart of the blur application process. Each process performed by the imaging device 100 shown in Figure 8 is executed by the CPU of the control unit 114 according to a program stored in the memory of the imaging device 100.

[0048] In step S800, the intention input unit 107 receives an input of the content genre, which is the shooting type by the user. The user selects a shooting style from the intention input options 500 displayed on the operation unit 116. The intention input unit 107 acquires, for example, the content genre selected by the user as the shooting style selected by the user. The intention input unit 107 acquires the shooting style selected in the intention input options 500 and outputs it to the correction target amount generation unit 108.

[0049] In step S801, the control unit 114 starts shooting in response to an instruction from the user. For example, the user starts imaging a video by pressing a video shooting start button provided on the operation unit 116 or a video shooting start button displayed on the touch panel provided on the operation unit 116. The control unit 114 starts shooting by detecting the pressing of the video shooting start button. When shooting is started, moving image data is input to the scene analysis unit 106 via the imaging element 102, A / D conversion unit 103, capture unit 104, and digital signal processing unit 105.

[0050] In step S802, the scene analysis unit 106 performs scene analysis. The scene analysis unit 106 performs scene analysis on the moving image data that has been shot and input to the scene analysis unit 106, and outputs the scene information, which is the analysis result, to the correction target amount generation unit 108. In step S803, the correction target amount generation unit 108 generates two or more blurs to be applied. The correction target amount generation unit 108 generates blurs to be applied based on the scene information of the image acquired as the analysis result from the scene analysis unit 106 in S802 and the shooting style selected by the user acquired from the intention input unit 107 in step S800. Further, the correction target amount generation unit 108 determines the type of blur corresponding to the generated blur to be applied. Then, the correction target amount generation unit 108 outputs the type of blur corresponding to each of the generated multiple blurs to be applied to the blur type notification unit 109.

[0051] In step S804, the assigned blur type notification unit 109 notifies the user of multiple types of assigned blur. The assigned blur type notification unit 109 displays the assigned blur type obtained from the correction target amount generation unit 108 in the assigned blur type notification area 501. In step S805, the intention input unit 107 determines whether a user operation to change the assigned blur type has been performed. The intention input unit 107 determines that a user operation to change the assigned blur type has been performed if a type of assigned blur is selected from the candidate types of assigned blur in the assigned blur type notification area 501. If a user operation to change the assigned blur type has been performed, the process in step S306 is performed to change the assigned blur to be applied to the video. On the other hand, if no user operation to change the assigned blur type has been performed, the process in step S307 is performed without changing the assigned blur to be applied to the video.

[0052] In step S806, the intent input unit 107 outputs the type of added blur selected by the user to the correction target amount generation unit 108. The correction target amount generation unit 108 sets the added blur corresponding to the type of added blur selected by the user as the added blur to be added to the video. In step S807, the correction target amount generation unit 108 calculates the correction target amount based on the scene information and added blur obtained in step S802. The correction target amount generation unit 108 outputs the calculated correction target amount to the correction amount calculation unit 110.

[0053] In step S808, the correction amount calculation unit 110 acquires the amount of camera shake, which is the shaking of the imaging device 100, based on the output of the gyro sensor 111. In step S809, the correction amount calculation unit 110 calculates the amount of shake required to drive the shift lens of the optical system unit 101, which is an image stabilization unit, based on the amount of camera shake obtained in step S808 and the correction target amount obtained in step S807. The correction amount calculation unit 110 then outputs the calculated amount of shake required to drive the shift lens to the image stabilization control unit 112.

[0054] In step S810, the image stabilization control unit 112 controls the drive of the shift lens of the optical system unit 101 based on the calculated amount of blur that drives the shift lens. In step S811, the control unit 114 determines whether to end shooting. The control unit 114 determines whether to end shooting depending on whether or not there is an instruction from the user to end shooting. The user ends the capture of video by pressing, for example, the video recording end button on the operation unit 116 or the video recording end button displayed on the touch panel of the operation unit 116. When the control unit 114 detects that the video recording end button has been pressed, it determines to end shooting. If it determines to end shooting, the control unit 114 ends video shooting and the blur application process. On the other hand, unless the control unit 114 detects that the video recording end button has been pressed, it returns to step S302 and continues video shooting and the blur application process. Through the above process, an imaging device that applies the effect of camera shake in real time and maximizes the visual effect can be realized.

[0055] In this embodiment, the blur effect was applied by driving a shift lens that constitutes part of the optical system unit, but it is not limited to this. The image stabilization unit may be configured using a sensor-shift type blur stabilization that drives the image sensor 102. Alternatively, the image stabilization unit may perform correction using both the driving of the shift lens and the driving of the image sensor.

[0056] In this embodiment, the scene analysis unit 106 performs scene analysis processing only on the input video; however, the target of scene analysis is not limited to video. For example, the imaging device 100 may be equipped with an audio input device, and the scene analysis unit 106 may add scene analysis processing based on ambient sound. By inputting and analyzing information about cheers and music playing, it becomes possible to add motion blur synchronized with the sound to the video. Furthermore, by inputting and analyzing cheers, it becomes possible to add motion blur to the video that corresponds to the excitement of the scene.

[0057] In this embodiment, the correction target amount generation unit 108 generates the blur information using a deep learning technology called generation AI, but it is not limited to this. For example, the correction target amount generation unit 108 may be configured to select and output pre-prepared blur information based on the scene analysis results. In this embodiment, the blur information to be added is generated based on scene information and the user's intention represented by the content genre, but the drive limit information of the optical system unit 101 may also be taken into consideration. Shift lenses that perform image stabilization have a drive limit position, and hitting this drive limit can result in jerky video. Therefore, the correction target amount generation unit 108 calculates the added blur based on the shooting style and analysis results, as well as the drive limit of the optical correction member (shift lens or image sensor 102), which is the blur stabilization unit. Specifically, when the correction target amount generation unit 108 generates the blur information to be added, the effect of adding blur near the drive limit position may be attenuated, and the output may be configured so as not to hit the edges.

[0058] Although the control processing implemented by the control unit 114 was described as being implemented by the CPU executing a computer program stored in memory, some or all of these processes may be implemented in hardware. Dedicated circuits (ASICs) or processors (reconfigurable processors, DSPs) can be used as hardware. Furthermore, functions implemented in hardware can also be implemented, for example, by generating circuits based on data read from memory by an FPGA (Field Programmable Gate Array). Alternatively, a gate array circuit can be formed in the same way as an FPGA and implemented as hardware, or by using an ASIC (Application Specific Integrated Circuit).

[0059] As described above, according to this embodiment, by intentionally inputting the video expression that the user wants to express and generating and adding various types of blur along with information about the scene being captured, it becomes possible to capture moving images that express visual effects to the fullest extent. Second Embodiment

[0060] The first embodiment described a method for optically applying a camera shake effect by driving a shift lens included in the optical system unit based on the user's intent and scene information during shooting. There are two methods for camera shake correction: optical correction and electronic correction. Although it will only be used for a portion of the shooting angle of view, it is also possible to apply the real-time camera shake effect of the present invention using electronic camera shake correction. By using electronic camera shake correction, post-shooting processing can be simplified, which is useful for users who are live streaming. Therefore, this embodiment will describe a configuration in which a camera shake effect is applied using electronic camera shake correction.

[0061] Figure 9 shows the configuration of the imaging device 900 in the second embodiment. In this embodiment, blocks assigned the same reference numerals as in the first embodiment perform the same processing and their explanation is omitted. The imaging device 900 includes an optical system unit 101, an image sensor 102, an A / D conversion unit 103, a capture unit 104, a digital signal processing unit 105, a control unit 114, an external recording device 113, a memory 115, and an operation unit 116. The control unit 114 includes a scene analysis unit 106, an intention input unit 107, a correction target amount generation unit 108, an assigned blur type notification unit 109, and a correction amount calculation unit 110. Furthermore, the control unit 114 in this embodiment includes a motion vector calculation unit 901 and an electronic image stabilization unit 902.

[0062] The motion vector calculation unit 901 calculates motion vectors for the inter-frame corresponding points of each feature point based on the output of the digital signal processing unit 105, similar to the motion vector calculation unit 200 of the first embodiment, using the feature extraction unit and the optical flow calculation unit. The motion vector calculation unit 901 extracts the amount of camera shake, which is the shake of the imaging device 100, by extracting the high-frequency components of the motion vector using an inter-frame high-pass filter, and outputs it to the correction amount calculation unit 110. The correction amount calculation unit 110 treats the amount of camera shake calculated based on the motion vector of the moving image as the shake of the imaging device 100. The correction amount calculation unit 110 calculates the amount of shake for the electronic camera shake correction unit 902 based on the applied shake and the amount of camera shake output by the correction target amount generation unit 108.

[0063] The electronic image stabilization unit 902 calculates a geometric deformation matrix for each frame based on the amount of blur output from the correction amount calculation unit 110, and performs image stabilization by applying the matrix to the image. The geometric deformation matrix is ​​a homography matrix, and highly accurate matrix calculations are possible by removing outliers using the RANSAC algorithm.

[0064] Furthermore, since geometric deformation due to electronic image stabilization requires a reference pixel area outside the recording pixel area, there are cases where a portion of the imaging angle of view is limited. In such cases, the area that has been cut off may be interpolated by inferring it from the previous frame. Alternatively, interpolation may be performed using an inpainting model (Generative Adversarial Networks, GANs) that uses deep learning, which is a technique that interpolates missing parts based on surrounding information. In addition, when blurring occurs that results in missing parts, the electronic image stabilization unit 902 may notify the user by displaying it on the operation unit 116, allowing the user to choose to reduce the blur effect. Furthermore, the correction target amount generation unit 108 may calculate the applied blur based on the shooting style and analysis results, as well as the blur application limit to avoid including the area where electronic correction is ineffective in the angle of view.

[0065] In the second embodiment, a method for applying the effect of camera shake using only electronic image stabilization was described, but the system is not limited to this, and processing may be performed in combination with optical image stabilization as described in the first embodiment. For example, in the first embodiment, there is a drive limit for the shift lens of the optical system unit that performs the correction, and it is not possible to move the shift lens beyond that limit. In cases where such blur is to be applied, electronic image stabilization may be used in combination. Conversely, when performing electronic image stabilization, if there is a loss of pixels, the system may be configured to drive the shift lens to compensate for it.

[0066] As described above, according to this embodiment, even in an imaging device 900 that performs electronic image stabilization, it is possible to maximize the visual effect by adding the image stabilization effect intended by the user.

[0067] The disclosure of this embodiment includes the following configuration of an imaging device: (Configuration 1) An imaging device for capturing moving images, comprising: an input unit that accepts the setting of a shooting style by a user; a scene analysis unit that analyzes the scene of a moving image being captured by an imaging unit; a generation unit that generates added blur, which is blur to be added to the moving image; a calculation unit that calculates the amount of blur based on the added blur obtained from the generation unit and the shaking of the imaging device; and a correction control unit that controls a blur correction unit based on the amount of blur, wherein the imaging device generates the added blur based on the shooting style and the analysis result by the scene analysis unit. (Configuration 2) The imaging device according to Configuration 1, wherein the generation unit generates two or more of the added blurs, notifies the user of the type of blur corresponding to each of the added blurs as an added blur candidate, and outputs the added blur corresponding to the type of blur selected by the user to the calculation unit. (Configuration 3) The imaging device according to Configuration 1 or 2, characterized in that the generation unit generates two or more of the added blurs and outputs the added blur corresponding to the type of blur that has been frequently used in the past to the calculation unit. (Configuration 4) The imaging device according to Configuration 1 or 2, characterized in that the generation unit generates two or more of the added blurs and outputs the added blur with the smallest integral value of the absolute value of the amount of blur of the added blur to the calculation unit. (Configuration 5) The imaging device according to any one of Configurations 1 to 4, characterized in that the generation unit has an inference machine based on an artificial intelligence model, and the artificial intelligence model is generated by machine learning using the shooting style and the analysis result as input, and the type of blur information and the blur to be added to the moving image as training data. (Configuration 6) The imaging device according to any one of Configurations 1 to 5, characterized in that the scene analysis unit analyzes at least one piece of information of the scene type, main object information, motion information, human body posture information and environmental information of the moving image being captured by the imaging unit in real time. (Configuration 7) The imaging apparatus according to any one of Configurations 1 to 6, characterized in that the blur correction unit is a combination of one or more correction units, which include a correction unit that performs electronic correction, a correction unit that performs optical correction by driving a shift lens, and a correction unit that performs optical correction by driving an image sensor.(Configuration 8) The imaging device according to any one of Configurations 1 to 7, characterized in that the generation unit calculates the implied blur based on the shooting style, the analysis results, and the drive limit of the optical correction member which is the blur correction unit. (Configuration 9) The imaging device according to any one of Configurations 1 to 8, characterized in that the shooting style includes at least one of the genre of content or the impression to be implied. (Configuration 10) The imaging device according to Configuration 9, characterized in that the genre of content includes at least one of documentary, live, martial arts, and sports. (Configuration 11) The imaging device according to Configuration 2, further comprising a display unit that displays on a display the options for the user to set the shooting style and the options for the user to select the type of blur corresponding to the implied blur. Other embodiments

[0068] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (for example, an ASIC) that implements one or more functions.

[0069] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist.

[0070] 100 Imaging device 101 Optical system unit 102 Image sensor 107 Input unit 108 Correction target amount generation unit 109 Notification unit 110 Correction amount calculation unit 112 Blur correction control unit 114 Control unit

Claims

1. An imaging device for capturing moving images, comprising: an input unit for receiving settings of a shooting style by a user; a scene analysis unit for analyzing the scene of a moving image being captured by an imaging unit; a generation unit for generating added blur, which is blur to be added to the moving image; a calculation unit for calculating the amount of blur based on the added blur obtained from the generation unit and the vibration of the imaging device; and a correction control unit for controlling a blur correction unit based on the amount of blur, wherein the added blur is generated based on the shooting style and the analysis results by the scene analysis unit.

2. The imaging apparatus according to claim 1, characterized in that the generation unit generates two or more of the assigned blurs, notifies the user of the type of blur corresponding to each of the assigned blurs as an assigned blur candidate, and outputs the assigned blur corresponding to the type of blur selected by the user to the calculation unit.

3. The imaging apparatus according to claim 1, characterized in that the generation unit generates two or more of the assigned blurs and outputs the assigned blur corresponding to the type of blur that has been frequently used in the past to the calculation unit.

4. The imaging apparatus according to claim 1, characterized in that the generation unit generates two or more of the applied blurs and outputs the applied blur with the smallest integral value of the absolute value of the amount of blur of the applied blurs to the calculation unit.

5. The imaging device according to claim 1, wherein the generation unit has an inference machine based on an artificial intelligence model, and the artificial intelligence model is generated by machine learning using the shooting style and the analysis results as input, and the type of blur information and the blur to be added to the moving image as training data.

6. The imaging apparatus according to claim 1, characterized in that the scene analysis unit analyzes in real time at least one of the following pieces of information: scene type, main object information, motion information, human body posture information, and environmental information of the moving image captured by the imaging unit.

7. The imaging apparatus according to claim 1, characterized in that the image stabilization unit is a combination of one or more correction units, which include an electronic correction unit, an optical correction unit that performs optical correction by driving a shift lens, and an optical correction unit that performs optical correction by driving an image sensor.

8. The imaging apparatus according to claim 1, characterized in that the generation unit calculates the implied blur based on the shooting style and the analysis results, as well as the drive limit of the optical correction member which is the blur correction unit.

9. The imaging device according to claim 1, characterized in that the shooting style includes at least one of the genre of content or the impression to be conveyed.

10. The imaging device according to claim 9, characterized in that the genre of the content includes at least one of the following: documentary, live performance, martial arts, and sports.

11. The imaging apparatus according to claim 2, further comprising a display unit that displays on a display the options for the user to set the shooting style and the options for the user to select the type of blur corresponding to the applied blur.

12. A method for controlling an imaging device that captures moving images, comprising: a step of receiving a shooting style setting from a user; a step of analyzing a scene of a moving image being captured by an imaging unit; a step of generating an added blur, which is a blur to be added to the moving image; a step of calculating the amount of blur based on the added blur and the vibration of the imaging device; and a step of adding blur by controlling a blur correction unit based on the amount of blur, wherein in the step of generating the added blur, the added blur is generated based on the shooting style and the analysis result.

13. A program that causes the computer of an imaging device to perform the steps described in claim 12.