Imaging device, control method for imaging device, display device, and imaging system

The imaging device addresses the challenge of capturing videos while looking away from the subject by using face direction detection to adjust the cutout range, ensuring continuous and uninterrupted recording.

JP7864534B2Active Publication Date: 2026-05-25CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2022-04-07
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Photographers face challenges in capturing desired videos while diverting their gaze from the subject during shooting, especially when operating a smartphone or mobile device, leading to difficulties in maintaining the imaging range and continuity of video capture.

Method used

An imaging device equipped with a face direction detection system that adjusts the cutout range in frame images based on the user's face direction, allowing continuous video capture even when the user looks away from the subject.

Benefits of technology

Enables seamless video capture by automatically adjusting the imaging range to maintain the desired field of view, ensuring continuous recording without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007864534000002
    Figure 0007864534000002
  • Figure 0007864534000003
    Figure 0007864534000003
  • Figure 0007864534000004
    Figure 0007864534000004
Patent Text Reader

Abstract

To capture a desired image even when a user looks away from a subject while capturing a moving image.SOLUTION: An imaging device includes imaging means, detection means for detecting the direction of the user's face with respect to the imaging device, setting means for setting a cropping range in each frame image of a moving image captured by the imaging means on the basis of the detected face direction, and generating means for generating a cropped moving image from the cropping range, and the generating means changes the cropping range set for the frame image in the operation period in which the user was turning his / her face toward a display device communicatively connected to the imaging device to the cropping range set for the frame image before the start of the operation period to generate the cropping moving image.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an imaging device, a control method for an imaging device, a display device, and an imaging system. [Background technology]

[0002] Traditionally, taking images with a camera requires the photographer to keep the camera pointed in the direction they want to capture the image. This means that the photographer's hands are tied to the act of taking the image, preventing them from doing anything else, or concentrating their attention on the act of taking the image prevents them from focusing on the experience of being in the moment.

[0003] For example, in terms of image capture operations, a parent who is taking the picture cannot play with their child while capturing images, and conversely, if they try to play with their child, they cannot capture images, which presents a challenge.

[0004] Furthermore, in terms of focusing attention on imaging, when imaging during a sporting event, the photographer may not be able to cheer or remember the details of the game, and focusing attention on watching the sport may prevent them from imaging. Similarly, when imaging during a group trip, the photographer may not be able to experience the emotions at the same level as the other members, and prioritizing the experience may lead to neglecting imaging.

[0005] One way to address these challenges is to use a head-mounting accessory to fix the action camera to the head and capture images in the direction of observation, allowing the photographer to capture images without being tied down by the camera's hands. Another method is to use a 360-degree camera to capture a wide area, allowing the participant to focus on the experience during the experience, and then, after the experience is over, to cut out and edit the necessary footage from the captured 360-degree video to preserve the video of the experience.

[0006] However, the former method requires the cumbersome act of attaching a head-mounting accessory to the head, to which the action camera 901 body is fixed, as shown in Figure 27(a). Furthermore, as shown in Figure 27(b), when the photographer attaches the action camera 901 to their head using the head-mounting accessory 902, it looks unsightly and causes problems such as messing up the photographer's hairstyle. In addition, the photographer was bothered by the weight and presence of the head-mounting accessory 902 and the action camera 901 attached to their head, and was also concerned about the unsightly appearance to third parties. As a result, in the state shown in Figure 27(b), the photographer was unable to concentrate on the experience, or felt resistance to being in the state shown in Figure 27(b), making it difficult to take images.

[0007] On the other hand, the latter method requires a series of operations such as image conversion and specifying the cropping position. For example, a 360-degree camera 903 equipped with a lens 904 and a shooting button 905 is known, as shown in Figure 28. The lens 904 is one of a pair of hemispherical fisheye lenses configured on both sides of the housing of the 360-degree camera 903, and the 360-degree camera 903 performs 360-degree photography using this pair of fisheye lenses. In other words, 360-degree photography is performed by combining the projected images from this pair of fisheye lenses.

[0008] Figure 29 shows an example of the conversion process for images captured by the 360-degree camera 903.

[0009] Figure 29(a) is an example of an image obtained by 360-degree imaging with the 360-degree camera 903, and includes the subjects: the photographer 906, the child 907, and the tree 908. This image is a hemispherical optical system image obtained by combining the projection images of a pair of fisheye lenses, therefore, the image is captured Person 906 is greatly distorted. Furthermore, the child 907, the subject that photographer 906 was trying to photograph, had its body positioned at the periphery of the hemispherical optical system, causing its body to be greatly distorted and stretched from side to side. On the other hand, the tree 908 was positioned directly in front of lens 904, and therefore was photographed without significant distortion.

[0010] To create an image representing the field of view that a person normally sees from the image in Figure 29(a), it is necessary to cut out a portion of it, transform it into a plane, and display it.

[0011] Figure 29(b) is an image cropped from the image in Figure 29(a) that is positioned directly in front of lens 904. In the image in Figure 29(b), the tree 908 is in the center, representing a field of view similar to that of a normal human. However, the child 907 that the photographer 906 was trying to capture is not included in Figure 29(b), so the cropping position must be changed. Specifically, the cropping position in Figure 29(a) must be to the left of tree 908 and 30° downward from the perspective of the drawing. After this cropping work is performed, the image is transformed into a plane and displayed as Figure 29(c). Thus, in order to obtain the image in Figure 29(c) that the photographer was trying to capture from the image in Figure 29(a), it is necessary to crop the required area and transform it into a plane (hereinafter referred to as "trimming"). Therefore, although the photographer can concentrate on the experience (while capturing images), the amount of work that follows becomes enormous.

[0012] Furthermore, when shooting video with a camera, if the user shifts their gaze from the camera to a smartphone or other mobile device to operate it, it becomes difficult to check the image being captured by the camera. It is preferable to be able to continuously capture the desired image even if the photographer temporarily looks at the smartphone. Patent Document 1 discloses a technology that, when it detects an abnormality in the walking of a user who is walking while operating a mobile device, transmits the acquired audio and image data to a management server for a predetermined period of time after the abnormality occurs. [Prior art documents] [Patent Documents]

[0013] [Patent Document 1] Japanese Patent Publication No. 2020-150355 [Overview of the project]

Problems to be Solved by the Invention

[0014] When a photographer diverts their line of sight from the subject to a smartphone or the like while shooting a video with a camera, even if various information is received from the smartphone, it is difficult to continuously capture a desired video with the camera unless a process for adjusting the imaging range or the like is executed on the camera side.

[0015] An object of the present invention is to provide a technique that enables a desired video to be captured even when a user diverts their line of sight from the direction of the subject during video shooting.

Means for Solving the Problems

[0016] An imaging device according to the present invention includes an imaging unit Shadow part and detection means for detecting the face direction of a user with respect to the imaging device, setting means for setting a cutout range in each frame image of a video captured by the imaging unit based on the detected face direction, Shadow part and generation means for generating a cutout video from the cutout range, and is characterized in that the generation means changes the cutout range set for frame images during an operation period in which a user is performing an operation of facing their face toward a display device communicably connected to the imaging device to the cutout range set for frame images before the start of the operation period, and generates the cutout video.

Advantages of the Invention

[0017] According to the present invention, a desired video can be captured even when a user diverts their line of sight from the direction of the subject during video shooting.

Brief Description of the Drawings

[0018] [Figure 1A]It is an external view of a camera body including a photographing / detecting unit as an imaging device according to Embodiment 1. [Figure 1B] It is a view showing how a user wears the camera body. [Figure 1C] It is a view of the battery unit in the camera body seen from the rear in FIG. 1A. [Figure 1D] It is an external view of a display device as a portable device according to Embodiment 1, which is configured separately from the camera body. [Figure 2A] It is a view of the photographing / detecting unit seen from the front. [Figure 2B] ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​This is a flowchart of the subroutine for determining the recording direction and range in step S300 of Figure 7A, according to Example 1. [Figure 7E] This is a flowchart of the subroutine for the recording range development process in step S500 of Figure 7A, according to Example 1. [Figure 7F] This diagram illustrates the process from steps S200 to S500 in Figure 7A in video mode. [Figure 8A] This diagram shows the user's image as seen through the face direction detection window. [Figure 8B] This diagram shows the case where a fluorescent light in the room is reflected as a background in the image of the user as seen through the face direction detection window. [Figure 8C] Figure 8B shows an image of the user and the fluorescent lamp in the background, when the infrared LED of the infrared detection processing device is not turned on, and the image is formed by the sensor of the infrared detection processing device through the face direction detection window. [Figure 8D] Figure 8B shows an image of the user and the fluorescent lamp in the background, when the infrared LED is turned on and the image is formed by the sensor of the infrared detection processing device through the face direction detection window. [Figure 8E] This figure shows the difference image calculated from the images in Figures 8C and 8D. [Figure 8F] This figure shows the case where the contrast of the difference image in Figure 8E is adjusted to match the light intensity of the infrared reflected light projected onto the user's face and neck. [Figure 8G] Figure 8F is a diagram in which symbols indicating different parts of the user's body, as well as double circles and black circles indicating the neck and chin positions, are superimposed. [Figure 8H] This figure shows the difference image calculated using the same method as in Figure 8E, when the user's face is turned to the right. [Figure 8I] Figure 8H shows the double circles and black circles superimposed to indicate the positions of the neck and chin. [Figure 8J] This diagram shows the user's image as seen through the face direction detection window when the user's face is turned 33° upward from the horizontal. [Figure 8K]This figure shows the difference image calculated using the same method as in Figure 8E, with double circles and black circles indicating the neck position and chin position superimposed on the image when the user's face is turned 33° above the horizontal. [Figure 9] This is a timing chart showing the timing of when the infrared LEDs light up. [Figure 10] This diagram illustrates the vertical movement of the user's face. [Figure 11A] This diagram shows the target field of view in an ultra-wide-angle image captured by the camera's shooting unit when the user is facing forward. [Figure 11B] This figure shows the image of the target field of view in Figure 11A, extracted from an ultra-wide-angle image. [Figure 11C] This diagram shows the target field of view in an ultra-wide-angle image when the user is observing subject A. [Figure 11D] This figure shows the image of the target field of view, extracted from ultra-wide-angle footage in Figure 11C, with distortion and shaking corrected. [Figure 11E] This figure shows the target field of view in ultra-wide-angle video when the user is observing subject A with a field of view setting smaller than that shown in Figure 11C. [Figure 11F] This figure shows the image of the target field of view, extracted from ultra-wide-angle footage in Figure 11E, with distortion and shaking corrected. [Figure 12A] This figure shows an example of the target field of view in ultra-wide-angle video. [Figure 12B] This figure shows an example of a target field of view in ultra-wide-angle video, where the field of view setting is the same as the target field of view in Figure 12A, but the observation direction is different. [Figure 12C] This figure shows another example of a target field of view in ultra-wide-angle video, with the same field of view setting as the target field of view in Figure 12A, but with a different observation direction. [Figure 12D] This figure shows an example of a target field of view in ultra-wide-angle video, where the observation direction is the same as the target field of view in Figure 12C, but the field of view setting value is smaller. [Figure 12E] This figure shows an example where a reserve area is added around the target field of view, as shown in Figure 12A. [Figure 12F] This figure shows an example where a reserve area with the same vibration isolation level as the reserve area in Figure 12E is added around the target field of view shown in Figure 12B. [Figure 12G] This figure shows an example where a reserve area with the same vibration isolation level as the reserve area in Figure 12E is added around the target field of view shown in Figure 12D. [Figure 13] This diagram shows the menu screen for various video mode settings, which is displayed on the display unit of the camera before image capture. [Figure 14] Figure 7A is a flowchart of the subroutine for the primary recording process in step S600. [Figure 15] This diagram shows the data structure of the video file generated by the primary recording process. [Figure 16] Figure 7A is a flowchart of the subroutine for the transfer process to the display device in step S700. [Figure 17] Figure 7A is a flowchart of the subroutine for the optical correction process in step S800. [Figure 18] This figure illustrates the case where distortion correction is performed in step S803 of Figure 17. [Figure 19] Figure 7A is a flowchart of the vibration isolation subroutine in step S900. [Figure 20] This is a flowchart of the preparation process according to Example 2. [Figure 21] This is a flowchart of the processing performed by the display device 800 according to Example 2. [Figure 22] This is a flowchart for determining the operating period required to view the display device. [Figure 23] This is a diagram illustrating frame management information. [Figure 24] This diagram illustrates the state of the flags during the operating period. [Figure 25] This is a flowchart of the recording range development process according to Example 2. [Figure 26] This is a flowchart for the recording range determination and development process. [Figure 27]This figure shows an example of a camera configuration that is fixed to the head using a conventional head-fixing accessory. [Figure 28] This figure shows an example configuration of a conventional 360-degree camera. [Figure 29] This figure shows an example of the conversion process for images captured by the 360-degree camera shown in Figure 28. [Modes for carrying out the invention]

[0019] Preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] (Example 1) Figures 1A to 1D illustrate a camera system according to this embodiment, consisting of a camera body 1 including an imaging / detection unit 10 and a display device 800 which is configured separately. In this embodiment, the camera body 1 and the display device 800 are shown as separate components, but they may be configured as an integrated unit. The user who wears the camera body 1 around their neck will be referred to as the user below.

[0021] Figure 1A is an external view of the camera body 1.

[0022] In Figure 1A, the camera body 1 includes a shooting / detection unit 10, a battery unit 90, and a connection unit 80 that connects the shooting / detection unit 10 and the battery unit 90 (power supply means).

[0023] The imaging and detection unit 10 includes a face direction detection window 13, a start switch 14, a stop switch 15, an imaging lens 16, an LED 17, and microphones 19L and 19R.

[0024] The face direction detection window 13 transmits infrared light and its reflected light emitted from the infrared LED lighting circuit 21 (Figure 5: infrared irradiation means), which is built into the shooting / detection unit 10 and is used to detect the position of various parts of the user's face.

[0025] The start switch 14 is a switch used to start imaging.

[0026] The stop switch 15 is a switch used to stop image acquisition.

[0027] The imaging lens 16 guides the light rays to be imaged to the solid-state image sensor 42 (Figure 5) inside the imaging / detection unit 10.

[0028] LED17 is an LED that indicates imaging is in progress or displays a warning.

[0029] Microphones 19R and 19L are microphones that pick up ambient sounds. Microphone 19L picks up sounds from the left side of the user's surroundings (right side in Figure 1A), and microphone 19R picks up sounds from the right side of the user's surroundings (left side in Figure 1A).

[0030] Figure 1B shows the camera body 1 being worn by the user.

[0031] When the battery unit 90 is attached to the user's back and the imaging / detection unit 10 is attached to the user's front, the imaging / detection unit 10 is biased and supported towards the chest by the connecting parts 80, which are connected to the left and right ends of the imaging / detection unit 10. As a result, the imaging / detection unit 10 is positioned in front of the user's collarbone. At this time, the face direction detection window 13 is positioned below the user's chin. Inside the face direction detection window 13 is an infrared focusing lens 26, which will be shown later in Figure 2E. The optical axis of the imaging lens 16 (imaging optical axis) and the optical axis of the infrared focusing lens 26 (detection optical axis) are oriented in different directions, and the face direction detection unit 20 (face direction detection means), which will be described later, detects the user's observation direction from the position of each part of the face. This enables imaging in that observation direction by the imaging unit 40 (imaging means), which will be described later.

[0032] Methods for adjusting the setting position due to individual differences in body shape and clothing will be discussed later.

[0033] Furthermore, by positioning the shooting / detection unit 10 on the front of the body and the battery unit 90 on the back, the weight is distributed, which reduces user fatigue and suppresses displacement caused by centrifugal force when the user moves.

[0034] In this embodiment, the image capture / detection unit 10 is shown to be mounted in a position near the user's collarbone, but this is not the only option. That is, as long as the camera body 1 can detect the user's viewing direction using the face direction detection unit 20 and capture images in that viewing direction using the image capture unit 40, the camera body 1 may be mounted anywhere on the user's body other than their head.

[0035] Figure 1C is a view of the battery unit 90 from the rear of Figure 1A.

[0036] In Figure 1C, the battery unit 90 includes a charging cable insertion port 91, adjustment buttons 92L and 92R, and a spine protection cutout 93.

[0037] The charging cable port 91 is an insertion port for a charging cable (not shown), and through this charging cable, the internal battery 94 is charged from an external power source, or power is supplied to the shooting / detection unit 10.

[0038] The adjustment buttons 92L and 92R are used to adjust the length of the band sections 82L and 82R of the connection section 80. Adjustment button 92L is for adjusting the band section 82L on the left side, and adjustment button 92R is for adjusting the band section 82R on the right side. In this embodiment, the lengths of the band sections 82L and 82R are adjusted independently using adjustment buttons 92L and 92R, but it is also possible to adjust the lengths of the band sections 82L and 82R simultaneously with a single button. Hereinafter, the band sections 82L and 82R will be collectively referred to as band section 82.

[0039] The spine-relief cutout 93 is a cutout designed to avoid contact with the user's spine, preventing the battery unit 90 from touching the spine. By avoiding the protruding part of the human spine, it reduces discomfort during wear and prevents the device from moving from side to side during use.

[0040] Figure 1D is an external view of the display device 800 as a portable device according to Embodiment 1, which is configured separately from the camera body 1.

[0041] In Figure 1D, the display device 800 includes button A802, display unit 803, button B804, front camera 805, face sensor 806, angular velocity sensor 807, and acceleration sensor 808. Although not shown in Figure 1D, it also includes a wireless LAN capable of high-speed connection with the camera body 1.

[0042] Button A802 is a button that functions as the power button for the display device 800. It accepts power ON and OFF operations by long-pressing and other processing timing instructions by short-pressing.

[0043] The display unit 803 allows users to view images captured by the camera body 1 and display menu screens necessary for settings. In this embodiment, a transparent touch sensor is also provided on the top surface of the display unit 803 to accept touch operations on the displayed screen (e.g., the menu screen).

[0044] Button B804 functions as a calibration button 854, which is used in the calibration process described later.

[0045] The front camera 805 is a camera capable of capturing images of a person observing the display device 800.

[0046] The face sensor 806 detects the face shape and viewing direction of a person observing the display device 800. The specific structure of the face sensor 806 is not particularly limited, but it can be implemented using various sensors such as a structural light sensor, a ToF sensor, or a millimeter-wave radar.

[0047] The angular velocity sensor 807 is located inside the display device 800 and is therefore shown with a dotted line in the perspective view. The display device 800 in this embodiment also has a calibrator function, which will be described later, and is therefore equipped with a gyro sensor in three directions: X, Y, and Z.

[0048] The accelerometer 808 detects the orientation of the display device 800.

[0049] In this embodiment, a standard smartphone is used as the display device 800, and the camera system according to the present invention can be implemented by making the firmware of the smartphone compatible with the firmware of the camera body 1. However, the camera system according to the present invention can also be implemented by making the firmware of the camera body 1 compatible with the application and OS of the smartphone used as the display device 800.

[0050] Figures 2A to 2F illustrate the imaging and detection unit 10 in detail. In subsequent figures, parts that have already been described will be numbered the same way to indicate the same function, and their explanation in this specification will be omitted.

[0051] Figure 2A is a front view of the imaging and detection unit 10.

[0052] The connection section 80 connects to the imaging / detection unit 10 at a right-side connection section 80R located on the right side of the user's body (left side in Figure 2A) and a left-side connection section 80L located on the left side of the user's body (right side in Figure 2A). More specifically, the connection section 80 is divided into a rigid angle-holding section 81 and a band section 82 that maintain the angle with respect to the imaging / detection unit 10. That is, the right-side connection section 80R has an angle-holding section 81R and a band section 82R, and the left-side connection section 80L has an angle-holding section 81L and a band section 82L.

[0053] Figure 2B shows the shape of the band portion 82 of the connecting portion 80. In this figure, the angle holding portion 81 is shown transparently to illustrate the shape of the band portion 82.

[0054] The band portion 82 includes a connection surface 83 and an electrical cable 84.

[0055] The connecting surface 83 is the connecting surface between the angle holding portion 81 and the band portion 82, and has a non-circular cross-sectional shape. Herein, it has an elliptical shape.Hereinafter, among the connection surfaces 83, the connection surfaces 83 that are positioned symmetrically on the right side (left side in Figure 2B) and left side (right side in Figure 2B) of the user's body when the camera body 1 is attached are referred to as the right connection surface 83R and the left connection surface 83L.The right connection surface 83R and the left connection surface 83L are shaped exactly like the Japanese katakana character "ハ".That is, the distance between the right connection surface 83R and the left connection surface 83L decreases as you move from the bottom to the top in Figure 2B.As a result, when the user puts the camera body 1 on, the long axis direction of the connection surface 83 of the connection part 80 is aligned with the user's body, so that the band part 82 is comfortable when it comes into contact with the user's body and the shooting / detection unit 10 does not move in the left, right, front, or back directions.

[0056] The electrical cable 84 (power supply means) is wired inside the band section 82L and is a cable that electrically connects the battery section 90 and the imaging / detection section 10. The electrical cable 84 connects the power supply of the battery section 90 to the imaging / detection section 10 and also transmits and receives electrical signals to and from the outside.

[0057] Figure 2C shows the imaging / detection unit 10 viewed from the back. Since Figure 2C is a view from the side that contacts the user's body, i.e., the opposite side of Figure 2A, the positional relationship between the right connection part 80R and the left connection part 80L is reversed compared to Figure 2A.

[0058] The imaging and detection unit 10 is equipped with a power switch 11, an imaging mode switch 12, and a chest connection pad 18 on its back side.

[0059] The power switch 11 is a power switch that switches the power of the camera body 1 ON / OFF. In this embodiment, the power switch 11 is a slide lever type switch, but is not limited to this. For example, the power switch 11 may be a push type switch, or it may be a switch integrated with a slide cover (not shown) of the imaging lens 16.

[0060] The imaging mode switch 12 (changing means) is a switch that changes the imaging mode and can change the mode related to imaging. In this embodiment, the imaging mode switch 12 can switch to still image mode, video mode, and a pre-setting mode set using the display device 800, which will be described later. In this embodiment, the imaging mode switch 12 is a switch in the form of a slide lever that allows selection of one of "Photo," "Normal," or "Pri" as shown in Figure 2C by sliding the lever. The imaging mode changes to still image mode by sliding to "Photo," to video mode by sliding to "Normal," and to pre-setting mode by sliding to "Pri." Note that the imaging mode switch 12 is not limited to the form of this embodiment as long as it is a switch that can change the imaging mode. For example, the imaging mode switch 12 may be composed of three buttons: "Photo," "Normal," and "Pri."

[0061] The chest connection pads 18 (fixing means) are the parts that come into contact with the user's body when the imaging / detection unit 10 is biased against the user's body. As shown in Figure 2A, the imaging / detection unit 10 is shaped so that its horizontal (left-right) length is longer than its vertical (up-down) length when attached, and the chest connection pads 18 are positioned near the left and right ends of the imaging / detection unit 10. This arrangement makes it possible to suppress lateral rotational blur during imaging with the camera body 1. In addition, the presence of the chest connection pads 18 prevents the power switch 11 and the imaging mode switch 12 from coming into contact with the body. Furthermore, the chest connection pads 18 also serve to prevent heat from being transferred to the user's body even if the temperature of the imaging / detection unit 10 rises during long-term imaging, and also play a role in adjusting the angle of the imaging / detection unit 10.

[0062] Figure 2D is a top view of the imaging and detection unit 10.

[0063] As shown in Figure 2D, a face direction detection window 13 is provided in the center of the upper surface of the imaging / detection unit 10, and the chest connection pad 18 protrudes from the imaging / detection unit 10.

[0064] Figure 2E shows the configuration of the infrared detection processing device 27, which is located inside the imaging and detection unit 10 and positioned below the face direction detection window 13.

[0065] The infrared detection and processing device 27 includes an infrared LED 22 and an infrared focusing lens 26.

[0066] The infrared LED 22 emits infrared light 23 (Figure 5) towards the user.

[0067] The infrared focusing lens 26 is a lens that focuses the reflected light rays 25 (Figure 5) reflected from the user when infrared rays 23 are emitted from the infrared LED 22 onto a sensor (not shown) of the infrared detection processing device 27.

[0068] Figure 2F shows the camera body 1 as seen from the left side of the user's body while the user is wearing it.

[0069] The angle adjustment button 85L is a button located on the angle holding unit 81L and is used to adjust the angle of the shooting / detection unit 10. Although not shown in this figure, an angle adjustment button 85R is also set inside the angle holding unit 81R on the opposite side, in a position symmetrical to the angle adjustment button 85L. Hereafter, when referring to the angle adjustment buttons 85R and 85L collectively, they will be referred to as the angle adjustment button 85.

[0070] The angle adjustment button 85 is visible in Figures 2A, 2C, and 2D, but it has been omitted for the sake of simplicity in this explanation.

[0071] The user can change the angle between the imaging / detection unit 10 and the angle holding unit 81 by pressing the angle adjustment button 85 and moving the angle holding unit 81 up and down in the direction of Figure 2F. In addition, the chest connection pad 18 can have its protrusion angle changed. Through the action of these two angle-changing members (angle adjustment button 85 and chest connection pad 18), the imaging / detection unit 10 can adjust the orientation of the imaging lens 16 to be horizontal to accommodate individual differences in the shape of the user's chest position.

[0072] Figure 3 is a diagram illustrating the details of the battery unit 90.

[0073] Figure 3(a) is a view of the battery unit 90 from the rear, with a partial perspective.

[0074] As shown in Figure 3(a), the battery unit 90 has two batteries, a left battery 94L and a right battery 94R (hereinafter also collectively referred to as battery 94), symmetrically mounted inside to balance its weight. By symmetrically arranging the batteries 94 in the center of the battery unit 90 in this way, the weight balance between the left and right sides is adjusted, preventing the camera body 1 from shifting position. The battery unit 90 may also be configured to house only one battery.

[0075] Figure 3(b) is a view of the battery unit 90 from above. In this figure as well, the battery 94 is shown in perspective.

[0076] As shown in Figure 3(b), the relationship between the spine-guard cutout 93 and the battery 94 can be seen. By symmetrically arranging the battery 94 on both sides of the spine-guard cutout 93 in this way, it is possible to attach the relatively heavy battery unit 90 to the user without causing any burden. ru.

[0077] Figure 3(c) is a view of the battery unit 90 from the back. Figure 3(c) is a view from the side that comes into contact with the user's body, that is, the opposite side from Figure 3(a).

[0078] As shown in Figure 3(c), the spine-guard cutout 93 is located in the center along the user's spine.

[0079] Figure 4 is a functional block diagram of the camera body 1. Details will be explained later, but here we will use Figure 4 to describe the general flow of processing performed by the camera body 1.

[0080] In Figure 4, the camera body 1 comprises a face direction detection unit 20, a recording direction / angle determination unit 30, a shooting unit 40, an image cropping / development processing unit 50, a primary recording unit 60, a transmission unit 70, and other control units 111. These functional blocks are executed under the control of the overall control CPU 101 (Figure 5), which controls the entire camera body 1.

[0081] The face direction detection unit 20 (observation direction detection means) is a functional block executed by the infrared LED 22 and infrared detection processing device 27 mentioned earlier, which detects the face direction, infers the observation direction, and passes this to the recording direction / angle determination unit 30.

[0082] The recording direction / angle determination unit 30 (recording direction determination means) performs various calculations based on the observation direction inferred by the face direction detection unit 20 to determine the position and range information for extracting the image from the shooting unit 40, and passes this information to the image extraction / development processing unit 50.

[0083] The shooting unit 40 converts light rays from the subject into an image and passes that image to the image extraction and development processing unit 50.

[0084] The image extraction and development processing unit 50 (development means) uses information from the recording direction and field of view determination unit 30 to extract and develop the video from the shooting unit 40, thereby passing only the video in the direction the user is looking to the primary recording unit 60.

[0085] The primary recording unit 60 is a functional block consisting of a primary memory 103 (Figure 5) and the like, which records video information and transmits it to the transmission unit 70 at the necessary timing.

[0086] The transmitting unit 70 (video output means) wirelessly connects to predetermined communication partners, namely the display device 800 (Figure 1D), the calibrator 850, and the simple display device 900, and communicates with them.

[0087] The display device 800 is a display device that can connect to the transmission unit 70 via a high-speed wireless LAN (hereinafter referred to as "high-speed wireless"). In this embodiment, wireless communication corresponding to the IEEE 802.11ax (WiFi 6) standard is used for the high-speed wireless, but wireless communication corresponding to other standards, such as the WiFi 4 standard or WiFi 5 standard, may also be used. Furthermore, the display device 800 may be a device developed specifically for the camera body 1, or it may be a general-purpose smartphone or tablet device.

[0088] Furthermore, the connection between the transmitter 70 and the display device 800 may use low-power wireless communication, or it may be connected using both high-speed wireless and low-power wireless communication, or switched between the two. In this embodiment, data with a large amount of data, such as video files of video footage described later, is transmitted using high-speed wireless communication, while lightweight data or data that can take longer to transmit is transmitted using low-power wireless communication. In this embodiment, Bluetooth is used for low-power wireless communication, but NFC (Near Field Communication) may also be used. Other short-range (short-range) wireless communication methods such as d Communication may also be used.

[0089] The calibrator 850 is a device used for initial setup of the camera body 1 and for individual settings. Like the display device 800, it can connect to the transmitter 70 via high-speed wireless communication. Further details about the calibrator 850 will be described later. The display device 800 may also incorporate the functions of the calibrator 850.

[0090] The simplified display device 900 is a display device that can only be connected to the transmitter 70 via, for example, low-power wireless communication.

[0091] The simplified display device 900 cannot transmit video footage to the transmission unit 70 due to time constraints, but it can transmit the timing of the start and stop of image capture, and perform basic image confirmation such as composition checks. Furthermore, the simplified display device 900, like the display device 800, may be a device developed specifically for the camera body 1, or it may be a smartwatch or similar device.

[0092] Figure 5 is a block diagram showing the hardware configuration of camera body 1. Furthermore, the same numbers are used for the configurations and functions described using Figures 1A to 1C, etc., and detailed explanations are omitted.

[0093] In Figure 5, the camera body 1 includes an overall control CPU 101, a power switch 11, an imaging mode switch 12, a face direction detection window 13, a start switch 14, a stop switch 15, an imaging lens 16, and an LED 17.

[0094] The camera body 1 also includes an infrared LED lighting circuit 21, an infrared LED 22, an infrared focusing lens 26, and an infrared detection processing device 27, which constitute a face direction detection unit 20 (Figure 4).

[0095] Furthermore, the camera body 1 includes an imaging unit 40 (Figure 4) consisting of an imaging driver 41, a solid-state image sensor 42, and an imaging signal processing circuit 43, and a transmitting unit 70 (Figure 4) consisting of a low-power wireless unit 71 and a high-speed wireless unit 72.

[0096] In this embodiment, the camera body 1 is provided with only one shooting unit 40, but two or more shooting units 40 may be provided to capture 3D images, capture images with a wider angle of view than that obtainable with one shooting unit 40, or capture images in multiple directions.

[0097] The camera body 1 also includes various types of memory, such as a large-capacity non-volatile memory 51, a built-in non-volatile memory 102, and a primary memory 103.

[0098] Furthermore, the camera body 1 includes an audio processing unit 104, a speaker 105, a vibrator 106, an angular velocity sensor 107, an acceleration sensor 108, and various switches 110.

[0099] The overall control CPU 101 is connected to the power switch 11 and other components as shown in Figure 2C, and controls the camera body 1. The recording direction / angle determination unit 30, image cropping / development processing unit 50, and other control units 111 shown in Figure 4 are all composed of the overall control CPU 101 itself.

[0100] The infrared LED lighting circuit 21 controls the on / off state of the infrared LED 22 as described above using Figure 2E, and controls the emission of infrared light 23 from the infrared LED 22 toward the user.

[0101] The face direction detection window 13 is composed of a visible light cut filter, which blocks almost all visible light but allows sufficient transmission of infrared light 23 and its reflected light 25, which are in the infrared region.

[0102] The infrared focusing lens 26 is a lens that focuses reflected light rays 25.

[0103] The infrared detection processing device 27 (infrared detection means) has a sensor that detects reflected light rays 25 focused by an infrared focusing lens 26. This sensor forms an image of the focused reflected light rays 25, converts it into sensor data, and passes it to the overall control CPU 101.

[0104] As shown in Figure 1B, when the user is wearing the camera body 1, the face direction detection window 13 is located below the user's chin. Therefore, as shown in Figure 5, the infrared light 23 emitted from the infrared LED lighting circuit 21 passes through the face direction detection window 13 and is irradiated onto the infrared irradiation surface 24 near the user's chin. The infrared light 23 reflected by the infrared irradiation surface 24 becomes reflected light 25, which passes through the face direction detection window 13 and is focused by the infrared condensing lens 26 onto the sensor in the infrared detection processing device 27.

[0105] The various switches 110 are not shown in Figures 1A to 1C, etc., and although details are omitted, they are switches that perform functions unrelated to this embodiment.

[0106] The imaging driver 41 includes a timing generator and other components, and generates and outputs various timing signals to each part involved in imaging, thereby driving the imaging process.

[0107] The solid-state image sensor 42 outputs a signal obtained by photoelectric conversion of the subject image projected from the imaging lens 16, as explained using Figure 1A, to the imaging signal processing circuit 43.

[0108] The imaging signal processing circuit 43 performs processing such as clamping and A / D conversion on the signal from the solid-state image sensor 42 and outputs the generated imaging data to the overall control CPU 101.

[0109] The built-in non-volatile memory 102 uses flash memory or the like and stores the startup program for the overall control CPU 101 and the settings for various program modes. In this embodiment, the observation field of view (angle of view) and the effectiveness level of vibration control can be set, so these settings are also recorded.

[0110] The primary memory 103 consists of RAM and other components, and temporarily stores video data being processed, as well as the calculation results of the overall control CPU 101.

[0111] The large-capacity non-volatile memory 51 records or reads primary image data. In this embodiment, for the sake of simplicity, the description assumes that the large-capacity non-volatile memory 51 is a semiconductor memory without a removable mechanism, but it is not limited to this. For example, the large-capacity non-volatile memory 51 may be configured as a removable recording medium such as an SD card, or it may be used in combination with the built-in non-volatile memory 102.

[0112] The low-power wireless unit 71 exchanges data with the display device 800, calibrator 850, and simple display device 900 using low-power wireless communication.

[0113] The high-speed wireless unit 72 exchanges data with the display device 800, calibrator 850, and simple display device 900 using high-speed wireless communication.

[0114] The audio processing unit 104 is equipped with a microphone 19L on the right side of Figure 1A and a microphone 19R on the left side of the same figure, which pick up external sound (analog signals). It processes the picked-up analog signals and generates audio signals.

[0115] The LED 17, speaker 105, and vibrator 106 communicate or warn the user about the status of the camera body 1 by emitting light, sound, or vibrating.

[0116] The angular velocity sensor 107 is a sensor that uses a gyroscope or the like, and detects the movement of the camera body 1 itself as gyro data.

[0117] The acceleration sensor 108 detects the orientation of the shooting / detection unit 10.

[0118] Furthermore, the angular velocity sensor 107 and acceleration sensor 108 are built into the imaging / detection unit 10, and a separate angular velocity sensor 807 and acceleration sensor 808 are also provided in the display device 800, which will be described later.

[0119] Figure 6 is a block diagram showing the hardware configuration of the display device 800. For the sake of simplicity, the same reference numerals are used for parts that were explained using Figure 1D, and their explanations are omitted.

[0120] In Figure 6, the display device 800 includes a display device control unit 801, button A 802, display unit 803, button B 804, in-camera 805, face sensor 806, angular velocity sensor 807, acceleration sensor 808, image capture signal processing circuit 809, and various switches 811.

[0121] The display device 800 also includes a built-in non-volatile memory 812, a primary memory 813, a large-capacity non-volatile memory 814, a speaker 815, a vibrator 816, an LED 817, an audio processing unit 820, a low-power wireless unit 871, and a high-speed wireless unit 872.

[0122] The display device control unit 801 is composed of a CPU and is connected to buttons A802 and face sensor 806, as explained using Figure 1D, and controls the display device 800.

[0123] The imaging signal processing circuit 809 performs the same functions as the imaging driver 41, solid-state image sensor 42, and imaging signal processing circuit 43 inside the camera body 1, but it is not very important for the explanation in this embodiment, so for the sake of simplicity, it is described as a single unit. The data output by the imaging signal processing circuit 809 is processed in the display device control unit 801. The details of this data processing will be described later.

[0124] The various switches 811 are not shown in Figure 1D, and details are omitted, but they are switches that perform functions unrelated to this embodiment.

[0125] The angular velocity sensor 807 is a sensor that uses a gyroscope or the like, and detects the movement of the display device 800 itself.

[0126] The accelerometer 808 detects the orientation of the display device 800 itself.

[0127] As mentioned above, the angular velocity sensor 807 and acceleration sensor 808 are built into the display device 800 and, although they have the same functions as the angular velocity sensor 107 and acceleration sensor 808 located in the camera body 1 described earlier, they are separate components.

[0128] The built-in non-volatile memory 812 uses flash memory or similar technology and stores the startup program for the display device control unit 801 and the settings for various program modes.

[0129] The primary memory 813 is composed of RAM or the like and temporarily stores video data being processed and the calculation results of the imaging signal processing circuit 809. In this embodiment, during video recording, gyro data detected by the angular velocity sensor 107 at the imaging time of each frame is associated with each frame and stored in the primary memory 813.

[0130] The high-capacity non-volatile memory 814 records or reads image data from the display device 800. In this embodiment, the high-capacity non-volatile memory 814 is configured as a removable memory, such as an SD card. Alternatively, it may be configured as a non-removable memory, such as the high-capacity non-volatile memory 51 located in the camera body 1.

[0131] The speaker 815, vibrator 816, and LED 817 communicate the status of the display device 800 to the user or provide warnings by emitting sound, vibrating, or emitting light.

[0132] The audio processing unit 820 is equipped with a left microphone 819L and a right microphone 819R that pick up external sound (analog signals), and processes the picked-up analog signals to generate audio signals.

[0133] The low-power wireless unit 871 exchanges data with the camera body 1 using low-power wireless communication.

[0134] The high-speed wireless unit 872 exchanges data with the camera body 1 using high-speed wireless communication.

[0135] The face sensor 806 (face detection means) includes an infrared LED lighting circuit 821, an infrared LED 822, an infrared focusing lens 826, and an infrared detection processing device 827.

[0136] The infrared LED lighting circuit 821 is a circuit that has the same function as the infrared LED lighting circuit 21 in Figure 5, and controls the lighting and extinguishing of the infrared LED 822 and controls the emission of infrared light 823 from the infrared LED 822 toward the user.

[0137] The infrared focusing lens 826 is a lens that focuses the reflected light rays 825 of the infrared rays 823.

[0138] The infrared detection and processing unit 827 has a sensor that detects reflected light rays focused by the infrared focusing lens 826. This sensor converts the focused reflected light rays 825 into sensor data and passes it to the display device control unit 801.

[0139] When the face sensor 806 shown in Figure 1D is pointed at the user, as shown in Figure 6, infrared light 823 emitted from the infrared LED lighting circuit 821 is irradiated onto the user's entire face, which is the infrared irradiation surface 824. The infrared light 823 reflected by the infrared irradiation surface 824 becomes reflected light 825, which is then focused by the infrared condensing lens 826 onto the sensor in the infrared detection processing device 827.

[0140] The other functional unit 830, whose details are omitted here, performs functions unrelated to this embodiment, such as telephone functions and other smartphone-specific functions like sensors.

[0141] The following explains how to use the camera body 1 and the display device 800.

[0142] Figure 7A is a flowchart showing an overview of the image recording process according to this embodiment, which is performed in the camera body 1 and the display device 800.

[0143] For further explanation, Figure 7A indicates on the right side of each step which device shown in Figure 4 is performing that step. Specifically, steps S100 to S700 in Figure 7A are performed by the camera body 1, and steps S800 to S1000 in Figure 7A are performed by the display device 800.

[0144] In step S100, when the power switch 11 is turned ON and power is supplied to the camera body 1, the overall control CPU 101 starts up and reads the startup program from the built-in non-volatile memory 102. After that, the overall control CPU 101 performs preparatory operations to configure the camera body 1 before image capture. Details of the preparatory operations will be described later using Figure 7B.

[0145] In step S200, the face direction detection unit 20 detects the face direction and performs a face direction detection process to infer the observation direction. Details of the face direction detection process will be described later using Figure 7C. This process is executed at a predetermined frame rate.

[0146] In step S300, the recording direction / angle determination unit 30 performs the recording direction / range determination process. Details of the recording direction / range determination process will be described later using Figure 7D.

[0147] In step S400, the imaging unit 40 performs imaging and generates imaging data.

[0148] In step S500, the image extraction and development processing unit 50 uses the recording direction and field of view information determined in step S300 to extract the image from the imaging data generated in step S400 and performs recording range development processing on that area. Details of the recording range development processing will be described later with reference to Figure 7E.

[0149] In step S600, the primary recording unit 60 (video recording means) performs a primary recording process in which the video developed in step S500 is saved as video data in the primary memory 103. Details of the primary recording process will be described later with reference to Figure 14.

[0150] In step S700, the transmission unit 70 performs a transfer process to the display device 800, in which it wirelessly transmits the video recorded in step S600 to the display device 800 at a specified timing. Details of the transfer process to the display device 800 will be described later with reference to Figure 16.

[0151] Steps from step S800 onward are executed on the display device 800.

[0152] In step S800, the display device control unit 801 performs optical correction processing on the video transferred from the camera body 1 in step S700. Details of the optical correction processing will be described later with reference to Figure 17.

[0153] In step S900, the display device control unit 801 performs vibration damping on the image that was optically corrected in step S800. Details of the vibration damping process will be described later with reference to Figure 19.

[0154] Furthermore, the order of steps S800 and S900 can be reversed. In other words, you can perform image stabilization first and then optical correction afterward.

[0155] In step S1000, the display device control unit 801 (video recording means) performs secondary recording, recording the video after the optical correction processing and vibration damping processing in steps S800 and S900 into the large-capacity non-volatile memory 814, and then terminates this process.

[0156] Next, using Figures 7B to 7F, we will explain in detail the subroutines for each step described in Figure 7A, along with the order of processing, using other diagrams as well.

[0157] Figure 7B is a flowchart of the subroutine for the preparation operation process in step S100 of Figure 7A. This process will be explained below using the parts illustrated in Figures 2 and 5.

[0158] In step S101, it is determined whether the power switch 11 is ON or OFF. If not, wait, and when it turns ON, proceed to step S102.

[0159] In step S102, the mode selected by the imaging mode switch 12 is determined. If the result of the determination is that the mode selected by the imaging mode switch 12 is video mode, the process proceeds to step S103.

[0160] In step S103, the various settings for the video mode are read from the built-in non-volatile memory 102 and saved to the primary memory 103, after which the process proceeds to step S104. Here, the various settings for the video mode include the field of view setting value ang (which is pre-set to 90° in this embodiment) and the vibration stabilization level specified as "Strong," "Medium," or "Off."

[0161] In step S104, the operation of the imaging driver 41 for video mode is started, and then the subroutine is exited.

[0162] If the result of the determination in step S102 is that the mode selected by the imaging mode switch 12 is still image mode, proceed to step S106.

[0163] In step S106, the various settings for still image mode are read from the built-in non-volatile memory 102 and saved in the primary memory 103, after which the process proceeds to step S107. Here, the various settings for still image mode include the field of view setting value ang (which is pre-set to 45° in this embodiment) and the image stabilization level specified as "Strong," "Medium," or "Off."

[0164] In step S107, the operation of the imaging driver 41 for still image mode is started, and then the subroutine is exited.

[0165] If the result of the determination in step S102 indicates that the mode selected by the imaging mode switch 12 is the pre-set mode, the process proceeds to step S108. Here, the pre-set mode is a mode in which the imaging mode is set for the camera body 1 from an external device such as the display device 800, and is one of the three imaging modes that can be switched using the imaging mode switch 12. The pre-set mode is, in other words, a mode for custom shooting. Here, since the camera body 1 is a small wearable device, there are no operation switches or setting screens on the camera body 1 to change its detailed settings, and the detailed settings of the camera body 1 are changed by an external device such as the display device 800.

[0166] For example, consider a scenario where you want to capture video with both a 90° and a 110° field of view consecutively. Since the standard video mode is set to a 90° field of view, this requires first capturing video in the standard mode, then ending the video capture, switching the display device 800 to the camera body's settings screen, and switching the field of view to 110°. However, such operations on the display device 800 can be cumbersome during events.

[0167] On the other hand, if the pre-setting mode is set in advance to capture video with a 110° field of view, after capturing video with a 90° field of view, the user can instantly switch to capturing video with a 110° field of view simply by sliding the capture mode switch 12 to "Pri". In other words, the user does not need to interrupt their current activity and perform the cumbersome operation described above.

[0168] Furthermore, the settings configured in pre-setting mode may include not only the field of view, but also the image stabilization level, which can be specified as "Strong," "Medium," or "Off," as well as voice recognition settings, which are not described in this embodiment.

[0169] In step S108, various settings for the pre-setting mode are read from the built-in non-volatile memory 102. After extracting the data and saving it to the primary memory 103, the process proceeds to step S109. Here, the various settings in the pre-setting mode include the field of view setting value ang and the image stabilization level specified by "Strong," "Medium," "Off," etc.

[0170] In step S109, the operation of the imaging driver 41 for the pre-setting mode is started, and then the subroutine is exited.

[0171] Here, we will explain the various video mode settings read in step S103 using Figure 13.

[0172] Figure 13 shows the menu screen for various video mode settings displayed on the display unit 803 of the display device 800 before image capture by the camera body 1. Note that the same reference numerals are used for parts identical to those in Figure 1D, and their explanation is omitted. The display unit 803 has a touch panel function, and the following explanation assumes that it functions using touch operations, including swiping.

[0173] In Figure 13, the menu screen includes a preview screen 831, a zoom lever 832, a recording start / stop button 833, a switch 834, a battery level indicator 835, a button 836, a lever 837, and an icon display unit 838.

[0174] The preview screen 831 allows you to check the image captured by the camera body 1, and to check the zoom level and field of view.

[0175] The zoom lever 832 is an operating unit that allows zoom settings to be adjusted by shifting it left or right. In this embodiment, we will describe a case where four values, 45°, 90°, 110°, and 130°, can be set as the angle of view setting value ang, but the zoom lever 832 may also be used to set values ​​other than these as the angle of view setting value ang.

[0176] The recording start / stop button 833 is a toggle switch that combines the functions of both the start switch 14 and the stop switch 15.

[0177] Switch 834 is a switch that toggles the vibration damping on and off.

[0178] Battery level indicator 835 displays the remaining battery level of the camera body 1.

[0179] Button 836 is the button to enter other modes.

[0180] Lever 837 is a lever for setting the vibration isolation strength. In this embodiment, only "strong" and "medium" vibration isolation strengths can be set, but other vibration isolation strengths, such as "weak," may also be set. Alternatively, the vibration isolation strength may be set steplessly.

[0181] The icon display unit 838 displays multiple thumbnail icons for previewing.

[0182] Figure 7C is a flowchart of the subroutine for the face direction detection process in step S200 of Figure 7A. Before explaining the details of this process, we will explain the method of detecting face direction using infrared projection with reference to Figures 8A to 8K.

[0183] Figure 8A shows the user's image as seen through the face direction detection window 13.

[0184] The image in Figure 8A shows that the face direction detection window 13 does not have a visible light cut filter component, and therefore does not have sufficient visible light. It is transparent and, if the infrared detection processing device 27 is a visible light image sensor, it is identical to the image captured by that visible light image sensor.

[0185] The video in Figure 8A shows the user's face, including the front of the neck above the collarbone 201, the base of the jaw 202, the tip of the chin 203, and the nose 204.

[0186] Figure 8B shows the case where a fluorescent light in the room is reflected as a background in the image of the user seen through the face direction detection window 13.

[0187] The image in Figure 8B shows multiple fluorescent lights 205 surrounding the user. As shown above, various backgrounds and other elements are reflected in the infrared detection processing unit 27 depending on the usage conditions, making it difficult for the face direction detection unit 20 and the overall control CPU 101 to isolate the face from the sensor data from the infrared detection processing unit 27. Nowadays, there are technologies that use AI to isolate such images, but these require high capabilities from the overall control CPU 101 and are not suitable for the camera body 1, which is a portable device.

[0188] In reality, the face direction detection window 13 is composed of a visible light cut filter, so almost no visible light is transmitted, and therefore the image from the infrared detection processing device 27 does not look like the images in Figures 8A and 8B.

[0189] Figure 8C shows the image obtained when the user and the fluorescent lamp in the background shown in Figure 8B are imaged by the sensor of the infrared detection processing device 27 via the face direction detection window 13, with the infrared LED 22 not illuminated.

[0190] In the image in Figure 8C, the user's neck and chin appear dark. On the other hand, fluorescent lamp 205 appears somewhat brighter because it contains not only visible light but also infrared components.

[0191] Figure 8D shows an image obtained when the user and the fluorescent lamp in the background shown in Figure 8B are imaged by the sensor of the infrared detection processing device 27 through the face direction detection window 13, with the infrared LED 22 turned on.

[0192] In the image in Figure 8D, the user's neck and chin are illuminated. On the other hand, unlike in Figure 8C, the brightness around fluorescent lamp 205 remains unchanged.

[0193] Figure 8E shows the difference image calculated from the images in Figures 8C and 8D. The user's face can be seen.

[0194] In this way, the overall control CPU 101 (image acquisition means) calculates the difference between the images formed by the infrared detection processing device 27's sensor when the infrared LED 22 is lit and when it is not, thereby obtaining a difference image (hereinafter also referred to as a face image) in which the user's face is extracted.

[0195] In this embodiment, the face direction detection unit 20 employs a method of acquiring face images by extracting infrared reflection intensity as a two-dimensional image using the infrared detection processing unit 27. The sensor of the infrared detection processing unit 27 employs a structure similar to that of a general image sensor and acquires face images one frame at a time. The vertical synchronization signal (hereinafter referred to as the V signal) for frame synchronization is generated by the infrared detection processing unit 27 and output to the overall control CPU 101.

[0196] Figure 9 is a timing chart showing the timing of the infrared LED 22 turning on and off.

[0197] Figure 9(a) shows the timing at which the V signal is generated by the infrared detection processing unit 27. When the V signal becomes Hi, the timing of frame synchronization and the on / off of the infrared LED 22 is determined.

[0198] In Figure 9(a), t1 represents the period for the first facial image acquisition, and t2 represents the period for the second facial image acquisition. Figures 9(a), (b), (c), and (d) are presented so that their horizontal time axes are identical.

[0199] Figure 9(b) shows the H position of the image signal output from the sensor of the infrared detection processing unit 27 on the vertical axis. The infrared detection processing unit 27 controls the movement of its sensor so that the H position of the image signal is synchronized with the V signal, as shown in Figure 9(b). As mentioned above, the sensor of the infrared detection processing unit 27 employs a structure similar to that of a general image sensor, and its movement is well known, so the detailed control is omitted.

[0200] Figure 9(c) shows the switching timing between Hi and Low of the IR-ON signal output from the overall control CPU 101 to the infrared LED lighting circuit 21. The switching between Hi and Low of the IR-ON signal is controlled by the overall control CPU 101 in synchronization with the V signal, as shown in Figure 9(c). Specifically, the overall control CPU 101 outputs a Low IR-ON signal to the infrared LED lighting circuit 21 during period t1, and outputs a Hi IR-ON signal to the infrared LED lighting circuit 21 during period t2.

[0201] Here, while the IR-ON signal is Hi, the infrared LED lighting circuit 21 lights up the infrared LED 22, and infrared light 23 is projected onto the user. On the other hand, while the IR-ON signal is Low, the infrared LED lighting circuit 21 turns off the infrared LED 22.

[0202] Figure 9(d) shows the imaging data output from the infrared detection and processing unit 27 sensor to the overall control CPU 101. The vertical axis represents the signal intensity and indicates the amount of reflected light 25 received. In other words, during period t1, the infrared LED 22 is off, so there is no reflected light 25 from the user's face, and imaging data like that shown in Figure 8C is obtained. On the other hand, during period t2, the infrared LED 22 is on, so there is reflected light 25 from the user's face, and imaging data like that shown in Figure 8D is obtained. Therefore, as shown in Figure 9(d), the signal intensity during period t2 is higher than the signal intensity during period t1 by the amount of reflected light 25 from the user's face.

[0203] Figure 9(e) shows the difference between the imaging data during periods t1 and t2 in Figure 9(d), resulting in imaging data in which only the component of reflected light 25 from the user's face is extracted, as shown in Figure 8E.

[0204] Figure 7C shows the face direction detection process in step S200, including the operations described using Figures 8C to 8E and Figure 9 above.

[0205] First, in step S201, when the V signal output from the infrared detection processing device 27 becomes timing V1, which is the start of period t1, the process proceeds to step S202.

[0206] Next, in step S202, the IR-ON signal is set to Low and output to the infrared LED lighting circuit 21. As a result, the infrared LED 22 is turned off.

[0207] In step S203, the imaging data for one frame output from the infrared detection processing device 27 during the period t1 is read out and temporarily stored as Frame1 in the primary memory 103.

[0208] In step S204, when the V signal output from the infrared detection processing unit 27 reaches timing V2, which marks the start of period t2, the process proceeds to step S203.

[0209] In step S205, the IR-ON signal is set to Hi and output to the infrared LED lighting circuit 21. As a result, the infrared LED 22 lights up.

[0210] In step S206, the imaging data for one frame output from the infrared detection processing unit 27 during the period t2 is read out and temporarily stored as Frame2 in the primary memory 103.

[0211] In step S207, the IR-ON signal is set to Low and output to the infrared LED lighting circuit 21. This turns off the infrared LED 22.

[0212] In step S208, Frame1 and Frame2 are read from the primary memory 103, and the difference obtained by subtracting Frame1 from Frame2 is used to calculate the light intensity Fn of the 25 components of the user's reflected light rays in Figure 9(e) (this is generally known as the blacking process).

[0213] In step S209, the neck position (center of neck rotation) is extracted from the light intensity Fn.

[0214] First, the overall control CPU 101 (division means) divides the face image into multiple distance areas, which will be explained using Figure 8F, based on the light intensity Fn.

[0215] Figure 8F shows the case where the density of the difference image in Figure 8E is adjusted to match the light intensity of the reflected light rays 25 of the infrared 23 projected onto the user's face and neck, in order to see the distribution of light intensity for each part of the user's face and neck.

[0216] Figure 8F(a) is a diagram showing the distribution of light intensity of reflected light rays 25 in the facial image of Figure 8E, divided into regions and shown in gray to simplify the explanation. For explanatory purposes, the Xf axis is taken in the direction from the center of the user's neck to the tip of the chin.

[0217] Figure 8F(i) shows the light intensity on the Xf axis of Figure 8F(a) on the horizontal axis, and the Xf axis on the vertical axis. The light intensity increases as you move to the right on the horizontal axis.

[0218] In Figure 8F(a), the facial image is divided into six regions (distance areas) 211-216 according to light intensity.

[0219] Region 211 is the region with the strongest light intensity and is shown in white as a gray area.

[0220] Region 212 is a region where the light intensity is slightly lower than that of region 211, and is shown as a gray area, specifically in a fairly light gray color.

[0221] Region 213 is a region where the light intensity is even lower than that of region 212, and is shown as a gray area, represented by a light gray color.

[0222] Region 214 is a region where the light intensity is even lower than that of region 213, and is shown as an intermediate gray color, representing a gray range.

[0223] Region 215 is a region where the light intensity is even lower than that of region 214, and is represented as a gray area. It is shown in a slightly dark gray color.

[0224] Region 216 is the region with the weakest light intensity and is the darkest shade of gray. Above region 216, there is no light intensity and it is black.

[0225] This light intensity will be explained in detail below using Figure 10.

[0226] Figure 10 illustrates the vertical movement of the user's face, showing the user's position as observed from the left side.

[0227] Figure 10(a) shows the user facing forward. The imaging and detection unit 10 is located in front of the user's collarbone. In addition, infrared light 23 from the infrared LED 22 is irradiated onto the lower part of the user's head from the face direction detection window 13 located above the imaging and detection unit 10. If we let Dn be the distance from the face direction detection window 13 to the base of the neck 200 above the user's collarbone, Db be the distance from the face direction detection window 13 to the base of the chin 202, and Dc be the distance from the face direction detection window 13 to the tip of the chin 203, then it can be seen that the distances increase in the order of Dn, Db, and Dc. Since light intensity is inversely proportional to the square of the distance, the light intensity when the reflected light 25 from the infrared irradiation surface 24 is imaged by the sensor of the infrared detection processing device 27 decreases in the order of the base of the neck 200, the base of the chin 202, and the tip of the chin 203. Furthermore, it can be seen that the light intensity of the face 204, including the nose, which is located at a distance greater than Dc from the face direction detection window 13, becomes even dimmer. In other words, in a case like Figure 10(a), it can be seen that an image with the light intensity distribution shown in Figure 8F is acquired.

[0228] Furthermore, the configuration of the face direction detection unit 20 is not limited to the configuration shown in this embodiment, as long as the direction of the user's face can be detected. For example, an infrared pattern may be irradiated from the infrared LED 22 (infrared pattern irradiation means), and the infrared pattern reflected from the irradiated object may be detected by the sensor (infrared pattern detection means) of the infrared detection processing device 27. In this case, the sensor of the infrared detection processing device 27 is preferably a structural light sensor. Alternatively, the sensor of the infrared detection processing device 27 may be a sensor that performs phase comparison between infrared rays 23 and reflected light rays 25 (infrared phase comparison means), for example, a ToF sensor.

[0229] Next, using Figure 8G, we will explain the extraction of the neck position in step S209 of Figure 7C.

[0230] Figure 8G(a) is a diagram in which the symbols indicating the various parts of the user's body in Figure 10(a), as well as the symbols of the double circle and black circle indicating the neck and chin positions, are superimposed on Figure 8F.

[0231] The white area 211 corresponds to the base of the neck 200 (Figure 10(a)), the fairly light gray area 212 corresponds to the front of the neck 201 (Figure 10(a)), and the light gray area 213 corresponds to the base of the chin 202 (Figure 10(a)). The medium gray area 214 corresponds to the tip of the chin 203 (Figure 10(a)), and the slightly darker gray area 215 corresponds to the lower part of the face 204 (Figure 10(a)), specifically the lips and the surrounding lower face. Furthermore, the darker gray area 216 corresponds to the upper part of the face 204 (Figure 10(a)), specifically the nose and the surrounding upper face.

[0232] Furthermore, as shown in Figure 10(a), the distance between Db and Dc is small compared to the distance from the face direction detection window 13 to other parts of the user, so the difference in reflected light intensity between the light gray region 213 and the medium gray region 214 is also small.

[0233] On the other hand, as shown in Figure 10(a), of the distances from the face direction detection window 13 to each part of the user, the distance Dn is the shortest, so the white area corresponding to the base of the neck 200 Region 211 is the area with the strongest reflectivity.

[0234] Therefore, the overall control CPU 101 (setting means) sets the position 206, indicated by a double circle in Figure 8G(a), which is the center of the left and right sides of region 211 and closest to the imaging / detection unit 10, as the neck rotation center position (hereinafter referred to as neck base position 206). This process is what is done in step S209 of Figure 7C.

[0235] Next, using Figure 8G, we will explain the extraction of the chin position in step S210 of Figure 7C.

[0236] The intermediate gray region 214, which is brighter than the region 215 corresponding to the lower part of the face including the lips inside the face 204 shown in FIG. 8G(A), is the region including the chin tip. As can be seen from FIG. 8G(B), the light intensity drops sharply in the region 215 adjacent to the region 214, and the change in distance from the face direction detection window 13 becomes large. The overall control CPU 101 determines that the region 214 in front of the region 215 where there is a sharp drop in light intensity is the chin tip region. Further, the overall control CPU 101 calculates (extracts) the position (the position indicated by the black circle in FIG. 8G(A)) that is the center of the left and right of the region 214 and the farthest from the neck base position 206 as the chin tip position 207.

[0237] For example, FIGS. 8H and 8I show the changes when the face is facing right.

[0238] FIG. 8H is a diagram showing a differential video image calculated in the same manner as in FIG. 8E when the user's face is facing right. FIG. 8I is a diagram in which the double circle and black circle symbols indicating the neck base position 206 and the chin tip position 207r, which is the center position of the neck operation, are superimposed on FIG. 8H.

[0239] Since the region 214 is on the left when looking up from the imaging / detection unit 10 side because the user is facing right, it moves to the region 214r shown in FIG. 8I. The region 215 corresponding to the lower part of the face including the lips inside the face 204 also moves to the region 215r on the left when looking up from the imaging / detection unit 10 side.

[0240] Therefore, the overall control CPU 101 determines that the region 214r in front of the 215r where there is a sharp drop in light intensity is the chin tip region. Further, the overall control CPU 101 calculates (extracts) the position (the position indicated by the black circle in FIG. 8I) that is the center of the left and right of the 214r and the farthest from the neck base position 206 as the chin tip position 207r.

[0241] After that, the overall control CPU 101 obtains a movement angle θr indicating how much the chin tip position 207r in FIG. 8I has moved to the right around the neck base position 206 from the chin tip position 207 in FIG. 8G(A). As shown in FIG. 8I, the movement angle θr is the angle in the left-right direction of the user's face.

[0242] In the above method, in step S210, the infrared detection device 27 of the face direction detection unit 20 (3D detection sensor) detects the position of the chin tip and the angle in the left - right direction of the user's face.

[0243] Next, the detection of the upward direction of the face will be described.

[0244] FIG. 10(b) is a diagram showing the state where the user is facing horizontally, and FIG. 10(c) is a diagram showing the state where the user is facing 33° above the horizontal direction.

[0245] In FIG. 10(b), the distance from the face direction detection window 13 to the chin tip 203 is denoted as Ffh, and in FIG. 10(c), the distance from the face direction detection window 13 to the chin tip 203u is denoted as Ffu.

[0246] As shown in FIG. 10(c), since the chin tip 203u also moves upward together with the face, it can be seen that the distance of Ffu is longer than that of Ffh.

[0247] FIG. 8J is a diagram showing the image of the user visible from the face direction detection window 1 when the user is facing 33° above the horizontal direction of the face. As shown in FIG. 10(c), since the user is looking upward, from the face direction detection window 13 located below the user's chin, the face 204 including the lips and nose cannot be seen, and only the chin tip 203 can be seen. The distribution of the light intensity of the reflected light ray 25 when the user is irradiated with infrared rays 23 at this time is shown in FIG. 8K. FIG. 8K is a diagram in which double - circle and black - circle signs indicating the neck base position 206 and the chin tip position 207u are superimposed on the differential image calculated in the same method as FIG. 8E.

[0248] The six regions 211u to 216u in Figure 8K, corresponding to light intensity, are regions with the same light intensity as the region shown in Figure 8F, but with "u" added to them. In Figure 8F, the light intensity of the user's chin 203 was in the intermediate gray region 214, but in Figure 8K, it has shifted towards the gray side and is in the slightly darker gray region 215u. Thus, as shown in Figure 10(c), the infrared detection and processing device 27 can detect that, as a result of Ffu being at a longer distance than Ffh, the light intensity of the reflected light 25 from the user's chin 203 is weakened inversely proportional to the square of the distance.

[0249] Next, we will explain the detection of the face in the downward direction.

[0250] Figure 10(d) shows the user with their face tilted 22° downward from the horizontal.

[0251] In Figure 10(d), Ffd is defined as the distance from the face direction detection window 13 to the chin tip 203d.

[0252] As shown in Figure 10(d), since the chin tip 203d moves downward along with the face, the distance of Ffd becomes shorter than that of Ffh, and the light intensity of the reflected light rays 25 from the chin tip 203 becomes stronger.

[0253] Returning to Figure 7C, in step S211, the overall control CPU 101 (distance calculation means) calculates the distance from the chin position to the face direction detection window 13 based on the light intensity of the chin position detected by the infrared detection processing device 27 of the face direction detection unit 20 (3D detection sensor). Based on this, the vertical angle of the face is also calculated.

[0254] In step S212, the angles of the face in the left-right direction (first detection direction) and the vertical direction perpendicular to it (second detection direction), acquired in steps S210 and S211 respectively, are stored in the primary memory 103 as the user's observation direction vi, which consists of three dimensions (i is an arbitrary sign). For example, if the user was observing the center of the front, the observation direction vo would be the vector information [0°,0°], since the left-right direction θh is 0° and the vertical direction θv is 0°. Also, if the user was observing 45° to the right, the observation direction vr would be the vector information [45°,0°].

[0255] In step S211, the vertical angle of the face was calculated by detecting the distance from the face direction detection window 13, but this method is not limited to this method. For example, the angle change may be calculated by comparing the level of variation in the light intensity of the chin tip 203. In other words, the angle change of the chin may be calculated based on the gradient change of the gradient CDu of the reflected light intensity from the chin base 202 to the chin tip 203 in Figure 8K(c), relative to the gradient CDh of the reflected light intensity from the chin base 202 to the chin tip 203 in Figure 8G(a).

[0256] Figure 7D shows the subroutine for determining the recording direction and recording range in step S300 of Figure 7A. This is a flowchart. Before explaining the details of this process, we will first use Figure 11A to explain the ultra-wide-angle video that is the target of determining the recording direction and recording range in this embodiment.

[0257] In this embodiment, the camera body 1 achieves the acquisition of an image in the observation direction by having the shooting unit 40 capture an ultra-wide-angle image around the shooting / detection unit 10 using an ultra-wide-angle imaging lens 16, and then cropping a portion of that image.

[0258] Figure 11A shows the target field of view 125 in the ultra-wide-angle image captured by the shooting unit 40 when the user is facing forward.

[0259] As shown in Figure 11A, the image-capable pixel area 121 of the solid-state image sensor 42 is a rectangular area. The effective projection area 122 (predetermined area) is the area where a circular hemispherical image projected onto the solid-state image sensor 42 by the imaging lens 16 is displayed. The imaging lens 16 is adjusted so that the center of the pixel area 121 and the center of the effective projection area 122 coincide.

[0260] The outermost edge of the circular effective projection area 122 indicates the position with an FOV (Field of View) angle of 180°. When the user is looking at the horizontal and vertical center, the target field of view 125, which is the area to be imaged and recorded, has an angle of 90° from the center of the effective projection area 122, which is half that angle. In this embodiment, the imaging lens 16 can also introduce light rays from outside the effective projection area 122, and can project light rays up to a maximum FOV angle of approximately 192° onto the solid-state image sensor 42 using a fisheye projection. However, beyond the effective projection area 122, the optical performance deteriorates significantly, with extreme drops in resolution, light intensity, and distortion. Therefore, in this embodiment, the recording area will be explained using an example where the image in the observation direction is extracted only from the image projected onto the pixel area 121 of the hemispherical image displayed on the effective projection area 122 (hereinafter simply referred to as the ultra-wide-angle image).

[0261] In this embodiment, the vertical size of the effective projection area 122 is larger than the shorter side size of the pixel area 121, so the images at the upper and lower edges of the effective projection area 122 are outside the pixel area 121, but this is not limited to this. For example, the configuration of the imaging lens 16 may be changed to design the effective projection area 122 so that the entire area of ​​the effective projection area 122 fits within the area of ​​the pixel area 121.

[0262] The invalid pixel region 123 is the pixel region of the pixel region 121 that was not included in the effective projection region 122.

[0263] The aiming visual field 125 is an area showing the range for cutting out the video in the observation direction of the user from the ultra-wide-angle image, and is defined by preset horizontal and vertical viewing angles (here, 45°, FOV angle 90°) centered on the observation direction. In the example of FIG. 11A, since the user is facing forward, the center of the aiming visual field 125 is the observation direction vo which is the center of the effective projection unit 122.

[0264] The ultra-wide-angle video shown in FIG. 11A includes a subject A131 who is a child, a subject B132 which is the staircase that the child who is the subject A is about to climb, and a subject C133 which is a play equipment in the shape of a locomotive.

[0265] Next, the recording direction / range determination process in step S300 executed to obtain the video in the observation direction from the ultra-wide-angle video described using FIG. 11A is shown in FIG. 7D. Hereinafter, this process will be described using FIGS. 12A to 12G which are specific examples of the aiming visual field 125.

[0266] In step S301, it is acquired by reading out a preset viewing angle setting value ang from the primary memory 103.

[0267] In this embodiment, all viewing angles at which the video in the observation direction can be cut out from the ultra-wide-angle image by the image cutting / development processing unit 50, namely 45°, 90°, 110°, 130° are stored in the built-in non-volatile memory 102 as the viewing angle setting value ang. Also, in any one of steps S103, S106, S108, one of the viewing angle setting values ang stored in the built-in non-volatile memory 102 is set and stored in the primary memory <0> 103. Also, in step S301, the observation direction vi determined in step S212 is determined as the recording direction, and the video of the aiming visual field 125 cut out from the ultra-wide-angle image with the acquired viewing angle setting value ang centered thereon is stored in the primary memory 103. <000>

[0268] <>

[0269] ​For example, if the field of view setting value ang is 90° and the observation direction vo (vector information [0°,0°]) is detected by the face direction detection process (Figure 7C), the target field of view 125 (Figure 11A) is set to a range of 45° to the left and right and 45° up and down, centered on the center O of the effective projection unit 122. In other words, the overall control CPU 101 (relative position setting means) sets the angle of the face direction detected by the face direction detection unit 20 to the observation direction vi, which is vector information indicating the relative position to the ultra-wide-angle image.

[0270] In this case, when the observation direction is vo, the effect of optical distortion due to the imaging lens 16 can be almost ignored, so the shape of the set target field of view 125 becomes the shape of the target field of view 125o (Figure 12A) after distortion conversion in step S303, which will be described later. Hereafter, the target field of view 125 after distortion conversion in the case of the observation direction vi will be called the target field of view 125i.

[0271] Next, in step S302, the pre-set vibration isolation level is obtained by reading it from the primary memory 103.

[0272] In this embodiment, as described above, the vibration isolation level is set in one of steps S103, S106, or S108 and stored in the primary memory 103.

[0273] Furthermore, in step S302, the amount of spare pixels Pis for vibration damping is set based on the vibration damping level obtained above.

[0274] In the vibration stabilization process, the image is acquired that tracks the amount of shake of the shooting / detection unit 10, and tracks the image in the opposite direction to the shake direction. For this reason, in this embodiment, a reserve area necessary for vibration stabilization is provided around the target field of view 125i.

[0275] In this embodiment, a table is stored in the built-in non-volatile memory 102 that holds the value of the vibration-damping reserve pixel count Pis associated with each vibration-damping level. For example, if the vibration-damping level is "medium", a reserve pixel area of ​​100 pixels, which is the vibration-damping reserve pixel count Pis read from the table, is set as the reserve area.

[0276] Figure 12E shows an example where a reserve area is added around the target field of view 125o shown in Figure 12A. Here, we will explain the case where the vibration stabilization level is "medium," that is, the vibration stabilization reserve pixel amount Pis is 100 pixels.

[0277] As shown in Figure 12E, the dotted lines in the upper, lower, left, and right directions of the target field of view 125o, each providing a margin (reserve area) of 100 pixels, which is the amount of spare image-stabilizing pixels Pis, represent the spare image-stabilizing pixel frame 126o.

[0278] Figures 12A and 12E illustrate the case where the observation direction vi coincides with the center O of the effective projection area 122 (the optical axis center of the imaging lens 16) for simplicity of explanation. However, as will be explained in the following steps, when the observation direction vi is at the periphery of the effective projection area 122, optical distortion occurs. It is affected by [something], so conversion is necessary.

[0279] In step S303, the shape of the target field of view 125 set in step S301 is corrected (distortion converted) considering the observation direction vi and the optical characteristics of the imaging lens 16 to generate the target field of view 125i. Similarly, the number of spare pixels Pi for vibration damping set in step S302 is also corrected considering the observation direction vi and the optical characteristics of the imaging lens 16.

[0280] For example, suppose the field of view setting value ang is 90° and the user is observing 45° to the right of the center o. In this case, the observation direction vi determined in step S212 is the observation direction vr (vector information [45°, 0°]), and the target field of view 125 is the range of 45° to the left and right and 45° up and down, centered on the observation direction vr. However, considering the optical characteristics of the imaging lens 16, the target field of view 125 is corrected to the target field of view 125r shown in Figure 12B.

[0281] As shown in Figure 12B, the target field of view 125r widens towards the periphery of the effective projection area 122, and the position of the observation direction vr is also slightly inward from the center of the target field of view 125r. This is because, in this embodiment, the imaging lens 16 has an optical design similar to that of a stereoscopic fisheye. Note that if the imaging lens 16 is designed as an equidistant fisheye, equisolid angle fisheye, or orthographic fisheye, this relationship will change, and corrections will be made to the target field of view 125 according to its optical characteristics.

[0282] Figure 12F shows an example in which a reserve area with the same vibration isolation level ("medium") as the reserve area in Figure 12E is added around the target field of view 125r shown in Figure 12B.

[0283] In the vibration-damping reserve pixel frame 126o (Figure 12E), a margin of 100 pixels, which is the vibration-damping reserve pixel number Pis, is set for each of the top, bottom, left, and right sides of the target field of view 125o. In contrast, in the vibration-damping reserve pixel frame 126r (Figure 12F), the vibration-damping reserve pixel number Pis is corrected and increases as you move towards the periphery of the effective projection area 122.

[0284] Thus, similar to the shape of the target field of view 125r, the shape of the spare area necessary for vibration damping, provided around it, also shows that the amount of correction increases towards the periphery of the effective projection section 122, as shown in the vibration damping spare pixel frame 126r in Figure 12F. This is because, in this embodiment, the imaging lens 16 has an optical design close to that of a stereoscopic fisheye. However, if the imaging lens 16 is designed as an equidistant fisheye, equisolid angle fisheye, or orthographic fisheye, the relationship will change, and correction will be applied to the vibration damping spare pixel frame 126r according to its optical characteristics.

[0285] The process performed in step S303, which sequentially switches the shape of the target field of view 125 and its reserve area considering the optical characteristics of the imaging lens 16, is a complex process. Therefore, in this embodiment, the process in step S303 is performed using a table stored in the built-in non-volatile memory 102 that holds the shape of the target field of view 125i and its reserve area for each observation direction vi. Depending on the optical design of the imaging lens 16 mentioned above, the calculation formula may be stored in the overall control CPU 101, and the optical distortion value may be calculated using that formula.

[0286] In step S304, the position and size of the video recording frame are calculated.

[0287] As described above, in step S303, a reserve area necessary for vibration isolation was targeted and placed around the field of view 125i, and this was calculated as the vibration isolation reserve pixel frame 126i. However, depending on the position in the observation direction vi, its shape becomes quite unusual, for example, as in the vibration isolation reserve pixel frame 126r.

[0288] The overall control CPU 101 can perform development processing only within this specially shaped range and extract the image. However, it is not possible to record it as video data in step S600. In step S700, when transferring images to the display device 800, it is not common to use images that are not rectangular. Therefore, in step S304, the position and size of the rectangular image recording frame 127i, which encompasses the entire vibration-damping spare pixel frame 126i, are calculated.

[0289] Figure 12F shows the video recording frame 127r, indicated by a dashed line, which was calculated in step S304 relative to the vibration-damping spare pixel frame 126r.

[0290] In step S305, the position and size of the video recording frame 127i calculated in step S304 are recorded in the primary memory 103.

[0291] In this embodiment, the upper left coordinates Xi,Yi of the video recording frame 127i in the ultra-wide-angle image are recorded as the position of the video recording frame 127i, and the width WXi and height WYi of the video recording frame 127i from coordinates Xi,Yi are recorded as the size of the video recording frame 127i. For example, for the video recording frame 127r shown in Figure 12F, the shown coordinates Xr,Yr, width WXr, and height WYr are recorded in step S305. Note that the coordinates Xi,Yi are XY coordinates with a predetermined reference point, specifically the optical center of the imaging lens 16, as the origin.

[0292] Once the vibration-damping spare pixel frame 126i and the video recording frame 127i have been determined in this way, the subroutine of step S300 shown in Figure 7D is exited.

[0293] Up to this point, in order to simplify the explanation of the complex optical distortion transformation, we have used observation directions vi that include horizontal 0°, i.e., observation direction vo (vector information [0°,0°]) and observation direction vr (vector information [45°,0°]) as examples of observation direction vi. However, in reality, the user's observation direction vi will be in various directions. Therefore, the recording range development process performed in such cases will be explained below.

[0294] For example, when the field of view setting value ang is 90° and the observation direction vl[-42°,-40°], the target field of view 125l is as shown in Figure 12C.

[0295] Furthermore, even with the same observation direction vl (vector information [-42°,-40°]) as the target field of view 125l, if the field of view setting value ang is 45°, the target field of view 128l will be slightly smaller than the target field of view 125l, as shown in Figure 12D. In addition, for the target field of view 128l, a vibration-damping reserve pixel frame 129l and a video recording frame 130l are set, as shown in Figure 12G.

[0296] Step S400 is a basic imaging operation and uses a general sequence for the imaging unit 40, so details are left to other literature and will be omitted here. In this embodiment, the imaging signal processing circuit 43 in the imaging unit 40 also performs processing to correct the signal output from the solid-state image sensor 42, which is in a specific output format (examples of standards: MIPI, SLVS), into imaging data using a general sensor readout method.

[0297] Furthermore, if the mode selected by the imaging mode switch 12 is video mode, the shooting unit 40 starts recording when the start switch 14 is pressed. Recording then ends when the stop switch 15 is pressed. On the other hand, if the mode selected by the imaging mode switch 12 is still image mode, the shooting unit 40 takes a still image each time the start switch 14 is pressed.

[0298] Figure 7E is a flowchart of the subroutine for the recording range development process in step S500 of Figure 7A.

[0299] In step S501, the raw data of the entire area of ​​the imaging data (ultra-wide-angle video) generated by the imaging unit 40 in step S400 is acquired and input to the video acquisition unit called the head unit (not shown) of the overall control CPU 101.

[0300] Next, in step S502, based on the coordinates Xi, Yi, width WXi, and height WYi recorded in the primary memory 103 in step S305, the portion of the video recording frame 127i is extracted from the ultra-wide-angle video acquired in step S501. After this extraction, the cropping and development process (Figure 7F), consisting of steps S503 to S508, is started only on the pixels within the vibration-damping reserve pixel frame 126i. This significantly reduces the amount of computation compared to performing development on the entire area of ​​the ultra-wide-angle video read in step S501, thereby reducing computation time and power consumption.

[0301] Furthermore, as shown in Figure 7F, if the mode selected by the imaging mode switch 12 is video mode, the processes in steps S200 and S300 and the process in step S400 are executed in parallel at the same or different frame rates. In other words, each time the Raw data for the entire area of ​​one frame generated by the imaging unit 40 is acquired, cropping and development processing is performed based on the coordinates Xi, Yi, width WXi, and height WYi recorded in the primary memory 103 at that time.

[0302] When the crop development process for pixels within the vibration-damping spare pixel frame 126i is started, first, in step S503, color interpolation is performed to complete the color pixel information arranged in the Bayer array.

[0303] After that, the white balance is adjusted in step S504, and then the color conversion is performed in step S505.

[0304] In step S506, gamma correction is performed to correct the gradation according to a pre-set gamma correction value.

[0305] In step S507, edge enhancement is performed according to the image size.

[0306] In step S508, the data is converted into a data format that can be temporarily stored by performing compression and other processing, recorded in the primary memory 103, and then the subroutine is exited. Details of this data format that can be temporarily stored will be described later.

[0307] Furthermore, the order and whether or not the cropping and development processes performed in steps S503 to S508 are performed may be adjusted according to the camera system and do not limit the present invention.

[0308] Furthermore, if video mode is selected, the process from steps S200 to S500 will be repeated until recording is finished.

[0309] This process significantly reduces the amount of computation required compared to processing the entire area read in step S501. As a result, an inexpensive and low-power microcontroller can be used as the overall control CPU 101, and heat generation in the overall control CPU 101 is suppressed, while the battery life of the battery 94 is also improved.

[0310] Furthermore, in this embodiment, in order to reduce the control load on the overall control CPU 101, the optical correction processing (step S800 in Figure 7A) and vibration damping processing (step S900 in Figure 7A) of the image are not performed on the camera body 1, but are transferred to the display device 800 and then performed by the display device control unit 801. Therefore, if only the image data partially extracted from the projected ultra-wide-angle image is sent to the display device 800, the optical correction processing and vibration damping processing cannot be performed. In other words, the extracted Since the video data alone does not contain positional information that can be used to substitute into formulas during optical correction processing or to reference from correction tables during image stabilization processing, these processes cannot be correctly executed on the display device 800. Therefore, in this embodiment, not only the extracted video data but also correction data, including information on the extraction position from the ultra-wide-angle video, is transmitted from the camera body 1 to the display device 800.

[0311] If the extracted image is a still image, the still image data and correction data are transmitted separately to the display device 800, but since there is a one-to-one correspondence between the still image data and the correction data, the display device 800 can correctly perform optical correction processing and image stabilization processing. On the other hand, if the extracted image is a video, transmitting the video data and correction data separately to the display device 800 makes it difficult to determine which correction data corresponds to which frame of the video. In particular, if the clock rate of the overall control CPU 101 in the camera body 1 and the clock rate of the display device control unit 801 in the display device 800 are slightly different, synchronization between the overall control CPU 101 and the display device control unit 801 becomes impossible after several minutes of video capture. As a result, the display device control unit 801 may correct the frame that should be processed with correction data different from the corresponding correction data, leading to problems such as this.

[0312] Therefore, in this embodiment, when transmitting video data extracted from the camera body 1 to the display device 800, correction data is appropriately added to the video data. The method for doing so will be described below.

[0313] Figure 14 is a flowchart of the subroutine for the primary recording process in step S600 of Figure 7A. This process will be explained below with reference to Figure 15. Figure 14 shows the process when the mode selected by the imaging mode switch 12 is video mode. If the selected mode is still image mode, this process starts from step S601 and ends when step S606 is completed.

[0314] In step S601a, the overall control CPU 101 reads an image of one frame from the video footage developed in the recording range development process (Figure 7E) that has not been processed in steps S601 to S606. The overall control CPU 101 (metadata generation means) also generates correction data, which is metadata for the read frame.

[0315] In step S601, the overall control CPU 101 attaches information about the image extraction position of the frame read in step S600 to the correction data. The information attached here is the coordinates Xi,Yi of the video recording frame 127i acquired in step S305. Alternatively, the information attached here may be vector information indicating the observation direction Vi.

[0316] In step S602, the overall control CPU 101 (optical correction value acquisition means) acquires an optical correction value. The optical correction value is the optical distortion value set in step S303. Alternatively, it may be a correction value corresponding to the lens optical characteristics, such as peripheral light intensity correction value or diffraction correction value.

[0317] In step S603, the overall control CPU 101 attaches the optical correction values ​​used for distortion conversion in step S602 to the correction data.

[0318] In step S604, the overall control CPU 101 determines whether or not the camera is in vibration damping mode. Specifically, if the pre-set vibration damping mode is "medium" or "strong," it determines that the camera is in vibration damping mode and proceeds to step S605. On the other hand, if the pre-set vibration damping mode is "off," it determines that the camera is not in vibration damping mode and proceeds to step S606. The reason for skipping step S605 when the vibration damping mode is "off" is that by skipping this step, the amount of calculation data for the overall control CPU 101 and the amount of data transmitted wirelessly can be reduced, and consequently, the camera itself This is because it also reduces power consumption and heat generation in unit 1. While this explanation focuses on reducing data used for vibration isolation, it is also possible to reduce data such as peripheral light correction values ​​and whether or not analytical correction is applied, which are included in the optical correction values ​​obtained in step S602.

[0319] In this embodiment, the vibration isolation mode is pre-set by user operation via the display device 800, but it may also be set as the initial setting of the camera body 1. Furthermore, if the camera system is configured to switch between vibration isolation and non-vibration isolation after data is transmitted to the display device 800, step S604 can be omitted, and the system proceeds directly from step S603 to step S605.

[0320] In step S605, the overall control CPU 101 (movement detection means) attaches the vibration isolation mode acquired in step S302 and the gyro data during video capture, which is linked to the frame read in step S600 and stored in the primary memory 813, to the correction data.

[0321] In step S606, the video file 1000 (Figure 15) is updated with data encoded from the image data of the frame read in step S600 and the correction data to which various data have been attached in steps S601 to S605. If the first frame of the video footage was read in step S601a, the video file 1000 is generated in step S606.

[0322] In step S607, it is determined whether the reading of all frames of the video developed in the recording range development process (Figure 7E) has finished. If not, the process returns to step S601a. ​​If it has finished, the subroutine is exited. The generated video file 1000 is saved in the built-in non-volatile memory 102. In addition to being saved in the primary memory 813 and the built-in non-volatile memory 102 as described above, it may also be saved in the large-capacity non-volatile memory 51. Alternatively, the generated video file 1000 may be immediately transferred to the display device 800 (step S700 in Figure 7A), and after being transferred to the display device 800, it may be saved in its primary memory 813.

[0323] In this embodiment, encoding refers to combining video data and correction data into a single file. However, the video data may be compressed, or the combined video data and correction data may be compressed.

[0324] Figure 15 shows the data structure of video file 1000.

[0325] The video file 1000 consists of a header 1001 and a frame 1002. Frame 1002 is composed of a frame dataset, which is a set of images for each frame that makes up the video and their corresponding frame metadata. In other words, frame 1002 contains as many frame datasets as there are frames in the video.

[0326] In this embodiment, the frame metadata is information encoded with correction data, which includes the cropping position (position information within the image), optical correction value, and gyro data as needed, but is not limited to this. For example, the amount of information in the frame metadata may be changed by adding other information to the frame metadata or deleting information in the frame metadata depending on the imaging mode selected by the imaging mode switch 12.

[0327] Header 1001 records the offset value or starting address to the frame dataset for each frame. Alternatively, metadata such as the time and size corresponding to the video file 1000 may be stored.

[0328] Thus, in the primary recording process (Figure 14), the display device 800 receives the recording range development process. As shown in Figure 7E), a video file 1000 is transferred, which is a set containing each frame of the developed video footage and its metadata. Therefore, even if the clock rate of the overall control CPU 101 of the camera body 1 and the clock rate of the display control unit 801 of the display device 800 are slightly different, the display control unit 801 can reliably perform the correction processing of the video footage developed by the camera body 1.

[0329] In this embodiment, the optical correction value was included in the frame metadata, but the optical correction value may be applied to the entire video.

[0330] Figure 16 is a flowchart of the subroutine for the transfer process to the display device 800 in step S700 of Figure 7A. Figure 16 shows the process when the mode selected by the imaging mode switch 12 is video mode. If the selected mode is still image mode, this process starts from step S702.

[0331] In step S701, it is determined whether the recording of video footage by the shooting unit 40 (step S400) has finished or is still in progress. If video footage is being recorded (video capture in progress), the recording range development process for each frame (step S500) and the updating of the video file 1000 in the primary recording process (step S600) (step S606) are performed sequentially. Wireless transfer has a high power load, so performing it in parallel with recording would require a large battery capacity for the battery 94 and separate measures to prevent overheating. Also, from the perspective of computing power, performing wireless transfer in parallel with recording increases the computing load, so it is necessary to prepare a high-spec overall control CPU 101, which also increases the cost. In this embodiment, taking these factors into consideration, the system waits for the video footage recording to finish (YES in step S701) before proceeding to step S702 to establish a connection with the display device 800. However, if the camera system of this embodiment has sufficient power supplied from the battery 94 and no additional heat dissipation measures are required, the camera body 1 may be connected to the display device 800 in advance, such as when it is started up or before recording begins.

[0332] In step S702, a connection to the display device 800 is established via the high-speed wireless unit 72 in order to transfer the video file 1000, which has a large amount of data, to the display device 800. The low-power wireless unit 71 is used for transferring low-resolution video (or video) to the display device 800 for checking the field of view, and for sending and receiving various setting values ​​with the display device 800, but it is not used for transferring the video file 1000 because it takes time to transmit.

[0333] In step S703, the video file 1000 is transferred to the display device 800 via the high-speed wireless unit 72. Once the transfer is complete, the process proceeds to step S704, where the connection with the display device 800 is closed, and then the subroutine is exited.

[0334] Up to this point, we have described the case of transferring a single video file containing images of all frames of a single video. However, for long video videos lasting several minutes, it is also acceptable to use multiple video files divided by time units. If the video file has the data structure shown in Figure 15, even if a single video is transferred to the display device 800 as multiple video files, the display device 800 can correct the video without any timing discrepancies with the correction data.

[0335] Figure 17 is a flowchart of the subroutine for the optical correction process in step S800 of Figure 7A. This process will be explained below with reference to Figure 18. As mentioned above, this process is executed by the display device control unit 801 of the display device 800.

[0336] In step S801, the display device control unit 801 (video file receiving means) first receives the video file from the camera body 1 that was transferred in the transfer process to the display device 800 (step S700). The device receives file 1000. Subsequently, the display device control unit 801 (first extraction means) obtains optical correction values ​​extracted from the received video file 1000.

[0337] Next, in step S802, the display device control unit 801 (second extraction means) acquires video (an image of one frame obtained by video capture) from the video file 1000.

[0338] In step S803, the display device control unit 801 (frame image correction means) performs optical correction of the image acquired in step S802 using the optical correction value acquired in step S801, and saves the corrected image to the primary memory 813. When performing optical correction, if cropping is performed from the image acquired in step S802, the image is cropped and processed in a range narrower than the development range (target field of view 125i) determined in step S303 (cropped development area).

[0339] Figure 18 is a diagram illustrating the case where distortion correction is performed in step S803 of Figure 17.

[0340] Figure 18(a) shows the position of the subject 1401 as seen by the user with the naked eye during imaging, and Figure 18(b) shows the image of the subject 1401 projected onto the solid-state image sensor 42.

[0341] Figure 18(c) shows the development region 1402 in the image of Figure 18(b). Here, the development region 1402 is the cropped development region explained earlier.

[0342] Figure 18(d) shows the cropped developed area, from which the image of the developed area 1402 has been extracted, and Figure 18(e) shows the image obtained by correcting the distortion of the cropped developed area in Figure 18(d). Since cropping is performed during the distortion correction of the cropped developed image, the field of view of the image shown in Figure 18(e) is even smaller than that of the cropped developed area shown in Figure 18(d).

[0343] Figure 19 is a flowchart of the vibration isolation subroutine in step S900 of Figure 7A. This process will be explained below with reference to Figure 25. As mentioned above, this process is executed by the display device control unit 801 of the display device 800.

[0344] In step S901, the gyro data for the current frame and the previous frame are obtained from the frame metadata of video file 1000, and the amount of shake V calculated for the previous frame in step S902 described below is obtained. n-1 Det Obtain the following information. Then, from this information, estimate the amount of deviation V n Pr e This is calculated. In this embodiment, the current frame is the frame currently being processed, and the previous frame is the frame immediately preceding the current frame.

[0345] Step S902 provides detailed information on the amount of blur V from the video. n Det This is how the amount of blur is determined. Blur detection is performed by calculating how much the feature points of the current frame's image have moved from the previous frame.

[0346] Known methods can be adopted for feature point extraction. For example, a luminance information image obtained by extracting only the luminance information of the frame image can be generated, and an image shifted by 1 to several pixels therefrom can be subtracted from the original image, and pixels whose absolute value is greater than or equal to a threshold can be extracted as feature points. Alternatively, an image obtained by applying a high-pass filter to the luminance information image can be subtracted from the original luminance information image, and the extracted edges can be extracted as feature points.

[0347] The movement amount is calculated by calculating the difference multiple times while shifting the luminance information images of the current frame and the previous frame by 1 to several pixels each, and calculating the position where the difference at the pixels of the feature points decreases.

[0348] Since a plurality of feature points are required as described later, it is preferable to divide the images of the current frame and the previous frame into a plurality of blocks and extract the feature points. The block division depends on the number of pixels and aspect ratio of the image, but generally 12 blocks of 4×3 to 96×64 blocks are preferable. If the number of blocks is small, corrections such as trapezoid due to the sway of the imaging unit 40 of the camera body 1 and rotation in the optical axis direction cannot be accurately performed. However, if the number of blocks is too large, the size of one block becomes small and the feature points become close to each other, resulting in errors. For these reasons, an optimal number of blocks is appropriately selected according to the number of pixels, ease of finding feature points, and angle of view of the subject.

[0349] To calculate the movement amount, it is necessary to shift the luminance information images of the current frame and the previous frame by 1 to several pixels each and perform multiple difference calculations, resulting in a large amount of calculation. Therefore, the actual movement amount is the blur amount V [[ID=E14]] n [[ID=E15]] Pre To calculate how many pixels it is shifted from, it is possible to significantly reduce the amount of calculation by performing difference calculations only in the vicinity thereof.

[0350] Next, in step S90 , after performing shake correction using the detailed blur amount V obtained in step S902, this subroutine is exited. n Det After performing shake correction using the detailed blur amount V obtained in step S902, this subroutine is exited.

[0351] Furthermore, conventional methods for vibration isolation include Euclidean transforms that allow rotation and translation, affine transforms that allow these actions, and projective transforms that allow trapezoidal correction.

[0352] While movement and rotation along the X and Y axes can be corrected using Euclidean transformations, the blur that occurs when actually capturing images with the camera body 1's shooting unit 40 also includes camera shake in the front-to-back direction and in the pan-tilt direction. Therefore, in this embodiment, vibration correction is performed using affine transformations, which can also correct for magnification and skew. In affine transformations, when the coordinates (x,y) of a reference feature point move to coordinates (x',y'), it is expressed by the following equation 100.

[0353]

number

[0354] Figure 18(f) shows the image after applying the image stabilization correction in step S903 to the distortion-corrected image shown in Figure 18(e). Because cropping is performed during image stabilization correction, the field of view of the image shown in Figure 18(f) is smaller than that of the image shown in Figure 18(e).

[0355] By performing this type of image stabilization, it is possible to obtain high-quality images with corrected blur.

[0356] The above describes a series of operations performed by the camera body 1 and the display device 800 included in the camera system of this embodiment.

[0357] After the user turns on the power switch 11 and selects the video mode with the imaging mode switch 12, and simply observes the front without turning their face up, down, left, or right, the face direction detection unit 20 first detects the observation direction vo (vector information [0°,0°]) (Figure 12A). Then, the recording direction / angle determination unit 30 extracts the image of the target field of view 125o shown in Figure 12A (Figure 11B) from the ultra-wide-angle image projected onto the solid-state image sensor 42.

[0358] Subsequently, without the user operating the camera body 1, if, for example, the user starts observing the child (subject A131) in Figure 11A, the face direction detection unit 20 first detects the observation direction vl (vector information [-42°, -40°]) (Figure 11C). Then, the recording direction / angle determination unit 30 extracts the image with the target field of view 125l (Figure 11C) from the ultra-wide-angle image captured by the shooting unit 40.

[0359] In this way, optical correction and vibration damping processing are performed on the display device 800 in steps S800 and S900 for the images cropped into various shapes according to the observation direction. As a result, even if the overall control CPU 101 of the camera body 1 has low specifications, even when cropping an image with significant distortion, such as the target field of view 125l (Figure 11C), it is possible to obtain an image with distortion and shaking corrected, centered on the child (subject A131), as shown in Figure 11D. In other words, the user can obtain an image captured in their observation direction without touching the camera body 1, other than turning on the power switch 11 and selecting a mode with the imaging mode switch 12.

[0360] Now, let's explain the pre-configuration mode. As mentioned above, the camera body 1 is a small wearable device, so it does not have any operation switches or setting screens for changing its detailed settings. Therefore, the detailed settings of the camera body 1 are changed using an external device such as the display device 800 (in this embodiment, the setting screen of the display device 800 (Figure 13)).

[0361] For example, consider a scenario where you want to capture video with both a 90° and a 45° field of view consecutively. In normal video mode, the field of view is set to 90°, so to perform this type of capture, you would first need to capture video in normal video mode, then stop the video capture, switch the display device 800 to the camera body 1 settings screen, and switch the field of view to 45°. However, performing such operations on the display device 800 during continuous capture is cumbersome, and you might miss capturing the footage you want.

[0362] On the other hand, if the pre-setting mode is set to capture video at a 45° field of view, after capturing video at a 90° field of view, simply sliding the capture mode switch 12 to "Pri" will instantly switch to zoomed-in video capture at a 45° field of view. In other words, the user does not need to interrupt the current capture process and perform the cumbersome operation described above.

[0363] The settings configured in pre-setting mode may include not only changes to the field of view, but also the image stabilization level, which can be specified as "strong," "medium," or "off," as well as changes to the voice recognition settings, which are not described in this embodiment.

[0364] For example, in the aforementioned imaging situation, if the user continues to observe the child (subject A131) and switches from video mode to pre-setting mode using the imaging mode switch 12, the field of view setting value ang changes from 90° to 45°. In this case, the recording direction / field of view determination unit 30 extracts an image of the target field of view 128l, shown by the dotted line frame in Figure 11E, from the ultra-wide-angle image captured by the shooting unit 40.

[0365] Even in pre-setting mode, optical correction and image stabilization processing are performed on the display device 800 in steps S800 and S900. As a result, even if the overall control CPU 101 of the camera body 1 has low specifications, distortion and shaking when zoomed in on the child (subject A131) as shown in Figure 11F can be corrected. This allows you to obtain corrected images. While the example of changing the field of view setting `ang` from 90° to 45° in video mode was used, the same applies to still image mode. Furthermore, the same is true when the video field of view setting `ang` is 90° and the still image field of view setting `ang` is 45°.

[0366] In this way, the user can obtain a zoomed-in image of their observation direction simply by switching modes using the imaging mode switch 12 on the camera body 1.

[0367] In this embodiment, the case in which the face direction detection unit 20 and the shooting unit 40 are integrally configured in the camera body 1 has been described. However, this is not limited to the case in which the face direction detection unit 20 is mounted on the user's body other than their head, and the shooting unit 40 is mounted on the user's body. For example, the shooting / detection unit 10 of this embodiment can also be installed on the shoulder or abdomen. However, in the case of the shoulder, if the shooting unit 40 is installed on the right shoulder, the subject on the left side may be obstructed by the head, so it is preferable to install multiple shooting means, including on the left shoulder, to compensate. Also, in the case of the abdomen, a spatial parallax occurs between the shooting unit 40 and the head, so it is desirable to be able to perform a correction calculation of the observation direction to correct for that parallax, as shown in Embodiment 3.

[0368] (Example 2) Embodiment 2 is an embodiment that enables the capture of desired images even if the user averts their gaze (face) from the direction of the subject while capturing video. In Embodiment 2, for example, even if the user turns their face towards the display device 800 that displays the status of the camera body 1, the camera body 1 can continue to capture images of the direction the user was looking before changing the direction of their face.

[0369] The imaging system according to Embodiment 2 will be described with reference to Figures 20 to 26. The imaging system includes a camera body 1 and a display device 800 that can communicate with the camera body 1. For the camera body 1 and display device 800 in Embodiment 2, the same reference numerals are used for components identical to those in Embodiment 1, and redundant explanations are omitted. For components with different processing compared to Embodiment 1, further details will be provided.

[0370] Since the camera body 1 is worn around the user's neck, the display unit (screen or LED lamp, etc.) that shows the status of the camera body 1 is not visible to the user during image capture. The display device 800 is connected to the camera body 1 in a communicative manner and displays a menu screen for various settings of the video mode, as explained in Figure 13, allowing the user to check the status of the camera body 1. The display device 800 can send and receive various data with the camera body 1 via an application or firmware that allows it to cooperate with the camera body 1. The status of the camera body 1 includes, for example, whether image capture is in progress or stopped, the remaining charge of the battery 94, and the remaining charge of the built-in non-volatile memory 102.

[0371] However, if the user turns their face towards the display device 800 during imaging, the camera body 1 records the area including the display device 800 in the direction of the user's face as the cropped area. As a result, the image in the direction that was originally intended to be captured is not captured. Furthermore, the screen of the display device 800 may display information other than the status of the camera body 1, such as credit card information, and the user may not want to record the image of the display device 800.

[0372] To prevent the display device 800 from being captured, the user must hold the display device 800 in a location where it is unlikely to be visible to the camera body 1, and direct only their gaze towards the display device 800 without changing the direction of their face. Checking the status of the camera body 1 displayed on the display device 800 while keeping the direction of the face unchanged during imaging is a tiring action for the user, and it is difficult to see the information on the camera body 1 that they want to check.

[0373] The camera body 1 in this embodiment is designed to continue recording the desired image even if the user turns their face towards the display device 800 to check the status of the camera body 1 without interrupting image capture. The purpose of the user turning their face towards the display device 800 is not limited to checking the status of the camera body 1, but may also be to respond to a phone call or email, etc. The following example will be explained assuming a flow in which the user turns their face towards the display device 800, displays the status of the camera body 1 on the screen of the display device 800, checks the status of the camera body 1, and then turns their face back.

[0374] It is believed that the user's face direction does not coincide with the direction to be recorded during the period from when the user begins to change their face direction toward the display device 800 until they return their face direction to its original position. In other words, during the period from when the user's face direction begins to change until the user's face is turned toward the display device 800 and the user's face direction changes again and stops (hereinafter referred to as the operation period), the user's face direction does not coincide with the direction the user wants to record.

[0375] Therefore, the image cropping and development processing unit 50 of the camera body 1 crops and records (develops) the face direction the user was facing when they began to change their face direction during the operation period for viewing the display device 800. In other words, the image cropping and development processing unit 50 changes the cropping range set for the frame image during the operation period in which the user was turning their face towards the display device 800 to the cropping range of the frame image before the operation period began.

[0376] Whether or not the user is facing the display device 800 is determined by the display device control unit 801 of the display device 800. If the display device control unit 801 determines that the user is facing the display device 800, it notifies the camera body 1 that the user is facing the display unit 803 of the display device 800.

[0377] The display unit control unit 801 can determine, for example, that the user has turned their face towards the display device 800 when the display unit 803 switches from an inactive state, such as being off, to an active state, such as being on, and becomes activated.

[0378] Furthermore, the display device control unit 801 is not limited to switching from an off state to an on state; it may also determine that the user has turned their face towards the display device 800 when the user is captured (detected) by the in-camera 805 provided by the display device 800.

[0379] Furthermore, the display device control unit 801 may determine that the user has turned their face towards the display device 800 when an application screen for checking the status of the camera body 1, such as the screen described in Figure 13, is displayed. The application screen may be displayed on the display unit 803 by user operation, or it may be displayed on the display unit 803 by instruction from the camera body 1.

[0380] The following describes the process in Figure 7A related to Example 2, which differs from that in Example 1. Figure 20 is a flowchart of the preparation process related to Example 2. The preparation process shown in Figure 20 exemplifies the detailed process of step S100 in Figure 7A. The preparation process related to Example 2 is the same as the preparation process in Figure 7B related to Example 1, with the addition of steps S110 and S111.

[0381] When the camera body 1 is powered on and the settings are switched based on the imaging mode (steps S101 to S109), in step S110, the overall control CPU 101 establishes a low-power wireless connection with the display device 800. In step S111, the overall control CPU 101 sends a low-power wireless instruction to turn off the screen of the display device 800. The display device control unit 801 of the display device 800 receives the instruction to turn off the screen and changes the screen (display unit 803) to the off state.

[0382] Referring to Figure 21, the processing flow of the display device 800 will be explained. Each step in Figure 21 is realized by the display device control unit 801 of the display device 800 controlling each component of the display device 800 according to a program stored in the built-in non-volatile memory 812.

[0383] In step S521, the display unit control 801 waits until the power switch is turned ON and power is supplied to the display device 800. In step S522, the display unit control 801 establishes a connection with the camera body 1. The display unit control 801 can connect to the camera body 1, for example, using low-power wireless communication. The process in step S522 corresponds to the process in step S110 of the camera body 1 as described in Figure 20.

[0384] In step S523, the display control unit 801 waits for a power-off command from the camera body 1 in step S111 of Figure 20. When the display control unit 801 receives the power-off command, it proceeds to step S524. In step S524, the display control unit 801 turns off the screen of the display device 800.

[0385] In step S525, the display control unit 801 determines whether the display device 800 has been activated by the user's operation. If the display control unit 801 determines that it has been activated, it proceeds to step S526; if it determines that it has not been activated, it proceeds to step S527. In step S526, the display control unit 801 notifies the camera body 1 that the display device 800 has been activated.

[0386] Here, the system determines whether the display device 800 has been activated, but the display device control unit 801 only needs to be able to determine whether the user is facing the display device 800. The display device control unit 801 may determine that the user is facing the display device 800 if the user is captured (detected) by the in-camera 805 provided by the display device 800. The method for determining whether the user has been captured can be any method that can detect a person's face or the user's face from the image captured by the in-camera 805. For example, the display device control unit 801 can determine that the user has been captured if the face or part of the face of the user registered for facial recognition is detected from the image captured by the in-camera 805.

[0387] Furthermore, the display device control unit 801 may determine that the user has turned their face towards the display device 800 when an application screen for checking the status of the camera body 1, as shown in Figure 13, is displayed.

[0388] In step S527, the display unit 801 determines whether or not it has received notification from the camera body 1 that image capture has ended. If it has not received notification that image capture has ended, the display unit 801 returns to step S525 and repeats the processes in steps S525 and S526 until it receives notification that image capture has ended. If it receives notification that image capture has ended, the display unit 801 proceeds to step S528.

[0389] In step S528, the display unit 801 establishes a high-speed wireless connection with the camera body 1. In step S529, the display unit 801 receives a video file from the camera body 1 in which the cropping area has been processed. The display unit 801 then performs the processing from steps S800 to S1000, as described in Figure 7A of Embodiment 1, on the received video file.

[0390] After the preparation operation process shown in Figure 20, the overall control CPU 101 of the camera body 1 executes the face direction detection process (step S200) and the recording direction / range determination process (step S300) as described in Figure 7A of Example 1, and then proceeds to the shooting process in step S400.

[0391] Referring to Figure 22, the process for determining the period during which the user was performing an action to view the display device 800 during the shooting process in step S400 will be explained. The process shown in Figure 22 is executed for each frame of the video captured by the shooting unit 40 in step S400. The frame to be processed in Figure 22 is referred to as the current frame. The processing of each step in Figure 22 is realized by the overall control CPU 101 controlling each component of the camera body 1 according to the program stored in the built-in non-volatile memory 102.

[0392] In step S401, the overall control CPU 101 determines whether the "Changing" flag is ON in the image capture of the frame preceding the current frame (hereinafter referred to as the "previous frame"). The "Changing" flag indicates whether the user is in the process of performing a series of operations to view the display device 800. Initially, the "Changing" flag is set to OFF, indicating that the user is not in the process of viewing the display device 800. In step S420, the value of the "Changing" flag is recorded in the primary memory 103 for each frame.

[0393] If the overall control CPU 101 determines in step S401 that the flag indicating the previous frame is being modified is OFF, it proceeds to step S402. In step S402, the overall control CPU 101 determines whether or not the user's face is in motion. Whether or not the user's face is in motion can be detected by a change in face direction. A change in face direction can be detected, for example, as a change in the face direction vector (observation direction vi) detected in Figure 7C, or as a change in the position and size (cropping range) of the video recording frame determined based on the direction vector. When detecting a change in the direction vector, the overall control CPU 101 can determine that the user's face is in motion if the amount of change in the direction vector is greater than or equal to a threshold angle.

[0394] When detecting a change in the cropping range, the overall control CPU 101 obtains the position and size (cropping range) of the video recording frame recorded in the primary memory 103 for the current frame and the previous frame in the recording direction / range determination process (step S300).

[0395] The overall control CPU 101 determines whether the difference in the cropping range between the current frame and the previous frame (for example, the distance between the position of the cropping range in the current frame and the position of the cropping range in the previous frame) is greater than or equal to a threshold. If the difference in the cropping range between the current frame and the previous frame is greater than or equal to the threshold, the overall control CPU 101 can determine that the user's face is in motion. The threshold can be, for example, the maximum or average value of the change in face direction when the user is not intentionally moving their face. If it is determined that the user's face is in motion, the process proceeds to step S403. If it is determined that the user's face is not in motion, the process proceeds to step S420.

[0396] In step S403, the overall control CPU 101 sets the "Changing" flag and the "Initial Operation" flag to ON, and sets the change frame count to 1. The initial operation is the action of turning the user's face away from the subject towards the display device 800 before looking at the display device 800. The "Initial Operation" flag is a flag that indicates the state in which the user's face direction has changed towards the display device 800 before looking at the display device 800, and the setting from the previous frame is carried over.

[0397] Even if the face direction changes, it is not necessarily an action to view the display device 800, so the "Changing" flag is temporarily set to ON. If it is determined that the change in face direction is not an action to view the display device 800, the "Changing" flag is changed to OFF in step S411.

[0398] The change frame count is a variable used to count the number of frames from when the user's face direction begins to change until an activation notification is received from the display device 800. If the user does not receive an activation notification from the display device 800 for a certain period of time, the display device 800 will... It is thought that the user changed the direction of their face to change the imaging direction, rather than trying to check the status of the camera body 1. If the overall control CPU 101 receives an activation notification from the display device 800 while the change frame count is below the threshold, it determines that the change in face direction is an action to look at the display device 800. Also, if the change frame count exceeds the threshold without receiving an activation notification from the display device 800, the overall control CPU 101 determines that the change in face direction is not an action to look at the display device 800. In step S403, since the current frame is the first frame in which the user's face began to move, the change frame count is set to 1.

[0399] If the overall control CPU 101 determines in step S401 that the flag indicating that the previous frame is being modified is ON, it proceeds to step S404. In step S404, the overall control CPU 101 determines whether the modified frame count is 0 or not. If it determines that the modified frame count is not 0, the process proceeds to step S405.

[0400] In step S405, the overall control CPU 101 determines whether the change frame count is above a threshold, and if it is below the threshold, proceeds to step S406. If the change frame count is below the threshold, it is determined that the change in face direction is an action for the user to look at the display device 800. If the change frame count is above the threshold, it is determined that the change in face direction is not an action for the user to look at the display device 800.

[0401] The threshold for the change frame count can be determined, for example, by multiplying the frame rate by the threshold time from the start of the operation period for the user to view the display device 800 until the user receives the activation notification. The threshold time may be, for example, around 3 to 5 seconds, or it may be set arbitrarily by the user.

[0402] Based on the determination in step S405, if the overall control CPU 101 receives an activation notification from the display device 800 within the threshold time, it can change the frame extraction range during the operation period to the extraction range before the face movement. If the overall control CPU 101 does not receive an activation notification from the display device 800 within the threshold time, it determines that the change in face direction was not an action to look at the display device 800, and does not change the extraction range.

[0403] In step S406, the overall control CPU 101 determines whether the user's face is active or not, similar to step S402. If the overall control CPU 101 determines that the user's face is not active, it changes the initial active flag to OFF in step S407. If it determines that the face is active, it proceeds to step S408 without changing the initial active flag.

[0404] In step S408, the overall control CPU 101 determines whether or not it received an activation notification from the display device 800 between the imaging of the previous frame and the imaging of the current frame. If the overall control CPU 101 receives an activation notification from the display device 800, it sets the changed frame count to 0 in step S410. If the overall control CPU 101 does not receive an activation notification from the display device 800, it increments the changed frame count in step S409. After setting the changed frame count in step S409 or step S410, the overall control CPU 101 proceeds to step S420.

[0405] In step S405, if the change frame count is above a threshold, the overall control CPU 101 determines that the change in face direction is not an action to view the display device 800, and the process proceeds to step S411. In step S411, the overall control CPU 101 sets the changing flag and the initial operation flag to OFF and resets the change frame count to 0.

[0406] In step S412, the overall control CPU 101 rewrites the "Changing" flag, which has been set to ON for consecutive frames up to the previous frame, to OFF. That is, if the display device 800 does not become active within a number of frames less than a threshold that have been captured since the user's face direction began to change, the overall control CPU 101 determines that the user's actions were not a series of actions to view the display device 800. Therefore, the overall control CPU 101 rewrites the "Changing" flag, which had been set to ON for each frame since the user's face direction began to change, to OFF.

[0407] If the change frame count is determined to be 0 in step S404, the overall control CPU 101 proceeds to step S413 and determines whether the user's face is active or not, similar to steps S402 and S406. If the overall control CPU 101 determines that the user's face is not active, it proceeds to step S414.

[0408] In step S414, the overall control CPU 101 determines whether the initial operation flag is ON or OFF. If the initial operation flag is ON, the overall control CPU 101 proceeds to step S417 and changes the initial operation flag to OFF. That is, when the initial operation is completed, in which the face direction starts to change and is seen by the display device 800 before the user looks at the display device 800, and the change in face direction stops, the overall control CPU 101 sets the initial operation flag to OFF.

[0409] If the system determines in step S413 that the user's face is not in operation, the overall control CPU 101 proceeds to step S418 to determine whether the initial operation flag is ON or OFF. If the initial operation flag is ON, the overall control CPU 101 proceeds to step S420; otherwise, it proceeds to step S419.

[0410] In step S419, the overall control CPU 101 sets the termination operation flag to ON and proceeds to step S420. The termination operation is the operation in which the user turns their face away from the display device 800 and towards the subject after looking at the display device 800. The termination operation flag is a flag that indicates the state from when the user's face direction starts moving again after looking at the display device 800 until it stops, and the setting from the previous frame is carried over.

[0411] If the initial operation flag is not ON in step S414, the overall control CPU 101 proceeds to step S415 to determine whether the termination operation flag is ON or OFF. If the termination operation flag is OFF, the overall control CPU 101 proceeds to step S420; if the termination operation flag is ON, it proceeds to step S416. In step S416, the overall control CPU 101 determines that the termination operation is complete, sets the termination operation flag and the current frame modification flag to OFF, and proceeds to step S420.

[0412] In step S420, the overall control CPU 101 records the information on the cropping range and the setting of the changing flag for each frame as frame management information in the primary memory 103. Figure 23 illustrates an image of the frame management information after video capture up to frame number j+2.

[0413] Frame management information includes frame number, cropping range, and modification flag information. The frame number indicates which frame in the video it is. The cropping range is indicated, for example, by the position and size of the video recording frame relative to the frame. The modification flag indicates whether the user is in the process of viewing the display device 800 (changing face direction) at the time the frame is captured.

[0414] Furthermore, frame management information is stored in primary memory 103 separately from the video file (video), but it may also be stored as metadata for the video file. In this case, the frame metadata explained in Figure 15 further includes a "modifying flag" item, and the cropping range (cropping) The position and the value of the changing flag are saved as metadata for the video file in step S420, along with the frame management information.

[0415] Referring to Figure 24, the states of the change frame count, change in progress flag, initial operation in progress flag, and completion operation in progress flag during the period when the user was facing the display device 800 will be explained. The states of the various flags will be explained assuming a flow in which the user faces the display device 800 to check the status of the camera body 1, displays the confirmation screen, and then returns the face direction to its original position after confirmation. Figure 24 shows the user's actions, face direction recording, the various flags used in the processing of Figure 22, and the state of the change frame count in chronological order.

[0416] The user starts shooting (A1), and for example, to check the status of the camera body 1, starts changing the face direction to view the display device 800 (A2). Since the changing flag is OFF until the face direction change operation is started, the process in Figure 22 proceeds from step S401 to step S402. In the frame in which the face direction change operation is started, the overall control CPU 101 sets the changing flag and the initial operation flag to ON in step S403, and sets the change frame count to 1.

[0417] In each frame from the start of the face direction change operation until the display device 800 is activated (A4), the change frame count is incremented by the processing from step S404 to step S409. During this time, the change in progress flag is in the ON state.

[0418] When the operation to change the face direction to view the display device 800 stops (A3), the initial operation flag is set to OFF in step S407. When a notification is received from the display device 800 that it has been activated (A4), the change frame count is initialized to 0 in step S410.

[0419] Furthermore, if the change frame count in step S405 is above a threshold, even if a notification is received that the display device 800 has been activated, the face direction change operation is determined not to be an operation for viewing the display device 800. In this case, in the frame management information of Figure 23, the changing flag recorded as ON for the frame after the face direction change operation has started (A2) is rewritten to OFF in step S412.

[0420] If the user has finished checking the screen of the display device 800 and has started the face direction change operation again (A5), the process in Figure 22 proceeds in the order of steps S404, S413, and S418. When the face direction change operation is started again (A5), the initial operation flag is set to OFF, so in step S419 the termination operation flag is set to ON.

[0421] When the face direction repositioning operation stops (A6), the process in Figure 22 proceeds in the order of steps S404, S413, S414, and S415. When the face direction repositioning operation stops (A6), the termination operation flag is set to ON, so in step S416, the termination operation flag and the change operation flag are set to OFF.

[0422] Referring to Figure 25, the recording range development process according to Example 2 will be described. The recording range development process according to Example 2 is a detailed process of step S500 in Figure 7A. In Example 2, the recording range development process shown in Figure 25 is performed instead of the recording range development process shown in Figure 7E, which was described in Example 1. The recording range development process according to Example 2 differs from the recording range development process in Figure 7E in that the process of cutting out the portion of the video recording frame 127i from the ultra-wide-angle video in step S502 is not performed.

[0423] In the recording range development process according to Example 2, the entire ultra-wide-angle video read in step S501 To perform development processing on the area, the amount of computation and the usage of primary memory 103 increase. If primary memory 103 becomes insufficient, the overall control CPU 101 may record the data in the built-in non-volatile memory 102.

[0424] Furthermore, the overall control CPU 101 may, while leaving the processing in step S502, crop a larger area than the video recording frame 127i described in Example 1, thereby limiting the amount of computation and the usage of the primary memory 103 compared to when developing the entire area. In this case, the overall control CPU 101 can change the cropping range within the range of the video cropped from a portion of the entire area during the operating period for viewing the display device 800.

[0425] If a larger area than that in Example 1 is cropped in step S502, the cropping area may be determined, for example, by pre-setting the expected face direction when the user views the display device 800, so that the image includes both the actual viewing direction Vi and the set face direction.

[0426] The processes from steps S200 to S500 in Figure 7A relating to Example 2 are repeatedly executed until recording is completed. Referring to Figure 26, the recording range confirmation and development process, which is executed after the recording range development process in step S500 is completed, will be described. The processes of each step in Figure 26 are realized by the overall control CPU 101 controlling each component of the camera body 1 according to a program stored in the built-in non-volatile memory 102.

[0427] In step S510, the overall control CPU 101 notifies the display device 800 that imaging has finished. In step S511, the overall control CPU 101 reads one frame's worth of frame management information, as explained in Figure 23, from the primary memory 103. The overall control CPU 101 reads the frame management information sequentially from the first frame and repeats the processing in steps S512 to S516 for the read frames.

[0428] In step S512, the overall control CPU 101 determines whether the "Changing" flag is ON or OFF. If the "Changing" flag is not ON, the overall control CPU 101 proceeds to step S513 and executes the extraction process using the extracted range that was read. In step S514, the overall control CPU 101 retains the extracted range (Xi, Yi, WXi, WYi) extracted in step S513. The retained extracted range is used as the extracted range immediately before the "Changing" flag is turned ON, and is used to change the extracted range during the period when the user is viewing the display device 800.

[0429] If the "Changing" flag is ON in step S512, the overall control CPU 101 proceeds to step S515 and changes the extraction range of the currently processed frame to the extraction range (Xi, Yi, WXi, WYi) that was saved in step S514. In step S516, the overall control CPU 101 executes the extraction process using the extraction range set in step S515.

[0430] In step S517, the overall control CPU 101 determines whether the extraction process has been completed for all frames. If it has not been completed, it repeats the processes from steps S511 to S516 until the last frame.

[0431] The case where the processing in Figure 26 is applied to the frame management information in Figure 23 will be explained. Since the changing flag is OFF for frames 1 to i, the overall control CPU 101 proceeds from step S512 to step S513. When the processing in step S514 is executed for frame i, the retained extraction range becomes (Xi, Yi, WXi, WYi).

[0432] In frame i+1, the changing flag is ON, so the overall control CPU 101 steps Proceed to step S515. The overall control CPU 101 sets the extraction range for frame i+1 to (Xi, Yi, WXi, WYi) and executes the extraction process. Since the flag for changing frames i+2 to j+2 is also ON, the overall control CPU 101 sets the extraction range for frames i+2 to j+2 to (Xi, Yi, WXi, WYi) in the same way as frame i+1 and executes the extraction process.

[0433] Once the recording range determination and development process shown in Figure 26 is completed, the overall control CPU 101 executes the processes from steps S600 to S1000 in Figure 7A, similar to the process in Example 1.

[0434] According to the above embodiment 2, during the series of actions taken by the user from the start of imaging to viewing the display device 800, the camera body 1 can capture the image at the face direction it was facing when the face direction began to change. Therefore, even if the user averts their gaze from the subject and turns their face towards the display device 800 to check the status of the camera body 1 while capturing video, the user can continue to record the desired image without interrupting imaging.

[0435] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence.

[0436] (Other embodiments) The present invention can also be implemented by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions. [Explanation of Symbols]

[0437] 1: Camera body, 10: Shooting / detection unit, 13: Face direction detection window, 16: Shooting lens, 20: Face direction detection unit, 30: Recording direction / angle determination unit, 40: Shooting unit, 50: Image cropping / development processing unit, 70: Transmission unit, 101: Overall control CPU, 102: Built-in non-volatile memory, 103: Primary memory, 800: Display device

Claims

1. An imaging device, The photography department, A detection means for detecting the direction of the user's face relative to the imaging device, A setting means for setting a cropping range in each frame image of the video captured by the imaging unit based on the detected face direction, A generation means for generating a cut video from the aforementioned cut range, It has, The generation means generates the extracted video by changing the cropping range set for the frame images during the period in which the user was facing the display device which is communicably connected to the imaging device, to the cropping range set for the frame images before the start of the period in which the user was facing. An imaging device characterized by the following features.

2. The aforementioned operating period is the period from when the user's face direction begins to change until the user's face is turned towards the display device, and then the user's face direction changes again and stops. The imaging apparatus according to feature 1.

3. The generation means, when it receives notification from the display device that the user has turned their face towards the display device within a threshold time from the start of the operation period, changes the cropping range set for the frame images of the operation period to the cropping range set for the frame images before the start of the operation period, and generates the cropped video. The imaging apparatus according to feature 1.

4. The generation means generates the cropped video by changing the cropping range set for the frame images during the operation period to the cropping range set for the frame images before the operation period started, if the number of frames captured from the start of the operation period until the display device receives notification from the display device that the user has turned their face towards the display device is less than a threshold. The imaging apparatus according to feature 1.

5. A method for controlling an imaging device, Imaging step, A detection step for detecting the direction of the user's face relative to the imaging device, A setting step in which, based on the detected face direction, a cropping range is set for each frame image of the video captured in the imaging step, A generation step of generating a cut video from the aforementioned cut range. It has, In the generation step, the cropping range set for the frame images during the period in which the user was facing the display device which is communicably connected to the imaging device is changed to the cropping range set for the frame images before the start of the period in which the user was facing the device, and the cropped video is generated. A control method for an imaging device, characterized by the following:

6. A program for causing a computer to function as each of the means of the imaging apparatus described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the imaging apparatus described in any one of claims 1 to 4.

8. A display device capable of communicating with an imaging device according to any one of claims 1 to 4, A determination means for determining whether the user turned their face towards the display device while the imaging unit was capturing images, If it is determined that the user has turned their face towards the display device, a notification means is provided to notify the imaging device that the user has turned their face towards the display device. A display device characterized by having the following features.

9. The determination means determines that the user has turned their face towards the display device if the display device switches from an off state to an on state, the user is captured by the in-camera provided in the display device, or the screen of an application for checking the status of the imaging device is displayed. The display device according to feature 8.

10. An imaging system comprising an imaging device and a display device capable of communicating with the imaging device, The imaging device is The photography department, A detection means for detecting the direction of the user's face relative to the imaging device, A setting means for setting a cropping range in each frame image of the video captured by the imaging unit based on the detected face direction, A generation means for generating a cut video from the aforementioned cut range, It has, The generation means generates the cropped video by changing the cropping range set for the frame images during the period in which the user was turning their face towards the display device which is communicably connected to the imaging device, to the cropping range set for the frame images before the start of the period in which the user was performing the action of turning their face towards the display device which is communicably connected to the imaging device. The aforementioned display device is A determination means for determining whether the user turned their face towards the display device while the imaging unit was capturing images, If it is determined that the user has turned their face towards the display device, a notification means is provided to notify the imaging device that the user has turned their face towards the display device. has An imaging system characterized by the following: