Camera device, control method thereof, and recording medium

By integrating the camera, search, authentication and control units in the camera device, the timing of automatic camera and authentication is optimized, the balance problem of automatic camera and authentication registration is solved, and the reliability and accuracy of the equipment are improved.

CN114500789BActive Publication Date: 2025-08-19CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202111248077.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-27
Filing Date
2021-10-26
Publication Date
2025-08-19
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

Existing cameras have difficulty finding a balance between automatic camera and automatic authentication registration, resulting in insufficient reliability of automatic camera and personal authentication accuracy, affecting the user experience.

Method used

An imaging device is designed, integrating an imaging unit, a search unit, an authentication registration unit and a control unit. Through the control unit, authentication registration determination and image determination are performed simultaneously during the search process, and the timing of automatic imaging and automatic authentication is optimized.

Benefits of technology

It improves the reliability of automatic camera equipment and the accuracy of personal authentication, ensuring that the images of users are interested in are accurately recorded and specific people are recognized during the automatic camera process, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114500789B_ABST
    Figure CN114500789B_ABST
Patent Text Reader

Abstract

The present invention provides an imaging device, a control method thereof, and a recording medium. The imaging device is capable of automatically capturing images of a subject and automatically registering and authenticating the subject. The imaging device includes a drive unit that rotates and moves the lens barrel in a panning direction and a tilting direction, and is capable of changing the imaging direction under the control of the drive unit. The imaging device is capable of performing automatic authentication registration to search for a subject detected in captured image data and authenticate and store the subject. A first control unit of the imaging device performs a determination regarding whether conditions for performing automatic authentication registration are met and a determination regarding whether conditions for performing automatic imaging are met. The first control unit performs automatic authentication registration determination processing and automatic imaging determination processing while performing a search for automatic imaging, and determines the timing for performing automatic authentication registration based on the determination result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an automatic photographing technology in a photographing device. Background Art

[0002] When using a camera to capture still or moving images, the photographer typically uses a viewfinder to identify the subject, check their camera status, and adjust the framing of the subject. Conventional technology includes mechanisms that detect user operational errors or external conditions, notify the user that the camera is unsuitable for capturing images, or control the camera in a manner that makes it suitable for capturing images.

[0003] Regarding a camera device that takes pictures by user operation, PCT Japanese Patent Publication No. 2016-536868 discloses a life log camera that takes pictures periodically and continuously without the user giving a shooting instruction. The life log camera is used when worn on the user by a belt or the like, and records the scenes that the user sees in daily life as videos at specified time intervals. When taking pictures using the life log camera, the pictures are taken at specified time intervals, rather than at the timing the user wants (such as when the shutter button is pressed, etc.). Therefore, it is possible to record videos of unexpected moments that the user would not normally take pictures of. There is a camera device that automatically takes pictures of objects. Japanese Patent Publication No. 2001-51338 discloses a device that automatically takes pictures when predetermined conditions are met.

[0004] On the other hand, Japanese Patent No. 4634527 discloses a camera device with a personal authentication function for determining a camera object by storing information related to the subject. The camera device preferentially focuses on the subject based on the stored information. Personal authentication is a process of identifying an individual by quantifying the number of features such as the face, but the number of facial features changes depending on the person's growth, the angle of the face, and the way light illuminates the face. Since it is difficult to identify an individual using only a single feature quantity data, the following method is used: by using multiple feature quantity data for the same person, the authentication accuracy is improved. Japanese Patent Laid-Open No. 2007-325285 discloses a method of separating a camera for storing information related to the subject and a camera for photographing the subject. The camera timing for personal authentication and the camera timing for photographing the subject can be controlled independently.

[0005] In the prior art, if the requirements for automatic video recording and automatic authentication and registration are different, it is difficult to simultaneously achieve these two requirements in one video recording. Summary of the Invention

[0006] The present invention is to control the timing of automatic authentication and registration of a subject in an image pickup device capable of automatic image pickup.

[0007] According to an embodiment of the present invention, a camera device is provided, which is capable of performing automatic camera shooting and automatic authentication registration, and the camera device includes: a camera unit, which is configured to shoot a subject; a search unit, which is configured to search for a subject detected in image data acquired by the camera unit; an authentication registration unit, which is configured to authenticate and store the detected subject; and a control unit, which is configured to perform an authentication registration determination related to whether a first condition for the authentication registration unit to perform automatic authentication registration is satisfied, and a camera determination related to whether a second condition for performing automatic camera shooting is satisfied, and control the timing of automatic camera shooting and automatic authentication registration, wherein the control unit determines the timing of automatic authentication registration by performing an authentication registration determination and a camera determination related to the detected subject while controlling the search in the search unit.

[0008] A control method performed in a camera device capable of performing automatic camera photography and automatic authentication registration, the control method comprising: searching for a subject detected in image data acquired by a camera unit; authenticating and registering the detected subject; and performing an authentication registration determination related to whether a first condition for performing automatic authentication registration is satisfied and a camera determination related to whether a second condition for performing automatic camera photography is satisfied, and controlling the timing of automatic camera photography and automatic authentication registration, wherein the control determines the timing of automatic authentication registration by performing an authentication registration determination and a camera determination related to the detected subject while controlling the search for the subject.

[0009] A non-temporary recording medium stores a control program for a camera device capable of performing automatic camera photography and automatic authentication and registration, the control program causing a computer to perform various steps of a control method for the camera device, the control method comprising: searching for a subject detected in image data acquired by a camera unit; authenticating and registering the detected subject; and performing an authentication and registration determination related to whether a first condition for performing automatic authentication and registration is satisfied and a camera determination related to whether a second condition for performing automatic camera photography is satisfied, and controlling the timing of automatic camera photography and automatic authentication and registration, wherein the control determines the timing of automatic authentication and registration by performing an authentication and registration determination and a camera determination related to the detected subject while controlling the search for the subject.

[0010] Further features of the present invention will become apparent from the following description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1A and Figure 1B Schematic diagram of the appearance and driving direction of the camera of the embodiment.

[0012] Figure 2 is a block diagram showing the overall configuration of a camera according to the embodiment.

[0013] Figure 3 is a diagram showing a configuration example of a wireless communication system between a camera and an external device.

[0014] Figure 4 It shows Figure 3 A block diagram of the construction of external devices in FIG.

[0015] Figure 5 is a diagram showing the configuration of a camera and external devices.

[0016] Figure 6 It shows Figure 5 A block diagram of the construction of external devices in FIG.

[0017] Figure 7 is a flowchart showing the operation of the first control unit.

[0018] Figure 8 is a flowchart showing the operation of the second control unit.

[0019] Figure 9A and Figure 9B : is a flowchart showing the image capture mode processing.

[0020] 10A to 10D 3 is a diagram showing the division of regions in a captured image.

[0021] Figure 11 : is a table showing execution determination based on automatic authentication registration determination and automatic imaging determination.

[0022] Figure 12A and Figure 12B is a diagram showing the arrangement of subjects during composition adjustment.

[0023] Figure 13 is a diagram showing a neural network.

[0024] Figure 14 is a diagram showing an image viewing state in an external device.

[0025] Figure 15 is a flowchart illustrating learning mode determination.

[0026] Figure 16 is a flowchart illustrating learning mode processing.

[0027] Figure 17 is a block diagram showing the configuration of an imaging apparatus.

[0028] Figure 18 is a table showing an example of character information.

[0029] Figure 19 : is a diagram showing a screen example of personal information displayed on an external device.

[0030] Figure 20A and Figure 20B It is a diagram showing image data and subject information.

[0031] Figure 21 is a flowchart showing an outline of periodic operations performed by the imaging apparatus.

[0032] Figure 22A and Figure 22B 1 and 2 are a flowchart and a table showing the temporary registration determination process.

[0033] Figure 23A and Figure 23B are a diagram and a table showing image data after the angle of view is adjusted due to the provisional registration determination.

[0034] Figure 24A and Figure 24B 1 and 2 are a flowchart and a table showing the main registration determination process.

[0035] Figure 25 is a flowchart illustrating a first primary registration count determination process.

[0036] Figure 26 is a flowchart illustrating the second primary registration count determination process.

[0037] Figure 27A and Figure 27B 1 and 2 are a flowchart and a table showing the imaging subject determination process.

[0038] Figure 28A and Figure 28B is a diagram showing an example of image data and subject information.

[0039] Figure 29 is a diagram showing an example of an image after the angle of view is adjusted due to imaging subject determination.

[0040] Figure 30 is a diagram showing an example of registered person information.

[0041] Figure 31A and Figure 31B is a diagram showing an example of image data and subject information.

[0042] Figure 32 is a flowchart showing an outline of periodic operations performed by the imaging apparatus.

[0043] Figure 33A and Figure 33B 1 and 2 are a flowchart and a table showing the importance determination process.

[0044] Figure 34A and Figure 34B 1 and 2 are a flowchart and a table showing the imaging subject determination process.

[0045] Figure 35A and Figure 35B is a diagram showing an example of image data and subject information related to a modification. DETAILED DESCRIPTION

[0046] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. First, the technical background of the present invention will be described. For example, in recording for the purpose of lifelogging, recording is performed periodically and continuously, which may result in recording image information that the user is not interested in. Therefore, there is a method that automatically performs panning or tilting operations on the camera to search for surrounding subjects and captures the image at a viewing angle that includes the detected subject. This increases the likelihood that image information that the user prefers will be recorded.

[0047] In an imaging device capable of automatically controlling the shooting direction, it is necessary to search for the subject to be imaged without missing the capture timing. The image composition must be adjusted using the pan, tilt, and zoom mechanisms, taking into account the number of subjects, their movement direction, and the background. Furthermore, the image capture operation must be performed quickly at the capture timing.

[0048] By using personal authentication information, it is possible to detect a subject to be searched with priority, and the personal authentication information can be used to determine a subject included in the angle of view during image capture. Therefore, it is possible to increase the possibility of recording an image that the user prefers.

[0049] However, in cameras capable of automatic recording, if personal authentication registration is not performed automatically, convenience could be significantly reduced. Personal identification processing in personal authentication is performed by quantifying feature quantities obtained from facial images. However, when these values change due to factors such as aging, slight changes in facial angle, and slight adjustments in the lighting on the face, there is a possibility that a person may not be recognized as the same person when they should be. In such cases, if a person is mistakenly identified as someone else due to misidentification during subject tracking control, the camera will track that person, resulting in the camera missing the intended person. Therefore, in cameras capable of automatic recording, the reliability of personal authentication is directly related to the reliability of automatic recording. Regarding the registration information for personal authentication of the same person, it is important to maintain and improve authentication accuracy by using multiple registration information by adding to it over time, and this information needs to be automatically updated. Automatic registration of personal authentication is crucial to achieving higher performance and more convenient automatic recording.

[0050] Highly accurate facial image data is required to register more accurate personal authentication. In other words, it is assumed that the composition is arranged at the optical center which is least affected by the aberration of the optical lens. It is necessary to capture an image of a larger facial area, and it is necessary to use the still image shooting function of the camera device to obtain a high-resolution image focused on the subject. However, in automatic camera shooting, composition adjustment is performed by taking into account multiple subjects and backgrounds so as not to miss the camera opportunity. Therefore, it may not be possible to meet the conditions required for automatic camera shooting and the composition adjustment conditions required for personal authentication registration at the same time. Therefore, in this embodiment, an example of a camera device will be described which is capable of controlling the timing to automatically register personal authentication without interfering with the camera opportunity of automatic camera shooting.

[0051] Figure 1A This figure schematically illustrates the appearance of the imaging apparatus of this embodiment. In addition to a power switch, the camera 101 is equipped with operating members for operating the camera. The lens barrel 102 integrally includes an imaging lens group and an imaging element as the imaging optical system for capturing an image of a subject, and is movably attached to the fixed portion 103 of the camera 101. Specifically, the lens barrel 102 is attached to the fixed portion 103 via a first rotation unit 104 and a second rotation unit 105 (the first and second rotation units 104, 105 are mechanisms capable of rotationally driving relative to the fixed portion 103), thereby enabling its imaging direction to be changed. The first rotation unit 104 drives the lens barrel 102 in the tilt direction (hereinafter referred to as the tilt rotation unit). The second rotation unit 105 drives the lens barrel 102 in the pan direction (hereinafter referred to as the pan rotation unit). An angular velocity meter 106 and an accelerometer 107 are arranged on the fixed portion 103 of the camera 101. For example, the angular velocity meter 106 has a gyro sensor, and the accelerometer 107 has an acceleration sensor.

[0052] Figure 1B This diagram shows the relationship between a three-dimensional orthogonal coordinate system (X-axis, Y-axis, and Z-axis) and three directions (pitch, yaw, and roll). The positions of the X-axis (horizontal axis), Y-axis (vertical axis), and Z-axis (depth axis) relative to the fixed portion 103 are defined. The direction around the X-axis is defined as the pitch direction, the direction around the Y-axis is defined as the yaw direction, and the direction around the Z-axis is defined as the roll direction.

[0053] The tilt rotation unit 104 includes Figure 1B The motor drive mechanism of the lens barrel 102 is shown in the pitch direction. The pan rotation unit 105 includes a Figure 1B The yaw direction shown is a motor drive mechanism that rotationally drives the lens barrel 102. That is, the camera 101 includes a mechanism that rotationally drives the lens barrel 102 in two axial directions.

[0054] The angular velocity meter 106 and the accelerometer 107 output angular velocity detection signals and acceleration detection signals, respectively. Based on the output signals from the angular velocity meter 106 or the accelerometer 107, vibration of the camera 101 is detected, and the tilt rotation unit 104 and the pan rotation unit 105 are rotationally driven. Thus, shake or tilt of the lens barrel 102 is corrected. Based on the output signals from the angular velocity meter 106 and the accelerometer 107, the movement of the camera 101 is detected by using measurement results over a certain period of time.

[0055] Figure 2 is a block diagram showing the overall configuration of the camera 101. The first control unit 223 includes a computing processing unit. The computing processing unit is, for example, a central processing unit (CPU), a microprocessor unit (MPU), or the like. The memory 215 includes, for example, dynamic random access memory (DRAM), static random access memory (SRAM), or the like. The first control unit 223 performs various processes according to programs stored in the non-volatile memory (EEPROM) 216, thereby controlling the various blocks of the camera 101 or controlling data transmission between the blocks. The non-volatile memory 216 is an electrically erasable and programmable memory and stores constants, programs, and the like used to operate the first control unit 223.

[0056] The zoom unit 201 includes a zoom lens that performs zooming (magnification / reduction of the formed subject image). The zoom drive control unit 202 drives and controls the zoom unit 201 and detects the focal length during drive control. The focus unit 203 includes a focus lens that adjusts the focus. The focus drive control unit 204 drives and controls the focus unit 203. The imaging unit 206 includes an imaging element that receives light incident through each lens group and outputs charge information corresponding to the amount of light to the image processing unit 207 as an analog image signal. The zoom unit 201, the focus unit 203, and the imaging unit 206 are arranged in the lens barrel 102.

[0057] The image processing unit 207 performs image processing on the digital image data obtained by applying A / D conversion to the analog image signal. Image processing includes distortion correction, white balance adjustment, color interpolation processing, etc., and the image processing unit 207 outputs the digital image data after image processing. The image recording unit 208 obtains the digital image data output from the image processing unit 207. The digital image data is converted to a recording format such as the Joint Photographic Experts Group (JPEG) format. The converted data is stored in the memory 215 and sent to the video output unit 217, which will be described later.

[0058] The lens barrel rotation drive unit 205 drives the tilt rotation unit 104 and the pan rotation unit 105 to rotationally move the lens barrel 102 in the tilt direction and the pan direction. The device shake detection unit 209 includes an angular velocity meter 106 that detects the angular velocity of the camera 101 in the three axial directions and an accelerometer 107 that detects the acceleration of the camera 101 in the three axial directions. The first control unit 223 calculates the rotation angle of the device, the displacement amount of the device, and the like based on the detection signal from the device shake detection unit 209.

[0059] The voice input unit 213 uses a microphone provided in the camera 101 to acquire a voice signal from around the camera 101, converts the voice signal into a digital voice signal, and transmits the digital voice signal to the voice processing unit 214. The voice processing unit 214 performs voice-related processing such as optimization processing on the input digital voice signal. The voice signal processed by the voice processing unit 214 is transmitted to the memory 215 via the first control unit 223. The memory 215 temporarily stores the image signal and voice signal obtained by the image processing unit 207 and the voice processing unit 214.

[0060] The image processing unit 207 and the voice processing unit 214 read the image signal and the voice signal temporarily stored in the memory 215, encode the image signal, encode the audio signal, etc., to generate a compressed image signal and a compressed voice signal. The first control unit 223 transmits the generated compressed image signal and compressed voice signal to the recording / reproducing unit 220.

[0061] The recording / reproducing unit 220 records the compressed image signal and compressed voice signal generated by the image processing unit 207 and the voice processing unit 214, control data related to image capture, and the like on the recording medium 221. If the voice signal is not compressed and encoded, the first control unit 223 transmits the voice signal generated by the voice processing unit 214 and the compressed image signal generated by the image processing unit 207 to the recording / reproducing unit 220 to be recorded on the recording medium 221.

[0062] The recording medium 221 is a recording medium built into the camera 101 or a recording medium that is detachable from the camera 101. The recording medium 221 can record various data such as compressed image signals, compressed audio signals, and audio signals generated by the camera 101. Typically, a medium having a capacity greater than that of the nonvolatile memory 216 is used as the recording medium 221. For example, as the recording medium 221, any type of recording medium can be used, such as a hard disk, an optical disk, a magneto-optical disk, a CD-R, a DVD-R, a magnetic tape, a nonvolatile semiconductor memory, and a flash memory.

[0063] The recording / reproducing unit 220 reads and reproduces the compressed image signal, compressed voice signal, voice signal, various data, and programs recorded on the recording medium 221. The first control unit 223 transmits the read compressed image signal and compressed voice signal to the image processing unit 207 and the voice processing unit 214, respectively. The image processing unit 207 and the voice processing unit 214 temporarily store the compressed image signal and compressed voice signal in the memory 215, decode the signals according to a predetermined process, and transmit the decoded signals to the video output unit 217.

[0064] Multiple microphones are arranged in the voice input unit 213 of the camera 101. The voice processing unit 214 can detect the direction of sound relative to the plane on which the multiple microphones are arranged, and the detection information is used for subject search or automatic recording, which will be described later. The voice processing unit 214 detects specific voice commands. Voice commands can be, for example, certain pre-registered commands, or, in embodiments where the user can register specific voices in the camera, commands based on registered voices. The voice processing unit 214 also identifies sound scenes. In sound scene recognition, sound scene determination processing is performed using a network that has previously undergone machine learning based on a large amount of voice data. For example, a network for detecting specific scenes such as "cheers," "applause," or "speech" is provided in the voice processing unit 214 to detect specific sound scenes or specific voice commands. When a specific sound scene or specific voice command is detected, the voice processing unit 214 outputs a detection trigger signal to the first control unit 223 or the second control unit 211.

[0065] The second control unit 211 is provided separately from the first control unit 223, which controls the entire camera system, and controls the power supplied to the first control unit 223. The first power supply unit 210 and the second power supply unit 212 respectively supply power for operating the first control unit 223 and the second control unit 211. Pressing the power switch on the camera 101 initially supplies power to both the first control unit 223 and the second control unit 211. As will be described later, the first control unit 223 also controls the cessation of power supply from the first power supply unit 210. Even when the first control unit 223 is not operating, the second control unit 211 is operating, thus receiving information from the device shake detection unit 209 and the voice processing unit 214. The second control unit 211 determines whether to activate the first control unit 223 based on the various input information. If it is determined that the first control unit 223 is to be activated, the second control unit 211 instructs the first power supply unit 210 to supply power to the first control unit 223.

[0066] The voice output unit 218 has a speaker built into the camera 101, and outputs voice with a preset pattern from the speaker, for example, during image capture. The LED control unit 224 controls a light emitting diode (LED) provided on the camera 101. The LED is controlled based on a preset lighting pattern or blinking pattern during image capture, etc.

[0067] The video output unit 217 has, for example, a video output terminal and outputs an image signal to display a video on an external display connected thereto. The voice output unit 218 and the video output unit 217 may be a combined terminal such as a High Definition Multimedia Interface (HDMI) terminal.

[0068] The communication unit 222 is a processing unit that performs communication between the camera 101 and external devices. For example, the communication unit 222 transmits and receives data such as voice signals, image signals, compressed voice signals, and compressed image signals. The communication unit 222 receives commands for starting or ending video recording and control signals related to video recording (such as pan, tilt, and zoom drive), and outputs the commands and control signals to the first control unit 223. Thus, the camera 101 can be driven according to instructions from the external device. The communication unit 222 transmits and receives information between the camera 101 and the external device, such as various parameters related to learning processed by the learning processing unit 219, which will be described later. The communication unit 222 includes a wireless communication module such as an infrared communication module, a Bluetooth communication module, a wireless LAN communication module, a wireless USB, or a GPS receiver.

[0069] The environment sensor 226 detects, in a predetermined cycle, the state of the surrounding environment of the camera 101. The environment sensor 226 is configured using, for example, the following sensors.

[0070] A temperature sensor that detects the temperature around the camera 101

[0071] An air pressure sensor that detects the air pressure around the camera 101

[0072] Illumination sensor, which detects the brightness around the camera 101

[0073] Humidity sensor, which detects the humidity around the camera 101

[0074] UV sensor, which detects the amount of ultraviolet light around the camera 101

[0075] In addition to various types of detection information (temperature, air pressure, illumination, humidity, and UV information), the rate of change at a predetermined time interval can also be calculated based on these various types of information. In other words, the amount of change in temperature, air pressure, illumination, humidity, and UV light can be used to determine automatic recording, etc.

[0076] Reference Figure 3 , communication between the camera 101 and the external device 301 will be described. Figure 3 3 is a diagram showing a configuration example of a wireless communication system between a camera 101 and an external device 301. The camera 101 is a digital camera having an imaging function, and the external device 301 is a smart device including a Bluetooth communication module and a wireless LAN communication module.

[0077] exist Figure 3 , communication between the camera 101 and the external device 301 is indicated by a first communication 302 (see the solid arrow) and a second communication 303 (see the dashed arrow). For example, the first communication 302 is wireless local area network (LAN) communication compliant with the IEEE 802.11 standard series. The second communication 303 is communication with a master-slave relationship between a control station and a subordinate station, such as Bluetooth Low Energy (hereinafter referred to as "BLE"). Wireless LAN and BLE are examples of communication methods. Each communication device has two or more communication functions. For example, one communication function used for communication between a control station and a subordinate station can be used to control other communication functions, but other communication methods can be used. However, the first communication 302 using a wireless LAN or the like can communicate at a higher speed than the second communication 303 using BLE or the like. It is assumed that the second communication 303 consumes less power and / or has a shorter communicable distance than the first communication 302.

[0078] Next, refer to Figure 4 , description will be made of the configuration of the external device 301. The external device 301 includes, for example, a wireless LAN control unit 401 for wireless LAN, a BLE control unit 402 for BLE, and a public wireless control unit 406 for public wireless communication.

[0079] The wireless LAN control unit 401 performs RF control and communication processing for wireless LAN, driver processing for various types of control for communication using wireless LAN that conforms to the IEEE 802.11 standard series, and protocol processing related to communication using wireless LAN. The BLE control unit 402 performs RF control and communication processing for BLE, driver processing for various types of control for communication using BLE, and protocol processing related to communication using BLE. The public wireless control unit 406 performs RF control and communication processing for public wireless communication, driver processing for various types of control for public wireless communication, and protocol processing related to public wireless communication. Public wireless communication is communication that conforms to, for example, the International Multimedia Telecommunications (IMT) standard or the Long Term Evolution (LTE) standard.

[0080] The external device 301 also includes a packet transmission / reception unit 403. The packet transmission / reception unit 403 performs processing for at least one of communication using wireless LAN and BLE, and transmission and reception of packets related to public wireless communication. The external device 301 of this embodiment will be described as performing at least one of transmission and reception of packets during communication, but other communication methods such as circuit switching may also be used in addition to packet switching.

[0081] The control unit 411 of the external device 301 includes a CPU and the like, and controls the entire external device 301 by executing a control program stored in the storage unit 404. The storage unit 404 stores, for example, the control program executed by the control unit 411 and various information such as parameters required for communication. Various operations to be described later are performed by the control unit 411 executing the control program stored in the storage unit 404.

[0082] The global positioning system (GPS) receiving unit 405 receives a GPS signal notified by an artificial satellite, analyzes the GPS signal, and estimates the current position (longitude and latitude information) of the external device 301. Alternatively, there is an embodiment in which the current position of the external device 301 is estimated based on information about wireless networks existing in the surroundings by using a Wi-Fi positioning system (WPS) or the like. For example, assume the following situation: the current GPS position information acquired by the GPS receiving unit 405 is within a preset position range (within a predetermined radius centered on the detection position), or a predetermined position change or a greater change occurs in the GPS position information. In this case, the movement information is notified to the camera 101 via the BLE control unit 402, and the movement information is used as a parameter for automatic shooting and automatic editing to be described later.

[0083] The display unit 407 has a function capable of outputting visually recognizable information (such as a liquid crystal display (LCD) or LEDs) or a function capable of outputting sound (such as a speaker), and presents various information. The operation unit 408 includes, for example, buttons for receiving user operations on the external device 301. The display unit 407 and the operation unit 408 can be configured using, for example, a touch panel.

[0084] The voice input / voice processing unit 409 acquires information related to the user's voice using, for example, a general-purpose microphone built into the external device 301. The user's operation commands can be recognized through voice recognition processing. There is a method for acquiring voice commands through the user's speech using a dedicated application of the external device 301. In this case, specific voice commands for recognition by the voice processing unit 214 of the camera 101 can be registered via the first communication 302 using wireless LAN. The power supply unit 410 supplies the necessary power to each unit of the external device 301.

[0085] The camera 101 and the external device 301 communicate using the wireless LAN control unit 401 and the BLE control unit 402 to transmit and receive data. For example, data such as voice signals, image signals, compressed voice signals, and compressed image signals are transmitted and received. The camera 101 can send shooting instructions, voice command registration data, notifications of predetermined location detection based on GPS location information, and notifications of location movement. Learning data can also be sent and received using a dedicated application on the external device 301.

[0086] Figure 5 1 is a diagram schematically illustrating an example configuration of an external device 501 capable of communicating with the camera 101. For example, the camera 101 is a digital camera having an image pickup function. The external device 501 is a wearable device including various sensing units capable of communicating with the camera 101 using a Bluetooth communication module or the like.

[0087] The external device 501 is configured to be wearable on a user's arm, etc. The external device 501 is equipped with sensors that detect biological information (such as the user's pulse, heart rate, blood flow, etc.) at a predetermined period, an acceleration sensor that can detect the user's motion state, and the like.

[0088] The biological information detection unit 602 of the external device 501 includes, for example, a pulse sensor, a heart rate sensor, and a blood flow sensor for respectively detecting the pulse, heart rate, and blood flow of the user, and a sensor for detecting a change in potential by contacting the skin with a conductive polymer. In this embodiment, the heart rate sensor included in the biological information detection unit 602 will be used for description. The heart rate sensor detects the heart rate of the user by irradiating the skin with infrared light using, for example, an LED, detecting the infrared light that has passed through body tissue using a light receiving sensor, and performing signal processing on the detected infrared light. The biological information detection unit 602 outputs the detected biological information signal to the control unit 607 (refer to Figure 6 ).

[0089] The shake detection unit 603 of the external device 501 detects the user's motion state. The shake detection unit 603 includes, for example, an acceleration sensor or a gyroscope sensor, and acquires movement information and motion detection information. Movement information indicates whether the user is moving based on acceleration information, motion speed, and the like. Motion detection information is information obtained by detecting motion, such as whether the user is swinging their arms and performing actions.

[0090] The external device 501 includes a display unit 604 and an operation unit 605. The display unit 604 outputs visually recognizable information, such as an LCD or LED, etc. The operation unit 605 receives an operation instruction for the external device 501 from a user.

[0091] Figure 6 5 is a block diagram showing the configuration of the external device 501. The external device 501 includes a control unit 607, a communication unit 601, a biological information detection unit 602, a shaking detection unit 603, a display unit 604, an operation unit 605, a power supply unit 606, and a storage unit 608.

[0092] The control unit 607 includes a CPU and the like, and controls the entire external device 501 by executing a control program stored in the storage unit 608. The storage unit 608 stores, for example, the control program executed by the control unit 607 and various information such as parameters required for communication. Various operations to be described later are performed by the control unit 607 executing the control program stored in the storage unit 608. The power supply unit 606 supplies power to each unit of the external device 501.

[0093] The operation unit 605 receives an operation instruction for the external device 501 from the user and notifies the control unit 607. The operation unit 605 obtains information related to the voice uttered by the user by using, for example, a general-purpose microphone built into the external device 501, recognizes the user's operation command through voice recognition processing, and notifies the control unit 607. The display unit 604 outputs visually recognizable information or outputs sound from a speaker or the like, and presents various information to the user.

[0094] The control unit 607 obtains detection information from the biometric information detection unit 602 and the shake detection unit 603 and processes the detection information. The various detection information processed by the control unit 607 is transmitted to the camera 101 via the communication unit 601. For example, the external device 501 can transmit the detection information to the camera 101 when it detects a change in the user's heart rate, or when it detects a change in movement state (such as walking, running, and stopping). The external device 501 can transmit the detection information to the camera 101 when it detects a preset arm swing motion, or when it detects movement of a preset distance.

[0095] Reference Figure 7 , the operation sequence of the camera 101 will be described. Figure 7 1 is a flowchart showing an example of processing performed by the first control unit 223 (main CPU) of the camera 101. When the user operates the power button provided on the camera 101, power is supplied from the first power supply unit 210 to the first control unit 223 and the various components of the camera 101. Power is supplied from the second power supply unit 212 to the second control unit 211. Figure 8 The flowchart of FIG. 1 describes the operation details of the second control unit 211.

[0096] After powering up the device, start Figure 7 In the process of , and in S701, read the start condition. In this embodiment, there are the following three cases related to the conditions for starting the power supply.

[0097] (1) When the power button is manually pressed to turn on the power

[0098] (2) When a startup instruction is sent from an external device (e.g., external device 301) via external communication (e.g., BLE communication) and the power supply is started

[0099] (3) When the power is turned on in response to an instruction from the second control unit 211

[0100] Here, in the case of (3), that is, if the power supply is started in response to an instruction from the second control unit 211, the start-up condition calculated in the second control unit 211 is read. Figure 8 The details are described below. The start condition read here is used as a parameter element during subject search or automatic imaging, which will also be described later. When the start condition is read in S701, the flow proceeds to the processing in S702.

[0101] In S702, detection signals from various sensors are read. The sensor signals read here are as follows.

[0102] A signal from a sensor that detects vibrations (such as a gyro sensor or an acceleration sensor of the device shake detection unit 209)

[0103] Signals related to the respective rotational positions of the tilt rotation unit 104 and the pan rotation unit 105

[0104] The voice signal detected by the voice processing unit 214, the detection trigger signal obtained by specific voice recognition, and the sound direction detection signal

[0105] Detection signal of environmental information from the environmental sensor 226

[0106] Detection signals in various sensors are read in S702 , and the flow proceeds to processing in S703 .

[0107] In S703, the first control unit 223 detects whether a communication instruction has been sent from an external device. If so, it controls communication with the external device. For example, it reads various information from the external device 301. This information includes information related to remote operation using wireless LAN or BLE, the transmission and reception of voice signals, image signals, compressed voice signals, compressed image signals, and the like, operating instructions for recording images from the external device 301, and the transmission of voice command registration data. This information also includes information related to notifications of predetermined location detection based on GPS location information, location movement notifications, and the transmission and reception of learning data. If the user's exercise information, arm movement information, and biometric information such as heart rate need to be updated from the external device 501, information reading using BLE is performed. While the example in which the environmental sensor 226 is mounted on the camera 101 has been described, the environmental sensor 226 can also be mounted on the external device 301 or the external device 501. In this case, in S703, environmental information reading using BLE is performed. Communication reading is performed in S703, and the flow proceeds to S704.

[0108] In S704, the mode setting is determined. Examples of "automatic camera mode" (S710), "automatic editing mode" (S712), "automatic image sending mode" (S714), "learning mode" (S716), and "automatic file deletion mode" (S718) will be described. In the following S705, a process is performed to determine whether the operation mode in S704 is set to the low power consumption mode. The low power consumption mode is a mode that is set when the operation mode is not any one of the "automatic camera mode", "automatic editing mode", "automatic image sending mode", "learning mode", and "automatic file deletion mode". If it is determined in S705 that the operation mode is set to the low power consumption mode, the process enters the process in S706, and if it is determined in S705 that the operation mode is not set to the low power consumption mode, the process enters the process in S709.

[0109] In S706, processing is performed to notify the second control unit 211 (sub-CPU) of various parameters related to the activation factor determined by the second control unit 211. These parameters include a shake detection parameter, a sound detection parameter, and a time lapse detection parameter. The parameter values are changed by learning the parameters in a learning process described later. When the processing in S706 ends, the flow proceeds to S707, and the power supply to the first control unit 223 (main CPU) is turned off, ending the series of processing.

[0110] In S709, processing is performed to determine whether the mode setting in S704 is the automatic camera mode. Subsequently, in S711, S713, S715, and S717, processing is performed to determine the corresponding modes. Here, the mode setting determination processing in S704 will be described. In the mode setting determination, a mode is selected from the modes shown below (1) to (5).

[0111] (1) Automatic camera mode

[0112] <Mode determination condition>

[0113] The conditions are: determination that automatic imaging is necessary based on various detection information set for learning, the time elapsed since the transition to automatic imaging mode, past imaging information, the number of captured images, etc. Various detection information refers to information such as images, sounds, time, vibrations, locations, physical changes, and environmental changes.

[0114] <Processing in mode>

[0115] If the mode setting is determined to be automatic shooting mode in S709, the process proceeds to automatic shooting mode processing (S710). Based on the detection information for the learning setting, pan / tilt or zoom drive is performed, and an automatic search for a subject is performed. When it is determined that the timing for shooting according to the photographer's preference has arrived, shooting is automatically performed.

[0116] (2) Automatic editing mode

[0117] <Mode determination condition>

[0118] The condition is that automatic editing needs to be performed based on the time that has elapsed since the time when the previous automatic editing was performed and past captured image information.

[0119] <Processing in mode>

[0120] If it is determined in S711 that the mode setting is the automatic editing mode, the process enters the automatic editing mode process (S712). A still image or moving image selection process based on learning is performed, and an automatic editing process based on the learning is performed to create a wonderful video combined into one video according to the image effect or the time of the edited video.

[0121] (3) Automatic image sending mode

[0122] <Mode determination condition>

[0123] The condition is that if the automatic image transmission mode is set in response to an instruction of a dedicated application using the external device 301, automatic transmission needs to be performed based on the time elapsed from the time of previous image transmission and past captured image information.

[0124] <Processing in mode>

[0125] If it is determined in S713 that the mode setting is the automatic image transmission mode, the flow proceeds to the automatic image transmission mode process (S714). The camera 101 automatically extracts images according to the user's preferences and automatically transmits images that appear to suit the user's preferences to the external device 301. The image extraction according to the user's preferences is performed based on a score (described later) added to each image to determine the user's preferences.

[0126] (4) Learning model

[0127] <Mode determination condition>

[0128] The conditions are: the need for automatic learning based on the time elapsed since the previous learning process was performed, information integrated with images available for learning, or the amount of learning data, etc. Alternatively, if there is an instruction to set the learning mode from the external device 301 through communication, the learning mode is also set.

[0129] <Processing in mode>

[0130] If it is determined in S715 that the mode setting is learning mode, the process enters learning mode processing (S716). Based on various operation information in the external device 301, notification of learning information from the external device 301, etc., learning is performed according to the user's preferences by using a neural network. Various operation information includes image acquisition information from the camera, information manually edited using a dedicated application, and user-defined value information input for the image in the camera. Simultaneously, learning is performed for detection such as registration of personal authentication, voice registration, sound scene registration, and general object recognition registration, as well as learning the conditions of the low power mode described above.

[0131] (5) Automatic file deletion mode

[0132] <Mode determination condition>

[0133] The conditions are that automatic file deletion needs to be performed based on the time that has passed since the time when the previous automatic file deletion was performed and the remaining capacity of the nonvolatile memory 216 in which the image data is recorded.

[0134] <Processing in mode>

[0135] If it is determined in S717 that the mode setting is the automatic file deletion mode, the flow proceeds to the automatic file deletion mode process (S718), which performs the following processing: among the images in the nonvolatile memory 216, based on the tag information and the shooting date and time of each image, a file to be automatically deleted is specified and deleted.

[0136] When completed Figure 7 When the processing in S710, S712, S714, S716 and S718 in the process, the process returns to S702 and continues processing. The details of the processing (S710 and S716) in each mode will be described later. Figure 7 If it is determined in S709 that the mode setting is not the automatic camera mode, the process enters the processing in S711. If it is determined in S711 that the mode setting is not the automatic editing mode, the process enters the processing in S713. If it is determined in S713 that the mode setting is not the automatic image sending mode, the process enters the processing in S715. If it is determined in S715 that the mode setting is not the learning mode, the process enters the processing in S717. If it is determined in S717 that the mode setting is not the automatic file deletion mode, the process returns to S702 and repeats the processing. The automatic editing mode, the automatic image sending mode and the automatic file deletion mode are not directly related to the spirit of the present invention, so their detailed descriptions will be omitted.

[0137] Figure 8 1 is a flowchart illustrating an example of processing performed by the second control unit 211 of the camera 101. When a user operates a power button provided on the camera 101, power is supplied from the first power supply unit 210 to the first control unit 223 and the various components of the camera 101. Power is supplied from the second power supply unit 212 to the second control unit 211.

[0138] After power is supplied, the second control unit (sub-CPU) 211 starts up and begins Figure 8 In S801, a process is performed to determine whether a predetermined sampling period has elapsed. The predetermined sampling period is set to, for example, 10 milliseconds (msec), and the flow proceeds to the process in S802 based on the determination result of the 10 msec period (when the predetermined sampling period has elapsed). If it is determined that the predetermined sampling period has not elapsed, the second control unit 211 waits until the determination process in S801 is performed again.

[0139] In S802, the learning information is read. Figure 7 The information is sent when information is sent to the second control unit 211 through communication in S706, and includes, for example, information used for the following determination.

[0140] (1) Information for determining specific vibration state detection (S804 described later)

[0141] (2) Information for determining specific sound detection (S805 described later)

[0142] (3) Information for determining time elapse detection (S807 described later)

[0143] After processing in S802, the process proceeds to S803, where a shake detection value is acquired. The shake detection value is the output value from the gyro sensor, acceleration sensor, etc. of the device shake detection unit 209. Next, the process proceeds to S804, where processing for detecting a preset specific shake state is performed. Here, several examples of changing the determination process based on the learning information read in S802 will be described.

[0144] <Tap Detection>

[0145] The tap state is, for example, a state in which the user taps the camera 101 with a fingertip or the like, and the tap state can be detected based on an output value from an acceleration sensor attached to the camera 101. The output from the three-axis acceleration sensor is processed by passing the output through a bandpass filter (BPF) set in a specific frequency region within a predetermined sampling period, and the component of the signal region in which the acceleration changes due to the tap is extracted. The number of times the acceleration signal after passing through the BPF exceeds a predetermined threshold (indicated by ThreshA) within a predetermined time (indicated by TimeA) is measured. A tap determination is made based on whether the measured number of times is a predetermined number of times (indicated by CountA). For example, in the case of a double-click, the value of CountA is set to 2, and in the case of a triple-click, the value of CountA is set to 3. The values of TimeA and ThreshA can also be changed based on learning information.

[0146] <Detection of vibration status>

[0147] The shake state of camera 101 can be detected based on the output value of a gyro sensor or accelerometer attached to camera 101. The output from the gyro sensor or accelerometer undergoes absolute value conversion after its high-frequency components are filtered by a high-pass filter (HPF) and its low-frequency components are filtered by a low-pass filter (LPF). The number of times the calculated absolute value exceeds a predetermined threshold (indicated by ThreshB) during a predetermined time (indicated by TimeB) is measured. Vibration is detected based on whether the measured number is equal to or greater than a predetermined number (indicated by CountB). For example, it is possible to determine whether the camera 101 is placed on a table or the like (i.e., a state with less shake) or whether the camera 101 is worn on the body as a wearable camera and the user is walking (i.e., a state with more shake). By setting multiple conditions in conjunction with the conditions for determining the threshold or the number of counts, detailed shake state detection can be performed based on the shake level. The values of TimeB, ThreshB, and CountB can be changed based on learning information.

[0148] In the above examples, a method for detecting a specific shake state by determining the detection value from the shake detection sensor has been described. Alternatively, there is a method in which data from the shake detection sensor, sampled within a predetermined time period, is input to a shake state determination device using a neural network (also referred to as an NN), thereby detecting a pre-registered specific shake state using a trained NN. In this case, in S802 (Reading of Learning Information), the weight parameters of the NN are read.

[0149] After the detection process in S804 is performed, the flow proceeds to the process of S805, and a process for detecting a preset specific sound is performed. Here, several examples of changing the detection determination process based on the learning information read in S802 will be described.

[0150] <Specific Voice Command Detection>

[0151] In the process of detecting the specific voice command, the specific voice command includes several commands registered in advance and commands based on a specific voice registered in the camera by the user.

[0152] <Specific Voice Recognition>

[0153] The network, which has previously undergone machine learning based on a large amount of speech data, detects sound scenes. For example, it can detect specific scenes such as "cheers," "applause," or "speech." The detection target scene changes through learning.

[0154] <Sound Level Determination>

[0155] The sound level is detected by determining whether the size of the voice level exceeds a predetermined size (threshold) within a predetermined time (threshold time). The threshold time, threshold value, etc. are changed through learning.

[0156] <Sound direction confirmed>

[0157] A sound direction is detected for a sound having a predetermined loudness by a plurality of microphones arranged on a plane.

[0158] The above-described determination processing is performed in the voice processing unit 214 , and it is determined in S805 whether a specific sound is detected based on the respective settings learned in advance.

[0159] After the detection processing in S805 is performed, the process enters the processing in S806, and the second control unit 211 determines whether the power supply to the first control unit 223 is turned off. If it is determined that the first control unit 223 (main CPU) is in the OFF state, the process enters the processing of S807, and if it is determined that the first control unit 223 (main CPU) is in the ON state, the process enters the processing of S811. In S807, a process is performed to detect whether a preset time has passed. Here, the detection determination process changes according to the learning information read in S802. The learning information is Figure 7 Information sent when information is transmitted to the second control unit 211 via communication in S706. The time elapsed from the ON state of the first control unit 223 to the OFF state is measured. If the measured elapsed time is equal to or longer than the predetermined time (indicated by TimeC), it is determined that the predetermined time has elapsed. If the measured elapsed time is shorter than TimeC, it is determined that the predetermined time has not elapsed. TimeC is a parameter that changes based on the learning information.

[0160] After the detection process of S807 is performed, the flow proceeds to the process of S808, and a process of determining whether the cancellation condition of the low power consumption mode is satisfied is performed. The cancellation of the low power consumption mode is determined based on the following conditions.

[0161] (1) Specific jitter is detected

[0162] (2) Detecting a specific sound

[0163] (3) The scheduled time has passed

[0164] For (1), it is determined in S804 (specific shaking state detection process) whether specific shaking is detected. For (2), it is determined in S805 (specific sound detection process) whether specific sound is detected. For (3), it is determined in S807 (time lapse detection process) whether a predetermined time has passed. If at least one of the conditions (1) to (3) is met, it is determined that the low power consumption mode is canceled. If it is determined in S808 that the low power consumption mode is canceled, the process proceeds to the process of S809, and if the conditions for canceling the low power consumption mode are not met, the process returns to S801 and continues processing.

[0165] In S809, the second control unit 211 turns on the power supply to the first control unit 223, and in S810 notifies the first control unit 223 of the condition (any one of vibration, sound, and time) for determining the low power consumption mode to be canceled. The flow returns to S801 and the processing continues.

[0166] On the other hand, if the transition from S806 to S811 occurs (if it is determined that the first control unit 223 is in the ON state), the flow enters the processing in S811. In S811, the information acquired in S803 to S805 is notified to the first control unit 223, and then the flow returns to S801 so that the processing is continued.

[0167] In the present embodiment, there is a configuration in which, even if the first control unit 223 is in the ON state, the second control unit 211 performs shake detection or specific sound detection and notifies the first control unit 223 of the detection result. The present embodiment is not limited to this example, and a configuration may be provided in which, if the first control unit 223 is in the ON state, the processing in S803 to S805 is not performed, and the processing in the first control unit 223 ( Figure 7 S702 in the embodiment performs jitter detection or specific sound detection.

[0168] As mentioned above, Figure 7 The processing in S704 to S707 and Figure 8 The processing in , and thereby learn the conditions for transitioning to low power consumption mode or the conditions for canceling low power consumption mode based on the user's operation. That is, the camera can be operated according to the usability of the user who owns the camera 101. The learning method will be described later.

[0169] In the above examples, a method for canceling low-power mode based on vibration detection, sound detection, and the passage of time has been described in detail, but low-power mode can be canceled based on environmental information. Cancellation can be determined based on whether the absolute value or change in temperature, air pressure, illuminance, humidity, or ultraviolet light as environmental information exceeds a predetermined threshold, and the threshold can be changed through learning, which will be described later. The detection information of vibration detection, sound detection, and the passage of time, or the absolute value or change in each piece of environmental information, can be determined based on a neural network, and the determination to cancel low-power mode can be made. In this determination process, the determination conditions can be changed through learning, which will be described later.

[0170] Reference Figure 9A and Figure 9B , will describe Figure 7 First, in S901 (image recognition processing), the image processing unit 207 performs image processing on the signal obtained by the imaging unit 206 to generate an image for subject detection. The generated image is subjected to subject detection processing for detecting a person, an object, etc.

[0171] If a person is detected as a subject, the subject's face or body is detected. In face detection processing, a pattern for identifying a person's face is pre-set, and the portion of the captured image that matches this pattern is detected as the person's facial area. A reliability level is also calculated, indicating the certainty of the subject's face. For example, the reliability level is calculated based on the size of the facial area in the captured image and the degree of agreement indicating the degree of match with the facial pattern. This also applies to object recognition, allowing objects matching a pre-registered pattern to be identified.

[0172] There is a method for extracting characteristic subjects by using histograms of hue, saturation, etc. in a captured image. For an image of a subject captured within a camera angle of view, the following processing is performed: a distribution derived from a histogram of hue, saturation, etc. is divided into a plurality of intervals, and the captured image is classified for each interval. For example, a histogram of a plurality of color components is created for the captured image, and the image is divided by a mountain-shaped distribution range. Images captured in areas belonging to the same interval combination are classified, and the image area of the subject is identified. Evaluation values of each identified image area of the subject are calculated, and thereby the image area of the subject with the highest evaluation value can be determined as the main subject area. According to the above method, each subject information can be obtained from the camera information.

[0173] In S902, image blur correction amount calculation processing is performed. Specifically, the absolute angle of camera shake is first calculated based on the information related to angular velocity and acceleration acquired by the device shake detection unit 209. The image blur correction amount is obtained by driving the tilt rotation unit 104 and pan rotation unit 105 in an angular direction that cancels out the absolute angles and obtaining an angle for correcting image blur. The calculation method used in this image blur correction amount calculation processing can be changed through the learning process described later.

[0174] In S903, the camera state is determined. The camera's current vibration / movement state is determined using the camera angle and camera movement detected based on angular velocity information, acceleration information, GPS location information, and the like. For example, assume that camera 101 is attached to a vehicle and capturing images. In this case, subject information, such as the surrounding scenery, changes significantly depending on the vehicle's travel distance. Therefore, a determination is made as to whether the vehicle is in a "vehicle moving state" where camera 101 is attached and moving at high speed. This determination result is used for automatic subject search, which will be described later. A determination is made as to whether the angle of camera 101 has significantly changed. A determination is made as to whether camera 101 is in a "stationary imaging state" with little to no vibration. If camera 101 is in a "stationary imaging state," it can be determined that the position of camera 101 itself has not changed. In this case, a subject search for stationary imaging can be performed. If the angle of camera 101 has significantly changed, it is determined that camera 101 is in a "handheld state." In this case, a subject search for handheld imaging can be performed.

[0175] In S904, a subject search process is performed. The subject search includes the following processes.

[0176] (1) Regional division

[0177] (2) Calculate the importance of each region

[0178] (3) Determine the search area

[0179] Hereinafter, each process will be described in sequence.

[0180] (1) Regional division

[0181] Reference 10A to 10D , the region division will be described. Set the origin O of the three-dimensional orthogonal coordinate system to the camera position. Figure 10A is a schematic diagram showing an example in which regions are divided around the entire circumference with the camera position (origin O) as the center. Figure 10AIn the example of , the area is divided into a plurality of areas at intervals of 22.5 degrees in each direction in the tilt direction and the pan direction. In the case of such division, as the tilt angle deviates from 0 degrees, the circumference in the horizontal direction becomes smaller and the area becomes smaller. On the other hand, Figure 10B is a diagram illustrating an example in which if the tilt angle is 45 degrees or more, the horizontal area range is set to be greater than 22.5 degrees. Figure 10C and Figure 10D : is a schematic diagram showing an example of area division areas within an imaging angle of view. Figure 10C The axis 1301 shown in FIG. 1 represents the direction of the camera 101 when initialized, and the area is divided based on the direction of the axis 1301 as a reference direction. The angle of view area 1302 for capturing an image is shown in FIG. Figure 10D An example of an image corresponding to this region is shown in FIG. In an image with a camera viewing angle, based on region division, such as Figure 10D An example of a plurality of divided regions 1303 to 1318 is shown.

[0182] (2) Calculate the importance of each region

[0183] For each divided area, an importance level indicating search priority is calculated based on the conditions of the subjects or scenes in that area. The importance level based on subject conditions is calculated based on factors such as the number of people in the area, the size of the person's face, the direction of the face, the certainty of face detection, the person's facial expression, and the person's personal authentication results. The importance level based on scene conditions is calculated based on factors such as general object recognition results, scene discrimination results (blue sky, backlight, evening scene, etc.), sound levels detected from the direction of the area or speech recognition results, and motion detection information within the area.

[0184] If in Figure 9A If the vibration of the camera is detected in the camera state determination (S903) in the image processing, the importance may be changed according to the shaking state. For example, it is assumed that the camera is determined to be in a "stationary imaging state". In this case, it is determined that the subject search is centered on a subject with a high priority among the subjects registered by using facial recognition (for example, the owner of the camera). For example, in the automatic imaging described later, priority is given to capturing the face of the owner of the camera. Therefore, even if the owner of the camera takes a long time to capture images while wearing the camera, many images of the owner can be recorded by removing the camera and placing the camera on a table or the like. In this case, since the face can be searched by panning or tilting, the user can record an image of the owner or a group photo of many faces simply by appropriately mounting the camera without considering the camera placement angle, etc.

[0185] Under these conditions alone, the areas with the highest importance are likely to remain the same unless there are changes in the respective areas. Consequently, the areas to be searched will never change. Therefore, a process is performed to change the importance based on past image capture information. Specifically, the importance of areas that have been continuously designated as search areas for a predetermined period of time is reduced, or the importance of areas captured in S910 (described later) is reduced for a predetermined period of time.

[0186] (3) Determine the search area

[0187] The following processing is performed: based on the importance of each area calculated as described above, an area with high importance is determined as a search target area. Search target angles for panning and tilting required to capture the search target area at a certain viewing angle are calculated.

[0188] exist Figure 9A In S905, pan and tilt drive is performed. Specifically, the pan and tilt drive amounts are calculated by adding the image blur correction amount for controlling the sampling frequency to the drive angle based on the search target angle for pan and tilt. The tilt rotation unit 104 and pan rotation unit 105 are driven and controlled by the lens barrel rotation drive unit 205.

[0189] In S906, zoom drive is performed by controlling the zoom unit 201. Specifically, zoom drive is performed based on the state of the search target object determined in S904. For example, assume that the search target object is a person's face. In this case, if the face size in the image is too small, the face may not be detected because it is smaller than the minimum detectable size, and the subject may be lost. In this case, control is performed as follows: zoom control is performed to increase the face size in the image by performing zoom control toward the telephoto side. On the other hand, if the face size in the image is too large, the subject may easily deviate from the angle of view due to movement of the subject or the camera itself. In this case, control is performed as follows: zoom control is performed to reduce the face size in the image by performing zoom control toward the wide-angle side. Zoom control is performed as described above, thereby maintaining a state suitable for tracking the subject. Zoom control includes optical zoom control performed by driving the lens and electronic zoom control for changing the angle of view through image processing. There are forms in which one control is performed or a form in which both controls are combined.

[0190] S907 is the process for determining automatic authentication registration. Based on the subject detection status, a determination is made as to whether automatic registration for personal authentication is possible. If the face detection reliability is high and remains high, a more detailed determination is made. Specifically, if the face is facing the front of the camera, not the side, and if the face size is equal to or larger than a predetermined value, the state is determined to be suitable for automatic registration for personal authentication.

[0191] The following S908 is an automatic image capture determination process. In the automatic image capture determination, whether to perform automatic image capture and the image capture method are determined (determining whether to perform still image capture, moving image capture, continuous image capture, panoramic image capture, etc.). The determination of whether to perform automatic image capture will be described later.

[0192] In S909, it is determined whether there is a manual camera instruction. Manual camera instructions include instructions given by pressing the shutter button, instructions given by tapping the camera housing with a finger or the like (tapping), instructions given by inputting a voice command, instructions from an external device, and the like. For example, when the user taps the camera housing, the camera instruction triggered by the tap operation is determined by detecting continuous high-frequency acceleration in a short period of time using the device shake detection unit 209. The voice command input method is a camera instruction method in which if the user speaks a password for giving a predetermined camera instruction (for example, "take a picture"), the voice processing unit 214 recognizes the voice and triggers the camera. The instruction method from an external device is a camera instruction method triggered by a shutter instruction signal, which is sent by using, for example, a smartphone connected to the camera via Bluetooth.

[0193] If it is determined in S909 that there is a manual camera instruction, the process proceeds to the processing in S910. If it is determined in S909 that there is no manual camera instruction, the process proceeds to the processing in S914. In S914, it is determined to perform automatic authentication registration. By using the determination result of the possibility of automatic authentication registration in S907 and the determination result of the possibility of automatic camera in S908, it is determined whether to perform automatic authentication registration. If it is determined in S914 that automatic authentication registration is to be performed, the process proceeds to the processing in S915, and if it is determined that automatic authentication registration is not to be performed, the process proceeds to the processing in S916. Reference will be made to Figure 11 A specific example is described.

[0194] Figure 11 This table shows the determination of automatic authentication registration and automatic image capture. The automatic authentication registration determination result is "Registration possible" or "Registration impossible," and the automatic image capture determination result is "Image capture possible" or "Image capture impossible." If personal authentication is determined to be suitable for registration, the personal authentication will be registered regardless of the automatic image capture determination result. If personal authentication is determined to be unsuitable for registration and the automatic image capture conditions are met ("Image capture possible"), automatic image capture will be performed.

[0195] The reason for prioritizing automatic authentication and registration is that it requires stable frontal facial information. Automatic recording can be determined based on factors such as the subject's silhouette, a temporary smile, and the time elapsed since the previous recording. However, the conditions for automatic authentication and registration are not often met. Therefore, in this embodiment, an algorithm is provided that prioritizes situations where the conditions suitable for automatic authentication and registration are met.

[0196] There may be a view that giving priority to automatic authentication registration will hinder the opportunity for automatic shooting. However, this view is wrong because automatic authentication registration improves the accuracy of personal authentication and the accuracy of searching and tracking priority subjects, which is very useful for finding cameras in automatic shooting. In this embodiment, if it is determined that personal authentication is suitable for registration, it will always take priority over the result of whether automatic shooting is possible. This embodiment is not limited to this. In automatic shooting, the priority can be changed according to the number of shots or shooting intervals within a predetermined time. For example, if the shooting frequency in automatic shooting is low, control can be performed so that automatic shooting is temporarily prioritized.

[0197] Figure 9B S915 in the figure is a personal authentication registration process. A series of processes are performed as follows, in which a camera process is performed by controlling the state to be suitable for personal authentication, and facial features are quantified and stored. Figure 12A and Figure 12B Describe its details.

[0198] Figure 12A and Figure 12B is a schematic diagram showing the arrangement of subjects during composition adjustment. Figure 12A shows the composition when shooting still images automatically. Figure 12B The composition of the image when shooting for personal authentication is shown. Figure 12B The state shown is suitable for personal authentication. In order to obtain facial features more accurately, it is important to place the subject in the center of the image that is less susceptible to optical aberrations and adjust the composition so that the face can be captured in a large size. On the other hand, if a still image is automatically captured in S910 to be described later, it is best to adjust the composition so that the face is captured in a large size. Figure 12A The main subject and background shown match, resulting in a more pleasing photo.

[0199] If a manual capture command is generated from the user during the personal authentication and registration process, the camera mode process can be terminated by pausing S915 and then re-executed. Composition adjustment control involves repeated panning, tilting, and zoom lens driving, as well as facial position checking based on face detection. Manual capture commands are constantly checked during this repetitive process, and if a break is detected, the user's intent can be promptly reflected by stopping the personal authentication and registration process.

[0200] Automatic video recording is a video recording that automatically records the image data output by the video recording unit. Figure 9B In S916, whether to perform automatic photographing is determined as follows. Specifically, it is determined that automatic photographing is to be performed in the following two cases. The first case is: based on the importance of each area obtained in S904, the importance exceeds a predetermined value. The second case is: using the determination result based on the neural network, which will be described later. Recording in automatic photographing is recording the image data in the memory 215 or recording the image data in the non-volatile memory 216. It is also assumed that the image data is automatically sent to the external device 301 and the image data is recorded in the external device 301.

[0201] In this embodiment, automatic imaging is controlled using a neural network-based automatic imaging determination process. Depending on the conditions of the imaging location and the camera, it may be desirable to change the parameters for automatic imaging. Unlike imaging at regular intervals, automatic imaging control based on condition determination is preferably performed in a manner that meets the following requirements.

[0202] (1) It is desirable to capture a large number of images including people and objects.

[0203] (2) Don’t expect to miss unforgettable scenes.

[0204] (3) Taking into account the remaining battery level and the remaining capacity of the recording medium, it is desirable to capture images with low power consumption.

[0205] If an evaluation value is calculated based on the state of the object, the evaluation value is compared with a threshold value, and if the evaluation value exceeds the threshold value, automatic imaging is performed. The evaluation value in the automatic imaging is determined based on determination using a neural network.

[0206] Next, determination based on a neural network (NN) will be described. As an example of NN, an example of a network using a multilayer perceptron is described in Figure 13 NN is used to predict output values from input values. Input values and output values used as a model for the input are learned in advance, and thus output values that follow the learned model can be estimated for new input values. The learning method will be described later.

[0207] Figure 13Node 1201 and the multiple nodes indicated by vertical circles in the diagram represent neurons in the input layer. Node 1203 and the multiple nodes indicated by vertical circles in the diagram represent neurons in the intermediate layer. Node 1204 represents a neuron in the output layer. Arrows 1202 indicate connections connecting the neurons to each other. In the NN-based determination, feature quantities based on the state of the subject, scene, or camera captured from the current viewing angle are input to the neurons in the input layer. The value output from the output layer is obtained through calculation based on the forward propagation rule of the multilayer perceptron. If the output value is equal to or greater than a threshold, it is determined that automatic recording is to be performed.

[0208] For example, the following information is used as the characteristics of the subject.

[0209] Information related to the recognition results of general objects at the current zoom factor and current viewing angle

[0210] Face detection results, number of faces captured in the current viewing angle, degree of facial smile, degree of eye closure, facial angle, facial authentication ID number, and subject’s gaze angle

[0211] Scene recognition results, time elapsed since the previous image was taken, current time, GPS location information, and change from the previous image position

[0212] Current voice level, speaker, applause, and information about whether cheers are being issued

[0213] Vibration information (acceleration information and camera status), environmental information (temperature, air pressure, illumination, humidity, ultraviolet radiation), etc.

[0214] If there is information notification from external device 501, the notification information (such as the user's exercise information, arm movement information, and biometric information such as heart rate) is also used as feature information. The feature information is converted into a numerical value within a predetermined range and assigned to each neuron in the input layer as a feature quantity. Therefore, the number of neurons required in the input layer is the same as the number of feature quantities to be used.

[0215] In the determination based on the neural network, the output value can be changed by changing the connection weights between the neurons in the learning process described later, so that the determination result can be adapted to the learning result.

[0216] The determination of automatic camera is also based on Figure 7 The first control unit 223 is changed based on the activation condition read in step S702. For example, if the activation is due to tap detection or a specific voice command, this operation is likely to be an instruction to currently record images in accordance with the user's intention. Therefore, the recording frequency is set to high.

[0217] In determining the imaging method, it is determined whether to perform imaging determined based on the state of the camera and the state of the surrounding subjects detected in S901 to S904. It is determined which of still image shooting, moving image shooting, continuous imaging, panoramic imaging, etc. is to be performed. For example, if the person serving as the subject is still, still image shooting is selected and performed. If the subject is moving, moving image shooting or continuous imaging is performed. If there are multiple subjects surrounding the camera, or if the location is determined to be a scenic area based on GPS information, panoramic imaging processing is performed. Panoramic imaging processing is a process of generating a panoramic image by synthesizing images captured sequentially while performing pan and tilt driving of the camera. In the same manner as the method of determining whether to perform automatic imaging, the imaging method can be determined by determining various information detected before imaging based on a neural network. In this determination process, the determination conditions can be changed by the learning process described later.

[0218] exist Figure 9B In S916, if it is determined in the automatic imaging determination process in S908 that automatic imaging is to be performed, the flow proceeds to the process in S910. If it is determined in S916 that automatic imaging is not to be performed, the imaging mode process ends. After S915 (automatic authentication registration process), the imaging mode process ends.

[0219] In S910, automatic imaging begins. That is, imaging begins according to the imaging method determined in S908. In this case, the focus drive control unit 204 performs automatic focus control. Furthermore, exposure control is performed using an aperture control unit, a sensor gain control unit, and a shutter control unit (not shown) to adjust the subject to appropriate brightness. After imaging, the image processing unit 207 performs various well-known image processing types (such as automatic white balance processing, noise reduction processing, and gamma correction processing) to generate image data.

[0220] If the predetermined condition is satisfied during the automatic image capturing in S910, the camera may notify the image capturing subject person to capture the image and then may capture the image. For example, the predetermined condition is set based on the following information.

[0221] The number of faces in the field of view, the degree of facial smile, the degree of eye closure, the subject's gaze angle or facial angle, and the facial authentication ID number

[0222] · Number of people registered for personal authentication and recognition results of general objects during video recording

[0223] Information on whether the current location is a scenic spot based on the scene recognition result, the time elapsed since the previous image capture, the image capture time, and GPS information

[0224] Information related to the sound level at the time of recording, the presence of speakers, applause, and cheers

[0225] Vibration information (acceleration information and camera status), environmental information (temperature, air pressure, illumination, humidity, and ultraviolet light), etc.

[0226] Notification methods include, for example, using sound emitted from the voice output unit 218 and lighting of the LED by the LED control unit 224. By performing image capture with notification based on these conditions, it is possible to record optimal images of the camera's line of sight in crucial scenes. For notification before image capture, information related to the captured image or various information detected before image capture can be determined using a neural network, and the notification method and timing can be determined. This determination process can be modified through the learning process described later.

[0227] In S911, editing processing such as processing the image generated in S910 and adding the processed image to a moving image is performed. Specifically, the image processing includes a cropping process based on a person's face or focus position, an image rotation process, a high dynamic range (HDR) effect process, a blur effect process, a color conversion filter effect process, and the like. In the image processing, a plurality of processed images are generated by a combination of the above processes based on the image data generated in S910. A process of storing the image data separately from the image data generated in S910 may be performed. For moving image processing, a process of adding a captured moving image or a still image to a generated edited moving image while performing special effect processing such as sliding, zooming, and fading is performed. With respect to the editing processing in S911, the image processing method can be determined by determining information related to the captured image based on a neural network or various information detected before shooting. In this determination process, the determination conditions can be changed by a learning process described later.

[0228] In S912, a learning information generation process for the captured image is performed. This process generates and records information used for the learning process described later. Specifically, for example, the following information exists.

[0229] The zoom ratio at the time of shooting, general object recognition results at the time of shooting, face detection results, the number of faces in the captured image, the degree of smiling, the degree of eye closure, the angle of the face, the face authentication ID number, and the angle of the subject's gaze in the image captured this time

[0230] Scene recognition results, time elapsed since the last image was taken, image capture time, GPS location information, and change from the last image capture location

[0231] Information about the sound level, speakers, applause, and whether cheers were given during filming

[0232] Vibration information (acceleration information and camera status), environmental information (temperature, air pressure, illumination, humidity, ultraviolet radiation), etc.

[0233] ・Motion image shooting time, information related to whether there is a manual shooting instruction, etc.

[0234] A score is calculated as the output of the neural network that quantifies the user's preference for the image. This information is generated and recorded as label information in the captured image file. Alternatively, there are methods of storing the information in non-volatile memory 216 or storing information related to each captured image in a list format as so-called catalog data in recording medium 221.

[0235] In S913, the past image capture information is updated. Specifically, this process updates the number of images captured for each area described in S908, the number of images captured for each person registered for personal authentication, the number of images captured for each subject identified through general object recognition, and the number of images captured for each scene identified. In other words, the count of the corresponding number of images captured this time is incremented by one. Simultaneously, the current capture time and the automatic capture evaluation value are stored and maintained as image capture history information. After S913, the series of processes ends.

[0236] Next, the learning based on the user's preferences will be described. In this embodiment, the learning processing unit 219 uses a machine learning algorithm based on the following example. Figure 13 The neural network (NN) shown is used to learn according to the user's preferences. The NN is used to predict the output value based on the input value, and by learning the actual value of the input value and the actual value of the output value in advance, the output value can be estimated for the new input value. By using the NN, learning can be performed according to the user's preferences for the above-mentioned automatic shooting, automatic editing, and subject search. It is also possible to change the registration of subject information (face recognition results, general object recognition results, etc.) as feature data to be input to the NN, as well as camera notification control, low power mode control, and automatic file deletion through learning.

[0237] In the present embodiment, an example of the operation of the application learning process is as follows.

[0238] (1) Automatic camera

[0239] (2) Automatic editing

[0240] (3) Subject Search

[0241] (4) Subject registration

[0242] (5) Camera notification control

[0243] (6) Low power mode control

[0244] (7) Automatic file deletion

[0245] (8) Image blur correction

[0246] (9) Automatic image sending

[0247] Among the operations of applying the learning process, (2) automatic editing, (7) automatic file deletion, and (9) automatic image sending are not directly related to the spirit of the present invention, and thus description thereof will be omitted.

[0248] <Automatic Camera>

[0249] The learning of automatic photography will be described. In automatic photography, learning of automatically taking pictures according to the user's preferences is performed. Figure 9B As described above, after imaging (after S910), learning information generation processing (S912) is performed. This is a process in which an image to be learned is selected according to a method described later, and learning is performed by changing the weights of the NN based on the learning information included in the image. Learning is performed by changing the NN used to determine the automatic imaging timing and the NN used to determine the imaging method (still image shooting, moving image shooting, continuous shooting, panoramic shooting, etc.).

[0250] <Subject Search>

[0251] The learning of the subject search will be described. In the subject search, learning is performed to automatically search for a subject according to the user's preference. Figure 9A In the subject search process (S904), the importance of each area is calculated, and the subject search is performed by panning, tilting, and zooming. Learning is performed based on the captured image or detection information during the search, and the learning is reflected as a learning result by changing the weight of the NN. By inputting various detection information during the search operation into the NN and determining the importance, a subject search can be performed that reflects the learning results. In addition to calculating the importance, the search method (speed or frequency of movement) is also controlled by panning and tilting.

[0252] <Subject Registration>

[0253] The following describes learning for subject registration. Subject registration involves learning to automatically register or rank subjects based on user preferences. Examples of learning include facial authentication, general object recognition, gesture or voice recognition, and scene recognition using sound. Authentication and registration of people and objects are performed, and ranking is determined based on the number or frequency of image acquisition, the number or frequency of manual capture, and the frequency of subject appearance during searches. Each piece of information is registered as input for determination using a neural network.

[0254] <Camera Notification Control>

[0255] The learning of the image capture notification will be described. Figure 9B As described in S910 of FIG. 2 , when a predetermined condition is satisfied immediately before recording, the camera notifies the subject person of recording to record the image, and then records the image. For example, the following processing is performed, wherein the subject's line of sight is visually guided by panning and tilting, or the subject's attention is attracted by using speaker sound emitted from the voice output unit 218 or LED light from the LED control unit 224. Whether the detection information of the subject (e.g., the degree of smile, line of sight detection, and gesture) is acquired immediately after the notification is determined, and learning is performed by changing the weight of the NN.

[0256] The detection information immediately before the image is captured is input to the NN, and a decision is made as to whether to provide a notification. The sound level, type, and timing of the notification sound, and the lighting duration, speed, and camera orientation (pan / tilt motion) of the notification light are determined.

[0257] <Low Power Mode Control>

[0258] As reference Figure 7 and Figure 8 As described, control for turning on / off the power supply to the first control unit 223 (main CPU) is performed. Conditions for returning from the low power consumption mode and conditions for transitioning to the low power consumption state are learned. First, learning of conditions for canceling the low power consumption mode will be described.

[0259] Sound detection

[0260] Learning can be performed by manually setting a specific voice of the user, a specific sound scene desired to be detected, or a specific sound level, for example, by communication using a dedicated application of the external device 301. There is a method in which a plurality of detection methods are pre-set in the voice processing unit and an image to be learned is selected according to a method to be described later. Learning can be performed by learning information related to the front and back sounds included in the selected image and setting sound determination (a specific sound command, or a sound scene such as "cheers" and "applause") as a starting factor.

[0261] Environmental information detection

[0262] Learning can be performed by manually setting changes in environmental information that the user desires to use as activation conditions, for example, through communication using a dedicated application on the external device 301. For example, specific conditions such as the absolute value or change in temperature, air pressure, illuminance, humidity, and ultraviolet light can be set, and the camera can be activated when the conditions are met. Determination thresholds based on various environmental information can be learned. If, based on camera detection information after activation based on environmental information, the factor is determined not to be a trigger factor, the parameters for each determination threshold are set to make it difficult to detect environmental changes.

[0263] Each of the above parameters also changes depending on the remaining battery level. For example, when the remaining battery level is low, it is difficult to make the various decisions, while when the remaining battery level is high, it is easy to make the various decisions. Specifically, even if neither the shaking state detection result nor the sound scene detection result indicates that the user wants to activate the camera, a high remaining battery level may still determine that the camera is to be activated.

[0264] The conditions for canceling the low power consumption mode may also be determined based on a neural network using vibration detection information, sound detection information, time lapse detection information, various environmental information, the remaining battery level, etc. In this case, an image to be learned is selected according to a method described later, and learning is performed by changing the weights of the neural network based on the learning information included in the image.

[0265] Next, the learning of the transition conditions to the low power consumption state will be described. Figure 7 As shown, in the mode setting determination in S704, if it is determined that the operation mode is not any of the "automatic camera mode", "automatic editing mode", "automatic image transmission mode", "learning mode" and "automatic file deletion mode", the operation mode is shifted to the low power consumption mode. The determination conditions of each mode are as described above, but the determination conditions of each mode also change through learning.

[0266] <Auto Camera Mode>

[0267] The importance of each area is determined, and automatic recording is performed while searching for a subject by panning and tilting. If it is determined that there is no subject to be recorded, automatic recording mode is canceled. For example, if the importance of all areas or the sum of the importance of each area is equal to or less than a predetermined threshold, automatic recording mode is canceled. In this case, the following setting is used, in which the predetermined threshold is reduced according to the time elapsed since the transition to automatic recording mode. The longer the time elapsed since the transition to automatic recording mode, the easier it is to transition to low power consumption mode.

[0268] By changing the predetermined threshold value according to the battery remaining amount, low power mode control can be performed taking into account the available time of the battery. For example, when the battery remaining amount is small, the threshold value is raised to facilitate the transition to low power mode, and when the battery remaining amount is large, the threshold value is lowered to make it difficult to transition to low power mode. Here, based on the time elapsed from the time of the previous transition to the automatic camera mode and the number of captured images, the parameters for the next low power mode cancellation condition (elapsed time threshold TimeC) are set in the second control unit 211. The above-mentioned various thresholds are changed by learning. Learning is performed by manually setting the camera frequency, startup frequency, etc. via communication using, for example, a dedicated application of the external device 301.

[0269] A configuration may be provided in which the average value or distribution data for each time period of the elapsed time from the time the camera 101 power button is turned on to the time the power button is turned off is accumulated, and various parameters are learned. In this case, for users whose elapsed time from power on to power off is short, the time interval for returning from low power mode or transitioning to a low power state is shortened through learning. Conversely, for users whose elapsed time from power on to power off is long, the time interval is increased through learning.

[0270] Learning is also performed based on detection information during subject search. If many important subjects are determined to be present, the time interval between returning from and transitioning to low power mode is shortened through learning. Conversely, if few important subjects are determined to be present, the time interval is lengthened through learning.

[0271] <Image Blur Correction>

[0272] The learning of image blur correction will be described. Figure 9A The image blur correction amount is calculated in S902, and based on the image blur correction amount, image blur correction is performed by panning and tilting driving in S905. In the image blur correction, learning is performed for correction based on the characteristics of user shaking. For example, the point spread function (PSF) is used to capture the image, so the direction and size of the blur can be estimated. Figure 9B In the learning information generation in S912 , information on the estimated blur direction and size is added to the image data.

[0273] exist Figure 7In the learning mode processing in S716, the weights of the neural network for image blur correction are learned based on predetermined input information and output (estimated blur direction and size). The predetermined input information is, for example, various detection information at the time of shooting (motion vector information in the image at a predetermined time before shooting, motion information of the detected subject (person or object), and vibration information (gyroscope output, acceleration output, and camera status)). Environmental information (temperature, air pressure, illuminance, and humidity), sound information (sound scene determination, specific voice detection, and sound level changes), time information (time elapsed since startup and time elapsed since the previous shooting), location information (GPS location information and position movement change), etc. can be added to the input.

[0274] When Figure 9A When calculating the image blur correction amount in S902, each piece of detection information is input into the neural network, allowing an estimate of the blur magnitude at the time the image was captured. If the estimated blur magnitude exceeds a threshold, control such as increasing the shutter speed can be performed. If the estimated blur magnitude exceeds the threshold, there is a possibility that a blurred image will be captured, so there is a method for prohibiting image capture.

[0275] Because the pan or tilt drive angle is limited, further image blur correction cannot be performed after reaching the drive end. In this embodiment, the magnitude and direction of blur during image capture are estimated, allowing the range of pan or tilt drive required to correct image blur during exposure to be estimated. Regarding the pan or tilt drive angle, if there is no margin within the movable range during exposure, processing is performed to set the drive angle within the movable range by increasing the cutoff frequency of the filter used to calculate the image blur correction amount. This allows for significant blur reduction. If it is predicted that the drive angle will exceed the movable range, the drive angle is changed immediately before exposure, rotating in the direction opposite to the direction in which the drive angle will exceed the movable range, and exposure is initiated. This allows for image capture with suppressed image blur while ensuring the movable range. By learning about image blur correction based on user characteristics during image capture or camera usage, image blur in captured images can be suppressed or prevented.

[0276] In this embodiment, panning determination processing can be performed when determining the imaging method. When panning, the image of a moving subject is not blurred, and the captured image appears to flow against the non-moving background. When determining whether to pan, the panning and tilting drive speeds required to capture the subject without blur are estimated based on pre-capture detection information, and the subject's image blur is corrected. In this case, the drive speed can be estimated by inputting information into a neural network that has learned the various detection information. The image is divided into multiple blocks, and the PSF of each block is estimated to estimate the direction and magnitude of blur in the block where the main subject is located. Learning is performed based on this information.

[0277] The amount of background blur can also be learned based on information about the user's selected image. In this case, the amount of blur in blocks (image areas) where the primary subject is absent is estimated, allowing the user's preference to be learned based on this information. Based on this learned preference, the shutter speed during recording is set based on the background blur, allowing for automatic recording that achieves a panning effect tailored to the user's preferences.

[0278] Next, we will describe the learning method. There are two learning methods: "in-camera learning" and "learning based on collaboration with external devices such as communication devices." First, we will describe the former learning method. The in-camera learning method of this embodiment includes the following methods.

[0279] (1) Learning based on detection information during manual photography

[0280] like Figure 9A As described in S907 to S913 in the manual shooting, the camera 101 can perform manual shooting and automatic shooting. If there is a manual shooting instruction in S907, information indicating that the image is manually shot is added to the shot image in S912. If it is determined in S916 that the automatic shooting is on and the image is shot, information indicating that the image is automatically shot is added to the shot image in S912. In the case of manual shooting, it is likely that the shooting is performed based on the subject, scene and location or time interval according to the user's preference. Therefore, learning is performed based on the learning data of each feature data or shot image obtained during manual shooting. Based on the detection information during manual shooting, learning is performed in connection with the extraction of feature quantities in the shot image, the registration of personal authentication, the registration of facial expressions of each person, and the registration of groups of people. For example, based on the detection information during the subject search, learning is performed to change the importance of nearby people or objects according to the facial expressions of the subjects registered by the individual.

[0281] (2) Learning based on detection information when searching for objects

[0282] During the subject search, it is determined what kind of people, objects, and scenes the subject registered for personal authentication is being photographed as at the same time, and the time ratio of the subjects being photographed at the same time within the field of view is also calculated. For example, the time ratio of person A, who is the subject registered for personal authentication, being photographed at the same time as person B, who is also the subject registered for personal authentication, is calculated. If person A and person B enter the field of view, various detection information is stored as learning data so that the score of automatic imaging determination is high, and learning is performed in the learning mode processing ( Figure 7 : S716). In another example, the ratio of time that person A, who is a subject registered for personal authentication, is simultaneously photographed as a "cat," which is a subject determined by general object recognition, is calculated. If person A and "cat" enter the field of view, various detection information is stored as learning data so that the score determined by automatic imaging is high, and learning is performed in the learning mode processing ( Figure 7 :S716).

[0283] If a high smile level is detected for person A, who is a subject registered for personal authentication, or if a facial expression such as "happy" or "surprised" is detected, the simultaneously photographed subject is learned to be important. Alternatively, if a facial expression such as "angry" or "expressionless" is detected in person A, it is determined that the simultaneously photographed subject is unlikely to be important, and learning is not performed.

[0284] Next, the following learning based on cooperation with an external device in the present embodiment will be described.

[0285] (1) Learning when an external device acquires an image

[0286] (2) Learning when a certain value is input to an image via an external device

[0287] (3) Learning when analyzing images stored in an external device

[0288] (4) Learning based on information uploaded to the social network service (SNS) server by external devices

[0289] (5) Learning when camera parameters are changed by external devices

[0290] (6) Learning based on information from manually edited images by external devices

[0291] The description will be given in order according to the given numbers.

[0292] <Learning when an external device acquires an image>

[0293] As reference Figure 3As described, the camera 101 and the external device 301 have a communication unit configured to perform a first communication 302 and a second communication 303. Typically, image data is sent and received via the first communication 302, and an image in the camera 101 is sent to the external device 301 using a dedicated application of the external device 301. Thumbnail images of the image data stored in the camera 101 can be viewed using the dedicated application of the external device 301. The user can send image data to the external device 301 by selecting and confirming an image that the user likes from the thumbnail images or by performing an image acquisition instruction operation. The image obtained by the user selecting an image is likely to be an image according to the user's preferences. Therefore, the image obtained is determined to be an image to be learned. Various types of learning according to the user's preferences can be performed based on the learning information of the acquired image.

[0294] Will refer to Figure 14 Describe the operation example. Figure 14 This figure shows an example of a user viewing images stored in the camera 101 using a dedicated application on the external device 301. Thumbnail images 1604 to 1609 of image data stored in the camera are displayed on the display unit 407. The user can select and retrieve their favorite images. Buttons 1601 to 1603 are operated to change the display method and constitute a display method changing unit.

[0295] When the first button 1601 is pressed, the mode changes to the date and time priority display mode, and images are displayed on the display unit 407 in the order of the shooting date and time of the images in the camera 101. For example, an image with a newer date and time is displayed at a position indicated by a thumbnail image 1604, and an image with an older date and time is displayed at a position indicated by a thumbnail image 1609. When the second button 1602 is pressed, the mode changes to the recommended image priority display mode. Based on the Figure 9B The images in the camera 101 are displayed on the display unit 407 in descending order of the scores calculated in S912 for determining the user's preference for each image. For example, images with high scores are displayed at positions indicated by thumbnail images 1604, while images with low scores are displayed at positions indicated by thumbnail images 1609. When the user presses the third button 1603, a subject such as a person or an object can be specified. When a subject such as a specific person or object is specified, only the specific subject can be displayed.

[0296] Buttons 1601 to 1603 can also be used to enable settings simultaneously. For example, if all settings are enabled, only the specified subjects will be displayed, images with the latest capture date and time will be prioritized, and images with high scores will be prioritized. As described above, since user preferences are also learned for captured images, only images that meet the user's preferences can be extracted from a large number of captured images with a simple confirmation operation.

[0297] <Learning When a Certain Value is Input to an Image via an External Device>

[0298] When viewing images stored in the camera 101, the user can rate each image. Images that the user finds pleasing can be given a high score (e.g., 5 points), while images that the user dislikes can be given a low score (e.g., 1 point). The camera is configured to learn the specific values of images based on user operations. The scores for each image are used together with the learning information in the camera for relearning. Learning is performed so that the output of a neural network, which uses feature data from a specified image as input, approaches the score specified by the user.

[0299] In addition to the configuration in which the user inputs a determination value to the captured image via the external device 301, there is also a configuration in which the user operates the camera 101 to directly input the determination value to the image. In this case, for example, the camera 101 includes a touch panel display. The user operates a graphical user interface (GUI) button displayed on the screen display unit of the touch panel display to set the mode for displaying the captured image. The user can perform the same learning as described above by inputting determination values to each image while checking the captured image.

[0300] <Learning when analyzing images stored in an external device>

[0301] The storage unit 404 of the external device 301 also stores images other than those captured by the camera 101. The images stored in the external device 301 are easily viewed by the user and easily uploaded to a shared server via the public wireless control unit 406, and therefore are likely to include many images according to the user's preferences.

[0302] It is assumed that the control unit 411 of the external device 301 can process the images stored in the storage unit 404 using a dedicated application with the same capabilities as the learning processing unit 219 of the camera 101. The processed learning data is sent to the camera 101 through communication for learning. Alternatively, a configuration may be adopted in which an image or data desired to be learned is sent to the camera 101 and learning is performed in the camera 101. A configuration may be adopted in which a user selects an image desired to be learned from the images stored in the storage unit 404 and learns by using a dedicated application.

[0303] <Learning from information uploaded to the SNS server from external devices>

[0304] The following method will be described, in which information from a social networking service (SNS) is used for learning. SNS is a service or website that can build a social network focused on connecting people. When an image is uploaded to the SNS, there is a technique that inputs tags related to the image from an external device 301 and then sends the tags along with the image. There is also a technique that inputs likes and dislikes for images uploaded by other users. It is also possible to determine whether an image uploaded by another user is a favorite photo of the user who owns the external device 301.

[0305] Images uploaded by a user and information related to the images can be obtained through a dedicated SNS application downloaded to the external device 301. It is also possible to obtain user-favorite images or tag information by inputting data on whether the user likes images uploaded by other users. Learning is performed in the camera 101 by analyzing such images or tag information.

[0306] The control unit 411 of the external device 301 acquires an image uploaded by the user or an image determined to be a favorite of the user, and can process the image with the same capabilities as the learning processing unit 219 of the camera 101. The processed learning data is sent to the camera 101 via communication for learning. Alternatively, a configuration is also possible in which image data desired for learning is sent to the camera 101, and learning is performed in the camera 101.

[0307] Based on the subject information set in the tag information (for example, subject information such as dogs and cats, scene information such as beaches, and facial expression information such as smiles), it is possible to estimate subject information that the user may like. Learning is performed by registering the subject as a subject to be input into the neural network and detected. A configuration may be as follows, in which image information currently popular in the world is estimated based on the statistical value of tag information (image filter information or subject information) in the SNS, and the image information is learned in the camera 101.

[0308] <Learning when camera parameters are changed by external devices>

[0309] The learning parameters currently set in the camera 101 (such as the weight of the NN and the selection of the subject to be input to the NN) can be sent to the external device 301 and stored in the storage unit 404 of the external device 301. By using the dedicated application of the external device 301, the learning parameters set in the dedicated server are obtained via the public wireless control unit 406. This can also be set as the learning parameters in the camera 101. The learning parameters can be returned by storing the parameters at a certain point in time in the external device 301 and setting the parameters in the camera 101. The learning parameters owned by other users can be obtained via the dedicated server and set in the owner's own camera 101.

[0310] Voice commands, authentication registrations, and gestures registered by the user can be registered by using a dedicated application of the external device 301, or important places can be registered. This information is used as a tool for processing in the camera mode ( Figure 9A and Figure 9B ) and automatic image capture determination described in

[15] . A configuration may be provided in which the image capture frequency, activation interval, ratio between still and moving images, preferred images, etc., can be set, and the activation interval, etc., described in

[15] low power consumption mode control can be set.

[0311] <Learning from information on manually edited images using external devices>

[0312] The dedicated application of the external device 301 may have a function that allows manual editing based on user operations, and the details of the editing work can be fed back to the learning. For example, image effects (cropping, rotation, sliding, zooming, fading, color conversion filter effects, time, still image / moving image ratio, and background music) can be edited. The automatic editing neural network is trained so that the manually edited image effects are determined based on the image learning information.

[0313] Next, the learning process sequence will be described. Figure 7 In the mode setting determination in S704 of FIG. 1 , it is determined whether or not to perform a learning process. If it is determined that a learning process is to be performed, the learning mode process in S716 is executed. The determination conditions of the learning mode will be described. Whether or not to transition to the learning mode is determined based on information such as the time elapsed since the previous learning process was performed, the amount of information available for learning, and whether or not there is an instruction to perform a learning process via the communication device. Figure 15 The learning mode determination process is described.

[0314] Figure 15 It is shown in Figure 7 Flowchart of the determination process of whether to change to the learning mode executed in S704 (mode setting determination process) of FIG. When the instruction to start the learning mode determination is given in the mode setting determination process in S704, the learning mode determination is started. Figure 15 In S1401, it is determined whether there is a registration instruction from the external device 301. The registration instruction is a registration instruction for learning (such as the above-mentioned <learning when an external device acquires an image>, <learning when a certain value is input to an image via an external device>, and <learning when an image stored in an external device is analyzed>).

[0315] If it is determined in S1401 that there is a registration instruction from the external device 301, the flow proceeds to the processing in S1408. In S1408, the learning mode determination flag is set to TRUE, so that the processing in S716 is set to proceed, and the learning mode determination processing ends. If it is determined in S1401 that there is no registration instruction from the external device 301, the flow proceeds to the processing in S1402.

[0316] In S1402, it is determined whether a learning instruction has been issued from the external device 301. This learning instruction is an instruction for setting learning parameters as described above in <Learning When Camera Parameters Are Changed by an External Device>. If it is determined in S1402 that a learning instruction has been issued from the external device 301, the process proceeds to S1408. If it is determined in S1402 that no learning instruction has been issued from the external device 301, the process proceeds to S1403.

[0317] In S1403, the time (indicated by TimeN) that has elapsed since the previous learning process (recalculation of the NN weights) was performed is acquired. The flow proceeds to S1404, where the amount of new data to be learned (indicated by DN) is acquired. The amount of data DN corresponds to the number of images specified to be learned during the time TimeN that has elapsed since the previous learning process was performed.

[0318] Next, the process enters the processing of S1405, and the threshold DT used to determine whether to transition to learning mode is calculated based on the elapsed time TimeN. As the threshold DT is set to a smaller value, it is easier to transition to learning mode. For example, the value of the threshold DT when TimeN is less than a predetermined value is indicated by DTa, and the value of the threshold DT when TimeN is greater than a predetermined value is indicated by DTb. DTa is set to be greater than DTb, and the threshold is set to become smaller as time passes. Therefore, even if the amount of learning data is small, if a long time has passed, it is easy to transition to learning mode and learn again. That is, the setting of the difficulty of the camera transitioning to learning mode changes according to the usage time.

[0319] After S1405, the process proceeds to S1406, where it is determined whether the number of data to be learned, DN, is greater than the threshold value, DT. If it is determined that the number of data, DN, is greater than the threshold value, DT, the process proceeds to S1407. If it is determined that the number of data, DN, is equal to or less than the threshold value, DT, the process proceeds to S1409. In S1407, the number of data, DN, is set to zero. Thereafter, S1408 is executed, and the learning mode determination process ends.

[0320] If the flow enters the processing in S1409, since there is no registration instruction or learning instruction from the external device 301 and the number of data DN is equal to or less than the threshold value DT, the learning mode determination flag is set to FALSE. After the processing in S716 is set to not be performed, the learning mode determination processing ends.

[0321] Next, the learning mode processing ( Figure 7 :Processing in S716). Figure 16 is a flowchart showing an example of learning mode processing. Figure 7 When the learning mode is determined in S715 and the process enters the processing in S716, the Figure 16 In S1501, it is determined whether there is a registration instruction from the external device 301. If it is determined in S1501 that there is a registration instruction from the external device 301, the flow proceeds to the process of S1502. If it is determined in S1501 that there is no registration instruction from the external device 301, the flow proceeds to the process of S1504.

[0322] In S1502, various registration processes are performed. These registration types include registration of features to be input into the neural network, such as facial recognition registration, general object recognition registration, voice information registration, and location information registration. After the registration process is completed, the process proceeds to S1503. In S1503, based on the information registered in S1502, the elements to be input into the neural network are changed. When S1503 is completed, the process proceeds to S1507.

[0323] In S1504, it is determined whether there is a learning instruction from the external device 301. If it is determined that there is a learning instruction from the external device 301, the flow proceeds to processing at S1505, whereas if it is determined that there is no learning instruction, the flow proceeds to processing at S1506.

[0324] In S1505, after the learning parameters transmitted from the external device 301 by communication are set in each determiner (such as the weight of NN), the flow enters the processing in S1507. In S1506, learning is performed (the weight of NN is recalculated). The case of transitioning to the processing in S1506 is as described in reference to Figure 15 The number of data DN described here exceeds the threshold DT and each determiner is relearned. The weights of the NN are recalculated by relearning using error backpropagation, gradient descent, etc., and the parameters in each determiner are changed. When the learning parameters are set, the process proceeds to the processing in S1507.

[0325] In S1507, the images in the file are re-scored. This embodiment provides a configuration in which all captured images in the file stored in recording medium 221 are scored based on the learning results, and automatic editing or file deletion is performed based on the scores. Therefore, if re-learning is performed or learning parameters are set from an external device, the scores of the captured images must also be updated. In S1507, the scores of the captured images stored in the file are recalculated, and when this process is completed, the learning mode processing ends.

[0326] The above description describes a method for extracting scenes assumed to be user-favorite scenes, learning their characteristics, and reflecting these characteristics in camera operations such as automatic shooting and automatic editing. The embodiments of the present invention are not limited to this application. For example, as described below, the present invention can be applied to an application for extracting videos that do not match the user's preferences.

[0327] <Method of using a neural network that has learned preferences>

[0328] According to the above method, the user's preferences are learned. Figure 9A In S908 of the NN, automatic image capture determination processing is performed. Automatic image capture is performed when the output value of the NN is a value indicating that the subject does not conform to the user's preferences as training data. For example, assume that an image preferred by the user is used as a training image, and learning is performed so that a high value is output when features similar to those of the training image are shown. In this case, automatic image capture is performed conversely when the output value is less than a predetermined threshold. Similarly, in the subject search processing or automatic editing processing, processing is performed so that the output value of the NN is a value indicating that the subject does not conform to the user's preferences as training data.

[0329] <Method using a neural network that has learned situations that are not in line with preferences>

[0330] During the learning process, the following process is performed: learning is performed on situations that do not conform to the user's preferences as training data. In the above example, it is assumed that manually captured images are scenes that the user voluntarily captured, and the learning method using this image as training data has been described. On the other hand, manually captured images are not used as training data, and the following process is performed: scenes that have not been manually captured for a predetermined time or longer are added as training data. Alternatively, if the training data includes scene data with characteristics similar to those of manually captured images, a process is performed to delete this data from the training data. A process is performed to add images with characteristics different from those of images acquired by an external device to the training data, or a process is performed to delete images with characteristics similar to those of the acquired images from the training data. In the above manner, data that does not conform to the user's preferences is accumulated in the training data, so that as a result of learning, the NN can discern situations that do not conform to the user's preferences. In automatic photography, photography is performed based on the output value of the NN, so that scenes that do not conform to the user's preferences can be photographed.

[0331] This method of suggesting images that don't suit the user's preferences reduces the number of missed shots by capturing scenes that the user wouldn't manually capture. Proposing shots in scenes not originally intended by the user also has the effect of increasing the user's awareness and broadening their tastes.

[0332] By combining the above methods, it is easy to propose situations that are somewhat similar to but partially different from the user's preferences, or to adjust the degree of compliance with the user's preferences. The degree of compliance with the user's preferences can be changed according to the mode settings, the status of various sensors, and the status of the detection information.

[0333] In this embodiment, the configuration for learning in the camera 101 has been described. On the other hand, if the external device 301 has a learning function, the data required for learning is sent to the external device 301, and learning is performed only by the external device 301. Even with this configuration, the same learning effect as described above can be achieved. For example, as described above in <Learning When Camera Parameters Are Changed by an External Device>, learning can be performed by setting parameters (such as the weights of the NN learned by the external device 301) in the camera 101 through communication.

[0334] In addition, there is an embodiment in which both the camera 101 and the external device 301 have a learning function. For example, a learning mode process ( Figure 7 : S716), the learning information stored in the external device 301 is sent to the camera 101, the learning parameters are merged, and learning is performed by using the merged parameters.

[0335] According to this embodiment, if a single camera is used for automatic photography and automatic authentication registration, it is possible to simultaneously realize the photography for automatic photography and the photography for automatic authentication registration. In particular, automatic authentication registration can improve the accuracy of automatic photography and can also realize control that does not interfere with automatic photography.

[0336] In the following, reference will be made to Figure 17 to Figure 3 The following example describes an example in which a person as a subject for image capture is identified and tracking control is performed. For example, in automatic image capture, the user registers characteristic information related to a primary person in the camera and specifies that the registered person should be tracked and captured with priority, thereby enabling image capture centered around that person (the priority person). If the priority person is not detected, or even if the priority person is detected but not recognized as the priority person, it is desirable to capture the primary person as closely as possible. Even if the priority person is detected, if other primary persons, such as family members or friends, are also detected, it is desirable to control the camera so that the person is also included in the field of view, and to minimize the inclusion of unrelated persons in the field of view.

[0337] As a subject recognition technology, there is a technique that analyzes image data frame by frame to identify detected subjects, extracts the frequency of appearance of the identified subjects, and selects the main subject from these subjects based on the frequency of appearance. With this technique, a specific number of subjects are always selected in descending order of frequency of appearance. Therefore, even in situations where the absolute number of people is small, a person can be identified as the main subject even if their frequency of appearance is significantly lower than that of the original main subject. Because the distance between the subject and the camera is not taken into account, there is a possibility that unrelated people far away from the camera may be included in the main subject.

[0338] Hereinafter, a technology will be described in which, in an automatic camera that regularly and continuously performs recording without the user giving a recording instruction, the frequency with which irrelevant persons are included in the recording angle of view is reduced while keeping the main person within the recording angle of view. Specifically, an example will be described in which the recording priority of a person is determined based on user settings and the facial size, facial position, facial reliability, and detection frequency of the detected person, and the person set as the recording subject is determined based on the recording priority of each person. Control is performed so that if a person with a high recording priority is detected, the person and persons with similar recording priorities are determined as recording subjects, and persons whose recording priorities differ by more than a certain amount are excluded from the recording subjects. By selecting the recording subject, the possibility of the user and persons with similar recording priorities to the user being recorded can be increased, and the possibility of irrelevant persons being recorded can be reduced.

[0339] Figure 171 is a block diagram illustrating an imaging apparatus including a lens barrel 102, a tilt rotation unit 104, a pan rotation unit 105, and a control box 1100. The control box 1100 includes a microcomputer and the like for controlling a group of imaging lenses included in the lens barrel 102, the tilt rotation unit 104, and the pan rotation unit 105. The control box 1100 is disposed in the fixed portion 103 of the imaging apparatus. The control box 1100 remains fixed even when the lens barrel 102 is being panned or tilted.

[0340] The lens barrel 102 includes a lens unit 1021 forming an imaging optical system and an imaging unit 1022 having an imaging element. The lens barrel 102 is controlled and driven to rotate in the tilt and pan directions by the tilt rotation unit 104 and the pan rotation unit 105, respectively. The lens unit 1021 includes a zoom lens for changing magnification, a focus lens for adjusting focus, and other components, and is driven and controlled by a lens drive unit 1113 in the control box 1100. The zoom mechanism unit consists of a zoom lens and a lens drive unit 1113 that drives the lens. The lens drive unit 1113 moves the zoom lens along the optical axis, thereby achieving a zoom function.

[0341] The imaging unit 1022 includes an imaging element, receives light incident through the various lens groups constituting the lens unit 1021, and outputs charge information corresponding to the amount of light as digital image data to the image processing unit 1103. The tilt rotation unit 104 and the pan rotation unit 105 rotate and drive the lens barrel 102 in accordance with a drive command input from the lens barrel rotation drive unit 1112 in the control box 1100.

[0342] Next, description will be given of the configuration of the control box 1100. The imaging direction in automatic imaging is controlled by the temporary registration determination unit 1108, imaging target determination unit 1110, drive control unit 1111, and lens barrel rotation drive unit 1112.

[0343] The image processing unit 1103 acquires the digital image data output from the imaging unit 1022. Image processing such as distortion correction, white balance adjustment, and color interpolation processing is applied to the acquired digital image data. The digital image data to which the image processing is applied is output to the image recording unit 1104 and the subject detection unit 1107. In response to an instruction from the provisional registration determination unit 1108, the image processing unit 1103 outputs the digital image data to the feature information extraction unit 1105.

[0344] Image recording unit 1104 converts the digital image data output from image processing unit 1103 into a recording format such as JPEG format and records the digital image data on a recording medium (non-volatile memory, etc.). Feature information extraction unit 1105 acquires a facial image located at the center of the digital image data output from image processing unit 1103. Feature information extraction unit 1105 extracts feature information from the acquired facial image and outputs the facial image and feature information to person information management unit 1106. Feature information is information indicating multiple facial feature points located in areas such as the eyes, nose, and mouth of the face, and is used to identify a person as the detected subject. Feature information may also be other information indicating facial features such as the outline of the face, facial color information, and facial depth information.

[0345] The character information management unit 1106 performs processing for storing and managing character information associated with each character in the storage unit. Figure 18 An example of describing person information. Person information consists of a person ID, facial image, feature information, registration status, priority setting, and name. The person ID is an ID (identification information) used to identify each piece of person information among multiple pieces of information. The same ID is not issued and a value of 1 or greater is set. Facial image data is facial image data input from feature information extraction unit 1105. Feature information is information input from feature information extraction unit 1105. Regarding registration status, it is assumed that two states are defined, namely "temporary registration" and "main registration." "Temporary registration" indicates a state in which a person is determined to be a main person through temporary registration. "Main registration" indicates a state in which a person is determined to be a main person through main registration or depending on whether a user operation is performed. The details of the temporary registration determination process and the main registration determination process will be described later. The priority setting is a setting that indicates whether or not to prioritize recording through user operation. The name is a name assigned to each person through user operation.

[0346] Upon receiving a facial image and feature information from the feature information extraction unit 1105, the character information management unit 1106 issues a new character ID, associates it with the input facial image and feature information, and adds the new character information. When adding new character information, the registration status is initially set to "Provisionally Registered," the priority setting is initially set to "Not Present," and the name is initially set to "None." Upon receiving the primary registration determination result (the character ID to be registered) from the primary registration determination unit 1109, the character information management unit 1106 changes the registration status of the character information corresponding to the character ID to "Primary Registered." If a user operation instructs the communication unit 1114 to change character information (priority setting or name), the character information management unit 1106 changes the character information according to the instruction. If the priority setting or name is changed for a character whose registration status is "Provisionally Registered," the character information management unit 1106 determines that the character is a primary character and changes the character's registration status to "Primary Registered." The importance determination unit 1514 will be described later.

[0347] Figure 19 1 is a diagram showing an example of a screen on a mobile terminal device (external device) communicating with the camera 101. The mobile terminal device acquires person information via the communication unit 1114 of the camera 101 and displays the person information on the screen in a list form. Figure 19 In the example shown, a face image, name, and priority setting are displayed on the screen. The user can change the name and priority setting. If the name or priority setting is changed, the mobile terminal device outputs an instruction to change the name or priority setting associated with the person ID to the communication unit 1114.

[0348] Subject detection unit 1107 ( Figure 17 ) detects a subject from the digital image data input from the image processing unit 1103, and extracts information related to the detected subject (subject information). An example in which the subject detection unit 1107 detects a human face as a subject will be described. The subject information includes, for example, the number of detected subjects, the position of the face, the face size, the orientation of the face, and the reliability of the face indicating the certainty of detection. The subject detection unit 1107 calculates the similarity by comparing the feature information of each person acquired from the person information management unit 1106 with the feature information of the detected subject. If the similarity is equal to or greater than the threshold value, a process of adding the person ID, registration status, and priority setting of the detected person to the subject information is performed. The subject detection unit 1107 outputs the subject information to the temporary registration determination unit 1108, the main registration determination unit 1109, and the imaging object determination unit 1110. This will be referred to later. Figure 20A and Figure 20BDescribes an example of subject information.

[0349] The temporary registration determination unit 1108 determines whether the subject detected by the subject detection unit 1107 is likely to be a principal person, that is, whether temporary registration is to be performed. If any subject is determined to be a person to be temporarily registered, the temporary registration determination unit 1108 calculates the pan drive angle, tilt drive angle, and target zoom position required to position the person to be temporarily registered at the center of the image frame at a specified size. A command signal based on the calculation results is output to the drive control unit 1111. The details of the temporary registration determination process will be described later with reference to FIG22.

[0350] The main registration determination unit 1109 determines a person similar to the user, that is, a person to be mainly registered, based on the subject information acquired from the subject detection unit 1107. If any person is determined to be a person to be mainly registered, the person ID of the person to be mainly registered is output to the person information management unit 1106. Figure 26 Details of the main registration determination process are described.

[0351] The imaging subject determination unit 1110 determines the imaging subject based on the subject information acquired from the subject detection unit 1107. Based on the determination result of the imaging subject person, the imaging subject determination unit 1110 calculates the pan drive angle, tilt drive angle, and target zoom position required to arrange the imaging subject person within the specified size within the angle of view. The command signal based on the calculation result is output to the drive control unit 1111. The details of the imaging subject determination processing will be described later with reference to FIG. 27.

[0352] Upon receiving a command signal from the temporary registration determination unit 1108 or the imaging object determination unit 1110, the drive control unit 1111 outputs control parameter information to the lens drive unit 1113 and the lens barrel rotation drive unit 1112. Parameters based on the target zoom position are output to the lens drive unit 1113. Parameters corresponding to the target position based on the pan drive angle and the tilt drive angle are output to the lens barrel rotation drive unit 1112.

[0353] If there is input from the temporary registration determination unit 1108, the drive control unit 1111 determines each target position (target zoom position and target position based on the drive angle) based on the input value from the temporary positioning determination unit 1108 without referring to the input from the imaging object determination unit 1110. The barrel rotation drive unit 1112 outputs a drive command to the tilt rotation unit 104 and the pan rotation unit 105 based on the target position and drive speed from the drive control unit 1111. The lens drive unit 1113 has a motor and a drive unit for driving the zoom lens, focus lens, etc. included in the lens unit 1021. The lens drive unit 1113 drives each lens based on the target position from the drive control unit 1111.

[0354] The communication unit 1114 transmits the personal information stored in the personal information management unit 1106 to an external device such as a mobile terminal device. When receiving an instruction to change the personal information from the external device, the communication unit 1114 outputs an instruction signal to the personal information management unit 1106. In this example, it is assumed that the change instruction from the external device is an instruction to change the priority setting and name of the personal information.

[0355] Figure 20A and Figure 20B is a diagram illustrating an example of image data and an example of subject information acquired by the subject detection unit 1107 . Figure 20A : is a schematic diagram showing an example of image data input to the subject detection unit 1107. For example, the image data is formed with a horizontal resolution of 960 pixels and a vertical resolution of 540 pixels. Figure 20B It is shown in Figure 20A 1 is a table showing an example of subject information extracted when the image data shown in FIG is input to the subject detection unit 1107. The illustrated subject information includes the number of subjects, the subject ID of each subject, the face size, the face position, the face orientation, the face reliability, the person ID, the registration status, and the priority setting.

[0356] The number of subjects indicates the number of detected faces. Figure 20B The example in the example shows four subjects, including face size, face position, face orientation, face reliability, person ID, registration status, and priority settings for each of the four subjects. The subject ID is a numerical value used to identify the subject and is published when a new subject is detected. The same subject ID is never published, and a new value is published each time a subject is detected. For example, if a specific subject moves out of view and becomes undetectable, then returns to view and is redetected, a new value is published even for the same subject.

[0357] Face size (w, h) is a numerical value indicating the size of the detected face, and the number of pixels of the width (w) and height (h) of the face is input. In this example, it is assumed that the width and height are the same value. Face position (x, y) is a numerical value indicating the relative position of the detected face within the camera range. If the upper left corner of the image data is defined as the starting point (0, 0) and the lower right corner of the screen is defined as the end point (960, 540), the number of horizontal pixels and the number of vertical pixels from the starting point to the center coordinates of the face are input. Face orientation is information indicating the orientation of the detected face, and any one of front, 45 degrees to the right, 90 degrees to the right, 45 degrees to the left, 90 degrees to the left, and unknown is input. Face reliability is information indicating the certainty of the detected face, and any value between 0 and 100 is input. The face reliability is calculated based on the similarity with the feature information of multiple pre-stored standard face templates.

[0358] The person ID is the same as the person ID managed by the person information management unit 1106. When a subject is detected, the subject detection unit 1107 calculates the similarity between the feature information of each person obtained from the person information management unit 1106 and the feature information of the subject. The person ID of the person whose similarity is equal to or greater than the threshold is input. If the feature information obtained from the person information management unit 1106 is not similar to the feature information of any person, zero is input as the ID value. The information related to the registration status and priority setting is the same as the information related to the registration status and priority setting managed by the person information management unit 1106. If the person ID is not zero (that is, it is determined that the person is one of the persons managed by the person information management unit 1106), the information related to the registration status and priority setting of the corresponding person obtained from the person information management unit 1106 is input.

[0359] In this example, we will refer to Figure 21 Describes the processing that is performed periodically. Figure 21 This is a flowchart illustrating the overall process of capturing images, registering and updating person information. When the camera is powered on, the camera unit 1022 begins periodic capturing (moving image capture) to acquire image data for various determinations (imaging subject determination, temporary registration determination, and main registration determination). At S500, repetitive processing begins.

[0360] Image data acquired through imaging is output to the image processing unit 1103. In S501, image data that has undergone various types of image processing is acquired. Because the acquired image data is used for various determinations, it is output from the image processing unit 1103 to the subject detection unit 1107. In other words, the acquired image data corresponds to image data displayed in live view in an imaging device where a user adjusts the composition and operates the shutter to capture an image. The periodic imaging used to acquire image data corresponds to live view imaging. The control box 1100 uses the acquired image data to adjust the composition or determine the timing for automatic imaging.

[0361] Next, in S502, the subject detection unit 1107 detects a subject based on the image data and acquires subject information (see Figure 20B After a subject is detected and subject information is acquired, primary registration determination is performed in step S503. In the primary registration determination, the person to be primarily registered is determined using information related to the detected subject. In this determination, the person information in the person information management unit 1106 is updated, but panning, tilting, and zooming are not performed.

[0362] In step S504, a temporary registration determination is performed. In this temporary registration determination, a subject to be temporarily registered is determined from among the detected subjects, and the pan drive angle and tilt drive angle are acquired based on the face position of the subject to be temporarily registered. A target zoom position is acquired based on the position and face size. The temporary registration determination unit 1108 instructs the image processing unit 1103 to output the image data to the feature information extraction unit 1105. Once the pan drive angle, tilt drive angle, and target zoom position have been acquired in the temporary registration determination, pan drive, tilt drive, and zoom drive are executed based on this information, thereby adjusting the composition for temporary registration.

[0363] After the process of S504, the flow proceeds to the process of S505, and it is determined whether the process of adjusting the composition for temporary registration is being executed. In S505, if the process of adjusting the composition for temporary registration is being executed, the flow proceeds to the process of S506, and if the process of adjusting the composition for temporary registration is not being executed, the flow proceeds to the process of S507.

[0364] In S506, the feature information extraction unit 1105 extracts feature information of the subject located in the image data center and outputs the extracted feature information to the person information management unit 1106. In S507, imaging subject determination is performed. The imaging subject determination unit 1110 determines the imaging subject from the detected subjects. Based on the facial position of the imaging subject, the pan drive angle and tilt drive angle are acquired. Based on the facial position and facial size, the target zoom position is acquired. Once the pan drive angle, tilt drive angle, and target zoom position have been acquired through imaging subject determination, pan drive, tilt drive, and zoom drive are executed based on this information, thereby adjusting the imaging composition.

[0365] After S506 and S507, the process enters the process of S508 and determines whether to end the repeated processing. If the processing is to be continued, the process returns to S500 and continues the processing. The processing in S501 to S507 is repeatedly executed according to the imaging cycle of the imaging unit 1022.

[0366] <Provisional Registration Process>

[0367] Will refer to Figure 22A and Figure 22B describe Figure 21 The temporary registration confirmation processing shown in S504 in . Figure 22A : is a flowchart showing the temporary registration determination process performed by the temporary registration determination unit 1108. This process is periodically executed and determines whether there is a possibility that a person may be the main person. Figure 22B This table shows temporary registration counts. The temporary registration count is associated with a subject ID, and if the temporary registration count reaches 50 or more, the corresponding subject is determined to be a person subject to temporary registration. Because temporary registration determination is performed over multiple cycles, the following process is performed: the current temporary registration count is stored during determination in the current cycle, and the temporary registration count added up to the previous cycle is referenced and taken over in the next cycle.

[0368] In S600, repeated processing corresponding to the number of detected subjects is started. When the subject information is acquired from the subject detection unit 1107, the temporary registration determination unit 1108 performs the processing in S601 to S609 for each subject, and performs the processing in S610 to S613 when any subject is determined to be a temporary registration object. In S601, a process of determining whether the subject is unregistered is performed. The temporary registration determination unit 1108 refers to the person ID of the subject information, and when it is determined that the subject is unregistered (the person ID is zero), the process enters the process of S602. If it is determined that the value of the person ID is 1 or greater, that is, the subject has been registered, the process enters the next subject determination process.

[0369] In S602, the temporary registration determination unit 1108 refers to the stored temporary registration count for the previous frame and, if a temporary registration count for the same subject ID exists, takes over the temporary registration count. Next, in S603, the temporary registration determination unit 1108 determines whether the face is facing forward. If it is determined that the face is facing forward, the process proceeds to S604. If it is determined that the face is not facing forward, the process proceeds to S607.

[0370] S604 is a process for determining whether the face size when zooming to wide angle is within the range of 100 to 200. If this condition is met, the process proceeds to S605, and if not, the process proceeds to S607. S605 is a process for determining whether the face reliability is equal to or greater than a threshold value of 80. If this condition is met, the process proceeds to S606, and if not, the process proceeds to S607.

[0371] When all the conditions shown in S603 to S605 are met, the process proceeds to the processing in S606. In S606, the temporary registration determination unit 1108 determines that the target person is likely to be a main person similar to the user, and adds 1 to the temporary registration count (increments). On the other hand, if any of the conditions shown in S603 to S605 are not met, the process proceeds to the processing in S607. In S607, the temporary registration determination unit 1108 determines that the target person is unlikely to be a main person, and sets the temporary registration count to 0.

[0372] After the processing in S606 and S607, in S608, the temporary registration determination unit 1108 compares the value of the temporary registration count of the subject with the threshold value 50. If it is determined that the value of the temporary registration count is less than 50, the flow proceeds to the flow in S609. If it is determined that the value of the temporary registration count is equal to or greater than 50, the flow proceeds to the processing in S611.

[0373] In S609, the temporary registration determination unit 1108 determines whether the value of the temporary registration count is greater than zero. If it is determined that the value of the temporary registration count is greater than zero, the process proceeds to S610. If this condition is not met (the value of the temporary registration count is zero), the process proceeds to S614 without storing the temporary registration count. In S610, the temporary registration determination unit 1108 stores the temporary registration count and then proceeds to the determination process in S614. In S614, it is determined whether to end the repeated processing. If the processing continues, the process returns to S600 and proceeds to the next subject determination process.

[0374] In S611, the temporary registration determination unit 1108 determines that the corresponding subject is likely a principal figure and sets the subject as a temporary registration target. In S612, the temporary registration determination unit 1108 calculates the pan drive angle, tilt drive angle, and zoom movement position so that the face of the temporary target subject is positioned at the center of the frame with an appropriate face size, and outputs a command based on the calculation results to the drive control unit 1111. For example, if the face center position is within 5% of the frame center and the face size is between 100 and 200, the feature information extraction unit 1105 can obtain feature information.

[0375] In this example, in order to obtain feature information, control is performed so that the subject of the image is positioned at the center of the screen. This example is not limited to this, and feature information can be extracted by performing image processing (such as cutting out a portion of the image data including the face of the subject without changing the position of the subject).

[0376] At S613, provisional registration determination unit 1108 instructs image processing unit 1103 to output the image data to feature information extraction unit 1105. Feature information extraction unit 1105 cuts out the facial image located at the center of the input image data, extracts feature information, and outputs the feature information to person information management unit 1106. Person information management unit 1106 adds new person information based on the input facial image and feature information. After S613, the series of processes ends.

[0377] Assume that the zoom position in the camera device of this example can be set to 0 to 100. The smaller the value of the zoom position, the closer the lens is to the wide-angle side, and the larger the value, the closer the lens is to the telephoto side. That is, the zoom wide angle shown in S604 indicates a state where the zoom position is zero and the viewing angle is the widest. In the camera device, if the face size at the time of zoom wide angle is 100 to 200, it is determined that the distance between the subject and the camera device can be predicted to be approximately 50cm to 150cm. That is, if the subject is neither too close nor too far from the camera device, it is determined that the subject may be the main person. In Figure 22A In the example of , a process of calculating the distance between the subject and the imaging device based on the face size has been described, but the distance to the subject may be measured according to other methods using a depth sensor, a compound eye lens, or the like.

[0378] Next, we will describe the input Figure 20B A specific example of temporary registration determination in the case of the subject information shown in FIG. Here, the zoom position is set to zero. Figure 20B Subject 1 and Subject 2 in the image are already in Figure 22A Since the person ID is not zero, the processing in S602 and subsequent steps is not performed.

[0379] because Figure 20B The person ID of subject 3 in Figure 22A In S601, the value is zero (not registered), so the processing in S602 and subsequent steps is executed. Figure 22B As shown in FIG, the temporary registration count of the subject ID 3 until the previous cycle is set to 30. Figure 22A In S602, the temporary registration counts up to the previous cycle are referred to, and if there is a temporary registration count for the subject ID 3, the information is taken over. Figure 20B The subject 3 in the image is facing forward, so the process starts from Figure 22A In S603, the process proceeds to S604. In S604, since the face size at the time of zoom wide angle is 120, the process proceeds to the process in S605, and in S605, the face reliability is 80, so the process proceeds to the process in S606. Figure 22A In S606, the temporary registration count is incremented by 1 to obtain 31. Since the temporary registration count is smaller than 50 in S608, the temporary registration count is stored in S609 and the flow proceeds to the next subject determination.

[0380] because Figure 20B The person ID of subject 4 in Figure 22A Since the temporary registration count in S601 is zero, the processing of S602 and thereafter is executed. In S602, the temporary registration count up to the previous cycle is referred to, and if there is a temporary registration count of the subject ID4, the information is taken over. Here, it is assumed that there is no temporary registration count of the subject ID up to the previous cycle. Figure 22A In S603, since the face orientation is 90 degrees to the left, the process proceeds to the process in S607, and the temporary registration count is set to 0. In S608, since the temporary registration count is less than 50, the process proceeds to the process in S609, and in S609, since the temporary registration count is 0, the temporary registration count is not stored, and the process ends.

[0381] Next, an example will be described in which Figure 22A In S608 of FIG, the temporary registration count becomes 50 or more, and the subject as the temporary registration target is arranged at the center of the angle of view by pan drive, tilt drive, and zoom drive. Figure 20B If Subject 3 in the image is a temporarily registered subject, the pan and tilt drive angles are calculated so that the subject's face position falls within a predetermined range. This predetermined range is within 5% of the center of the frame, meaning that the x-position coordinate values are between 432 and 528, and the y-position coordinate values are between 513 and 567. Since Subject 3's face size is between 100 and 200, the zoom position is not changed.

[0382] Figure 23A is shown relative to Figure 20A A diagram showing an example of image data when the pan position and tilt position are changed. Figure 23B It is shown that Figure 23A is a table showing an example of subject information extracted when image data is input to subject detection unit 1107. In this example, the face is positioned at the center of the frame at an appropriate size, so feature information extraction unit 1105 can acquire feature information. In the provisional registration determination process, unregistered persons who meet specific conditions over multiple periods are likely to be determined as key persons and added to person information management unit 1106.

[0383] <Main Registration>

[0384] Next, we will refer to Figure 24A and Figure 24B ,describe Figure 21 The main registration determination processing shown in S503 in . Figure 24A : is a flowchart showing the principal registration determination process performed by the principal registration determination unit 1109. Like the provisional registration determination, this determination process is executed in a plurality of cycles, and the principal person is determined from among persons who have been provisionally registered.

[0385] Figure 24B This table shows Count A, Count B, and the main registration count associated with a person ID. Count A and Count B are added under different conditions. If either Count A or Count B is 50 or greater, the main registration count is added. If the main registration count reaches 100, the corresponding subject is determined to be the main registration target person. Assume that the following processing is performed: the current Count A, Count B, and main registration count are stored at the time of determination for each cycle, and the various counts added up to the previous cycle are referenced and taken over in the next cycle.

[0386] In S1700, repeated processing corresponding to the number of detected subjects is started. When subject information is acquired from the subject detection unit 1107, the main registration determination unit 1109 performs Figure 24A The main registration determination unit 1109 determines "temporary registration" in S1701. If the registration status of the subject information is determined to be "temporary registration," the flow proceeds to S1702. If the registration status is not "temporary registration," the flow proceeds to the next subject determination process.

[0387] In S1702, the main registration determination unit 1109 refers to the counts stored until the previous frame, and in the case where various counts exist for the same person ID, the registration determination unit 1109 takes over the various counts. Then, the main registration determination unit 1109 performs a first main registration count determination (S1703), and further performs a second main registration count determination (S1704). The first main registration count determination is a determination based on the subject information of a single person. The following processing is performed: count A is added and the main registration count is added according to the distance and reliability between the object person and the camera device. The second main registration count determination is a determination based on the correlation with the "main registration" person who has been determined as the main person. Specifically, multiple "main registration persons" are detected at the same time, and the following processing is performed: count B is added and the main registration count is added according to whether the distance from the camera device is the same. The details of the first main registration count determination processing and the second main registration count determination processing will be described later.

[0388] Following S1704, in S1705, the principal registration determination unit 1109 compares the principal registration count value of the corresponding person with a threshold value of 100. If the principal registration count value is determined to be greater than 100, the process proceeds to S1706. If the principal registration count value is determined to be less than 100, the process proceeds to S1707. In S1706, the principal registration determination unit 1109 instructs the person information management unit 1106 to change the registration status of the corresponding person to "principal registration." Furthermore, in S1707, the principal registration determination unit 1109 stores various current counts. After S1706 and S1707, the process proceeds to S1708, where it is determined whether to terminate the repetitive processing. If the processing is to continue, the process returns to S1700 and continues processing for the next detected subject.

[0389] Then, refer to Figure 25 The flowchart is described Figure 24A The processing in S1703 (first principal registration count determination) in

[1701] is repeated. In S1801, the principal registration determination unit 1109 determines whether the face size when zooming wide is within the range of 100 to 200. If this condition is satisfied, the flow proceeds to the processing in S1802, and if not, the flow proceeds to the processing in S1804.

[0390] In S1802, the main registration determination unit 1109 determines whether the face reliability is equal to or greater than a threshold value of 80. If this condition is met, the process proceeds to S1803; if not, the process proceeds to S1804. If all of the conditions in S1801 and S1802 are met, the process proceeds to S1803, and a value corresponding to "zoom wide-angle face size / 10" is added to the count A. In S1804, the main registration determination unit 1109 sets the count A to zero and ends the process.

[0391] Following S1803, in S1805, the primary registration determination unit 1109 compares the value of Count A with a threshold value of 50. If it is determined that the value of Count A is equal to or greater than 50, the process proceeds to S1806. If it is determined that Count A is less than 50, the process ends. The primary registration determination unit 1109 increments the primary registration count by 1 in S1806 and sets Count A to zero in S1807. After S1807, the process ends.

[0392] Will refer to Figure 26 The flowchart is described Figure 24A The processing in S1704 (Second Principal Registration Count Determination) in the above flow proceeds to S1901. In S1902, the principal registration determination unit 1109 refers to the subject information and determines whether a person with a registration status of "principal registration" (i.e., multiple persons determined to be principal persons) is detected simultaneously. If it is determined that a principal registered person is detected simultaneously, the flow proceeds to S1902. If it is determined that a principal registered person is not detected simultaneously, the flow proceeds to S1905.

[0393] In S1902, the main registration determination unit 1109 refers to the facial size of the subject information and determines whether the facial size is similar to the facial size of any of the simultaneously detected main registration persons. Specifically, for example, if the facial size of the subject information is within the range of "±10% of the facial size of the main registration person" as the determination condition, the facial size is considered similar. If the condition of S1902 is met, the process proceeds to S1903. If not, the process proceeds to S1905.

[0394] In S1903, the principal registration determination unit 1109 compares the face reliability with a threshold value of 80. If the face reliability is determined to be equal to or greater than 80, the process proceeds to S1904. If the face reliability is determined to be less than 80, the process proceeds to S1905. In S1904, the principal registration determination unit 1109 adds a value corresponding to "face size when zooming wide angle / 10" to count B. In S1905, the principal registration determination unit 1109 sets count B to zero and ends the process.

[0395] Following S1904, in S1906, the primary registration determination unit 1109 compares the value of Count B with the threshold value 50. If it is determined that the value of Count B is equal to or greater than the threshold value 50, the process proceeds to S1907. If it is determined that the value of Count B is less than the threshold value 50, the process ends. In S1907, the primary registration determination unit 1109 increments the primary registration count by 1, and in S1908, sets Count B to zero, and then ends the process.

[0396] Subsequently, the process of obtaining the information in the main registration determination unit 1109 will be described. Figure 20B A specific example of the main registration determination in the case of the subject information shown. The zoom position is set to zero. Figure 20B The registration status of subject 1, subject 3 and subject 4 in Figure 24A Since S1701 is not "temporary registration", S1702 and subsequent processes are not executed. Figure 20B The registration status of subject 2 in Figure 24A Since S1701 is "temporary registration", the processing after S1702 is executed.

[0397] exist Figure 24A In S1702 of , count A, count B and main registration count up to the previous cycle are referred to, and if there are various counts of person ID 4, the information is taken over. Figure 24B As shown, the count A, count B, and main registration count of the person ID 4 until the previous cycle are set to 30, 40, and 70, respectively. The sum of the values of the count A and count B is the value of the main registration count. Figure 24A In S1703, the first primary registration count determination is performed. Figure 25 The face size when zooming wide angle in S1801 is 110, so the process enters the process in S1802, and since the face reliability is 90 in S1802, the process enters the process in S1803. Figure 25 In S1803 of FIG, since the face size when zooming wide is 110, 11 (=110 / 10) is added to the count A, and the count A becomes 41 (=30+11). Figure 25 In S1805 , since the value of the count A is smaller than the threshold value 50, the first primary registration count determination process ends.

[0398] Later, in Figure 24A In S1704, the second primary registration count determination is performed. Figure 26 In S1901, the subject information is referred to, and it is determined that the registration status of the subject 1 detected at the same time is "main registration". It is determined that the main registration person is detected at the same time, and the flow proceeds to the processing in S1902. Figure 26 In S1902, the facial dimensions of Subject 1 and Subject 2, the main registered person, are compared. Since Subject 1's facial dimension is 120, if the facial dimensions are 120 ± 10%, i.e., 108 to 132, the facial dimensions are determined to be similar. Since Subject 2's facial dimension is 110, the facial dimensions are determined to be similar to that of the main registered person, and the process proceeds to S1903. Since the facial reliability is 90 in S1903, the process proceeds to S1904.

[0399] exist Figure 26 In S1904 of , since the face size when zooming wide is 110, 11 (=110 / 10) is added to the count B, and the count B becomes 51 (=40+11). Figure 26 In S1906 of FIG, since the count B is equal to or greater than 50, the flow proceeds to the processing in S1907. In S1907, 1 is added to the value of the primary registration count 70 to obtain 71. In S1908, after the count B is set to zero, the second primary registration count determination processing is ended. Subsequently, in Figure 24A In S1705, since the value of the main registration count is less than the threshold value 100, the flow proceeds to processing in S1707. For person ID 4, count A is set to 41, count B is set to 0, and the main registration count is set to 71, and processing for storing the various counts is performed.

[0400] A person who has been temporarily registered and who has continuously satisfied the following conditions for multiple cycles through the principal registration determination process: the distance to the camera is within a predetermined range, or the distance to a person already determined as a principal is short, is determined as a principal person. The person information management unit 1106 can update the information based on the determination result.

[0401] <Impression Subject Determination>

[0402] Will refer to Figure 27A and Figure 27B describe Figure 21 Details of the imaging subject determination processing shown in S507 are described. Figure 27AThis is a flowchart showing the processing performed by the imaging subject determination unit 1110. This processing is executed every cycle, and the imaging subject person is determined from the detected persons. Upon receiving subject information from the subject detection unit 1107, the imaging subject determination unit 1110 executes the processing in S1001 to S1008 to determine the imaging subject. Based on the determination results, the pan drive angle, tilt drive angle, and zoom movement position are calculated in the processing in S1009 and S1010.

[0403] In S1001 , the imaging target determination unit 1110 refers to the subject information and determines whether a person with a priority set to “exist” is detected. If the corresponding person is detected, the flow proceeds to the process of S1002 , and if the corresponding person is not detected, the flow proceeds to the process of S1005 .

[0404] In S1002, the imaging target determination unit 1110 adds the person whose priority is set to "Present" as the imaging target person and proceeds to S1003. In S1003, the imaging target determination unit 1110 refers to the subject information and determines whether a person with a registration status of "Primary Registration" has been detected. If the corresponding person is detected, the process proceeds to S1004. If the corresponding person is not detected, the process proceeds to S1009. In S1004, the imaging target determination unit 1110 adds the person with a registration status of "Primary Registration" as the imaging target person and proceeds to S1009.

[0405] If a person with a priority of "Present" is detected, the person with a priority of "Present" and a registration status of "Primary Registration" is determined as the imaging target person in the processes of S1001 to S1004. In S1005, the imaging target determination unit 1110 refers to the subject information and determines whether a person with a registration status of "Primary Registration" has been detected. If the corresponding person is detected, the process proceeds to S1006. If the corresponding person is not detected, the process proceeds to S1009. In S1006, the imaging target determination unit 1110 adds the person with a registration status of "Primary Registration" as the imaging target person and proceeds to S1007.

[0406] In S1007, the imaging target determination unit 1110 refers to the subject information and determines whether a person whose registration status is "temporarily registered" has been detected. If the corresponding person has been detected, the process proceeds to S1008. If the corresponding person has not been detected, the process proceeds to S1009. In S1008, the imaging target determination unit 1110 adds the person whose registration status is "temporarily registered" as the imaging target person and proceeds to S1009.

[0407] If a person whose priority is set to "Present" is not detected and a person whose registration status is "Main Registration" is detected, the person to be imaged is determined through the processes in S1006 to S1008. That is, the person whose registration status is "Main Registration" and the person whose registration status is "Provisional Registration" are determined to be the person to be imaged.

[0408] At S1009, the imaging subject determination unit 1110 determines the number of imaging subject persons. If it is determined that there are one or more imaging subject persons, the process proceeds to S1010. If it is determined that the number of imaging subject persons is zero, the process ends. At S1010, the imaging subject determination unit 1110 calculates the pan drive angle, tilt drive angle, and zoom movement position so that the imaging subject is included in the angle of view, and outputs the calculation results to the drive control unit 1111.

[0409] Figure 27B : is a table showing the importance of people according to the registration status of subject information and priority settings. The imaging priority is represented by a numerical value from 1 to 4, where 1 is the highest imaging priority and 4 is the lowest imaging priority.

[0410] A person whose imaging priority is 1 is a person whose registration status is “primary registration” and whose priority is set to “existence”.

[0411] The person with a camera priority of 2 is a person whose registration status is "Main Registration" and whose priority is set to "None".

[0412] The person with a camera priority of 3 is a person whose registration status is "temporarily registered".

[0413] ·People with a camera priority of 4 are unregistered people.

[0414] according to Figure 27A In the processing in , if a person with an imaging priority of 1 is detected, the imaging subject determination unit 1110 sets the persons with imaging priorities of 1 and 2 as imaging subjects, and does not set the persons with imaging priorities of 3 and 4 as imaging subjects. If no person with an imaging priority of 1 is detected but a person with an imaging priority of 2 is detected, the imaging subject determination unit 1110 sets the persons with imaging priorities of 2 and 3 as imaging subjects, and does not set the person with imaging priority of 4 as imaging subjects. If no person with an imaging priority of 1 or 2 is detected, the determination result is that no subject is set as the imaging subject.

[0415] Figure 28A and Figure 28B is a diagram showing an example of image data and subject information. Figure 28A is a schematic diagram illustrating an example of image data input to the subject detection unit 1107 . Figure 28B It is shown in Figure 28A 1 is a table showing an example of subject information extracted when image data is input to the subject detection unit 1107 . Figure 28B The example in FIG is an example of information in which the number of subjects is 6 and includes subject ID, face size, face position, face orientation, face reliability, person ID, registration status, and priority setting for six subjects. Figure 28B A specific example of determining an imaging target in the case of the subject information shown is shown in FIG. The zoom position is set to zero.

[0416] exist Figure 27A S1001, refer to Figure 28B , and since the priority of subject 2 is set to "exist", the process enters the processing of S1002 and subject 2 is added as the imaging target. In S1003, referring to Figure 28B , and since the registration status of subject 1 is “main registration”, the flow proceeds to the process of S1004, and subject 1 is added as an imaging target.

[0417] exist Figure 27A In S1009, since the number of people being imaged is two, the process proceeds to S1010. In S1010, the pan drive angle, tilt drive angle, and zoom movement position are calculated so that subjects 1 and 2 are included in the angle of view. The description of specific numerical methods for calculating angles or positions will be omitted. There are methods for specifying angles or positions as absolute values, and methods for setting a minimum value for the drive angles and positions that can be specified and gradually changing the minimum value to the target angle or position over multiple cycles.

[0418] Figure 29 is a schematic diagram showing an example of image data as a result of the drive control unit 1111 controlling the respective drive units when the calculated pan drive angle, tilt drive angle, and zoom movement position are input. Figure 29 In the example, the pan drive, tilt drive and zoom position movement are controlled so that the center of gravity of the face positions of the right subject 1 and the left subject 2 is arranged at the center of the screen, and the face size of each subject is between 150 and 200.

[0419] Through the above control, it is possible to perform video recording so that subjects 1 and 2, which are the subject of the video and have been determined to have high video priority, are included in the field of view, while subjects 3 to 6, which are not the subject of the video and have been determined to have low video priority, are not included in the field of view. If a person with a video priority equal to or greater than a certain level is detected, the following processing is performed, in which people with similar video priorities are set as the subject of the video, while people with video priorities far from the main person are not set as the subject of the video. As a result, it is possible to perform video recording in which the main person is set as the subject of the video and people with low relationship are excluded from the subject of the video as much as possible.

[0420] Next, we will refer to Figure 17 、 Figure 30 to Figure 3 4 describes an example in which the importance determination unit 1514 is added. In this example, an example is shown in which the person information used to determine the imaging priority is further subdivided and the importance is changed according to the detection interval of each person to improve the discrimination accuracy of the main person.

[0421] Reference Figure 17 , the details of the processing performed by the control box 1100 will be described focusing on the differences from the above example. The character information management unit 1106 stores and manages character information associated with each character. Figure 30 Describe the character information.

[0422] Figure 30 This table shows an example of person information including importance. Items other than importance are the same as in the above example, so their descriptions are omitted. Importance is set on a scale of 1 to 10, with 1 indicating the lowest importance and 10 indicating the highest importance. The lower limit for importance is "0" if the name is left blank and "5" if a name is entered.

[0423] Upon receiving a facial image and feature information from the feature information extraction unit 1105, the character information management unit 1106 issues a new character ID, associates the character ID with the input facial image and feature information, and adds the new character information. When adding new character information, the initial value of the registration status is "Provisional Registration," the importance is "0" (not set), the initial value of the priority setting is "Not Present," and the initial value of the name is blank. Upon receiving the primary registration determination result (the character ID to be registered) from the primary registration determination unit 1109, the character information management unit 1106 changes the registration status of the character information corresponding to the corresponding character ID to "Primary Registration" and sets the importance to "1." If a user operation instructs the communication unit 1114 to change character information (information related to priority settings or names), the character information management unit 1106 changes the character information according to the instruction. If the priority setting or name of a character whose registration status is "Provisional Registration" is changed, the character information management unit 1106 changes the corresponding character's registration status to "Primary Registration" and sets the importance to "5" if the name has been changed.

[0424] When receiving an instruction to increase or decrease the importance of a person ID from the importance determination unit 1514, the person information management unit 1106 increases or decreases the importance of the person information corresponding to the person ID of the corresponding person. The subject detection unit 1107 detects a subject from the digital image data from the image processing unit 1103 and extracts information related to the detected subject. An example in which the subject detection unit 1107 detects a person's face as a subject will be described. The subject information includes, for example, the number of detected subjects, the position of the face, the size of the face, the orientation of the face, and the reliability of the face indicating the certainty of detection. This will be referred to later. Figure 31A and Figure 31B Describes an example of subject information.

[0425] The subject detection unit 1107 calculates the degree of similarity by comparing the characteristic information of each person acquired from the person information management unit 1106 with the characteristic information of the detected subject. If the degree of similarity is equal to or greater than a threshold, the subject detection unit 1107 adds the person ID, registration status, importance, and priority setting of the detected person to the subject information. The subject detection unit 1107 outputs the subject information to the temporary registration determination unit 1108, the main registration determination unit 1109, the imaging target determination unit 1110, and the importance determination unit 1514.

[0426] The imaging subject determination unit 1110 determines the imaging subject based on the subject information acquired from the subject detection unit 1107. The imaging subject determination unit 1110 also calculates the pan drive angle, tilt drive angle, and target zoom position required to arrange the imaging subject person at a specified size within the angle of view based on the determination result of the imaging subject person. A command based on the calculation result is output to the drive control unit 1111. This will be referred to later. Figure 34A The details of the imaging subject determination processing are described.

[0427] Figure 31A and Figure 31B is a diagram showing an example of image data and subject information. Figure 31A is a schematic diagram illustrating an example of image data input to the subject detection unit 1107 . Figure 31B It is shown in Figure 31A is a table showing an example of subject information extracted when the image data shown in FIG is input to the subject detection unit 1107. FIG shows an example in which the subject information includes the number of subjects, the subject ID of each subject, the face size, face position, face orientation, face reliability, person ID, registration status, importance, and priority setting. Items other than importance are the same as in the above example, and their description will be omitted.

[0428] The importance is the same as that managed by the person information management unit 1106. If the person ID is not zero, that is, if it is determined that the person is any person managed by the person information management unit 1106, the importance of the corresponding person acquired from the person information management unit 1106 is acquired.

[0429] Figure 32 2800 is a flowchart showing the overall process of recording, registering, and updating person information in this example, and the following processing is performed periodically. When the camera device is powered on, the camera unit 1022 begins periodic recording (moving image capture) to obtain image data for various determinations. The various determinations include image subject determination, temporary registration determination, primary registration determination, and importance determination. In S2800, repeated processing begins.

[0430] In S2801, image data acquired by video recording is output to the image processing unit 1103, and image data subjected to various types of image processing is acquired. When a subject is detected and subject information is acquired in S2802, main registration determination is performed in S2803, importance determination is performed in S2804, and temporary registration determination is performed in S2805. Descriptions of the temporary registration determination processing and the main registration determination processing will be omitted. In S2804, the importance determination unit 1514 determines the importance of the person by using information related to the detected subject. In the importance determination, the person information in the person information management unit 1106 is updated, but pan drive, tilt drive, and zoom drive are not performed.

[0431] S2806 is a process for determining whether temporary registration composition adjustment processing is currently being executed. If it is determined that temporary registration composition adjustment processing is currently being executed, the process proceeds to S2807. If it is determined that temporary registration composition adjustment processing is not currently being executed, the process proceeds to S2808. In S2807, the feature information extraction unit 1105 extracts feature information of the subject located in the image data center and outputs the feature information to the person information management unit 1106. In S2807, the imaging subject is determined.

[0432] After S2807 and S2808, the flow proceeds to the process of S2809, and it is determined whether the repetitive processing is to be ended. If the processing is to be continued, the flow returns to S2800. The processes in S2801 to S2808 are repeatedly executed according to the imaging cycle of the imaging unit 1022.

[0433] Next, refer to Figure 33A and Figure 33B ,describe Figure 32 Importance determination processing shown in S2804. Figure 33A : is a flowchart showing the processing performed by the importance determination unit 1514. The importance determination processing is performed in a plurality of cycles, and the importance of the person who has been mainly registered is determined and updated. Figure 33B This table shows the last detection date and time and the last update date and time associated with a person ID. The last detection date and time indicates the last detection of a primary registered person. The last update date and time indicates the last update of the primary registered person's importance. The last detection date and time and the last update date and time are stored in memory using data corresponding to the number of primary registered people and are referenced when determining each period.

[0434] Upon acquiring subject information from the subject detection unit 1107, the importance determination unit 1514 executes processing in S2901, then executes processing in S2902 to S2906 for the detected subject, and also executes processing in S2907 to S2909 for the primarily registered person. In S2901, the importance determination unit 1514 acquires the current date and time from the system time of the camera 101. In ST, repeated processing corresponding to the number of detected subjects begins. In S2902, the importance determination unit 1514 refers to the subject information and determines whether the registration status is "primary registration." If it is determined that the registration status is "primary registration," the process proceeds to processing in S2903. If the registration status is not "primary registration," the process proceeds to processing in STB.

[0435] In S2903, the importance determination unit 1514 sets the current date and time as the last detected date and time of the detected person. In S2904, the importance determination unit 1514 determines whether the current date and time has elapsed for more than 30 minutes and is within 24 hours of the last updated date and time. If this condition is met, the process proceeds to S2905. If not, the process proceeds to the STB.

[0436] The importance determination unit 1514 instructs the person information management unit 1106 to increase the importance by 1 in S2905 and sets the current date and time as the last updated date and time in S2906. It is determined whether the repeated processing is ended in the STB. If the processing continues, the flow returns to the STA and processes the next subject.

[0437] Next, the following processing is performed for each main registered person. In STC, repetitive processing begins for the number of main registered subjects. In S2907, the importance determination unit 1514 refers to the current date and time and determines whether there is a gap of more than one week between the last detected date and time and the last updated date and time. If it is determined that no detection or update has been performed for more than one week, the process proceeds to S2908. If it is determined that a detection or update has been performed within one week, the process proceeds to STD.

[0438] The importance determination unit 1514 instructs the person information management unit 1106 to decrement the importance by 1 in S2908 and sets the current date and time as the last updated date and time in S2909. In STD, it is determined whether to end the repetitive processing. If the processing continues, the flow returns to STC and continues processing the next main registered subject.

[0439] By using the importance determination process, the importance of people who are rediscovered every other day is increased, and the importance of subjects who have not been detected for more than a week is reduced. In other words, the importance of main characters who appear frequently can be increased, and the importance of irrelevant characters who are rarely seen or have been mainly registered can be reduced.

[0440] Will refer to Figure 34A and Figure 34B describe Figure 32 The imaging subject determination processing shown in S2808 in . Figure 34A 3 is a flowchart showing the processing performed by the imaging subject determination unit 1110. This processing is executed every cycle, and an imaging subject person is determined from among detected persons. Figure 34B This is a table showing the imaging priority of people according to the registration status, importance, and priority of subject information (imaging priority table). The imaging priority is represented by a numerical value from 1 to 13, where 1 indicates the highest imaging priority and 13 indicates the lowest imaging priority.

[0441] The person whose imaging priority is 1 is a person whose registration status is "Main Registration" and whose priority is set to "Present".

[0442] The persons with imaging priorities 2 to 11 are persons whose registration status is "Main registration" and whose priority is set to "Not present", and the higher the importance, the higher the imaging priority.

[0443] The person with a camera priority of 12 is a person whose registration status is "temporarily registered".

[0444] ·The person with camera priority 13 is an unregistered person.

[0445] Upon receiving the subject information from the subject detection unit 1107, the imaging subject determination unit 1110 performs processing in S3001 to S3004 to determine the imaging subject. Based on the determination result, processing in S3005 and S3006 is performed to calculate the pan drive angle, tilt drive angle, and zoom movement position.

[0446] In S3001, the imaging target determination unit 1110 refers to the subject information and Figure 34BThe imaging priority table shown in FIG3001 is used to obtain the imaging priority of each subject. In S3002, the imaging subject determination unit 1110 determines whether the imaging priority of the subject with the highest imaging priority among all detected subjects is equal to or less than a threshold value of 10. If this condition is met, the process proceeds to the processing in STE. If not, it is determined that there is no imaging subject and the processing ends. In STE, repeated processing corresponding to the number of detected subjects is started. In S3003, the imaging subject determination unit 1110 determines whether the imaging priority of each subject is less than the value obtained by adding 2 to the highest imaging priority of all subjects. If this condition is met, the process proceeds to S3004. If not, the process proceeds to STF. In STF, it is determined whether to terminate the repeated processing. If processing continues, the process returns to STE and continues processing for the next detected subject.

[0447] In S3004, the imaging subject determination unit 1110 adds the determined detected subject as an imaging subject. For example, if the imaging priority of the subject with the highest imaging priority is "4," subjects with imaging priorities of "4," "5," or "6" are determined as imaging subjects. If the imaging priority of the subject with the highest imaging priority is "7," subjects with imaging priorities of "7," "8," or "9" are determined as imaging subjects. After S3004, the process enters the processing in STF, where it is determined whether to end the repeated processing. If the processing is to continue, the process returns to STE and continues processing the next detected subject. When the repeated processing ends, the process enters the processing in S3005.

[0448] At S3005, the imaging subject determination unit 1110 determines whether one or more imaging subject persons exist. If this condition is met, the process proceeds to S3006; if not, the process ends. At S3006, the imaging subject determination unit 1110 calculates the pan drive angle, tilt drive angle, and zoom movement position so that the imaging subject is included in the angle of view, and outputs the calculation results to the drive control unit 1111. This completes the series of processes.

[0449] Through the above control, it is possible to perform video recording so that the subject of the video recording (i.e., the subject determined to have a high video recording priority) is included in the field of view, and the subject that is not the subject of the video recording (i.e., the subject determined to have a low video recording priority) is not included in the field of view. If a person with a high video recording priority is detected, multiple people with similar video recording priorities are determined as the video recording subjects, and people with different video recording priorities are determined not to be the video recording subjects. The following video recording can be performed, in which the main person is set as the video recording subject and people with low relationship are excluded from the video recording subject as much as possible.

[0450] (Variation)

[0451] The following describes a variation of the above example. In the above example, information related to features of a person's face is used as subject information. In this variation, feature information related to a subject other than a person (such as an animal or an object) can be used as subject information.

[0452] Figure 35A and Figure 35B An example is shown in which facial information of animals can be detected in addition to people. Figure 35A is a schematic diagram illustrating an example of image data input to the subject detection unit 1107 . Figure 35B is shown with Figure 35A A table of subject information corresponding to the image data in the image is displayed. When recording animals or objects, provisional registration and primary registration are performed separately from human subjects. Alternatively, if animals or objects are mixed with human subjects, the following process is performed to determine the recording subject by weighting their importance according to the subject type.

[0453] The above example is an example in which the lens barrel 102 including the imaging unit 1022 rotates about both the X-axis and the Y-axis, thereby enabling panning and tilting. Even if the lens barrel 102 cannot rotate about both the X-axis and the Y-axis, the present invention is still applicable as long as the lens barrel 102 can rotate about any axis. For example, in the case of a configuration in which the lens barrel 102 can rotate about the Y-axis, panning can be performed based on the orientation of the subject.

[0454] In the above example, the following camera device has been described, in which a lens barrel including a camera optical system and a camera element, and a camera control device for controlling the camera direction by using the lens barrel are integrated. The present invention is not limited to this. For example, the camera device can have a structure in which the lens device is replaceable. There is a structure in which the camera device is attached to a platform provided with a rotating mechanism driven in a pan direction and a tilt direction. The camera device can have a camera function and other functions. For example, there is a structure in which a platform capable of fixing a smartphone with a camera function and the smartphone are combined. The lens barrel, the rotating mechanism (tilt rotation unit and pan rotation unit) and the control box do not need to be physically connected. For example, the rotating mechanism or the zoom function can be controlled by wireless communication such as Wi-Fi.

[0455] The example of acquiring the characteristic information of a person by an imaging device has been described above. The present invention is not limited to this, and, for example, there may be a configuration in which a facial image or characteristic information in the person information is acquired from another imaging device for face registration or an external device such as a mobile terminal device, and the facial image or characteristic information is registered or added.

[0456] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes can be made within the scope of the gist of the present invention.

[0457] (Other embodiments)

[0458] The embodiments of the present invention can also be implemented by the following method, that is, providing software (program) that performs the functions of the above-mentioned embodiments to a system or device through a network or various storage media, and the computer or central processing unit (CPU) or microprocessing unit (MPU) of the system or device reads and executes the program.

[0459] While the present invention has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments. The scope of the following claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.

[0460] This application claims priority from Japanese Patent Application No. 2020-179882, filed on October 27, 2020, which is hereby incorporated by reference herein in its entirety.

Claims

1. A camera device capable of automatic camera shooting and automatic authentication and registration, the camera device comprising: a camera unit configured to capture a subject; a search unit configured to search for a subject detected in the image data acquired by the imaging unit; an authentication and registration unit configured to authenticate and store the detected subject; as well as a control unit configured to perform an authentication registration determination regarding whether a first condition for the authentication registration unit to perform automatic authentication registration is satisfied and an imaging determination regarding whether a second condition for performing automatic imaging is satisfied, and to control timings of the automatic imaging and the automatic authentication registration, wherein the control unit determines the timing of automatic authentication registration by performing authentication registration determination and imaging determination related to the detected subject while controlling the search in the search unit, and If the first condition is satisfied as a result of the authentication registration determination and the imaging determination, the authentication registration unit registers the detected subject, and if the first condition is not satisfied but the second condition is satisfied, the control unit controls automatic imaging.

2. The imaging device according to claim 1, in, The control unit prioritizes the authentication registration determination over the image capturing determination.

3. The imaging device according to claim 1 , further comprising: a first changing unit configured to change a camera direction; as well as a second changing unit configured to change a camera viewing angle, The control unit controls the timing at which the first changing unit or the second changing unit changes the shooting direction or the shooting angle of view in automatic shooting and automatic authentication registration.

4. The imaging device according to claim 3, in, The first changing unit includes a driving unit configured to rotationally move the imaging unit in a plurality of directions, and The second changing unit changes the viewing angle in automatic photography by driving a lens or performing image processing.

5. The imaging device according to claim 3, in, If it is determined that the first condition is satisfied, the control unit performs control so that the first changing unit arranges the face of the subject at the center of the imaging angle of view.

6. The imaging device according to claim 3, in, If it is determined that the first condition is satisfied, the control unit controls so that the second changing unit changes the face size of the subject to a preset size.

7. The imaging device according to claim 3, in, If it is determined that the second condition is satisfied and the detected subject is a person, the control unit controls so that the second changing unit changes the imaging angle of view to an imaging angle of view including the subject.

8. The imaging device according to claim 1, in, The control unit performs control to stop the process of automatic authentication registration if it is determined that the first condition is satisfied and there is an image capturing instruction from an external device.

9. The imaging device according to claim 1, in, The first condition based on the acquired facial information of the subject is: the reliability of facial detection is greater than a threshold and the subject's face faces the front of the camera device, or the state where the reliability is greater than the threshold continues and the subject's face faces the front of the camera device.

10. The imaging device according to claim 1, in, The control unit acquires information about the detected subject and image capturing history information, and calculates a score of image capturing and a determination threshold value, and The second condition is that the score exceeds the determination threshold.

11. The imaging device according to claim 3, in, If it is determined that the first condition is satisfied, the control unit controls the second changing unit to adjust the camera angle of view before automatic authentication registration.

12. The imaging device according to claim 1, further comprising: an acquisition unit configured to acquire information calculated or changed by machine learning of the image data, The control unit performs registration determination of the subject or imaging determination based on the second condition by using the information acquired by the acquisition unit.

13. The imaging device according to claim 12, in, The control unit makes a determination as to whether a condition for transitioning to a low power consumption state or a condition for canceling a low power consumption state is satisfied by using the information acquired by the acquisition unit, and controls power supply based on a result of the determination.

14. The imaging apparatus according to any one of claims 1 to 13, in, In automatic photography, the control unit obtains information related to the distance and detection frequency of the subject, determines the photography priority of each subject, and determines the subject with a priority within a preset range among the multiple detected subjects as the photography target subject.

15. The imaging device according to claim 14, in, The control unit determines a first subject having a first priority level and a second subject having a second priority level as imaging target subjects, the second priority level being within a preset range from the first priority level.

16. The imaging device according to claim 15, in, The control unit performs control of automatic imaging without setting a subject having a priority lower than the second priority as an imaging target.

17. The imaging device according to claim 15, in, The control unit determines an imaging priority of each subject by using information on distances from the imaging apparatus to the first subject and the second subject.

18. The imaging device according to claim 14, in, The control unit performs processing for storing and managing feature information of a subject in a storage unit, and determines whether the feature information of the detected subject is consistent with the feature information stored in the storage unit.

19. The imaging device according to claim 18, in, The storage unit stores feature information of the subject and the priority in association with each other.

20. The imaging device according to claim 19, in, If a subject corresponding to the feature information stored in the storage unit is detected, the control unit performs processing for updating the priority level stored in the storage unit according to the priority level of the detected subject.

21. The imaging device according to claim 18, in, If feature information of the detected object is acquired, the control unit performs a process for storing the feature information of the object whose priority is a preset value or within a preset range in the storage unit.

22. The imaging device according to claim 14, in, The control unit determines a priority level of the subject based on an elapsed time from a last detection date and time of the detected subject.

23. The imaging device according to claim 14, further comprising: a changing unit configured to change a camera direction, The control unit controls the image capturing of the determined image capturing object by controlling the changing unit.

24. The imaging device according to claim 14, further comprising: a changing unit configured to change a camera angle of view, The control unit controls the image capturing in a state where the determined image capturing target object is included in the image capturing angle of view by controlling the changing unit.

25. The imaging device according to claim 24, in, The control unit determines the priority of the subject by using information on the face orientation of the subject or reliability indicating certainty of the face.

26. The imaging device according to claim 25, in, The control unit performs control for outputting priority levels and image data of a face of a subject.

27. A camera device capable of automatic camera shooting and automatic authentication and registration, the camera device comprising: a camera unit configured to capture a subject; a search unit configured to search for a subject detected in the image data acquired by the imaging unit; an authentication and registration unit configured to authenticate and store the detected subject; as well as a control unit configured to perform an authentication registration determination regarding whether a first condition for the authentication registration unit to perform automatic authentication registration is satisfied and an imaging determination regarding whether a second condition for performing automatic imaging is satisfied, and to control timings of the automatic imaging and the automatic authentication registration, wherein the control unit determines the timing of automatic authentication registration by performing authentication registration determination and imaging determination related to the detected subject while controlling the search in the search unit, and The control unit sets the result determined by the authentication registration to take priority over the result determined by the photographing according to the number of photographing times or the time interval between photographing.

28. The imaging device according to claim 27, in, In automatic photography, the control unit obtains information related to the distance and detection frequency of the subject, determines the photography priority of each subject, and determines the subject with a priority within a preset range among the multiple detected subjects as the photography target subject.

29. The imaging device according to claim 28, in, The control unit determines a first subject having a first priority level and a second subject having a second priority level as imaging target subjects, the second priority level being within a preset range from the first priority level.

30. The imaging device according to claim 29, in, The control unit performs control of automatic imaging without setting a subject having a priority lower than the second priority as an imaging target.

31. The imaging device according to claim 29, in, The control unit determines an imaging priority of each subject by using information on distances from the imaging apparatus to the first subject and the second subject.

32. The imaging device according to claim 28, in, The control unit performs processing for storing and managing feature information of a subject in a storage unit, and determines whether the feature information of the detected subject is consistent with the feature information stored in the storage unit.

33. The imaging device according to claim 32, in, The storage unit stores feature information of the subject and the priority in association with each other.

34. The imaging device according to claim 33, in, If a subject corresponding to the feature information stored in the storage unit is detected, the control unit performs processing for updating the priority level stored in the storage unit according to the priority level of the detected subject.

35. The imaging device according to claim 32, in, If feature information of the detected object is acquired, the control unit performs a process for storing the feature information of the object whose priority is a preset value or within a preset range in the storage unit.

36. The imaging device according to claim 28, in, The control unit determines a priority level of the subject based on an elapsed time from a last detection date and time of the detected subject.

37. The imaging device according to claim 28, further comprising: a changing unit configured to change a camera direction, The control unit controls the image capturing of the determined image capturing object by controlling the changing unit.

38. The imaging device according to claim 28, further comprising: a changing unit configured to change a camera angle of view, The control unit controls the image capturing in a state where the determined image capturing target object is included in the image capturing angle of view by controlling the changing unit.

39. The imaging device according to claim 38, in, The control unit determines the priority of the subject by using information on the face orientation of the subject or reliability indicating certainty of the face.

40. The imaging device according to claim 39, in, The control unit performs control for outputting priority levels and image data of a face of a subject.

41. A control method executed in a camera device capable of automatic camera shooting and automatic authentication and registration, the control method comprising: searching for a subject detected in image data acquired by the imaging unit; Authenticate and register the detected subject; as well as performing an authentication registration determination regarding whether a first condition for performing automatic authentication registration is satisfied and an imaging determination regarding whether a second condition for performing automatic imaging is satisfied, and controlling the timing of automatic imaging and automatic authentication registration, wherein the control determines the timing of automatic authentication registration by performing authentication registration determination and imaging determination related to a detected subject while controlling a search for the subject, and If the results of the authentication registration determination and the photographing determination are that the first condition is met, the detected subject is registered through the authentication and registration, and if the first condition is not met but the second condition is met, automatic photographing is controlled through the control.

42. A non-transitory recording medium storing a control program for an imaging device capable of automatic imaging and automatic authentication and registration, the control program causing a computer to perform steps of a method for controlling the imaging device, the method comprising: searching for a subject detected in image data acquired by the imaging unit; Authenticate and register the detected subject; as well as performing an authentication registration determination regarding whether a first condition for performing automatic authentication registration is satisfied and an imaging determination regarding whether a second condition for performing automatic imaging is satisfied, and controlling the timing of automatic imaging and automatic authentication registration, wherein the control determines the timing of automatic authentication registration by performing authentication registration determination and imaging determination related to a detected subject while controlling a search for the subject, and If the results of the authentication registration determination and the photographing determination are that the first condition is met, the detected subject is registered through the authentication and registration, and if the first condition is not met but the second condition is met, automatic photographing is controlled through the control.

43. A computer program product comprising a control program for an imaging device capable of automatic imaging and automatic authentication and registration, the control program causing a computer to perform the steps of a method for controlling the imaging device, the method comprising: searching for a subject detected in image data acquired by the imaging unit; Authenticate and register the detected subject; as well as performing an authentication registration determination regarding whether a first condition for performing automatic authentication registration is satisfied and an imaging determination regarding whether a second condition for performing automatic imaging is satisfied, and controlling the timing of automatic imaging and automatic authentication registration, wherein the control determines the timing of automatic authentication registration by performing authentication registration determination and imaging determination related to a detected subject while controlling a search for the subject, and If the results of the authentication registration determination and the photographing determination are that the first condition is met, the detected subject is registered through the authentication and registration, and if the first condition is not met but the second condition is met, automatic photographing is controlled through the control.

Citation Information

Patent Citations

  • Camera

    JP2001051338A

  • Automatic photographing system, and picture management server

    JP2007325285A

  • Roll product package

    JP2020179882A

  • Zoom control device and control method of zoom control device

    CN105744148A

  • Search apparatus, imaging apparatus including the same, and search method

    CN107862243A