Camera device, control method thereof, and storage medium

By integrating status detection and control functions in the camera device, the subject search range is optimized, and the problem of meaningless search in the camera device is solved, and the quality of image capture and battery life are improved.

CN114827458BActive Publication Date: 2025-05-02CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210385296.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-28
Filing Date
2018-12-04
Publication Date
2025-05-02
Estimated Expiration
2038-12-04

AI Technical Summary

Technical Problem

When the user wears the camera device on his body, the process of automatically searching for the subject by the camera device may lead to meaningless searches, consume battery power, and fail to capture images that the user likes.

Method used

The imaging equipment design is adopted that includes an imaging component, a subject detection component, a state detection component and a control component. The state detection component detects the moving state of the imaging device, and the control component controls the search range of the subject detection component based on the detected state information to reduce meaningless searches.

Benefits of technology

Effectively reduce meaningless subject searches, improve the probability of capturing users' favorite images, and save battery power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114827458B_ABST
    Figure CN114827458B_ABST
Patent Text Reader

Abstract

The present invention provides a camera device, a control method thereof, and a storage medium. In automatic camera shooting using a mobile camera device, the present invention eliminates meaningless searches for a subject and increases the probability of obtaining images that a user likes. The camera device includes: a camera unit for capturing an image of a subject; a subject detection unit for detecting a subject from image data captured by the camera unit; a state detection unit for detecting information related to the state of movement of the camera device itself; and a control unit for controlling the range of the subject detection unit to search for a subject based on the state information of the camera device detected by the state detection unit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] (This application is a divisional application of an application filed on December 4, 2018, with application number 201880081863.2 and invention name “Camera device and control method thereof, program and storage medium”.) Technical Field

[0002] The invention relates to an automatic photographing technology used in a photographing device. Background Art

[0003] A life-logging camera that periodically takes continuous pictures without requiring a shooting instruction from the user is known (Patent Document 1). The life-logging camera is used in a state of being attached to the user's body with a belt or the like, and records scenes from the user's daily life as images at set time intervals. The life-logging camera does not take pictures at a time specified by the user pressing a shutter button or the like. Rather, the camera automatically takes pictures at each set time interval, making it possible to take images of unexpected moments that would not normally be photographed.

[0004] Patent document 2 discloses a technique for automatically searching for and photographing a subject, which is applied in a camera configured to be able to change the shooting direction. Even in automatic photography, composing the picture based on the detected subject makes it possible to increase the chance of photographing an image that the user will like.

[0005] Prior art literature

[0006] Patent Literature

[0007] Patent Document 1: Japanese Patent Application No. 2016-536868

[0008] Patent Document 2: Japanese Patent 05453953 Summary of the invention

[0009] Problem that the invention aims to solve

[0010] When capturing images for the purpose of life recording, images may also be recorded that the user is not very interested in. Automatically panning and tilting the camera to search for surrounding subjects and taking pictures at a viewing angle that includes the detected subjects can increase the chance of recording images that will be liked by the user.

[0011] However, in the case where the user searches for a subject while wearing the camera on his or her body, the camera itself is moving. Thus, even if the camera is pointed at the detected subject again to shoot the subject after the search operation is performed, the subject may no longer be visible. There are also cases where the subject has moved away and is too small, making the subject search meaningless. This situation is problematic in that not only does the user fail to obtain an image he or she likes, but battery power will be consumed to re-search for the subject, which reduces the amount of time that can be used to shoot images.

[0012] The present invention is achieved in view of the above-mentioned problems, and eliminates meaningless searches for subjects and increases the probability that images that users like can be obtained.

[0013] Solutions for solving problems

[0014] A camera device according to the present invention is characterized in that it includes: a camera component for capturing an image of a subject; a subject detection component for detecting a subject from image data captured by the camera component; a state detection component for detecting information related to the state of movement of the camera device itself; and a control component for controlling the range in which the subject detection component searches for a subject based on the state information of the camera device detected by the state detection component.

[0015] Effects of the Invention

[0016] According to the present invention, meaningless searches for objects can be eliminated, and the probability of being able to obtain an image that a user likes can be increased.

[0017] Other features and advantages of the present invention will be apparent from the following description taken in conjunction with the accompanying drawings. Note that throughout the accompanying drawings, the same reference numerals represent the same or similar components. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0019] Figure 1A is a diagram schematically showing the appearance of a camera serving as a first embodiment of an image pickup apparatus according to the present invention.

[0020] Figure 1B is a diagram schematically showing the appearance of a camera serving as a first embodiment of an image pickup apparatus according to the present invention.

[0021] Figure 2 is a block diagram showing the overall structure of a camera according to the first embodiment.

[0022] Figure 3 is a diagram showing an example of the structure of a wireless communication system between a camera and an external device.

[0023] Figure 4 It is a diagram showing the structure of an external device.

[0024] Figure 5 is a diagram showing the configuration of a camera and external devices.

[0025] Figure 6 It is a diagram showing the structure of an external device.

[0026] Fig. 7A is a flowchart showing operations performed by the first control unit.

[0027] Figure 7B is a flowchart showing operations performed by the first control unit.

[0028] Figure 8 is a flowchart showing operations performed by the second control unit.

[0029] Fig. 9 is a flowchart showing operations performed in the image capture mode process.

[0030] Fig. 10A It is a diagram showing the area division in the captured image.

[0031] Fig. 10B It is a diagram showing the area division in the captured image.

[0032] Fig. 10C It is a diagram showing the area division in the captured image.

[0033] Fig. 10D It is a diagram showing the area division in the captured image.

[0034] Fig. 10E It is a diagram showing the area division in the captured image.

[0035] Fig.11 is a diagram showing a neural network.

[0036] Fig.12 It is a diagram showing browsing of images in an external device.

[0037] Fig.13 : is a flowchart showing the learning mode determination.

[0038] Fig.14 is a flowchart showing the learning process.

[0039] Fig.15A is a diagram showing an example of attaching an accessory.

[0040] Fig. 15B is a diagram showing an example of attaching an accessory.

[0041] Fig. 15C is a diagram showing an example of attaching an accessory.

[0042] Fig.15D is a diagram showing an example of attaching an accessory.

[0043] Fig.16A is a diagram showing a subject search range when a handheld accessory is attached.

[0044] Fig. 16B is a diagram showing a subject search range when a handheld accessory is attached.

[0045] Fig.17A : is a diagram showing a subject search range when the desktop accessory is attached.

[0046] Fig. 17B : is a diagram showing a subject search range when the desktop accessory is attached.

[0047] Fig.18 is a diagram showing controls for each accessory type. DETAILED DESCRIPTION

[0048] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0049] (First embodiment)

[0050] <Camera Structure>

[0051] Figure 1A and Figure 1B is a diagram schematically showing the appearance of a camera serving as a first embodiment of an image pickup apparatus according to the present invention. Figure 1A The camera 101 shown is provided with a power switch, an operation member capable of performing camera operations, and the like. A lens barrel 102 is movably attached to a fixed portion 103 of the camera 101, and the lens barrel 102 includes an imaging lens group and an image sensor, etc. in an integrated manner as an optical imaging system for capturing an image of a subject. Specifically, the lens barrel 102 is attached to the fixed portion 103 via a pitch rotation unit 104 and a pan rotation unit 105 as mechanisms capable of being rotationally driven relative to the fixed portion 103.

[0052] The pitch rotation unit 104 includes a Figure 1B The motor drive mechanism for driving the lens barrel 102 to rotate in the tilt direction shown in FIG. Figure 1B1 and 2. The motor driving mechanism for rotationally driving the lens barrel 102 in the yaw direction shown in FIG. 1. In other words, the camera 101 has a mechanism for rotationally driving the lens barrel 102 in two axis directions. Figure 1B The axes shown are defined relative to the position of the fixed portion 103. An angular velocity meter 106 and an accelerometer 107 are arranged in the fixed portion 103 of the camera 101. The camera 101 detects vibration based on output signals from the angular velocity meter 106, the accelerometer 107, and the like, and can correct shake and tilt, and the like in the lens barrel 102 by rotationally driving the tilt rotation unit 104 and the pan rotation unit 105. The angular velocity meter 106, the accelerometer 107, and the like also detect movement of the camera based on measurement results obtained at set intervals.

[0053] Figure 2 1 is a block diagram showing the overall structure of the camera 101 according to the present embodiment. Figure 2 , the first control unit 223 includes, for example, a CPU (MPU) and a memory (DRAM, SRAM), etc. The first control unit 223 controls the respective blocks of the camera 101 and controls data transfer between the respective blocks, etc., by executing various processes according to programs stored in the nonvolatile memory (EEPROM) 216. The nonvolatile memory 216 is an electrically erasable / recordable memory that stores operation constants and programs, etc. used by the first control unit 223 as described above.

[0054] exist Figure 2 , the zoom unit 201 includes a zoom lens for performing magnification (enlarging and reducing the formed subject image). The zoom drive control unit 202 controls the drive of the zoom unit 201 and detects the focal length at this time. The focus unit 203 includes a focus lens for adjusting the focus. The focus drive control unit 204 controls the drive of the focus unit 203. The image pickup unit 206 includes an image sensor. The image pickup unit 206 receives incident light via each lens group and outputs information on the charge generated by the amount of light as an analog image signal to the image processing unit 207. Note that the zoom unit 201, the focus unit 203, and the image pickup unit 206 are arranged in the lens barrel 102.

[0055] The image processing unit 207 applies image processing such as distortion correction, white balance adjustment, and color interpolation to the digital image data obtained by A / D conversion of the analog image signal, and outputs the processed digital image data. The digital image data output from the image processing unit 207 is converted into a format for recording such as JPEG by the image recording unit 208, and then stored in the memory 215 or sent to the image output unit 217 (described later) or the like.

[0056] The lens barrel rotation drive unit 205 rotates the lens barrel 102 in the pitch direction and the pan direction by driving the pitch rotation unit 104 and the pan rotation unit 105. The device oscillation detection unit 209 includes an angular velocity meter (gyro sensor) 106 for detecting the angular velocity of the camera 101 in the three-axis direction, and an accelerometer (accelerometer) 107 for detecting the acceleration of the camera 101 in the three-axis direction. The rotation angle and offset of the device are calculated based on the signals detected by these sensors.

[0057] The audio input unit 213 obtains an audio signal from the periphery of the camera 101 via a microphone provided in the camera 101, converts the audio into a digital audio signal, and sends the signal to the audio processing unit 214. The audio processing unit 214 performs audio-related processing such as optimization on the input digital audio signal. The audio signal processed by the audio processing unit 214 is sent to the memory 215 by the first control unit 223. The memory 215 temporarily stores the image signal and the audio signal obtained from the image processing unit 207 and the audio processing unit 214.

[0058] The image processing unit 207 and the audio processing unit 214 read out the image signal and the audio signal temporarily stored in the memory 215, and encode the image signal and the audio signal to generate a compressed image signal and a compressed audio signal. The first control unit 223 sends the compressed image signal and the compressed audio signal to the recording / playback unit 220.

[0059] The recording / playback unit 220 records the compressed image signal and the compressed audio signal generated by the image processing unit 207 and the audio processing unit 214, and other control data related to imaging, etc., in the recording medium 221. Without compressing and encoding the audio signal, the first control unit 223 sends the audio signal generated by the audio processing unit 214 and the compressed image signal generated by the image processing unit 207 to the recording / playback unit 220, and causes these signals to be recorded in the recording medium 221.

[0060] The recording medium 221 may be a recording medium built into the camera 101, or a removable recording medium, and can record various data such as a compressed image signal, a compressed audio signal, and an audio signal generated by the camera 101. A medium having a larger capacity than the nonvolatile memory 216 is generally used for the recording medium 221. For example, the recording medium 221 may be any type of recording medium such as a hard disk, an optical disk, a magneto-optical disk, a CD-R, a DVD-R, a magnetic tape, a nonvolatile semiconductor memory, a flash memory, or the like.

[0061] The recording / playback unit 220 reads out (or plays back) the compressed image signal, compressed audio signal, audio signal, various data and programs, etc. recorded in the recording medium 221. Then, the first control unit 223 sends the read compressed image signal and compressed audio signal to the image processing unit 207 and the audio processing unit 214. The image processing unit 207 and the audio processing unit 214 temporarily store the compressed image signal and the compressed audio signal in the memory 215, decode the signals through a predetermined process, and send the decoded signals to the image output unit 217.

[0062] The audio input unit 213 is provided with a plurality of microphones. The audio processing unit 214 can detect the direction of the sound relative to the plane on which the plurality of microphones are arranged, so that it can perform the search for the subject and automatic video recording, which will be described later. In addition, the audio processing unit 214 detects a specific voice command. The structure can be as follows: in addition to several commands registered in advance, the user can also register a specific voice as a voice command in the camera. The audio processing unit 214 also recognizes the sound scene. In the sound scene recognition, a network pre-trained by machine learning based on a large amount of audio data is used to determine the sound scene. For example, a network for detecting specific scenes such as cheering, applause and speaking of the audience is set in the audio processing unit 214, and the network is used to detect specific sound scenes and specific voice commands, etc. When detecting a specific sound scene or a specific voice command, the audio processing unit 214 outputs a detection trigger signal to the first control unit 223 or the second control unit 211, etc.

[0063] The camera 101 is provided with a second control unit 211 for controlling power supply to the first control unit 223, in addition to the first control unit 223 of the main system that controls the entirety of the camera 101. The first power supply unit 210 and the second power supply unit 212 supply power for operation to the first control unit 223 and the second control unit 211, respectively. In response to pressing of a power button provided in the camera 101, power is first supplied to the first control unit 223 and the second control unit 211. However, as will be described later, the first control unit 223 itself may perform control for disconnecting power supply to the first power supply unit 210. The second control unit 211 operates even when the first control unit 223 is not operating, and takes information from the device oscillation detection unit 209 and the audio processing unit 214, etc. as input. The second control unit 211 determines whether the first control unit 223 is operating based on various input information, and instructs the first power supply unit 210 to supply power to the first control unit 223 when it is determined that the first control unit 223 is operating.

[0064] The audio output unit 218 outputs a preset audio pattern from a speaker built into the camera 101, for example, during image capture, etc. The LED control unit 224 causes an LED provided in the camera 101 to light up based on a preset lighting pattern or blinking pattern, for example, during image capture, etc. The image output unit 217 is constituted by, for example, an image output terminal, and outputs an image signal so that an image is displayed in a connected external display, etc. The audio output unit 218 and the image output unit 217 may be a single integrated terminal, for example, a High-Definition Multimedia Interface (HDMI; registered trademark) terminal.

[0065] The communication unit 222 is a portion for communicating between the camera 101 and an external device, and, for example, transmits and receives data such as an audio signal, an image signal, a compressed audio signal, and a compressed image signal. The communication unit 222 also receives commands for starting and stopping image capture, and control signals related to image capture such as pan, tilt, and zoom drive, and the like, and drives the camera 101 based on instructions from the external device. The communication unit 222 also transmits and receives information such as various parameters related to learning processed by the learning processing unit 219 (described later) between the camera 101 and the external device. For example, the communication unit 222 may include an infrared communication module, a Bluetooth (registered trademark) communication module, a wireless LAN communication module such as a wireless LAN module, a wireless USB (registered trademark) or a GPS receiver, or the like.

[0066] The environment sensor 226 detects the state of the surrounding environment of the camera 101 at every predetermined period. The environment sensor 226 includes a temperature sensor for detecting the temperature around the camera 101, an air pressure sensor for detecting the change of the air pressure around the camera 101, and an illumination sensor for detecting the brightness around the camera 101. The environment sensor 226 also includes a humidity sensor for detecting the humidity around the camera 101, and a UV sensor for detecting the amount of ultraviolet light around the camera 101, etc. In addition to the detected temperature information, air pressure information, brightness information, humidity information, and UV information, the temperature change amount, air pressure change amount, brightness change amount, humidity change amount, ultraviolet light change amount, etc. obtained by calculating the change rate of the detected various information at predetermined time intervals are used to determine automatic camera shooting, etc.

[0067] <Communication with External Devices>

[0068] Figure 3 is a diagram showing an example of a configuration of a wireless communication system between the camera 101 and the external device 301. The camera 101 is a digital camera having an imaging function, and the external device 301 is a smart device including a Bluetooth communication module and a wireless LAN communication module.

[0069] The camera 101 and the external device 301 can communicate using a first communication 302, which is performed, for example, through a wireless LAN conforming to the IEEE802.11 standard series, and a second communication 303, such as Bluetooth Low Energy (hereinafter referred to as "BLE"), which has, for example, a master / slave relationship including a control station and a slave station. Note that the wireless LAN and BLE are merely examples of communication methods, and, for example, other communication methods may be used as long as the communication device has two or more communication functions and one of the communication functions can control the other communication functions in the communication performed according to the relationship between the control station and the slave station. However, it is assumed that the first communication 302, which is a wireless LAN or the like, can communicate at a higher speed than the second communication 303, which is BLE or the like, and the second communication 303 consumes less power, has a shorter communication range, or both, than the first communication 302.

[0070] Will refer to Figure 4 301. The structure of the external device 301 is described. In addition to the wireless LAN control unit 401 for wireless LAN and the BLE control unit 402 for BLE, the external device 301 also includes a public wireless control unit 406 for public wireless communication. The external device 301 also includes a packet transmission / reception unit 403. The wireless LAN control unit 401 performs RF control and communication processing of the wireless LAN, driver processing for implementing various controls of communication through the wireless LAN that conforms to the IEEE 802.11 standard series, and protocol processing related to communication through the wireless LAN. The BLE control unit 402 performs RF control and communication processing of the BLE, driver processing for implementing various controls of communication through the BLE, and protocol processing related to communication through the BLE. The public wireless control unit 406 performs RF control and communication processing of public wireless communication, driver processing for implementing various controls of public wireless communication, and protocol processing related to public wireless communication. Public wireless communication conforms to, for example, the IMT (International Multimedia Telecommunications) standard or the LTE (Long Term Evolution) standard. The packet transmission / reception unit 403 performs processing for performing at least one of transmission and reception of packets related to wireless LAN and BLE communication and public wireless communication. Although the present embodiment describes the external device 301 as performing at least one of transmission and reception of packets in communication, it should be noted that a communication format other than packet switching, such as line switching, may be used instead.

[0071] The external device 301 further includes a control unit 411, a storage unit 404, a GPS receiving unit 405, a display unit 407, an operation unit 408, an audio input / audio processing unit 409, and a power supply unit 410. The control unit 411 controls the entire external device 301 by, for example, executing a control program stored in the storage unit 404. The storage unit 404 stores, for example, the control program executed by the control unit 411, and various information such as parameters required for communication. Various operations (described later) are realized by the control unit 411 executing the control program stored in the storage unit 404.

[0072] The power supply unit 410 supplies power to the external device 301. The display unit 407 has a function that enables it to output information that is visually recognizable using an LCD or LED, etc., and to perform audio output using a speaker, etc., and displays various information. The operation unit 408 includes, for example, a button that accepts an operation performed by a user on the external device 301, etc. Note that the display unit 407 and the operation unit 408 may be configured by a common member such as a touch panel, etc., for example.

[0073] The audio input / audio processing unit 409 obtains the voice uttered by the user using, for example, a general-purpose microphone built into the external device 301, and may be configured to recognize an operation command from the user using voice recognition processing. In addition, using a dedicated application in the external device 301, voice commands uttered by the user may be obtained, and these voice commands may be registered as specific voice commands to be recognized by the audio processing unit 214 of the camera 101 via the first communication 302 using the wireless LAN.

[0074] The GPS (Global Positioning System) receiving unit 405 receives a GPS signal from satellite communication, analyzes the GPS signal, and estimates the current position (longitude / latitude information) of the external device 301. Alternatively, the current position of the external device 301 can be estimated by using information based on wireless networks existing in the surrounding area, such as WPS (Wi-Fi Positioning System). In the case where the current GPS position information obtained is within a preset position range (within a range with a predetermined radius centered on the detection position), and in the case where the GPS position information has changed by a predetermined amount or more, etc., the movement information is communicated to the camera 101 via the BLE control unit 402. Then, this information is used as a parameter in automatic shooting and automatic editing, etc., which will be described later.

[0075] As described above, the camera 101 and the external device 301 exchange data by communication using the wireless LAN control unit 401 and the BLE control unit 402. For example, data such as an audio signal, an image signal, a compressed audio signal, and a compressed image signal are transmitted and received. In addition, an image capturing instruction, etc., voice command registration data, a notification of detection of a predetermined position based on GPS position information, and a notification of movement of a place, etc. are transmitted from the external device 301 to the camera 101. Training data used in a dedicated application in the external device 301 is also transmitted and received.

[0076] <Structure of accessories>

[0077] Figure 5 is a diagram showing an example of a structure of an external device 501 capable of communicating with the camera 101. The camera 101 is a digital camera having an imaging function, and the external device 501 is, for example, a wearable device including various sensing units, which can communicate with the camera 101 using a Bluetooth communication module or the like.

[0078] The external device 501 is configured to be attachable to the user's arm or the like, and is equipped with a sensor for detecting biological information such as the user's pulse, heartbeat, and blood flow at a predetermined cycle, and an accelerometer capable of detecting the user's movement state.

[0079] The biological information detection unit 602 includes, for example, a pulse sensor for detecting a pulse, a heartbeat sensor for detecting a heartbeat, a blood flow sensor for detecting a blood flow, and a sensor for detecting a change in electric potential caused by skin contact using a conductive polymer. This embodiment describes a heartbeat sensor as being used as the biological information detection unit 602. The heartbeat sensor detects the heartbeat of the user by irradiating the user's skin with infrared light using an LED or the like, detecting the infrared light that has passed through body tissue using a light receiving sensor, and processing the signal thus obtained. The biological information detection unit 602 outputs the detected biological information as a signal to the control unit 607 (see Figure 6 ).

[0080] The shake detection unit 603 for detecting the moving state of the user includes, for example, an accelerometer and a gyro sensor, etc., and can detect motion based on acceleration information, such as whether the user is moving, or performing an action such as waving his or her arm, etc. An operation unit 605 for accepting the user's operation of the external device 501, and a display unit 604 for outputting visually recognizable information, such as an LCD or LED monitor, are also provided.

[0081] Figure 6is a diagram showing the structure of the external device 501. As described above, the external device 501 includes, for example, the control unit 607, the communication unit 601, the biological information detection unit 602, the shaking detection unit 603, the display unit 604, the operation unit 605, the power supply unit 606, and the storage unit 608.

[0082] The control unit 607 controls the entire external device 501 by, for example, executing a control program stored in the storage unit 608. The storage unit 608 stores, for example, the control program executed by the control unit 607 and various information such as parameters required for communication. For example, various operations (described later) are implemented by the control unit 607 executing the control program stored in the storage unit 608.

[0083] The power supply unit 606 supplies power to the external device 501. The display unit 604 has an output unit capable of outputting visually recognizable information using an LCD or LED, etc., and an output unit capable of outputting audio using a speaker, etc., and displays various information. The operation unit 605 includes, for example, a button that accepts an operation performed by a user on the external device 501, etc. Note that the display unit 604 and the operation unit 605 can be configured by a common component such as a touch panel, etc. The operation unit 605 uses, for example, a general-purpose microphone built into the external device 501 to obtain a voice uttered by the user, and can be configured to recognize an operation command from the user using a voice recognition process.

[0084] Various detection information obtained by the biological information detection unit 602 and the shaking detection unit 603 and processed by the control unit 607 is sent to the camera 101 by the communication unit 601. For example, the detection information can be sent to the camera 101 at the timing of detecting a change in the user's heartbeat; or the detection information can be sent at the timing of a change in the movement state (state information) indicating walking movement, running movement, or standing still. In addition, the detection information can be sent at the timing of detecting a preset arm waving action; and the detection information can be sent at the timing of detecting a movement equivalent to a preset distance.

[0085] <Camera Operation Sequence>

[0086] Fig. 7A and 7B : is a flowchart showing an example of the operation processed by the first control unit 223 of the camera 101 according to the present embodiment.

[0087] When the user operates the power button provided on the camera 101, power is supplied from the first power supply unit 210 to the first control unit 223 and each block in the camera 101. Similarly, power is supplied from the second power supply unit 212 to the second control unit 211. Figure 8The operation of the second control unit 211 is described in detail with reference to the flowchart of FIG.

[0088] When power is supplied, Fig. 7A and 7B The process starts. In step S701, the startup condition is loaded. In this embodiment, the following three situations are used as conditions for starting the power supply.

[0089] (1) The power button is manually pressed and the power is turned on;

[0090] (2) A case where a start instruction is sent from an external device (eg, external device 301 ) through external communication (eg, BLE communication) and the power is turned on; and

[0091] (3) A case where the power is turned on in response to an instruction from the second control unit 211 .

[0092] Here, in the case of (3), that is, when the power is turned on in response to an instruction from the second control unit 211, the start condition calculated in the second control unit 211 is loaded; this will be referred to later. Figure 8 The start condition loaded here is used as a single parameter during subject search and automatic imaging, etc., and this will also be described later. Once the start condition is loaded, the sequence proceeds to step S702.

[0093] In step S702, detection signals are loaded from various sensors. One of the sensor signals loaded here is a signal from a sensor that detects oscillations (such as a gyro sensor or an accelerometer in the device oscillation detection unit 209). Another signal is a signal indicating the rotation position of the pitch rotation unit 104 and the pan rotation unit 105, etc. In addition, an audio signal detected by the audio processing unit 214, a detection trigger signal for specific voice recognition, a sound direction detection signal, and a detection signal for environmental information detected by the environmental sensor 226, etc. are other such signals. Once the detection signals are loaded from various sensors in step S702, the sequence enters step S703.

[0094] In step S703, it is detected whether a communication instruction is sent from an external device, and in the case where such a communication instruction is sent, communication is performed with the external device. For example, the following contents are loaded: remote operation from the external device 301 via wireless LAN or BLE; transmission and reception of audio signals, image signals, compressed audio signals, compressed image signals, etc.; operation instructions such as for video recording from the external device 301; transmission of voice command registration data; and transmission and reception of predetermined position detection notifications, place movement notifications, and training data based on GPS position information; etc. In addition, in the case where there is an update of user movement information, arm motion information, and biometric information such as heartbeat, the information is loaded from the external device 501 via BLE. Although the above-mentioned environmental sensor 226 can be built into the camera 101, it can also be built into the external device 301 or the external device 501. In this case, environmental information is loaded via BLE in step S703. Once communication with the external device and loading from the external device are performed in step S703, the sequence enters step S704.

[0095] In step S704, a mode setting judgment is performed, after which the sequence proceeds to step S705. In step S705, it is judged whether the operation mode is set to the low power mode in step S704. In the case where the operation mode is not the automatic camera mode, automatic editing mode, automatic image transfer mode, learning mode, and automatic file deletion mode to be described later, it is judged that the operation mode is the low power mode. In the case where it is judged in step S705 that the operation mode is the low power mode, the sequence proceeds to step S706.

[0096] In step S706, various parameters (jitter detection judgment parameter, voice detection judgment parameter, and elapsed time detection parameter) related to the start trigger judged in the second control unit 211 are communicated to the second control unit 211 (sub-CPU). As a result of learning performed in the learning process to be described later, the values ​​of the various parameters change. Once the processing of step S706 ends, the sequence proceeds to step S707, in which the first control unit 223 (main CPU) is turned off, and the processing ends.

[0097] In the case where it is determined in step S705 that the operation mode is not the low power mode, it is determined whether the mode setting in step S704 is the automatic imaging mode. Here, the processing for determining the mode setting in step S704 will be described. The mode to be subjected to determination is selected from the following modes.

[0098] (1) Automatic camera mode

[0099] <Mode judgment condition>

[0100] Automatic camera mode is set when it is determined that automatic camera mode is to be performed based on various detection information (image, audio, time, oscillation, location, body changes, environmental changes) that have been learned and set, the amount of time that has passed since the transition to automatic camera mode, and past camera information / number of captured images.

[0101] <Processing in mode>

[0102] In the automatic camera mode processing (step S710), the subject is automatically searched by driving the pan, tilt and zoom operations based on various detection information (image, sound, time, vibration, location, body change, environmental change). Then, when it is determined that an image matching the user's preference can be captured, the image is automatically captured.

[0103] (2) Automatic editing mode

[0104] <Mode judgment condition>

[0105] In the event that it is determined that automatic editing should be performed based on the amount of time that has passed since the previous automatic editing and the past captured image information, the automatic editing mode is set.

[0106] <Processing in mode>

[0107] In the automatic editing mode processing (step S712), processing is performed for selecting still images and moving images based on learning, and then automatic editing processing is performed based on the learning to create a wonderful video that aggregates these images into a single moving image based on the image effects and the time of the edited moving images.

[0108] (3) Image transmission mode

[0109] <Mode judgment condition>

[0110] The automatic image transfer mode is set in response to an instruction using a dedicated application in the external device 301 and it is determined that images are to be automatically transferred based on the amount of time that has passed since the previous image transfer and past captured image information.

[0111] <Processing in mode>

[0112] In the automatic image transfer mode processing (step S714), the camera 101 automatically extracts an image assumed to match the user's preference, and automatically transfers the image assumed to match the user's preference to the external device 301. The image matching the user's preference is extracted based on a score for judging the user's preference added to the image as will be described later.

[0113] (4) Learning Mode

[0114] <Mode judgment condition>

[0115] The automatic learning mode is set when it is determined that automatic learning should be performed based on the amount of time that has passed since the previous learning process, the amount of information integrated with the image and the amount of training data that can be used in learning, etc. The mode is also set when an instruction for setting the learning mode is given via communication from the external device 301.

[0116] <Processing in mode>

[0117] In the learning mode processing (step S716), learning based on user preferences is performed using a neural network based on various operation information in the external device 301 (image acquisition information from the camera, information manually edited via a dedicated application, judgment value information input by the user for the image in the camera) and notification of training information from the external device 301. At the same time, learning related to detection (such as personal authentication registration, voice registration, sound scene registration, and general object recognition registration), and learning of the conditions of the above-mentioned low power mode, etc. are performed.

[0118] (5) Automatic file deletion mode

[0119] <Mode judgment condition>

[0120] The automatic file deletion mode is set when it is determined that a file should be automatically deleted based on the amount of time that has passed since the last automatic file deletion and the remaining capacity of the nonvolatile memory 216 in which images are recorded.

[0121] <Processing in mode>

[0122] In the automatic file deletion mode processing (step S718), a file to be automatically deleted is specified from the images in the nonvolatile memory 216 based on the tag information of the image and the date / time when the image was captured, and then the file is deleted.

[0123] The processing performed in the above-mentioned mode will be described in detail later.

[0124] Return to Fig. 7A and Figure 7B If it is determined in step S705 that the operation mode is not the low power mode, the sequence proceeds to step S709, in which it is determined whether the mode setting is the automatic camera mode. If it is determined that the operation mode is the automatic camera mode, the sequence proceeds to step S710, in which the automatic camera mode processing is performed. Once the processing is completed, the sequence returns to step S702 and the processing is repeated. If it is determined in step S709 that the operation mode is not the automatic camera mode, the sequence proceeds to step S711.

[0125] In step S711, it is determined whether the mode setting is the automatic editing mode; if the operating mode is the automatic editing mode, the sequence proceeds to step S712, and the automatic editing mode processing is performed. Once the processing is completed, the sequence returns to step S702, and the processing is repeated. If it is determined in step S711 that the operating mode is not the automatic editing mode, the sequence proceeds to step S713. Note that the automatic editing mode is not directly related to the main concept of the present invention, and therefore will not be described in detail.

[0126] In step S713, it is determined whether the mode setting is the automatic image transfer mode; if the operation mode is the automatic image transfer mode, the sequence proceeds to step S714, and the automatic image transfer mode processing is performed. Once the processing is completed, the sequence returns to step S702, and the processing is repeated. If it is determined in step S713 that the operation mode is not the automatic image transfer mode, the sequence proceeds to step S715. Note that the automatic image transfer mode is not directly related to the main concept of the present invention, and therefore will not be described in detail.

[0127] In step S715, it is determined whether the mode setting is the learning mode; if the operation mode is the learning mode, the sequence proceeds to step S716 and the learning mode processing is performed. Once the processing is completed, the sequence returns to step S702 and the processing is repeated. If it is determined in step S715 that the operation mode is not the learning mode, the sequence proceeds to step S717.

[0128] In step S717, it is determined whether the mode setting is the automatic file deletion mode; if the operation mode is the automatic file deletion mode, the sequence enters step S718 and the automatic file deletion mode processing is performed. Once the processing is completed, the sequence returns to step S702 and the processing is repeated. If it is determined in step S717 that the operation mode is not the automatic file deletion mode, the sequence returns to step S702 and the processing is repeated. Note that the automatic file deletion mode is not directly related to the main concept of the present invention and will not be described in detail.

[0129] Figure 8 : is a flowchart showing an example of the operation processed by the second control unit 211 of the camera 101 according to the present embodiment.

[0130] When the user operates a power button provided on the camera 101, power is supplied from the first power supply unit 210 to the first control unit 223 and each block in the camera 101. Likewise, power is supplied from the second power supply unit 212 to the second control unit 211.

[0131] When power is supplied, the second control unit (sub-CPU) 211 starts up, and Figure 8The process shown starts. In step S801, it is determined whether a predetermined sampling period has passed. The predetermined sampling period is set to 10 ms, for example, and the sequence proceeds to step S802 every 10 ms. If it is determined that the predetermined sampling period has not passed, the second control unit 211 stands by.

[0132] In step S802, training information is loaded. Fig. 7A The information transmitted when communicating information to the second control unit 211 in step S706, and includes, for example, the following information.

[0133] (1) Determination of Detection of Specific Oscillation (to be used in step S804 described later)

[0134] (2) Determination of Detection of Specific Sound (to be used in step S805 described later)

[0135] (3) Determination of the amount of time that has passed (used in step S807 described later)

[0136] Once the training information is loaded in step S802, the sequence proceeds to step S803, in which an oscillation detection value is obtained. The oscillation detection value is an output value from a gyro sensor or an accelerometer or the like of the device oscillation detection unit 209.

[0137] Once the oscillation detection value is obtained in step S803, the sequence proceeds to step S804, in which a process for detecting a preset specific oscillation state is performed. Here, the judgment process is changed according to the training information loaded in step S802. Several examples will be explained.

[0138] <Tap Detection>

[0139] A state (tap state) in which a user taps the camera 101 with his or her fingertip or the like can be detected based on an output value from the accelerometer 107 attached to the camera 101. By passing the output of the three-axis accelerometer 107 through a bandpass filter (BPF) set to a specific frequency range at each predetermined sampling period, a signal range corresponding to the acceleration change caused by the tap can be extracted. A tap is detected based on whether the number of times the acceleration signal obtained after bandpass filtering exceeds a predetermined threshold value TreshA within a predetermined time TimeA is a predetermined number of times CountA. For a double tap, CountA is set to 2, and for a triple tap, CountA is set to 3. Note that TimeA and ThreshA can also be changed according to training information.

[0140] <Oscillation status detection>

[0141] The oscillation state of the camera 101 can be detected based on the output values ​​from the gyro sensor 106 and the accelerometer 107, etc. attached to the camera 101. A high-pass filter (HPF) is used to cut off the high-frequency components of the outputs from the gyro sensor 106 and the accelerometer 107, etc., and a low-pass filter (LPF) is used to cut off the low-frequency components, after which the output is converted into an absolute value. Oscillation is detected based on whether the number of times the calculated absolute value exceeds a predetermined threshold ThreshB in a predetermined time TimeB is greater than or equal to a predetermined number CountB. This makes it possible to judge the state of low oscillation (where, for example, the camera 101 is placed on a table, etc.) and the state of high oscillation (where, the camera 101 has been attached to the user's body as a wearable camera, etc. and the user is walking). It is also possible to detect a fine oscillation state based on the oscillation level by setting multiple judgment thresholds and conditions for judging the number of counts used, etc. Note that TimeB, ThreshB, and CountB can also be changed according to training information.

[0142] The above describes a method for detecting a specific oscillation state by judging the detection value from the oscillation detection sensor. However, it is also possible to detect a pre-registered specific oscillation state using a trained neural network by inputting data sampled by the oscillation detection sensor during a predetermined time into an oscillation state judger using a neural network. In this case, the training information loaded in step S802 is the weight parameter of the neural network.

[0143] Once the process for detecting a specific oscillation state is performed in step S804, the sequence proceeds to step S805, in which the process for detecting a preset specific oscillation state is performed. Here, the detection judgment process is changed according to the training information loaded in step S802. Several examples will be described below.

[0144] <Specific Voice Command Detection>

[0145] Specific voice command detected. In addition to several commands registered in advance, users can also register specific voices as voice commands in the camera.

[0146] <Specific Sound Scene Recognition>

[0147] The sound scene is judged using a network pre-trained by machine learning based on a large amount of audio data. For example, specific scenes such as audience cheering, applause, and talking are detected. The detected scenes are changed through learning.

[0148] <Sound Level Judgment>

[0149] The sound level is detected by determining whether the volume of the audio level exceeds a predetermined volume and continues for a predetermined amount of time. The predetermined amount of time and the predetermined volume, etc., are changed through learning.

[0150] <Sound direction determination>

[0151] The direction of the sound is detected for the sound of a predetermined volume using a plurality of microphones arranged in a plane.

[0152] The determination processing described is performed in the audio processing unit 214, and it is determined in step S805 whether a specific sound is detected using various settings learned in advance.

[0153] Once the process for detecting the specific sound is performed in step S805, the sequence proceeds to step S806, in which it is determined whether the power of the first control unit 223 is turned off. If the first control unit 223 (main CPU) is turned off, the sequence proceeds to step S807, in which a process for detecting the passage of a preset amount of time is performed. Here, the detection judgment process is changed according to the training information loaded in step S802. The training information is in Fig. 7A The information transmitted when the information is communicated to the second control unit 211 in step S706 of the embodiment. The amount of time that has passed since the first control unit 223 was turned from on to off is measured; if the amount of time is greater than or equal to the predetermined time TimeC, it is determined that the amount of time has passed, and if the amount of time is less than TimeC, it is determined that the amount of time has not passed. TimeC is a parameter that changes according to the training information.

[0154] Once the process for detecting the amount of time that has elapsed is performed in step S807, the sequence proceeds to step S808, where it is determined whether a condition for canceling the low power mode is satisfied. Whether to cancel the low power mode is determined according to the following conditions.

[0155] (1) Whether specific oscillation is detected

[0156] (2) Whether a specific sound is detected

[0157] (3) Has the predetermined amount of time passed?

[0158] Regarding (1), it is determined whether a specific oscillation is detected by a specific oscillation state detection process performed in step S804. Regarding (2), it is determined whether a specific sound is detected by a specific sound detection process performed in step S805. Regarding (3), it is determined whether a predetermined amount of time has passed by a process for detecting the passage of an amount of time performed in step S807. If at least one of (1) to (3) is satisfied, it is determined that the low power mode is canceled.

[0159] Once it is determined in step S808 that the low power mode is canceled, the sequence proceeds to step S809, in which the power of the first control unit 223 is turned on; then, in step S810, the condition (oscillation, sound, or time) for determining that the low power mode is canceled is communicated to the first control unit 223. Then, the sequence returns to step S801, and the processing loops. If no condition is satisfied in step S808 and it is determined that there is no condition for canceling the low power mode, the sequence returns to step S801, and the processing loops.

[0160] On the other hand, if it is determined in step S806 that the first control unit 223 is turned on, the sequence proceeds to step S811, in which the information obtained in steps S803 to S805 is communicated to the first control unit 223; then, the sequence returns to step S801, and the processing loops.

[0161] In the present embodiment, the structure is as follows: even when the first control unit 223 is turned on, the second control unit 211 performs oscillation detection and specific sound detection, etc., and communicates the detection result to the first control unit 223. However, the structure may be as follows: when the first control unit 223 is turned on, the processing of steps S803 to S805 is not performed, and the oscillation detection and specific sound detection, etc. are performed by the processing in the first control unit 223 ( Fig. 7A Step S702).

[0162] As mentioned above, by executing Fig. 7A The processing of steps S704 to S707 and Figure 8 The conditions for transitioning to the low power mode and the conditions for canceling the low power mode are learned based on the user operation, such as the processing of the camera 101. This makes it possible to perform camera operations that are more user-friendly for the user who owns the camera 101. The method used for learning will be described later.

[0163] Although the method for canceling the low power mode in response to oscillation detection, sound detection, or time lapse is described in detail above, the low power mode may be canceled based on environmental information. Environmental information may be determined based on whether the absolute amount or change in the temperature, air pressure, brightness, humidity, and ultraviolet light amount exceeds a predetermined threshold, and the threshold may also be changed by learning as will be described later.

[0164] In addition, detection information related to oscillation detection, sound detection or time lapse, and absolute values ​​or changes in various environmental information can be determined based on the neural network and used to determine whether to cancel the low power mode. The judgment conditions used in the judgment process can be changed through learning as described later.

[0165] <Automatic camera mode processing>

[0166] Will refer to Fig. 9 The automatic imaging mode processing will be described below. First, in step S901, the image processing unit 207 performs image processing on the signal obtained from the imaging unit 206, and generates an image for subject detection. The generated image is subjected to subject detection processing for detecting a person or an object, etc.

[0167] In the case of detecting a person, the face or body of the subject is detected. In the face detection process, a pattern for determining the face of a person is set in advance, and the place in the captured image that matches the pattern can be detected as the facial area of ​​the person. In addition, a reliability level indicating the degree of certainty that the subject is a face is calculated at the same time. The reliability level is calculated based on, for example, the size of the facial area in the image or the degree to which the area matches the facial pattern. The same applies to object recognition that identifies an object that matches a pre-registered pattern.

[0168] There is also a method of extracting feature subjects using a histogram of hue or saturation, etc. within a captured image. For an image of a subject that appears within the shooting angle of view, a distribution is derived from a histogram of hue or saturation, etc., and the distribution is divided into multiple intervals; then a process for classifying the captured image for each of these intervals is performed. For example, histograms are created for multiple color components of a captured image, and then these histograms are divided into distribution ranges corresponding to peaks; then, the captured image area is identified by classifying the captured image according to areas belonging to the same interval combination. An evaluation value is calculated for each identified subject image area, and the subject image area with the highest evaluation value can be determined as the main subject area. The above method can be used to obtain various subject information from the camera information.

[0169] In step S902, the image blur correction amount is calculated. Specifically, first, the absolute angle of the camera's oscillation is calculated based on the angular velocity and acceleration information obtained by the device oscillation detection unit 209. Then, an angle for correcting the image blur by moving the pitch rotation unit 104 and the pan rotation unit 105 in an angle direction that cancels the absolute angle is obtained, and the angle is regarded as the image blur correction amount. Note that the calculation method used in the image blur correction amount calculation process described here can be changed by the learning process described later.

[0170] In step S903, the (holding) state of the camera is determined. The current oscillation / movement state of the camera is determined based on the camera angle and the camera movement amount detected based on the angular velocity information, acceleration information, GPS position information, etc. For example, in the case of capturing an image in a state where the camera 101 is mounted on a vehicle, subject information such as surrounding scenery will change significantly depending on the distance traveled. Therefore, it is determined whether the state is a "vehicle moving state" in which the camera is mounted on a vehicle, etc. and is moving at a high speed, and is used in the automatic subject search to be described later. It is also determined whether the camera angle is changing significantly to determine whether the state is a "stationary shooting state" in which the camera 101 undergoes little oscillation. In the stationary shooting state, it can be assumed that the position of the camera 101 itself will not change, and thus a subject search for stationary shooting can be performed. In the case where the camera angle undergoes a relatively large change, the state can be determined as a "handheld state", and a subject search for the handheld state can be performed.

[0171] In step S904, a subject search process is executed. The subject search is composed of the following processes.

[0172] (1) Region segmentation

[0173] (2) Calculate the importance level of each area

[0174] (3) Determine the search area

[0175] These processes will be described below in sequence.

[0176] (1) Region segmentation

[0177] Reference FIG. 10A to FIG. 10E To illustrate the region segmentation. Fig. 10A As shown, the entire periphery is divided into regions using the position of the camera (when the camera position is represented by the origin O) as the center. Fig. 10A In the example shown, the division is performed every 22.5 degrees in both the pitch direction and the pan direction. Fig. 10A When segmentation is performed as shown, as the angle in the pitch direction moves away from 0 degrees, the circle in the horizontal direction becomes smaller, and thus the area becomes smaller. Fig. 10B As shown, when the pitch angle is greater than or equal to 45 degrees, the range of the area in the horizontal direction is set to be greater than 22.5 degrees.

[0178] Fig. 10C and Fig. 10D 1301 shows an example of regions obtained by region segmentation within the shooting angle of view. Axis 1301 represents the orientation of the camera 101 in the initial state, and the region segmentation is performed using this direction as a reference position. 1302 represents the angle of view region of the captured image, and Fig. 10D An example of the image obtained at this time is shown. Based on image segmentation, the image within the shooting angle of view is segmented into Fig. 10D The images are represented by reference numerals 1303 to 1318.

[0179] (2) Calculate the importance level of each area

[0180] For each area obtained by the above segmentation, an importance level representing the priority ranking of the search is calculated based on the condition of the subject existing in the area and the condition of the scene, etc. The importance level based on the condition of the subject is calculated based on, for example, the number of people existing in the area, the size of each person's face, the direction of the face, the certainty of face detection, the expression of the person, and the result of personal authentication of the person, etc. In addition, the importance level based on the condition of the scene is calculated based on, for example, the result of general object recognition, the result of scene judgment (blue sky, backlight or night scene, etc.), the level of sound from the direction of the area, the result of voice recognition, and the motion detection information from within the area, etc.

[0181] In addition, Fig. 9 In the case where camera oscillation is detected in the camera state judgment (step S903) shown, the importance level can also be changed according to the oscillation state. For example, in the case where the "stationary shooting state" is judged, it can be judged that a subject search focusing on a subject registered for facial authentication and having a high priority level (for example, the owner of the camera) is performed. For example, automatic photography to be described later can also be performed by giving priority to the face of the owner of the camera. As a result, even in the case where the owner of the camera often takes images while walking with the camera attached to him or her, the owner can obtain many images in which he or she appears by removing the camera and placing the camera on a table or the like. At this time, a face search can be performed by panning and tilting, so that images in which the owner appears and group photos showing many faces can be obtained simply by placing the camera as desired without paying special attention to the placement angle of the camera, etc.

[0182] Note that only under the above conditions, as long as there is no change in each area, the same area will have the highest importance level, so the searched area will remain the same indefinitely. Therefore, the importance level is changed according to the past camera information. Specifically, the importance level of an area that is continuously designated as a search area during a predetermined amount of time may be lowered, or the importance level of an area whose image is captured in step S910 to be described later may be lowered during a predetermined amount of time, and so on.

[0183] Furthermore, when the camera is moving (such as when the owner of the camera wears the camera on his or her body, or when the camera is attached to a vehicle, etc.), there is a case where even if a subject in the surroundings is searched for by panning and tilting, the subject is no longer visible when the image is captured. There is also a case where the subject has left and is too small, which makes the subject search meaningless. Therefore, the moving direction and moving speed of the camera are calculated based on the angular velocity information, acceleration information, and GPS position information of the camera detected in step S903, and in addition, based on the motion vector calculated for each coordinate from the captured image. Based on these, it can be assumed from the beginning that an area far from the traveling direction does not have a subject, or conversely, the search time interval can be changed according to the moving speed (such as by shortening the subject search time interval during high-speed movement, etc.) to ensure that important subjects are not lost.

[0184] Specifically, reference will be made to Fig. 10E To illustrate the state where the camera is hung from the neck. Fig. 10E 1320 denotes a person (owner of the camera), 1321 denotes a camera, and 1322 and 1323 denote subject search ranges, respectively. The subject search range is set, for example, to an angle range that is substantially left-right symmetrical with respect to a traveling direction 1324 of the camera. The subject search range 1322 denotes a subject search range in a state where the person is completely stopped. In order to limit the viewing angle to a viewing angle where the body and clothing of the owner of the camera do not occupy more than the set area, a 360-degree search is not performed to prevent the body of the camera owner from appearing in the image.

[0185] The subject search range 1323 indicates the search range when the person is moving in the direction shown in the figure (travel direction 1324). Thus, by changing the subject search range according to the moving speed (for example, by narrowing the range when the moving speed is high and expanding the range when the moving speed is low), it is possible to perform a wasteful subject search in an adaptive manner. Fig. 10E The subject search range is shown as changing only in the horizontal direction, but the processing can also be performed in the same manner in the vertical direction. In addition, the "travel direction" described here is calculated based on the measurement results obtained by the angular velocity meter 106 and the accelerometer 107 during the set time period. This makes it possible to prevent the search range from changing frequently even when the movement is unstable.

[0186] Furthermore, in order to prevent the subject search range from becoming uncertain due to a sudden change in the travel direction, the sensitivity may be reduced by taking the past travel direction into consideration. Fig. 10EThe case where the camera is hung from the neck is shown, but if it can be determined that the camera is placed on a table, the processing for changing the subject search range according to the moving speed can be omitted. In this way, the subject search processing can be changed according to the change of the state of the camera (such as whether the camera is in a handheld state, hanging from the neck, in a wearable state, placed on a table, or attached to a moving object). Changing the subject search range according to the movement information eliminates waste of subject search and also helps reduce battery power consumption.

[0187] (3) Determine the search area

[0188] Once the importance level is calculated for each area as described above, the area with a high importance level is set as the search target area. Then, the pan / tilt search target angle required to capture the search target area within the angle of view is calculated.

[0189] Return to Fig. 9 In step S905, the pan / tilt drive is performed. Specifically, the pan / tilt drive amount is calculated by adding the image blur correction amount at the control sampling frequency to the drive angle based on the pan / tilt search object angle. Then, the drive of the tilt rotation unit 104 and the pan rotation unit 105 is controlled by the lens barrel rotation drive unit 205.

[0190] In step S906, zoom drive is performed by controlling the zoom unit 201. Specifically, zoom drive is performed according to the state of the search object subject determined in step S904. For example, in the case where the search object subject is the face of a person, if the face is too small in the image, the face may be below the minimum size required for detection, which makes it impossible to detect the face; there is a risk that the face will be missed as a result. In this case, control is performed to increase the size of the face in the image by zooming toward the telephoto side. On the other hand, if the face is too large in the image, the subject is more likely to move out of the angle of view due to the movement of the subject and the camera itself, etc. In this case, control is performed to reduce the size of the face in the image by zooming toward the wide-angle side. Controlling the zoom in this way makes it possible to maintain a state suitable for tracking the subject.

[0191] In step S907, it is determined whether a manual imaging instruction is performed, and if a manual imaging instruction is performed, the sequence proceeds to step S910. At this time, the manual imaging instruction may be pressing a shutter button, tapping (tap) the camera housing with a fingertip or the like, inputting a voice instruction, or an instruction from an external device or the like. The imaging instruction using the tapping operation as a trigger is determined by detecting high-frequency acceleration that lasts for a short period of time using the device oscillation detection unit 209 when the user taps the camera housing. The voice command input is an imaging instruction method that uses the audio processing unit 214 to recognize the voice when the user issues a predetermined phrase (e.g., "take a photo" or the like) that instructs the taking of an image and uses the voice as a trigger for taking an image. The use of an instruction from an external device is an imaging instruction method that uses a shutter instruction signal sent from, for example, a dedicated application, which is connected to the camera via Bluetooth or the like.

[0192] If there is no manual image capture instruction in step S907, the sequence proceeds to step S908, in which automatic image capture determination is performed. In the automatic image capture determination, determinations are made as to whether to perform automatic image capture and as to the shooting method (whether to shoot still images, shoot moving images, perform continuous shooting, or perform panoramic shooting, etc.).

[0193] <Determine whether to execute automatic recording>

[0194] The judgment on whether to perform automatic photography is performed as follows. Specifically, the judgment to perform automatic photography is performed in the following two cases. In one case, when the importance level obtained based on the importance level obtained for each area in step S904 is greater than a predetermined value, it is judged that automatic photography is performed. In another case, the judgment is based on a neural network.

[0195] Fig.11 An example of a network composed of a multilayer perceptron is shown as an example of a neural network. A neural network is used to predict output values ​​from input values, and by training the network in advance using input values ​​and output values ​​used as a model for these inputs, an output value that conforms to the learned model can be estimated for new input values. Note that the learning method will be described later. Fig.11, 1201 and the circles arranged vertically below it represent neurons of the input layer; 1203 and the circles arranged vertically below it represent neurons of the middle layer; and 1204 represent neurons of the output layer. Arrows such as the arrow represented by 1202 represent connections between neurons. In the judgment based on the neural network, the subject appearing in the current viewing angle, or the feature quantity based on the scene or camera state, etc. are supplied as input to the neurons of the input layer, and the value output from the output layer is obtained after performing calculation based on the forward propagation of the multilayer perceptron. If the output value is greater than or equal to the threshold value, it is judged that automatic photography is performed. Note that the following are used as features of the subject: current zoom ratio; common object recognition result at the current viewing angle; face detection result; the number of faces appearing in the current viewing angle; the degree to which the face is smiling; the degree to which the eyes are closed; the angle of the face; the face authentication ID number; the angle of sight of the person used as the subject; the result of scene judgment; the amount of time that has passed since the last image capture; the current time, GPS position information, and the amount of change from the last image capture position; the current audio level; the person using his or her voice; whether people are clapping or cheering; vibration information (acceleration information, camera status); environmental information (temperature, air pressure, illumination, humidity, and ultraviolet light amount); and the like. In addition, in the case where information is communicated from the external device 501, the communicated information (user movement information, arm motion information, and biological information such as heartbeat, etc.) is also used as a feature. The feature is converted into a numerical value within a predetermined range and is supplied to the neurons of the input layer as a feature quantity. Therefore, the neurons of the input layer need to use an equal number of feature quantities.

[0196] Note that in judgment based on a neural network, the output value can be changed by changing the weights of connections between neurons using a learning process to be described later, and then the result of the judgment can be applied to the learning result.

[0197] In addition, according to Fig. 7A The judgment of automatic video recording is changed according to the activation condition of the first control unit 223 loaded in step S702. For example, in the case where the unit is activated in response to detecting a tap or a specific voice command, it is very likely that the operation indicates that the user currently wants to capture an image. Therefore, settings are made to increase the frequency of video recording.

[0198] <Determining the Camera Method>

[0199] When judging the shooting method, it is judged whether to shoot a still image, shoot a moving image, perform continuous shooting, or shoot a panoramic image, etc. based on the camera state detected in steps S901 to S904 and the state of the surrounding subjects. For example, a still image is shot when the subject (person) is stationary, and a moving image or a continuous image is shot when the subject is moving. In addition, when there are multiple subjects around the camera, or when the place is judged to be a scenic spot based on the above-mentioned GPS information, a panoramic image shooting process can be performed, which generates a panoramic image by synthesizing the images captured in sequence while performing a pan / tilt operation. As with the judgment method used in "Judging whether to perform automatic shooting", various information detected before shooting can be judged based on a neural network, and then the shooting method can be set. The judgment conditions used in this judgment process can be changed by the learning process to be described later.

[0200] Return to Fig. 9 As explained above, if in step S909, the automatic imaging determination performed in step S908 is determined to be performed for automatic imaging, the sequence proceeds to step S910; however, if it is not determined to be performed for automatic imaging, the automatic imaging mode processing ends.

[0201] In step S910, automatic imaging is started. At this time, imaging is started using the imaging method determined in step S908. At this time, automatic focus control is performed using the focus drive control unit 204. In addition, exposure control is performed using an aperture control unit, a sensor gain control unit, a shutter control unit, etc. (not shown) so that the subject is photographed with appropriate brightness. In addition, after imaging, the image processing unit 207 performs various known image processing such as white balance processing, noise reduction processing, and gamma correction processing to generate an image.

[0202] Note that during this image capture, if a predetermined condition is satisfied, the person whose image is to be captured by the camera may be notified of this before the image is captured. For example, as a method for making such a notification, sound may be emitted from the audio output unit 218, and the LED may be lit using the LED control unit 224. The predetermined conditions are, for example: the number of faces in the current viewing angle; the degree to which the face is smiling; the degree to which the eyes are closed; the angle of sight or the angle of the face of the person serving as the subject; the face authentication ID number; the number of persons registered for personal authentication; the result of ordinary object recognition at the time of image capture; the result of scene judgment; the amount of time that has passed since the previous image capture; the time of image capture; whether the current position based on GPS information is a scenic spot; the audio level at the time of image capture; whether there is a person making a sound; whether there is applause or cheering; vibration information (acceleration information, camera status); and environmental information (temperature, air pressure, illuminance, humidity, amount of ultraviolet light); and the like. By performing image capture based on these conditions for notification, an image in which a person is viewing the camera in a favorable manner can be obtained in a scene with high importance.

[0203] Also regarding such notification before shooting, information of the captured image or various information detected before shooting can be determined based on a neural network, and then a notification method and timing can be set. The determination conditions used in this determination process can be changed by a learning process to be described later.

[0204] In step S911, an editing process for processing the image generated in step S910 and adding a moving image is performed. "Processed image" specifically refers to: cropping based on the face and focus position of a person, etc.; image rotation processing; HDR (high dynamic range) effect processing; bokeh effect processing; and color conversion filter effect processing; etc. In image processing, a plurality of processed images generated by a combination of the above-mentioned processes can be generated based on the image generated in step S910, and these processed images can be stored separately from the image generated in step S910. Regarding moving image processing, the following process can be performed, which is used to add a captured moving image or still image while giving special effect processing such as slide, zoom, and fade in and out to the generated edited moving image. Through this editing in step S911, the information of the captured image or various information detected before shooting can be judged based on a neural network, and then the image processing method can be set. The judgment condition used for this judgment process can be changed by the learning process to be described later.

[0205] In step S912, captured image training information generation processing is performed. Here, information used in the learning processing to be described later is generated and recorded. Specifically, the following information of the current captured image is used: zoom ratio at the time of shooting; general object recognition result at the time of shooting; face detection result; the number of faces appearing in the captured image; the degree to which the face is smiling; the degree to which the eyes are closed; the face angle; the face authentication ID number; the angle of the line of sight of the person who is the subject; the scene judgment result; the amount of time that has passed since the last shooting; the shooting time; GPS position information and the amount of change from the last shooting position; the audio level at the time of shooting; the person using his or her voice; whether people are clapping or cheering; vibration information (acceleration information, camera status); environmental information (temperature, air pressure, illumination, humidity, ultraviolet light amount); motion image shooting time; and whether the shooting instruction was manually performed; etc. In addition, a score can also be calculated, which is a neural network output that represents the user's image preference as a numerical value. This information is generated and recorded as label information in the captured image file. Alternatively, the information may be written to the nonvolatile memory 216, or the information of each captured image may be stored in the recording medium 221 in a list format known as "bibliographic data".

[0206] In step S913, the past camera information is updated. Specifically, regarding the number of images captured for each area as described in step S908, the number of images captured for each person who has undergone personal authentication registration, the number of images captured for each subject identified in general object recognition, and the number of images captured for each scene in scene judgment, the count of the number of images captured at this time is increased by 1.

[0207] <Learning Process>

[0208] Next, the learning based on the user's preference according to the present embodiment will be described. In the present embodiment, the learning processing unit 219 uses a method such as Fig.11 The neural network shown in the figure and the machine learning algorithm are used to perform learning based on user preferences. The neural network is used to predict the output value based on the input value, and by training the network in advance with the actual value of the input value and the actual value of the output value, the output value can be estimated for the new input value. By using the neural network, learning based on user preferences is performed for the above-mentioned automatic camera, automatic editing, and subject search. In addition, the following operations are also performed: using learning to change the registration of subject information (results of facial authentication, and general object recognition, etc.) used as feature data for input to the neural network, controlling camera notification, controlling low power mode, and automatically deleting files, etc.

[0209] In the present embodiment, the operation to which the learning process is applied is the following operation.

[0210] (1) Automatic camera

[0211] (2) Automatic editing

[0212] (3) Subject Search

[0213] (4) Subject Registration

[0214] (5) Camera notification control

[0215] (6) Low power mode control

[0216] (7) Automatic file deletion

[0217] (8) Image blur correction

[0218] (9) Automatic image transmission

[0219] Among the above-mentioned operations to which the learning process is applied, automatic editing, automatic file deletion, and automatic image transfer have no direct relation to the main concept of the present invention and thus will not be described.

[0220] <Automatic Camera>

[0221] Here, learning for automatic photography is described. In automatic photography, learning for automatically photographing an image that matches the user's preference is performed. Fig. 9 As described in the flowchart, after the image is captured (after step S910), training information generation processing (step S912) is performed. The image to be learned is selected by the method to be described later, and based on the training information included in the image, the neural network is trained by changing the weight of the neural network.

[0222] The training is performed by changing the neural network for determining the timing of automatic imaging, and changing the neural network for determining the imaging method (photographing a still image, photographing a moving image, continuous shooting, panoramic image shooting, etc.).

[0223] <Subject Search>

[0224] Here, learning for subject search will be described. In subject search, learning is performed to automatically search for a subject that matches the user's preference. Fig. 9As described in the flowchart, in the subject search process (step S904), the subject search is performed by calculating the importance level of each area and then performing pan, tilt and zoom drive. Learning is performed based on the captured image and the detection information obtained during the search, and the result is obtained as the learning result by changing the weight of the neural network. Various detection information is input into the neural network during the search operation, and a subject search reflecting the learning is performed by judging the importance level. In addition to calculating the importance level, for example, the pan / tilt search method (speed and frequency of movement) is controlled, and the subject search area is controlled according to the movement speed of the camera, etc. In addition, the optimal subject search is performed by setting different neural networks for each of the above-mentioned camera states and applying a neural network suitable for the current camera state.

[0225] <Subject Registration>

[0226] Learning for subject registration will be described here. In subject registration, learning is performed for automatically registering subjects according to user preferences and ranking the subjects. For example, as learning, facial authentication registration, registration of general object recognition, registration of gesture and voice recognition, and sound-based scene recognition are performed. Authentication registration is performed on people and objects, and then these people and objects are ranked based on the number and frequency of obtaining images, the number and frequency of manually taking images, and the frequency of the subject appearing in the search. The registered information is registered as an input for judgment using the corresponding neural network.

[0227] <Camera notification control>

[0228] Here we will explain the learning of camera notification. Fig. 9 As described in step S910, immediately before capturing an image, if a predetermined condition is satisfied, a notification indicating that an image will be captured is provided to the person to be captured by the camera, and then the image is captured. For example, the subject's line of sight may be visually guided by a pan / tilt drive operation, or the subject's attention may be attracted by using a speaker sound emitted by the audio output unit 218, or by using an LED control unit 224 to emit light from an LED, etc. Whether to use the subject's detection information in learning is determined based on whether detection information (e.g., the degree of a smile, whether a person is looking at the camera, or a gesture) is obtained immediately after the above notification, and training is performed by changing the weights in the neural network.

[0229] Various detection information immediately before the image is taken is input into the neural network, and then judgments are made regarding whether to notify, various operations (sound (sound level / sound type / timing), light (luminous time, speed), camera direction (pan / tilt movement), etc.) are made.

[0230] <Low Power Mode Control>

[0231] As reference Fig. 7A , Figure 7B and Figure 8 However, conditions for canceling the low power mode and conditions for transitioning to the low power state are also learned. Here, the conditions for canceling the low power mode will be described.

[0232] <Tap Detection>

[0233] As described above, the predetermined time TimeA and the predetermined threshold ThreshA are changed by learning. Preliminary tap detection is performed even in the state where the threshold of tap detection has been lowered, and the parameters of TimeA, ThreshA, etc. are set to make detection easier depending on whether the preliminary tap detection is determined before the tap is detected. In addition, after the tap is detected, if it is determined that the tap is not a start trigger based on the camera detection information, the parameters of TimeA, ThreshA, etc. are set to make the tap detection more difficult.

[0234] <Oscillation status detection>

[0235] As described above, the predetermined time TimeB, the predetermined threshold ThreshB, the predetermined number of times CountB, etc. are changed by learning. In the case where the oscillation state judgment result corresponds to the start condition, the start is performed; however, in the case where the result is judged not to be a start trigger in the predetermined amount of time after the start according to the camera detection information, the learning is performed so that the start is more difficult to occur in response to the oscillation state judgment. In addition, in the case where it is judged that the shooting frequency is high in the state of high oscillation, the start is set to be more difficult to occur in response to the oscillation state judgment.

[0236] <Sound Detection>

[0237] For example, learning can be performed by the user manually setting a specific voice, a specific sound scene to be detected, or a specific sound level, etc. via communication using a dedicated application in the external device 301. In addition, learning can also be performed by setting a plurality of detection methods in advance in the audio processing unit, so that an image to be learned is selected by the method described later, audio information before and after the image is learned, and a sound to be judged as a start trigger (a specific voice command, and a sound scene such as cheering or applause, etc.) is set.

[0238] <Environmental information detection>

[0239] For example, learning may be performed by manually setting a change in environmental information to be used as a start-up condition by the user via communication using a dedicated application in the external device 301. For example, start-up may be performed under specific conditions such as the absolute amount or change amount of temperature, air pressure, brightness, humidity, or ultraviolet light amount. Judgment thresholds based on various environmental information may also be learned. If, after start-up performed in response to environmental information, it is determined that the environmental information is not a start-up trigger based on camera detection information, parameters of various judgment thresholds are set to make detection of environmental changes more difficult.

[0240] In addition, the above parameters change according to the remaining battery power. For example, when the remaining battery power is low, it becomes more difficult to make various judgments, and when the remaining battery power is high, it becomes easier to make various judgments. Specifically, there is a case where even in the case of an oscillation state detection result and a sound scene detection result that are not necessarily a trigger for the user to start the camera, when the remaining battery power is high, it is determined that the camera is started.

[0241] In addition, the conditions for canceling the low power mode can be determined based on the neural network according to information such as oscillation detection, sound detection, elapsed time detection, various environmental information, and remaining battery power. In this case, an image to be learned is selected by a method to be described later, and the neural network is trained by changing the weight of the neural network based on the training information included in the image.

[0242] Next, the learning of the conditions for transitioning to the low power state will be described. Fig. 7A As shown, if the mode setting judgment performed in step S704 indicates that the operation mode is not the automatic camera mode, the automatic editing mode, the automatic image transfer mode, the learning mode, and the automatic file deletion mode, the camera enters the low power mode. The conditions for judging each mode are as described above, and the conditions based on which each mode is judged also change in response to learning.

[0243] <Automatic Camera Mode>

[0244] As described above, the importance level is determined for each area, and automatic video recording is performed while using pan / tilt to search for a subject; however, if it is determined that there is no subject to be photographed, the automatic video recording mode is canceled. For example, in the case where the importance level of all areas or a value obtained by adding the importance levels of these areas together becomes less than or equal to a predetermined threshold, the automatic video recording mode is canceled. At this time, as time passes after the transition to the automatic video recording mode, the predetermined threshold also decreases. As more time passes after the transition to the automatic video recording mode, it is easier to transition to the low power mode.

[0245] Low power mode control taking into account battery life can be performed by changing a predetermined threshold value according to the remaining battery power. For example, when the remaining power is small, the threshold value is increased to make it easier to switch to low power mode, and when the remaining power is large, the threshold value is reduced to make it more difficult to switch to low power mode. Here, the parameters (elapsed time threshold TimeC) of the conditions for canceling the low power mode next time are set for the second control unit 211 (sub-CPU) according to the amount of time that has passed since the last switch to the automatic camera mode and the number of images captured. The above threshold values ​​change as a result of learning. For example, learning is performed by manually setting the camera frequency and the startup frequency, etc. via communication using a dedicated application of the external device 301.

[0246] A structure may be adopted in which each parameter is learned by accumulating the average value and distribution data of each time period thereof for the time elapsed from turning on the power button of the camera 101 until the power button is turned off. In this case, learning is performed so that the return from the low power mode and the transition to the low power state, etc. occur at shorter time intervals for a user whose time from turning on the power until the power is turned off is shorter, and the time interval is longer for a user whose time from turning on the power until the power is turned off is longer.

[0247] Learning is also performed based on the detection information during the search. Learning is performed so that when it is determined that there are many subjects that have been set as important through learning, the return from the low power mode and the transition to the low power state occur at shorter time intervals, and when there are fewer important subjects, the time intervals are longer.

[0248] <Image Blur Correction>

[0249] Here we will explain the learning for image blur correction. Fig. 9 The image blur correction is performed by calculating the correction amount in step S902 and then performing the pan / tilt drive operation in step S905 based on the correction amount. In the image blur correction, learning for correction according to the characteristics of the user's oscillation is performed. The direction and size of the blur can be estimated by using, for example, a PSF (point spread function) for the captured image. Fig. 9 In the learning information generation performed in step S912, the estimated blur direction and size are added to the image as information.

[0250] exist Figure 7BIn the learning mode processing performed in step S716 of , the estimated direction and size of blur are used as outputs, and various detection information when the image is captured (motion vector information of the image a predetermined amount of time before the image is captured, movement information of the detected subject (person or object, etc.), oscillation information (gyro sensor output, acceleration output, camera status)) is used as input to train the weights of the neural network used for image blur correction. It is also possible to make a judgment by adding other information to the input, such as environmental information (temperature, air pressure, illumination, and humidity), sound information (sound scene judgment, specific audio detection, sound level change), time information (time elapsed since startup, time elapsed since the last image was captured), and location information (GPS location information, amount of change in location movement), etc.

[0251] When calculating the image blur correction amount in step S902, the size of blur when the image is captured at the moment can be estimated by inputting the above-mentioned various detection information into the neural network. When the size of blur is estimated to be high, control for increasing the shutter speed or the like can be performed. In addition, a method can also be used in which, when the size of blur is estimated to be high, the image becomes blurred and the capturing is prohibited.

[0252] Since there is a limit on the pan / tilt drive angle, additional correction cannot be performed once the end of the drive range is reached; however, the range required for the pan / tilt drive for correcting blur in the image being exposed can be estimated by estimating the size and direction of the blur when the image is captured. In the case where there is no margin in the range of motion during exposure, a larger amount of blur can be suppressed by increasing the cutoff frequency of the filter used to calculate the image blur correction amount so that the range of motion is not exceeded. In the case where the range of motion seems to be exceeded, exposure is started after first turning the pan / tilt angle in the direction opposite to the direction in which the range of motion is to be exceeded, which makes it possible to ensure the range of motion and capture a blur-free image. Therefore, it is possible to learn image blur correction that conforms to the characteristics of the user when capturing an image, how the user uses the camera, etc., so that the captured image can be prevented from becoming blurred.

[0253] In addition, in the above-mentioned "imaging method judgment", it can be judged whether panning is performed, in which the moving subject is not blurred but the stationary background appears blurred due to the movement. In this case, subject blur correction can be performed by estimating the pan / tilt drive speed for shooting a blur-free subject based on the detection information obtained until the image is captured. At this time, the drive speed can be estimated by inputting the above-mentioned various detection information into the trained neural network. Learning is performed by dividing the image into blocks, estimating the PSF of each block, estimating the direction and size of the blur in the block where the main subject is located, and then performing learning based on the information.

[0254] It is also possible to learn the amount of blur in the background from information of an image selected by the user. In this case, the size of blur is estimated in a block where the main subject is not located, and the user's preference can be learned based on this information. By setting the shutter speed during imaging based on the learned preferred amount of blur in the background, imaging that provides the user's desired panning effect can be automatically performed.

[0255] Next, the learning method will be described. "Learning within the camera" and "learning by association with a communication device" can be given as the learning method.

[0256] The following will describe a method for in-camera learning. In this embodiment, the following method for in-camera learning is given.

[0257] (1) Learning from detection information during manual shooting

[0258] (2) Learning based on the detection information when searching for the subject

[0259] <Learning from Detection Information During Manual Image Recording>

[0260] As reference Fig. 9 As described in steps S907 to S913 of the embodiment, in the present embodiment, the camera 101 can capture images in two ways (i.e., by manual capture and automatic capture). When a manual capture instruction is given in step S907, information indicating that the image is captured manually is added to the captured image in step S912. If the image is captured in a state where the automatic capture is determined to be on in step S909, information indicating that the image is captured automatically is added to the captured image in step S912.

[0261] Here, when manually shooting an image, it is likely that the image is shot based on the user's preferred subject, preferred scene, preferred place, and time interval. Therefore, learning is performed based on various feature data obtained during manual shooting and training information of the shot image, etc. Learning is also performed based on the detection information obtained during manual shooting for extraction of feature quantities in the shot image, personal authentication registration, registration of expressions of individual persons, and registration of combinations of persons, etc. In addition, learning is performed so that the importance of nearby people and objects, etc. is changed based on the detection information obtained during the subject search (for example, based on the expression of the subject registered as an individual). In addition, different training data and neural networks can be set for each of the above-mentioned camera states, and training data consistent with the state of the camera when the image is shot can be added.

[0262] <Learning from detection information when searching for a subject>

[0263] During the subject search operation, it is determined with respect to the subject registered for personal authentication which person, object, and scene the subject appears at the same time, and the time ratio of the subject appearing at the same time in the viewing angle is calculated. For example, the time ratio of person A, who is a subject registered for personal authentication, and person B, who is also a subject registered for personal authentication, appearing at the same time is calculated. Various detection information is saved as learning data so that when person A and person B are within the same viewing angle, the score for determining automatic photography is increased, and then learning is performed through learning mode processing (step S716).

[0264] As another example, the ratio of the time when person A who has performed personal authentication registration and the subject "cat" determined by general object recognition appear at the same time is calculated. Various detection information is saved as learning data so that when person A and the cat are within the same viewing angle, the score for automatic camera determination increases, and then learning is performed through learning mode processing (step S716).

[0265] In addition, in the case where a high degree of smile or an expression indicating "joy" or "surprise" is detected for person A as a subject for which personal authentication registration is performed, the subject appearing at the same time is learned as important. Alternatively, in the case where an expression indicating "anger" or "seriousness" is detected, the subject appearing at the same time is unlikely to be important, and the processing may be performed so that learning is not performed.

[0266] Next, learning by association with an external device according to the present embodiment will be described. According to the present embodiment, the following method can be given as a method for learning by association with an external device.

[0267] (1) Learning by obtaining images from external devices

[0268] (2) Learning by inputting judgment values ​​for images via an external device

[0269] (3) Learning by analyzing images stored in external devices

[0270] (4) Learning from information uploaded to the SNS server by an external device

[0271] (5) Learning by changing camera parameters using external devices

[0272] (6) Learning based on information obtained by manually editing images in an external device

[0273] <Learning by obtaining images from external devices>

[0274] As reference Figure 3As described above, the camera 101 and the external device 301 have communication means for performing a first communication 302 and a second communication 303. The first communication 302 is mainly used to send and receive images, and images in the camera 101 can be sent to the external device 301 via a dedicated application in the external device 301. In addition, thumbnail images of image data stored in the camera 101 can be browsed using a dedicated application in the external device 301. The user can select an image he or she likes from the thumbnail images, confirm the image, and issue an instruction to obtain the image, so that the image is sent to the external device 301.

[0275] At this time, the user selects and obtains an image, so it is very likely that the obtained image is an image that matches the user's preference. Therefore, the obtained image can be judged as an image to be learned, and various user preferences can be learned by performing training based on training information of the obtained image.

[0276] An example of the operation will be described here. Fig.12 An example is shown in which images in the camera 101 are being browsed using a dedicated application of the external device 301. Thumbnail images (1604 to 1609) of image data stored in the camera are displayed in the display unit 407, and the user can select and obtain his or her favorite image. At this time, buttons 1601, 1602, and 1603 constituting a display method change unit for changing the display method are provided.

[0277] When the button 1601 is pressed, the display method changes to a date / time priority display mode in which images within the camera 101 are displayed in the order of the date / time at which they were captured in the display unit 407. For example, images with newer dates / times are displayed at positions indicated by 1604, and images with older dates / times are displayed at positions indicated by 1609.

[0278] When button 1602 is pressed, the mode changes to the recommended image priority display mode. Fig. 9 In order to determine the score calculated for the user's preference for each image in step S912, the images in the camera 101 are displayed in the display unit 407 in order from the image with the highest score. For example, images with higher scores are displayed at the position indicated by 1604, and images with lower scores are displayed at the position indicated by 1609.

[0279] When button 1603 is pressed, a subject such as a person or an object can be specified, and when a specific person or object is then specified, only the specific subject can be displayed. Buttons 1601 to 1603 can also be turned on at the same time. For example, when all buttons are turned on, only the specified subject is displayed, with images captured at a newer date / time being displayed preferentially, and images with higher scores being displayed preferentially. In this way, the user's preferences are also learned for the captured images, so that only images that match the user's preferences can be extracted from a large number of captured images by performing a simple confirmation task.

[0280] <Learning by inputting judgment values ​​for images via external devices>

[0281] As described above, the camera 101 and the external device 301 include a communication component, and images stored in the camera 101 can be browsed using a dedicated application within the external device 301. Here, the structure may be as follows: The user adds a score to each image. The user can add a high score (e.g., 5 points) to an image that matches his or her preference, and a low score (e.g., 1 point) to an image that does not match his or her preference, so the structure is as follows: The camera learns in response to user operations. The score of each image is used together with the training information for retraining within the camera. Learning is performed so that the output of the neural network that takes the feature data from the specified image information as input is close to the score specified by the user.

[0282] Although the present embodiment describes a configuration in which a user inputs a judgment value for a captured image via the external device 301, a configuration may be such that a judgment value is directly input for an image by operating the camera 101. In this case, for example, the camera 101 is provided with a touch panel display, and a mode is set to a mode in which a captured image is displayed when a user presses a GUI button displayed in a screen display section of the touch panel display. The same type of learning can be performed by a method in which a user inputs a judgment value for each captured image while confirming the image.

[0283] <Learning by analyzing images stored in external devices>

[0284] The external device 301 includes a storage unit 404, and is configured to record images other than the images captured by the camera 101 in the storage unit 404. At this time, it is easy for the user to browse the images stored in the external device 301, and it is also easy to upload these images to the shared server via the public wireless control unit 406, so it is likely that many images matching the user's preference are included.

[0285] The control unit 411 of the external device 301 is configured to be able to process the images stored in the storage unit 404 using a dedicated application with performance equivalent to that of the learning processing unit 219 in the camera 101. Learning is performed by communicating the processed training data to the camera 101. Alternatively, a structure may be such that an image and data to be learned, etc. are sent to the camera 101, and learning is performed in the camera 101. A structure is also possible in which a user selects an image to be learned from the images stored in the storage unit 404 using a dedicated application, and then performs learning.

[0286] <Learning from information uploaded to the SNS server by an external device>

[0287] Next, a method will be described in which information from a social network service (SNS) that is a service or website that can build a social network focusing on connections between people is used in learning. There is a technology in which, when an image is uploaded to the SNS, the image is sent from the external device 301 together with tag information input for the image. There is also a technology in which likes or dislikes are input for images uploaded by other users, so that it can be determined whether the images uploaded by other users are images that match the preferences of the user who owns the external device 301.

[0288] Images uploaded by the user himself or herself and information related to the images as described above can be obtained via a dedicated SNS application downloaded to the external device 301. In addition, images matching the user's preferences and tag information, etc. can also be obtained based on the user's input of whether he or she likes images uploaded by other users. By analyzing these images and tag information, etc., learning can be performed in the camera 101.

[0289] The control unit 411 of the external device 301 is configured to be able to obtain images uploaded by the user, images judged to match the user's preference, etc. as described above, and process these images with performance equivalent to that of the learning processing unit 219 in the camera 101. Learning is performed by communicating the processed training data to the camera 101. Alternatively, the structure may be as follows: an image to be learned is sent to the camera 101, and learning is performed in the camera 101.

[0290] In addition, subject information assumed to match the user's preference is estimated based on subject information set in the tag information (for example, subject information indicating a subject such as a dog or a cat, scene information indicating a beach, and expression information indicating a smile, etc.) Then, learning is performed by registering the information as a subject to be detected by inputting it into the neural network.

[0291] In addition, a configuration may be adopted in which image information currently popular in the world is estimated based on statistical values ​​of tag information (image filter information, subject information, and the like) in the above-described SNS, and then learning may be performed in the camera 101 .

[0292] <Learning by changing camera parameters using external devices>

[0293] As described above, the camera 101 and the external device 301 have a communication component. The learning parameters currently set in the camera 101 (neural network weights, and the selection of the subject to be input to the neural network, etc.) can be communicated to the external device 301 and stored in the storage unit 404 of the external device 301. In addition, the learning parameters set in the dedicated server can be obtained via the public wireless control unit 406 using a dedicated application in the external device 301, and then these learning parameters can be set as learning parameters in the camera 101. Therefore, by storing the parameters at a given point in time in the external device 301 and then setting these parameters in the camera 101, the learning parameters can also be restored. In addition, the learning parameters maintained by other users can also be obtained via the dedicated server and set in the user area camera 101.

[0294] In addition, the structure may be as follows: a dedicated application of the external device 301 may be used for voice commands, authentication registration, gesture registration, etc. registered by the user, or may be used to register important places. This information is processed as in the automatic camera mode ( Fig. 9 ) and the input data for determining automatic photography, etc. In addition, the structure can be as follows: the photography frequency, the start-up interval, the ratio of still images to moving images, and the preferred image can be set, and then the settings such as the start-up interval, etc. described in "Low Power Mode Control" can be set.

[0295] <Learning from Information Obtained by Manually Editing Images in External Devices>

[0296] The dedicated application in the external device 301 may be provided with a function capable of manual editing via user operation, and then the details of the editing task are fed back to the learning. For example, editing for adding image effects (e.g., cropping, rotation, slideshow, zoom, fade in and fade out, color conversion filter effect, time, still image to moving image ratio, BGM) may be performed. Then, the neural network for automatic editing is trained, thereby judging the image effects added by manual editing for the training information of the image.

[0297] Next, the sequence of the learning process will be described. Fig. 7AIn the mode setting judgment performed in step S704, it is judged whether the learning process should be executed, and if it is judged that the learning process should be executed, the learning mode processing of step S716 is executed.

[0298] The conditions for determining the learning mode will be described here. Whether to shift to the learning mode is determined based on the amount of time since the previous learning process was executed, the amount of information that can be used in learning, and whether an instruction to execute the learning process is given via the communication device. Fig.13 The flow of processing for determining whether to shift to the learning mode is shown, and this determination is performed in the mode setting determination processing of step S704.

[0299] If an instruction to start learning mode determination is given in the mode setting determination process of step S704, Fig.13 The sequence shown in FIG. 14 starts. In step S1401, it is determined whether a registration instruction has been issued from the external device 301. The determination here is related to whether a registration instruction has been issued for the above-mentioned learning (for example, "learning by obtaining an image through an external device", "learning by inputting a judgment value for an image via an external device", or "learning by analyzing an image stored in an external device", etc.).

[0300] If a registration instruction is given from the external device 301 in step S1401, the sequence proceeds to step S1408, in which the learning mode determination is set to "true", the processing of step S716 is set to be executed, and the learning mode determination processing ends. If there is no registration instruction from the external device in step S1401, the sequence proceeds to step S1402.

[0301] In step S1402, it is determined whether a learning instruction is given from an external device. The determination here is made based on whether an instruction for setting learning parameters such as "learning to change camera parameters by using an external device" is given. If a learning instruction is given from an external device in step S1402, the sequence proceeds to step S1408, in which the learning mode determination is set to "true", the processing of step S716 is set to be executed, and the learning mode determination processing ends. If there is no learning instruction from an external device in step S1402, the sequence proceeds to step S1403.

[0302] In step S1403, the elapsed time TimeN since the previous learning process (recalculation of the weights of the neural network) is obtained, and the sequence proceeds to step S1404. In step S1404, the number of new data used for learning DN (the number of images designated for learning during the elapsed time TimeN since the previous learning process was performed) is obtained, and the sequence proceeds to step S1405. In step S1405, a threshold DT for judging whether to enter the learning mode after the elapsed time TimeN is calculated. The structure is as follows: as the value of the threshold DT decreases, it is easier to enter the learning mode. For example, DTa, which is the value of the threshold DT when TimeN is less than a predetermined value, is set to be greater than DTb, which is the value of the threshold DT when TimeN is greater than a predetermined value, and the threshold is set to decrease as time passes. Therefore, even if there is little training data, it is easier to enter the learning mode when a large amount of time has passed; by performing learning again, the camera is more likely to change through learning according to the use time.

[0303] Once the threshold DT is calculated in step S1405, the sequence proceeds to step S1406, in which it is determined whether the number of data DN used for learning is greater than the threshold DT. If the number of data DN is greater than the threshold DT, the sequence proceeds to step S1407, in which DN is set to 0. Then, the sequence proceeds to step S1408, in which the learning mode determination is set to "true", and step S716 ( Figure 7B ) is set to be executed, and the learning mode judgment processing ends.

[0304] If DN is less than or equal to the threshold value DT in step S1406, the sequence proceeds to step S1409. There is no registration instruction or restriction instruction from the external device, and the amount of data used for learning is less than or equal to the predetermined value; in this case, the learning mode judgment is set to "false", the processing of step S716 is set not to be executed, and the learning mode judgment processing ends.

[0305] Next, the processing performed in the learning mode processing (step S716) will be described. Fig.14 is a flowchart illustrating in detail the operations performed in the learning mode process.

[0306] exist Figure 7B When it is determined in step S715 that the learning mode is in progress and the sequence enters step S716, Fig.14 The sequence starts. In step S1501, it is determined whether a registration instruction is given from the external device 301. If there is no registration instruction from the external device 301 in step S1501, the sequence proceeds to step S1502. Various registration processes are performed in step S1502.

[0307] The various registrations are registrations of features to be input to the neural network, for example, facial authentication registration, general object recognition registration, sound information registration, and place information registration, etc. Once the registration process is completed, the sequence proceeds to step S1503, and the elements to be input to the neural network are changed based on the information registered in step S1502. Once the process of step S1503 is completed, the sequence proceeds to step S1507.

[0308] If there is no registration instruction from the external device 301 in step S1501, the sequence proceeds to step S1504, in which it is determined whether a learning instruction is given from the external device 301. If there is a learning instruction from the external device 301, the sequence proceeds to step S1505, in which the learning parameters communicated from the external device 301 are set in various judgements (neural network weights, etc.), and then the sequence proceeds to step S1507.

[0309] If there is no learning instruction from the external device 301 in step S1504, learning is performed (recalculation of the neural network weights) in step S1506. Fig.13 When the amount of data DN used for learning exceeds the threshold value DT and each judgement device is to be retrained, the process of step S1506 is performed. Retraining is performed by a method such as error back propagation or gradient descent, the weight of the neural network is recalculated, and the parameters of each judgement device are changed. Once the learning parameters are set, the sequence enters step S1507.

[0310] In step S1507, the images in the file are re-scored. In the present embodiment, the structure is as follows: scores are assigned to all captured images stored in the file (recording medium 221) based on the learning results, and automatic editing and automatic file deletion, etc. are performed according to the assigned scores. Therefore, if retraining is performed or learning parameters from an external device are set, the scores of the captured images also need to be updated. Therefore, in step S1507, recalculation is performed to assign new scores to the captured images stored in the file, and once the processing is completed, the learning mode processing is also completed.

[0311] The present embodiment describes a structure in which learning is performed within the camera 101. However, the same learning effect can be achieved even with a structure in which a learning function is set in the external device 301, and learning is performed only on the external device side by communicating data required for learning to the external device 301. In this case, as described above in “Learning by changing camera parameters using an external device”, the structure may be as follows: learning is performed by setting parameters such as neural network weights learned on the external device side in the camera 101 via communication.

[0312] In addition, the structure may be as follows: both the camera 101 and the external device 301 are provided with a learning function; for example, the structure may be as follows: the training information maintained by the external device 301 is communicated to the camera 101 at a timing when the learning mode processing (step S716) is performed within the camera 101, and learning is performed by merging learning parameters.

[0313] (Second embodiment)

[0314] The structure of the camera according to the second embodiment is the same as that in the first embodiment; thus, only portions different from those in the first embodiment will be described below, and the structure of the same processing will not be described.

[0315] In the present embodiment, the type of the accessory attached to the camera 101 can be detected using an accessory detection unit (not shown). For example, a method is used in which information on the type of the attached accessory is transmitted from the camera 101 to the camera 101 using a non-contact communication component or the like. FIG. 15A to FIG. 15D The accessories 1501 to 1504 shown are sent to the camera 101. It is also possible to use the connectors provided in the camera 101 and the accessories 1501 to 1504 to send and receive information and perform detection. However, in the case where the camera 101 includes a battery, there is a case where it is not necessary to provide a battery connector in the accessory. In this case, purposefully providing a connector will require also including components such as for adding a waterproof function to the connecting portion, which increases the size and cost of the device, etc. Therefore, it is preferred to use a contactless communication component, etc. Bluetooth Low Energy (BLE) or Near Field Communication (NFC) BLE, etc. can be used as a contactless communication component, or other methods can be used instead.

[0316] In addition, the radio wave transmission source in accessories 1501 to 1504 can be a compact power source with a low capacity; for example, a button battery, or a component that generates very small electricity from the force used to press an operating member (not shown), etc. can be used.

[0317] By detecting the type of accessory, the state of the camera can be determined in a limited manner according to the type of accessory (e.g., whether the camera is in a handheld state, hanging from the neck, in a wearable state, placed on a table, or attached to a moving object, etc., i.e., state information). The attachment of the accessory can be detected using existing methods such as detecting a change in voltage or detecting an ID.

[0318] FIG. 15A to FIG. 15D It is a diagram showing a usage example when an accessory is attached. Fig.15A Shows the handheld state. Fig. 15B showing a camera hanging from the neck, Fig. 15C shows the wearable status, and Fig.15D1501 shows a fixed placement state; here, various accessories are attached to the camera 101. 1501 denotes a handheld accessory; 1502 denotes an accessory for hanging around the neck; 1503 denotes a wearable accessory; and 1504 denotes an accessory for fixed placement. Instead of a head-mounted form, it is also conceivable that the wearable accessory 1503 is attached to a person's shoulder or belt, etc.

[0319] When attaching accessories in this manner, there are situations where the state of the camera is restricted, which improves the accuracy of judging the camera state; this in turn makes it possible to more appropriately control the timing of automatic photography, the subject search range, and the timing of turning low-power mode on and off, etc.

[0320] In addition, the state of the camera can be further limited by combining the camera state judgment performed by the type of accessory used with subject information, camera movement information, oscillation status, etc. For example, when a handheld accessory is detected, the subject search range control, automatic camera control, and low power mode control are changed according to whether the user is walking or standing still in a handheld state. Similarly, in the case of detecting a fixed-place accessory, it is determined based on movement information and oscillation status whether the camera is stationary on a table or attached to a vehicle or drone, etc. and moving, and various controls are changed.

[0321] Specifically, Fig.18 A summary of examples of subject search ranges, automatic imaging control, and low power mode control for each type of attachment is shown.

[0322] First, in Fig.15A In the case of attaching the handheld accessory 1501 as shown, Fig.16A As shown, this is regarded as a high probability that the user is pointing the camera at a given subject. Therefore, even if the subject search range 1601 is set to a narrow range, the subject 1602 can be found. Fig. 16B 1 is a diagram showing only the camera 101 attached to the handheld accessory 1501. The subject search range can be set to an area such as a subject search range 1604 that is narrow relative to the shootable range 1603. In the automatic shooting performed in this case, the user intentionally points the camera at the subject, so the frequency of shooting can be increased (shooting frequency determination), and the shooting direction can be limited to the direction the camera is facing. In addition, since the image is to be captured when the camera is pointing at the subject, there is no need to set the camera to low power mode. This example is an example in which the subject is in the direction the camera is facing, so when this is not the case, there is no need to narrow the subject search range or increase the shooting frequency, etc.

[0323] Then, in Fig. 15B In the case of the neck suspension attachment 1502 as shown, as by FIG. 10A to FIG. 10EAs shown in 1322 in , the subject search range can be limited to a range that avoids showing the user's body as much as possible. When the camera is hung from the neck, it can be imagined that the camera is being used as a life recording camera, so the shooting frequency can be set to a constant interval, and the forward direction can be prioritized as the shooting direction. However, the structure can be as follows: in response to the user's movement or voice, etc., the changes in the surrounding environment are detected, the shooting frequency is increased, and the restrictions on the shooting direction are eliminated, so that more images can be recorded when an event occurs. In addition, once a predetermined number of images of the same scene are captured, the camera can be switched to a low-power mode until a trigger such as a large movement of the user is detected, thereby avoiding taking many similar photos.

[0324] In such Fig. 15C With the wearable accessory 1503 attached as shown, the control can be greatly varied depending on the intended use. Fig. 15C A head-mounted accessory such as the one shown can perform the same control as that applied when using the neck-hanging accessory in a situation where the user is hiking or mountain climbing, etc. Alternatively, when a technician is using a camera to record his or her work, a subject search can be performed so that the technician's hands are shown as much as possible, and control can be performed to record at an increased camera frequency.

[0325] Will use Fig.17A and Fig. 17B To illustrate how Fig.15D The condition in which the attachment 1504 for fixing and placing is attached is shown. Fig.17A The following situation is shown as viewed from the front side: the camera 101 is placed on a table using the fixed placement attachment 1504. As shown by 1702, the subject search range can be slightly reduced relative to the photographable range 1701 so that the table (ground, floor) does not occupy a large part of the image. Fig. 17B 1703 is a diagram showing a state in which the camera is viewed from above. In this state, there are no specific obstacles in the surroundings, so the subject search range 1703 applies to all directions. Since the camera does not move at this time, in automatic camera control, the camera frequency can be reduced each time an image is captured to avoid taking similar photos, and then the camera frequency can be increased once a new person or change in the environment is detected. The camera direction is also set to cover all directions to avoid taking similar photos. In addition, once a predetermined number of images are captured, the camera can be switched to low power mode.

[0326] In addition, when the camera is attached to a moving object using a fixed placement accessory, it is assumed that the camera will be moving toward the captured subject; therefore, the subject search can be preferentially performed in the forward direction, and the shooting direction can also be preferentially set to the forward direction. In this case, the shooting frequency is changed according to the movement of the moving object. For example, if the direction of travel changes frequently, the shooting frequency is increased. However, if the direction of travel does not change and the speed does not change significantly, control can be performed to switch to low power mode at set intervals.

[0327] In this way, using the attached information makes it possible to limit the camera state judgment, which in turn makes it possible to more accurately judge the state. Therefore, subject search control, automatic camera control, and low power mode control can be performed more accurately, thereby increasing the possibility that the user can capture images according to his or her expectations.

[0328] (Other embodiments)

[0329] The present invention may also be implemented as a process performed in the following manner: a program that implements one or more functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and then one or more processors of a computer of the system or device read out and execute the program. The present invention may also be implemented by a circuit (e.g., ASIC) that implements one or more functions.

[0330] Examples of embodiments of the present invention will be described below.

[0331] (Example 1)

[0332] A camera device, characterized in that it includes: a camera component for capturing an image of a subject; a subject detection component for detecting a subject from image data captured by the camera component; a state detection component for detecting information related to the state of movement of the camera device itself; and a control component for controlling the range in which the subject detection component searches for a subject based on state information of the camera device detected by the state detection component.

[0333] (Example 2)

[0334] The image pickup apparatus according to Embodiment 1 is characterized in that the state detection section detects a traveling direction and a moving speed of the image pickup apparatus.

[0335] (Example 3)

[0336] The camera device according to embodiment 1 or 2 is characterized in that the state detection component detects the travel direction and moving speed of the camera device based on at least one of the angular velocity information, acceleration information, GPS position information and motion vectors calculated for each coordinate from the captured image of the camera device.

[0337] (Example 4)

[0338] The image pickup apparatus according to Embodiment 3 is characterized in that the state detection section detects the traveling direction and the moving speed of the image pickup apparatus based on the measurement result in the set time period.

[0339] (Example 5)

[0340] The image pickup apparatus according to any one of Embodiments 1 to 4 is characterized in that the object detection section changes a time interval for searching for the object based on a detection result from the state detection section.

[0341] (Example 6)

[0342] The image pickup apparatus according to any one of Embodiments 1 to 5 is characterized in that the control means narrows the range for searching for an object when the moving speed of the image pickup apparatus detected by the state detection means increases.

[0343] (Example 7)

[0344] According to any one of embodiments 1 to 6, the imaging device is characterized in that, when the imaging device is determined to be stationary through detection performed by the state detection component, the control component widens the range for searching the subject compared to when the imaging device is moving.

[0345] (Example 8)

[0346] A camera device, characterized in that it includes: a camera component for capturing an image of a subject; a subject detection component for detecting a subject from image data captured by the camera component; a state detection component for detecting information related to the state in which the camera device is being maintained; and a control component for controlling the range in which the subject detection component searches for a subject based on the state information of the camera device detected by the state detection component.

[0347] (Example 9)

[0348] The imaging device according to Embodiment 8 is characterized in that the state in which the imaging device is being held includes at least one of a handheld state, a state suspended from the neck, a wearable state, a state placed on a table, and a state placed on a moving object.

[0349] (Example 10)

[0350] The image pickup apparatus according to any one of Embodiments 1 to 9 is characterized in that the control means sets the range for searching for the object to a left-right symmetrical angle range with respect to a traveling direction of the image pickup apparatus.

[0351] (Example 11)

[0352] The imaging device according to Embodiment 1 or 8 is characterized in that the state detection component detects information of an accessory attached to the imaging device, and the control component controls the range in which the subject detection component searches for the subject based on the information of the attached accessory.

[0353] (Example 12)

[0354] The imaging device according to Embodiment 8 is characterized in that, when the imaging device is being held in a state of being suspended from the neck, the control unit limits the range in which the subject detection unit searches for the subject so that the user's body is not visible.

[0355] (Example 13)

[0356] The imaging device according to any one of Embodiments 1 to 12 is characterized by further comprising: a changing component for changing the orientation of the imaging component so that the imaging component faces the direction of the subject.

[0357] (Example 14)

[0358] The image pickup apparatus according to Embodiment 13 is characterized in that the changing section causes the image pickup section to rotate in a pan direction or a tilt direction.

[0359] (Example 15)

[0360] According to the image pickup apparatus of Embodiment 13 or 14, it is characterized in that the range for searching for the subject is a range of change of the orientation of the image pickup apparatus by the changing means.

[0361] (Example 16)

[0362] According to any one of Embodiments 1 to 15, the image pickup apparatus is characterized in that a different neural network is set for each state of the image pickup apparatus, and a neural network suitable for the state of the image pickup apparatus is applied.

[0363] (Example 17)

[0364] The imaging device according to any one of embodiments 1 to 16 is characterized in that it further comprises: an imaging frequency determination component for determining the imaging frequency of the automatic imaging component, wherein the imaging frequency is determined based on status information of the imaging device.

[0365] (Example 18)

[0366] The imaging apparatus according to any one of Embodiments 1 to 17 is characterized by further comprising: a low power mode control section, wherein low power mode control is performed based on state information of the imaging apparatus.

[0367] (Example 19)

[0368] The imaging device according to any one of embodiments 1 to 18 is characterized in that it also includes: an automatic imaging component for enabling the imaging component to capture an image based on information of the subject detected by the subject detection component, and recording the captured image data.

[0369] (Example 20)

[0370] A control method for a camera device, the camera device includes a camera component for capturing an image of a subject, and the control method is characterized in that it includes: a subject detection step for detecting a subject from image data captured by the camera component; a state detection step for detecting information related to the state of movement of the camera device itself; and a control step for controlling a range for searching a subject in the subject detection step based on state information of the camera device detected in the state detection step.

[0371] (Example 21)

[0372] A control method for a camera device, the camera device includes a camera component for capturing an image of a subject, and the control method is characterized in that it includes: a subject detection step for detecting a subject from image data captured by the camera component; a state detection step for detecting information related to the state in which the camera device is being maintained; and a control step for controlling a range for searching a subject in the subject detection step based on state information of the camera device detected in the state detection step.

[0373] (Example 22)

[0374] A program for causing a computer to execute the steps of the control method according to embodiment 20 or 21.

[0375] (Example 23)

[0376] A computer-readable storage medium storing a program for causing a computer to execute the steps of the control method according to Embodiment 20 or 21.

[0377] The present invention is not limited to the above-described embodiments, and various changes and modifications can be made within the spirit and scope of the present invention. Therefore, in order to inform the public of the scope of the present invention, the following claims are added.

[0378] This application claims the benefit of Japanese Patent Application No. 2017-242228, filed on December 18, 2017, and Japanese Patent Application No. 2017-254402, filed on December 28, 2017, which are hereby incorporated by reference herein in their entirety.

Claims

1. A camera device, comprising: A camera component, used for capturing an image of a subject; as well as a subject detection component, used to detect the subject from the image data captured by the camera component, Characterized in that the camera device also includes: a state detection unit for detecting information related to a moving state of the imaging device and information related to a state in which the imaging device is installed; and A control component for controlling a time interval for a rotation operation based on information related to a moving state of the camera device and information related to a state in which the camera device is installed, detected by the state detection component, wherein in the rotation operation, the camera component is rotated in at least one of a pan direction and a tilt direction.

2. The imaging device according to claim 1, wherein: The state detection component detects information about the moving speed of the camera device as information about the moving state of the camera device, and the control component controls the time interval for performing the rotation operation based on the information about the moving speed of the camera device and the information about the state in which the camera device is installed.

3. The imaging device according to claim 2, wherein: The control section shortens the time interval for performing the rotation operation when the state detection section detects that the imaging apparatus is moving at the first speed, compared to when the state detection section detects that the imaging apparatus is moving at a second speed lower than the first speed.

4. The imaging device according to claim 1, wherein: The rotation operation is performed to automatically search for a subject.

5. The imaging device according to claim 1, wherein: The state detection section detects information on a moving state of the imaging device based on at least one of angular velocity information, acceleration information, GPS position information, and a motion vector calculated for each coordinate from a captured image of the imaging device.

6. The imaging device according to claim 5, wherein: The state detection section detects a moving direction and a moving speed based on a detection result during a constant time interval.

7. The imaging device according to claim 1, wherein: The state in which the imaging device is installed includes at least one of a handheld state, a state suspended from the neck, a wearable state, a state placed on a table, and a state placed on a moving body.

8. The imaging device according to claim 1, wherein: The control section also controls the range in which the subject detection section searches for the subject by controlling the range in which the orientation of the imaging section is changed based on information on the state in which the imaging device is installed.

9. The imaging device according to claim 8, wherein: When the imaging device is mounted in a state of being suspended from the neck, the control section limits the range in which the subject detection section searches for the subject so that the user's body is not visible.

10. The imaging device according to claim 1, wherein A different neural network is set for each state of the camera device.

11. The imaging device according to claim 1, wherein: The control section controls the range in which the subject detection section searches for the subject by controlling the range in which the orientation of the imaging section is changed based on information on the movement state of the imaging device.

12. The imaging device according to claim 1, wherein: The state detection section detects information about a moving speed of the imaging device as information about a moving state of the imaging device, and the control section controls a range in which the subject detection section searches for a subject based on the information about the moving speed.

13. The imaging device according to claim 12, wherein: When the state detection component detects that the imaging device is moving at the third speed, the control component narrows the range in which the subject detection component searches for the subject, compared to when the state detection component detects that the imaging device is moving at a fourth speed lower than the third speed.

14. The imaging device according to claim 12, wherein: The control section causes the subject detection section to search for a subject in a wider range when the state detection section detects that the imaging apparatus is stationary than when the state detection section detects that the imaging apparatus is moving.

15. The imaging device according to claim 1, wherein: The state detection section detects information about a moving direction of the image pickup device, and the control section controls a range in which the object detection section searches for an object based on the information about the moving direction.

16. The imaging device according to any one of claims 1 to 15, further comprising an automatic imaging component configured to cause the imaging component to perform imaging based on information of the subject detected by the subject detection component, and to record the captured image data.

17. The imaging device according to claim 16, further comprising: a photographing frequency determining component, used to determine the photographing frequency of the automatic photographing component, The imaging frequency is determined based on information related to a moving state of the imaging device.

18. The imaging device according to any one of claims 1 to 15, further comprising: Low power mode control components, Here, low power mode control is performed based on information related to a movement state of the imaging device.

19. The imaging device according to claim 16, wherein: The automatic imaging component performs control to automatically perform an imaging operation using parameters generated by machine learning.

20. The imaging device according to claim 19, wherein: The imaging operation is changed by updating the parameters based on machine learning using data output by the imaging component.

21. A control method for an imaging device, the imaging device comprising an imaging component for capturing an image of an object, and a subject detection component for detecting the object from image data captured by the imaging component, characterized in that: The control method comprises: a state detection step for detecting information related to a moving state of the imaging device and information related to a state in which the imaging device is installed; and A control step for controlling a time interval for performing a rotation operation based on information related to a moving state of the imaging device and information related to a state in which the imaging device is installed, which are detected in the state detection step, wherein in the rotation operation, the imaging component is rotated in at least one of a pan direction and a tilt direction.

22. A computer-readable storage medium storing a program for causing a computer to execute the steps of the control method according to claim 21.

Citation Information

Patent Citations

  • Operation method of attachable life logging device

    JP2016536868A

  • Electronic camera

    US20070211161A1