Information processing device, control method, and program

The information processing device controls multiple imaging devices to prevent overlapping targets and ranges by comparing and adjusting their shooting directions, enhancing image capture efficiency and reducing redundancy.

JP2026059616APending Publication Date: 2026-04-07CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies for automatic shooting with multiple cameras often result in overlapping shooting targets and ranges, leading to redundant or biased image capture.

Method used

An information processing device that communicates with multiple imaging devices, acquires and compares images, and controls their shooting directions to avoid overlap by using a system with rotatable imaging units, angular velocity and acceleration sensors, and neural networks for image processing.

Benefits of technology

Enables imaging without target or range overlap between multiple devices, optimizing image capture efficiency and reducing redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026059616000001_ABST
    Figure 2026059616000001_ABST
Patent Text Reader

Abstract

Capture images in a way that avoids overlap between the target objects and shooting ranges of multiple imaging devices. [Solution] The information processing device includes communication means for communicating with a plurality of imaging devices whose shooting direction can be changed, acquisition means for acquiring images taken from the plurality of imaging devices for each shooting direction, comparison means for comparing the images acquired from the plurality of imaging devices, and control means for controlling the shooting directions of the plurality of imaging devices so that they do not overlap, based on the results of the comparison.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a control technology for a plurality of imaging devices.

Background Art

[0002] An automatic shooting camera having a function of automatically detecting a subject from a captured image and automatically performing shooting is known. Patent Document 1 describes selecting a camera that captures an image optimal for image processing on an overlapping imaging area when the imaging areas of a plurality of cameras overlap. Patent Document 2 describes notifying an information processing device of a shooting area that can be covered and a shooting area that cannot be covered based on the shooting ranges of a plurality of cameras when shooting is performed by interlocking a plurality of cameras.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In Patent Documents 1 and 2, when automatically shooting with a plurality of cameras, the shooting targets and shooting ranges may overlap between the cameras, and similar images may be captured by the plurality of cameras, or shooting may be biased toward the same subject.

[0005] The present invention has been made in view of the above problems, and an object thereof is to realize a technique for performing shooting so that the shooting targets and shooting ranges do not overlap between a plurality of imaging devices.

Means for Solving the Problems

[0006] To solve the above problems and achieve the objective, the information processing device of the present invention includes communication means for communicating with a plurality of imaging devices whose shooting direction can be changed, acquisition means for acquiring images captured from the plurality of imaging devices for each shooting direction, comparison means for comparing the images acquired from the plurality of imaging devices, and control means for controlling the shooting directions of the plurality of imaging devices so that they do not overlap, based on the results of the comparison. [Effects of the Invention]

[0007] According to the present invention, it is possible to perform imaging without overlapping imaging targets or imaging ranges between multiple imaging devices. [Brief explanation of the drawing]

[0008] [Figure 1] This figure schematically illustrates the external appearance of the imaging device of this embodiment. [Figure 2] This is a block diagram illustrating the internal configuration of the imaging device of this embodiment. [Figure 3] This diagram illustrates the connection relationship between the imaging device and the external device in this embodiment. [Figure 4] This is a block diagram illustrating the internal configuration of the external device of this embodiment. [Figure 5] This figure schematically illustrates the external appearance of the imaging device and external device of this embodiment. [Figure 6] This is a block diagram illustrating the internal configuration of the external device of this embodiment. [Figure 7] This is a flowchart illustrating the first control process of this embodiment. [Figure 8] This is a flowchart illustrating the second control process of this embodiment. [Figure 9] This is a flowchart illustrating the automatic shooting mode processing of this embodiment. [Figure 10] This is a diagram illustrating the shooting area of ​​this embodiment. [Figure 11] This diagram illustrates the configuration of the neural network in this embodiment. [Figure 12]This is a diagram for explaining the image selection method of the present embodiment. [Figure 13] This is a flowchart exemplifying the learning mode determination process of the present embodiment. [Figure 14] This is a flowchart exemplifying the learning mode process of the present embodiment. [Figure 15] This is a diagram exemplifying the configuration of a system including a plurality of imaging devices and an external device in the present embodiment. [Figure 16] This is a diagram exemplifying the shooting ranges of a plurality of imaging devices in the present embodiment. [Figure 17] This is a diagram exemplifying an omnidirectional image by a plurality of imaging devices in the present embodiment. [Figure 18] This is a flowchart exemplifying the control process of a plurality of imaging devices in the present embodiment. [Figure 19] This is a flowchart exemplifying the control process of an information processing device to which a plurality of imaging devices are connected in the present embodiment. [Figure 20] This is a flowchart exemplifying the shooting area determination process of the present embodiment. [Figure 21] This is a diagram exemplifying the shooting area determination result of the present embodiment. [Figure 22] This is a flowchart exemplifying the shooting permission determination process based on image similarity in the present embodiment. [Figure 23] This is a flowchart exemplifying the image similarity determination process of the present embodiment. [Figure 24] This is a diagram exemplifying the image similarity determination result of the present embodiment. [Figure 25] This is a diagram exemplifying the shooting permission determination result based on image similarity in the present embodiment. [Figure 26] This is a flowchart exemplifying the shooting permission determination process based on subject bias in the present embodiment. [Figure 27] This is a diagram exemplifying the subject determination result of the present embodiment. [Figure 28] This is a diagram exemplifying the shooting permission determination result based on subject bias in the present embodiment. [Figure 29]This figure illustrates the final settings for enabling or disabling photography in this embodiment. [Figure 30] This figure illustrates a modified example of the image similarity-based photo capture feasibility determination process of this embodiment. [Modes for carrying out the invention]

[0009] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0010] The following describes embodiments in which the image processing apparatus of the present invention is applied to an imaging device such as a digital camera. The digital camera in this embodiment is an automatic shooting camera that has the function of automatically detecting a subject from an captured image and automatically taking a picture. In this embodiment, an example of controlling multiple cameras so that the shooting target and shooting range do not overlap between cameras when automatic shooting is performed will be described.

[0011] <Device configuration> First, the configuration of the system and apparatus of this embodiment will be described with reference to Figure 1.

[0012] Figure 1(a) is a schematic diagram illustrating the external appearance of the imaging device of this embodiment.

[0013] The imaging device (hereinafter referred to as camera) 101 of this embodiment includes an imaging unit 102 and a support unit 103. The camera 101 of this embodiment is also provided with operating parts such as a power button and a shutter button (not shown).

[0014] The imaging unit 102 includes an optical system for capturing an image of a subject. The optical system includes a lens group that forms an optical image of the subject onto the image sensor 206, and an image sensor 206 which is composed of a CCD or CMOS that converts the optical image of the subject into an electrical signal. The imaging unit 102 is rotatably mounted on the support unit 103. More specifically, the imaging unit 102 is rotatably mounted on the support unit 103 by a first rotation mechanism 104 and a second rotation mechanism 105, and the shooting direction (hereinafter referred to as the shooting angle) of the imaging unit 102 can be changed.

[0015] The first rotation mechanism 104 is a tilt unit that rotates the imaging unit 102 in the tilt direction. The second rotation mechanism 104 is a pan unit that rotates the imaging unit 102 in the pan direction.

[0016] Furthermore, the support section 103 is provided with an angular velocity meter 106 and an accelerometer 107. For example, a gyro sensor is used for the angular velocity meter 106, and for example, an acceleration sensor is used for the accelerometer 107.

[0017] Figure 1(b) illustrates the relationship between the three-dimensional Cartesian coordinate system of camera 101 and the rotation direction of camera 101.

[0018] The X-axis (horizontal axis), Y-axis (vertical axis), and Z-axis (depth axis) of the three-dimensional Cartesian coordinate system are defined with respect to the position of the support part 103. In this embodiment, the direction around the X-axis is defined as the pitch direction, the direction around the Y-axis is defined as the yaw direction, and the direction around the Z-axis is defined as the roll direction.

[0019] The tilt unit 104 includes a drive mechanism such as a motor that allows the imaging unit 102 to rotate in the pitch direction shown in Figure 1(b). The pan unit 105 includes a drive mechanism such as a motor that allows the imaging unit 102 to rotate in the yaw direction shown in Figure 1(b). The camera 101 has a drive mechanism that allows the imaging unit 102 to rotate around at least two axes of a three-dimensional Cartesian coordinate system.

[0020] The angular velocity meter 106 outputs an angular velocity detection signal. The accelerometer 107 outputs an acceleration detection signal. Based on the detection signals from the angular velocity meter 106 and the accelerometer 107, vibration of the camera 101 is detected, and the tilt unit 104 and the pan unit 105 are rotated. This corrects the shake and tilt of the imaging unit 102. In addition, the direction and distance of movement of the camera 101 are detected based on the detection results from the angular velocity meter 106 and the accelerometer 107 over a predetermined period.

[0021] Figure 2 is a block diagram illustrating the internal configuration of the camera 101 in this embodiment.

[0022] The first control unit 223 includes a processor (main processor) such as a CPU (Central Processing Unit) or MPU (Micro-Processing Unit) that performs calculation and control processing for the camera 101.

[0023] The work memory 215 includes dynamic RAM and static RAM, and is loaded with constants and variables for the operation of the first control unit 223, as well as control programs read from the non-volatile memory 216.

[0024] The non-volatile memory 216 includes flash ROM and stores constants and control programs for the operation of the first control unit 223.

[0025] The first control unit 223 controls each component of the camera 101 and controls data transfer between each component by loading a control program stored in the non-volatile memory 216 into the work memory 215 and executing it.

[0026] The zoom unit 201 includes a zoom lens that performs magnification (enlargement or reduction of the image of the subject). The zoom drive control unit 202 drives the zoom unit 201 and detects the focal length during drive control.

[0027] The focus unit 203 includes a focus lens for adjusting the focus. The focus drive control unit 204 drives and controls the focus unit 203.

[0028] The image sensor 206 converts charge information corresponding to the amount of light incident through the lens group into an analog image signal and outputs it to the image processing unit 207. The zoom unit 201, focus unit 203, and image sensor 206 are located in the imaging unit 102.

[0029] The image processing unit 207 performs image processing on the image data obtained by converting the analog image signal into a digital signal. Image processing includes distortion correction, white balance adjustment, and color interpolation, and the image processing unit 207 outputs the image data after image processing.

[0030] The image recording unit 208 acquires image data output from the image processing unit 207. The image data is converted to a recording format such as JPEG (Joint Photographic Experts Group) format, temporarily stored in the work memory 215, and then transmitted to the video output unit 217, which will be described later.

[0031] The pan / tilt drive control unit 205 drives the tilt unit 104 and the pan unit 105 to rotate the imaging unit 102 in the tilt and pan directions.

[0032] The camera shake detection unit 209 includes an angular velocity meter 106 that detects the angular velocity of the camera 101 in three axes and an accelerometer 107 that detects the acceleration of the camera 101 in three axes. The first control unit 223 calculates the rotation angle of the imaging unit 102 and the amount of camera shake of the camera 101 based on the detection signals from the camera shake detection unit 209.

[0033] The audio input unit 213 acquires audio signals from the area around the camera 101 using a microphone provided on the camera 101, converts them into digital audio signals, and transmits them to the audio processing unit 214. The audio processing unit 214 performs audio-related processing, such as optimizing the digital audio signals received from the audio input unit 213. The audio signals processed by the audio processing unit 214 are transmitted to the work memory 215 by the first control unit 223. The work memory 215 temporarily stores the image signals and audio signals obtained by the image processing unit 207 and the audio processing unit 214.

[0034] The image processing unit 207 and the audio processing unit 214 read the image signal and audio signal temporarily stored in the work memory 215, encode the image signal and the audio signal, and generate a compressed image signal and a compressed audio signal. The first control unit 223 transmits the compressed image signal and the compressed audio signal to the recording and playback unit 220.

[0035] The recording and playback unit 220 records the compressed image signal and compressed audio signal generated by the image processing unit 207 and the audio processing unit 214, as well as control data related to shooting, onto the recording medium 221. If the audio signal is not compressed, the first control unit 223 transmits the audio signal generated by the audio processing unit 214 and the compressed image signal generated by the image processing unit 207 to the recording and playback unit 220 for recording on the recording medium 221.

[0036] The recording medium 221 is either built into the camera 101 or is removable from the camera 101, and can record various data such as compressed image signals, compressed audio signals, and audio signals generated by the camera 101. The recording medium 221 uses a medium with a larger capacity than the non-volatile memory 216. The recording medium 221 can be, but is not limited to, a hard disk, optical disk, magneto-optical disk, CD-R, DVD-R, magnetic tape, non-volatile semiconductor memory, or flash memory, and any type of recording medium can be used.

[0037] The recording and playback unit 220 reads and plays back the compressed image signal, compressed audio signal, and audio signal recorded on the recording medium 221. The first control unit 223 transmits the compressed image signal and compressed audio signal read from the recording medium 221 to the image processing unit 207 and the audio processing unit 214. The image processing unit 207 and the audio processing unit 214 temporarily store the compressed image signal and compressed audio signal in the work memory 215, decode them according to a predetermined procedure, and transmit the decoded signals to the video output unit 217.

[0038] The audio input unit 213 includes multiple microphones. The audio processing unit 214 can detect the direction of sound relative to the plane on which the multiple microphones are installed, and the detected information is used for subject search and automatic shooting, which will be described later. The audio processing unit 214 detects specific voice commands. Voice commands are, for example, several commands that have been registered in advance, or commands based on registered voices that the user can register to the camera. The audio processing unit 214 also performs sound scene recognition. In sound scene recognition, the sound scene determination process is performed by a network that has undergone machine learning based on a large amount of audio data in advance. For example, a learning model for detecting specific scenes such as "cheers are being made," "applause is being made," and "voices are being spoken" is set in the audio processing unit 214, and specific sound scenes and specific voice commands are detected. When the audio processing unit 214 detects a specific sound scene or specific voice command, it outputs a trigger signal based on the specific voice recognition to the first control unit 223 and the second control unit 211.

[0039] The second control unit 211 is a separate processor (sub-processor) from the first control unit 223 and controls the power supply to the first control unit 223. The first power supply unit 210 and the second power supply unit 212 supply power to operate the first control unit 223 and the second control unit 211. When the power button on the camera 101 is pressed, power is first supplied to both the first control unit 223 and the second control unit 211. As will be described later, the first control unit 223 also controls the power supply from the first power supply unit 210 to be turned off. Even when the first control unit 223 is not operating, the second control unit 211 is operating and receives detection signals from the camera shake detection unit 209 and the audio processing unit 214. Based on various input information, the second control unit 211 determines whether or not to start the first control unit 223. If it determines that the first control unit 223 should be activated, the second control unit 211 instructs the first power supply unit 210 to supply power to the first control unit 223.

[0040] The audio output unit 218 includes a speaker built into the camera 101 and outputs pre-set audio patterns from the speaker, for example, when shooting.

[0041] The light emission control unit 224 controls the light emission of the LEDs (light-emitting diodes) provided on the camera 101. In addition, the light emission control unit 224 controls the light emission of the LEDs based on a preset lighting pattern or flashing pattern when taking pictures or performing other actions.

[0042] The video output unit 217 includes, for example, a video output terminal and outputs an image signal to display video on an external display connected to the camera 101. The audio output unit 218 and the video output unit 217 may be combined into a single terminal, such as an HDMI® (High-Definition Multimedia Interface) terminal.

[0043] The learning processing unit 219 executes the learning process using the learning model and learning parameters. The learning process can be executed by an image processing processor such as a GPU (Graphics Processing Unit). A GPU is a processor capable of performing a large number of multiply-accumulate operations and has the computational processing power to perform matrix operations of neural networks and other operations in a short amount of time.

[0044] The communication unit 222 includes an interface for communication between the camera 101 and an external device. The communication unit 222 transmits and receives data such as audio signals, image signals, compressed audio signals, and compressed image signals to and from the external device. The communication unit 222 receives commands for starting and ending shooting, as well as control signals related to shooting such as pan, tilt, and zoom, and outputs them to the first control unit 223. This allows the operation of the camera 101 to be controlled based on instructions from the external device. The communication unit 222 also transmits and receives information such as learning models and learning parameters used for learning processing by the learning processing unit 219 between the camera 101 and the external device. The communication unit 222 includes wireless communication modules such as an infrared communication module, a Bluetooth® communication module, a wireless LAN communication module, a WirelessUSB®, and a GPS receiver.

[0045] The environmental sensor 226 detects the state of the surrounding environment of the camera 101 at predetermined intervals. The environmental sensor 226 includes, for example, the following sensors: • Temperature sensor that detects the ambient temperature around camera 101 • Barometric pressure sensor that detects the atmospheric pressure around camera 101 • Illuminance sensor that detects ambient light around camera 101 • Humidity sensor that detects ambient humidity around camera 101 • UV sensor that detects the amount of ultraviolet light around camera 101 The environmental sensor 226 can collect various information detected by each sensor (temperature information, atmospheric pressure information, illuminance information, humidity information, ultraviolet information), and can also calculate the rate of change at predetermined time intervals from this information. The temperature change, atmospheric pressure change, illuminance change, humidity change, and ultraviolet change can be used for judgments such as automatic shooting.

[0046] Next, with reference to Figure 3, the connection relationship between the camera 101 and the external device 301 will be explained.

[0047] Figure 3 illustrates a system configuration in which the camera 101 and the external device 301 are connected wirelessly.

[0048] The camera 101 in this embodiment is, for example, an automatic camera installed at any location. The external device 301 in this embodiment is an information processing device capable of controlling the shooting angles of multiple cameras 101. The information processing device is, for example, a smart device including a wireless communication module such as Bluetooth® or Wi-Fi, but is not limited to this, and may also be a personal computer (notebook PC or tablet PC), a cloud server, etc. Furthermore, the external device 301 in this embodiment may be any of the multiple cameras 101.

[0049] In the example shown in Figure 3, the camera 101 and the external device 301 perform a first communication 302 (solid arrow) and a second communication 303 (dotted arrow). The first communication 302 is, for example, wireless LAN (Local Area Network) communication compliant with the IEEE 802.11 standard. The second communication 303 is a communication with a master-slave relationship, such as a control station and a slave station, such as Bluetooth® Low Energy (BLE). Note that wireless LAN and BLE are just examples of communication methods. If the camera 101 and the external device 301 have two or more communication functions, and it is possible to control one communication function by communicating in a relationship between a control station and a slave station, for example, other communication methods may be used. However, the first communication 302, such as wireless LAN, is capable of faster communication than the second communication 303, such as BLE. Also, the second communication 303 consumes less power or has a shorter communication range than the first communication 302, or at least one of the above.

[0050] Next, the internal configuration of the external device 301 will be described with reference to Figure 4.

[0051] Figure 4 is a block diagram illustrating the internal configuration of the external device 301.

[0052] The wireless LAN control unit 401 performs RF control of the wireless LAN, communication processing, driver processing to control various aspects of wireless LAN communication in accordance with the IEEE 802.11 standard series, and protocol processing related to wireless LAN communication.

[0053] The BLE control unit 402 performs RF control of BLE, communication processing, driver processing to control various aspects of BLE communication, and protocol processing related to BLE communication.

[0054] The public radio control unit 406 performs RF control of public radio communications, communication processing, driver processing for various controls of public radio communications, and protocol processing related to public radio communications. Public radio communications are communications that comply with standards such as IMT (International Multimedia Telecommunications) and LTE (Long Term Evolution).

[0055] The packet transceiver 403 performs at least one of the following: transmission and reception of packets related to wireless LAN and BLE communication, as well as public wireless communication. In this embodiment, the external device 301 is described as performing at least one of the following in communication: transmission and reception of packets. However, other communication formats, such as circuit switching, may be used in addition to packet switching.

[0056] The control unit 411 includes a processor such as a CPU or MPU that performs calculation and control processing for the external device 301, and controls each component of the external device 301 by executing a control program stored in the memory unit 404.

[0057] The memory unit 404 is a memory that stores various information such as the control program executed by the control unit 411 and parameters necessary for communication. The various operations described later are realized by the control unit 411 executing the control program stored in the memory unit 404.

[0058] The GPS (Global Positioning System) receiver 405 receives GPS signals from satellites, analyzes the GPS signals, and estimates the current position (longitude and latitude information) of the external device 301. Alternatively, it estimates the current position of the external device 301 based on information from surrounding wireless networks, such as using WPS (Wi-Fi Positioning System). For example, consider the case where the current GPS information acquired by the GPS receiver 405 is located within a preset position range (within a predetermined radius centered on the detected position), or where there has been a position change greater than a predetermined amount in the GPS information. In this case, the BLE control unit 402 notifies the camera 101 of the movement information, which is then used as parameters for automatic shooting and automatic editing, as described later.

[0059] The display unit 407 has the function of outputting visible information such as an LCD (liquid crystal display device) or LED, or outputting sound such as from a speaker, and presents various information.

[0060] The control unit 408 includes buttons and the like for receiving user input on the external device 301. The display unit 407 and the control unit 408 may be configured as, for example, a touch panel.

[0061] The voice processing unit 409 acquires user voice information, for example, using a microphone built into the external device 301. It may also be configured to identify user-spoken instructions through voice recognition processing. Alternatively, it may be configured to acquire voice commands from user speech using a dedicated application on the external device 301. In this case, specific voice commands for recognition by the camera 101's voice processing unit 214 can be registered via the first wireless LAN communication 302. The power supply unit 410 supplies the necessary power to each part of the external device 301.

[0062] Camera 101 and external device 301 transmit and receive data via wireless LAN control unit 401 and BLE control unit 402. For example, data such as audio signals, image signals, compressed audio signals, and compressed image signals are transmitted and received. In addition, the external device 301 transmits instructions such as shooting commands to camera 101, transmits voice command registration data, transmits notifications of predetermined location detection based on GPS signals, and transmits notifications of location movement. Furthermore, training data is transmitted and received using a dedicated application on the external device 301.

[0063] Figure 5 is a schematic diagram illustrating the appearance of an external device 501 that can communicate with the camera 101.

[0064] The camera 101 in this embodiment is a wearable camera that can be attached to the user's neck, for example. The external device 501 in this embodiment is a wearable device that can be attached to the user's arm, for example. The external device 501 in this embodiment is an information processing device that includes sensors for detecting the user's biometric information and movement state, and can communicate with the camera 101 via a Bluetooth® communication module or the like.

[0065] The external device 501 includes a biometric information detection unit 602. The biometric information detection unit 602 includes a pulse sensor, a heart rate sensor, and a blood flow sensor that detect the user's pulse, heart rate, and blood flow, respectively, and a sensor that detects changes in potential due to skin contact using a conductive polymer. In this embodiment, an example in which the biometric information detection unit 602 is a heart rate sensor is described. The heart rate sensor detects the user's heart rate by irradiating the skin with infrared light using, for example, an LED, and processing the detection signal from a sensor that receives infrared light transmitted through body tissue. The biometric information detection unit 602 outputs the detected biometric information to the control unit 607 (see Figure 6).

[0066] The external device 501 includes a vibration detection unit 603. The vibration detection unit 603 detects the user's movement state. The vibration detection unit 603 is equipped with, for example, an acceleration sensor and a gyroscope sensor, and acquires user movement information and motion detection information. Movement information includes information indicating whether the user is moving or not, and the speed of movement, based on acceleration information. Motion detection information is information indicating that the user is making a motion, such as swinging their arms.

[0067] The external device 501 includes a display unit 604 and an operation unit 605. The display unit 604 is equipped with a display device such as an LCD or LED and outputs visible information. The operation unit 605 receives operation instructions for the external device 501 from the user.

[0068] Figure 6 is a block diagram illustrating the internal configuration of the external device 501.

[0069] The external device 501 includes a control unit 607, a communication unit 601, a biometric information detection unit 602, a vibration detection unit 603, a display unit 604, an operation unit 605, a power supply unit 606, and a storage unit 608.

[0070] The control unit 607 includes a processor such as a CPU or MPU that performs calculation and control processing for the external device 501, and controls each component of the external device 501 by executing a control program stored in the memory unit 608.

[0071] The memory unit 608 is a memory that stores various information such as the control program executed by the control unit 607 and parameters necessary for communication. The various operations described later are realized when the control unit 607 executes the control program stored in the memory unit 608.

[0072] The power supply unit 606 supplies power to each part of the external device 501.

[0073] The operation unit 605 receives operation instructions from the user for the external device 501 and notifies the control unit 607. The operation unit 605 also acquires user voice information from a microphone built into the external device 501, identifies the user's operation commands through voice recognition processing, and notifies the control unit 607. The display unit 604 presents various information to the user through the output of visually recognizable information or through audio output from a speaker or other means.

[0074] The control unit 607 processes the detection signals acquired from the biometric information detection unit 602 and the vibration detection unit 603, and transmits them to the camera 101 via the communication unit 601. For example, the external device 501 transmits detection information to the camera 101 when a change in the user's heart rate is detected. The external device 501 also transmits detection information to the camera 101 when a change in movement state occurs, such as walking, running, or standing still. The external device 501 also transmits detection information to the camera 101 when a preset arm swing motion is detected. The external device 501 also transmits detection information to the camera 101 when a preset distance of movement is detected.

[0075] Next, with reference to Figure 7, the first control process by camera 101 will be described.

[0076] Figure 7 is a flowchart illustrating the first control process performed by the first control unit 223 (main processor) of the camera 101.

[0077] When the user operates the power button on the camera 101, power is supplied from the first power supply unit 210 to the first control unit 223 and each component of the camera 101. Power is also supplied from the second power supply unit 212 to the second control unit 211. Details of the operation of the second control unit 211 will be described later in Figure 8.

[0078] In step S701, the first control unit 223 acquires the startup conditions and proceeds to the process in step S702. In this embodiment, the startup conditions are as follows: (1) When the power is turned on by manually pressing the power button (2) When a startup command is sent from an external device (e.g., external device 301) via communication (e.g., BLE) and the power is started. (3) When the power supply is turned on by an instruction from the second control unit 211 Here, (3) when the power is turned on by instruction from the second control unit 211, the startup condition calculated by the second control unit 211 is used. The startup condition is also used as one of the parameters when performing subject search or automatic shooting. Details of the determination condition will be described later in Figure 8. After the processing of step S701 is performed, the process proceeds to step S702.

[0079] In step S702, the first control unit 223 acquires detection signals from various sensors. The detection signals from the various sensors are as follows: • Detection signals from vibration detection sensors such as gyro sensors and accelerometers in the camera shake detection unit 209. • Detection signals for the rotational positions of the tilt unit 104 and the pan unit 105 • Audio signal detected by the audio processing unit 214, trigger signal from specific speech recognition, and sound direction detection signal • Detection signal of environmental information by environmental sensor 226 After the processing in step S702 is completed, the process proceeds to step S703.

[0080] In step S703, the first control unit 223 determines whether a communication instruction has been sent from the external device, and if a communication instruction has been received, it controls communication with the external device. For example, it executes the process of acquiring various information from the external device 301. This information includes remote operation via wireless LAN or BLE, voice signals, image signals, compressed voice signals, compressed image signals, etc., as well as shooting control information (shooting instructions) from the external device 301, voice command registration data, a predetermined location detection notification based on GPS information, a location movement notification, and learning data. If it is necessary to update biometric information such as the user's movement information, arm action information, and heart rate acquired from the external device 501, the process of acquiring information via BLE is executed. Although an example in which the environmental sensor 226 is mounted on the camera 101 has been described, it may also be mounted on the external device 301 or the external device 501. In that case, in step S703, the process of acquiring environmental information via BLE is performed. After the processing in step S703 is completed, the process proceeds to step S704.

[0081] In step S704, the first control unit 223 determines the operating mode of the camera 101. The operating modes include "automatic shooting mode" (step S710), "automatic editing mode" (step S712), "automatic image transfer mode" (step S714), "learning mode" (step S716), and "automatic file deletion mode" (step S718).

[0082] In step S705, the first control unit 223 determines whether the operating mode determined in step S704 is the low-power mode. The low-power mode is set when the operating mode is not one of the following: "automatic shooting mode", "automatic editing mode", "automatic image transfer mode", "learning mode", or "automatic file deletion mode". If it is determined in step S705 that the operating mode is the low-power mode, the process proceeds to step S706; otherwise, the process proceeds to step S709.

[0083] In step S706, the first control unit 223 notifies the second control unit 211 (subprocessor) of various parameters related to the activation conditions to be determined by the second control unit 211. These parameters include parameters for vibration detection, sound detection, and time elapsed detection, and their values ​​change as a result of the learning process described later.

[0084] When the process in step S706 is completed, the process proceeds to step S707, where the power to the first control unit 223 (main processor) is turned off, and the process ends.

[0085] In steps S709, S711, S713, S715, and S717, the first control unit 223 determines whether the mode determined in step S704 is the automatic shooting mode, automatic editing mode, automatic image transfer mode, learning mode, or automatic file deletion mode.

[0086] In the mode determination process of step S704, one of the following modes (1) to (5) is determined. (1) Automatic shooting mode <Mode determination conditions> The system is designed to determine whether to perform automatic shooting based on the detection information set in the training data, the elapsed time since switching to automatic shooting mode, past shooting information, and the number of photos taken. The detection information includes information such as images, sounds, time, vibrations, location, physical changes, and environmental changes.

[0087] <In-mode processing> If automatic shooting mode is determined in step S709, the process proceeds to automatic shooting mode processing (step S710). Based on the detection information set in the training data, pan, tilt, and zoom are driven, and automatic subject search is performed. When it is determined that the timing is suitable for shooting according to the user's preferences, shooting is performed automatically. (2) Automatic editing mode <Mode determination conditions> The system is designed to determine if automatic editing is appropriate based on the elapsed time since the last automatic editing was performed and the information from previously captured images.

[0088] <In-mode processing> If automatic editing mode is determined in step S711, the process proceeds to automatic editing mode processing (step S712). Based on the learned data, still images and moving images are selected, and based on the learned data, an automatic editing process is performed to generate a highlight video that combines the images into a single video, taking into account image effects and the length of the edited video. (3) Automatic image transfer mode <Mode determination conditions> If the system is set to automatic image transfer mode by instruction from a dedicated application on the external device 301, the condition is that it is determined to perform automatic transfer based on the elapsed time since the last image transfer and the information of previously captured images.

[0089] <In-mode processing> If the automatic image transfer mode is determined in step S713, the process proceeds to the automatic image transfer mode processing (step S714). The camera 101 automatically extracts images preferred by the user and automatically transfers the extracted images to the external device 301. The extraction of images preferred by the user is performed based on a score that determines the user's preference, which is attached to each image as described later. (4) Learning mode <Mode determination conditions> The system is set to automatic learning if it is determined to do so based on the elapsed time since the last learning process, the amount of information integrated into the images that can be used for learning, and the amount of learning data. Alternatively, it is set to learning mode if it receives instructions to set the learning mode via communication with the external device 301.

[0090] <In-mode processing> If the system is determined to be in learning mode in step S715, it proceeds to the learning mode processing (step S716). In learning mode, the system uses a learning model such as a neural network (NN) as input, along with operation information from the external device 301 and notifications of learning information from the external device 301, to perform learning processing tailored to the user's preferences. In this embodiment, the learning processing is machine learning using an NN such as deep learning, and the NN is a Convolutional Neural Network (CNN). The operation information includes image acquisition information from the camera, information manually edited by a dedicated application, and judgment information entered by the user regarding the camera image. Simultaneously, learning processing related to detection, such as registration of personal authentication, voice registration, sound scene registration, and general object recognition registration, as well as learning processing for the conditions of the low-power mode described above, are also performed. (5) Automatic file deletion mode <Mode determination conditions> The condition for performing automatic file deletion is that the system determines whether to delete files based on the elapsed time since the last automatic file deletion and the remaining capacity of the non-volatile memory 216 that stores the image data.

[0091] <In-mode processing> If the system determines in step S717 that it is in automatic file deletion mode, it proceeds to the automatic file deletion mode process (step S718). From the images in the non-volatile memory 216, a process is executed to specify and delete files that should be automatically deleted based on the tag information and date and time the image was taken for each image.

[0092] Once steps S710, S712, S714, S716, and S718 in Figure 7 are completed, the process returns to step S702 and continues. Details of the automatic shooting mode processing in step S710 and the learning mode processing in step S716 will be described later.

[0093] If it is determined in step S709 of Figure 7 that the system is not in automatic shooting mode, the process proceeds to step S711. If it is determined in step S711 that the system is not in automatic editing mode, the process proceeds to step S713. If it is determined in step S713 that the system is not in automatic image transfer mode, the process proceeds to step S715. If it is determined in step S715 that the system is not in learning mode, the process proceeds to step S717. If it is determined in step S717 that the system is not in automatic file deletion mode, the process returns to step S702 and is repeated. Detailed explanations of automatic editing mode, automatic image transfer mode, and automatic file deletion mode are omitted.

[0094] Next, with reference to Figure 8, the second control process by camera 101 will be described.

[0095] Figure 8 is a flowchart illustrating the second control process performed by the second control unit 211 (subprocessor) of the camera 101.

[0096] When the user operates the power button on the camera 101, power is supplied from the first power supply unit 210 to the first control unit 223 and each component of the camera 101. Power is also supplied from the second power supply unit 212 to the second control unit 211.

[0097] After power is supplied, the second control unit 211 starts up, and the process shown in Figure 8 begins.

[0098] In step S801, the second control unit 211 waits until a predetermined sampling period has elapsed, and once the predetermined sampling period has elapsed, it proceeds to the process in step S802. The predetermined sampling period is set to, for example, a period of 10 milliseconds.

[0099] In step S802, the second control unit 211 acquires learning information. The learning information is the information transmitted from the first control unit 223 to the second control unit 211 in step S706 of Figure 7, and includes, for example, information used for the following determinations. (1) Information for determining specific shaking state (step S804) (2) Information for determining specific sound detection (step S805) (3) Information for determining the passage of time (step S807) In step S803, the second control unit 211 acquires shake detection information. The shake detection information consists of detection signals from the gyro sensor and acceleration sensor of the camera shake detection unit 209.

[0100] In step S804, the second control unit 211 detects a predetermined specific shaking state. Here, we will describe several examples in which the determination process is modified based on the learning information acquired in step S802.

[0101] <Tap detection> A tap state is, for example, when a user taps the camera 101 with their fingertip, and can be detected from the output value of an accelerometer attached to the camera 101. The output of the 3-axis accelerometer is processed by passing it through a bandpass filter (BPF) set to a specific frequency range at a predetermined sampling period, and the signal domain component of the acceleration change due to the tap is extracted. The number of times the acceleration signal after passing through the bandpass filter (BPF) exceeds a predetermined threshold ThreshA within a predetermined time TimeA is measured. A tap is determined based on whether the measured number is a predetermined number CountA or not. For example, in the case of a double tap, the value of the predetermined number CountA is set to 2, and in the case of a triple tap, the value of the predetermined number CountA is set to 3. The values ​​of the predetermined time TimeA and the predetermined threshold ThreshA can also be changed based on learned information.

[0102] <Detection of shaking state> The shaking state of camera 101 can be detected from the output values ​​of the gyro sensor and accelerometer attached to camera 101. The output of the gyro sensor and accelerometer is filtered by a high-pass filter (HPF) to remove high-frequency components and a low-pass filter (LPF) to remove low-frequency components, and then converted to an absolute value. The number of times the calculated absolute value exceeds a predetermined threshold ThreshB during a predetermined time TimeB is measured. Vibration detection is performed based on whether the number of measured counts is greater than or equal to a predetermined number CountB. For example, it is possible to determine whether camera 101 is placed on a table or the like, i.e., the shaking is small, or whether camera 101 is worn on the user's body as a wearable camera and walking, i.e., the shaking is large. Furthermore, by setting multiple conditions regarding the judgment threshold and the judgment count, it is possible to detect a detailed shaking state according to the shaking level. The values ​​of the predetermined time TimeB, predetermined threshold ThreshB, and predetermined count CountB can be changed using learned information.

[0103] The above example describes a method for detecting a specific shaking state by determining the detected value of a shaking sensor. Alternatively, there is a method for detecting a specific shaking state that has been pre-registered by inputting data from a shaking sensor acquired at a predetermined sampling period into a shaking state determination device using a neural network (NN) and training a model. In this case, step S802 (acquisition of training information) only involves acquiring the weight coefficients of the NN.

[0104] In step S805, the second control unit 211 performs a detection process for a pre-set specific sound. In this embodiment, several examples of modifying the detection process based on the learning information acquired in step S802 will be described.

[0105] <Detection of specific voice commands> In the process of detecting specific voice commands, these specific voice commands include several pre-registered commands and commands based on specific voices registered by the user with the camera.

[0106] <Specific Sound Scene Recognition> Based on a large amount of pre-programmed audio data, a machine learning network determines the sound scene. For example, it can detect specific scenes such as "cheers," "applause," or "speaking." The scenes targeted for detection change through learning.

[0107] <Sound level determination> Sound level detection is performed by determining whether the volume of the sound exceeds a predetermined level (threshold) for a predetermined period of time (threshold time). The threshold time and threshold may change through learning.

[0108] <Sound direction determination> Multiple microphones arranged on a plane detect the direction of a sound of a predetermined magnitude.

[0109] The audio processing unit 214 performs the above determination process, and based on the settings learned in advance, the second control unit 211 determines in step S805 whether or not a specific sound has been detected.

[0110] In step S806, the second control unit 211 determines whether the power supply to the first control unit 223 is off. If it is determined that the first control unit 223 is off, the process proceeds to step S807. If it is determined that the first control unit 223 is on, the process proceeds to step S811.

[0111] In step S807, the second control unit 211 performs a detection process to detect the passage of a preset time. In this embodiment, the detection process is modified based on the learning information acquired in step S802. The learning information is the information transmitted from the first control unit 223 to the second control unit 211 in step S706 of Figure 7. Here, the elapsed time from the point when the power supply of the first control unit 223 transitioned from the ON state to the OFF state is measured, and if the elapsed time is greater than or equal to a predetermined time TimeC, it is determined that the predetermined time has elapsed. Also, if the measured elapsed time is shorter than the predetermined time TimeC, it is determined that the predetermined time has not elapsed. The predetermined time TimeC is a parameter that changes based on the learning information.

[0112] In step S808, the second control unit 211 determines whether the conditions for disabling the low-power mode have been met. Disabling the low-power mode is determined by the following conditions. (1) A specific type of tremor was detected. (2) A specific sound was detected. (3) The prescribed time has elapsed. (1) In step S804 (Specific vibration state detection process), it is determined whether a specific vibration has been detected. (2) In step S805 (Specific sound detection process), it is determined whether a specific sound has been detected. (3) In step S807 (Time elapsed detection process), it is determined whether a predetermined time has elapsed. If at least one of the conditions shown in (1) to (3) is met, it is determined that the conditions for disabling low power mode have been met. If it is determined in step S808 that the conditions for disabling low power mode have been met, the process proceeds to step S809. If it is determined that the conditions for disabling low power mode have not been met, the process returns to step S801 and continues.

[0113] In step S809, the second control unit 211 turns on the power to the first control unit 223.

[0114] In step S810, the second control unit 211 notifies the first control unit 223 of the conditions (vibration, sound, or time) that have been determined to cause the low-power mode to be deactivated, and returns to step S801 to continue processing.

[0115] On the other hand, if the process moves from step S806 to step S811 (when it is determined that the power supply to the first control unit 223 is turned on), the process proceeds to step S811.

[0116] In step S811, the second control unit 211 notifies the first control unit 223 of the detection information acquired in steps S803 to S805, and returns to step S801 to continue processing.

[0117] In this embodiment, even when the power supply of the first control unit 223 is turned on, the second control unit 211 performs vibration detection and specific sound detection and notifies the first control unit 223 of the detection results, but this is not the only example. If the power supply of the first control unit 223 is turned on, the processing in steps S803 to S805 may be omitted, and the first control unit 223 may perform the detection of specific vibration or specific sound in its processing (step S702 in Figure 7).

[0118] As described above, by performing steps S704 to S707 in Figure 7 and the processing shown in Figure 8, the conditions for switching to low-power mode and the conditions for deactivating low-power mode are learned based on user operation. In other words, it becomes possible to operate the camera in a way that suits the user's preferences. The learning method will be described later.

[0119] In the example shown in Figure 8, a method for disabling low-power mode is described based on vibration detection, sound detection, and time elapsed. However, low-power mode may also be disabling based on environmental information. Disabling can be determined based on whether the absolute values ​​or changes in temperature, atmospheric pressure, illuminance, humidity, and ultraviolet radiation exceed predetermined thresholds. These thresholds can be changed through learning, as described later. Alternatively, the detection information from vibration detection, sound detection, and time elapsed, as well as the absolute values ​​and changes in each environmental information, can be determined through a learning process to decide whether to disabling low-power mode. In this determination process, the determination conditions can be changed through a learning process, as described later.

[0120] Next, with reference to Figure 9, the automatic shooting mode processing in step S710 of Figure 7 will be described.

[0121] In step S901, the first control unit 223, using the image processing unit 207, performs image processing on the image signal generated by the image sensor 206 to generate image data for subject detection. The image processing unit 207 also performs subject recognition processing to detect people, general objects, etc., from the generated image data.

[0122] When detecting a person, the subject's face and body are detected. In face detection processing, areas in image data that match a predetermined pattern for determining a person's face can be detected as the face region. At the same time, a confidence score indicating the likelihood that the subject is a face is calculated. The confidence score is calculated, for example, from the size of the face region in the image data and the degree of agreement with the face pattern. General object recognition is performed in a similar manner, and objects that match a pre-registered pattern can be detected.

[0123] Another method involves extracting characteristic information about a subject using histograms of hue and saturation from image data. For images of subjects within the field of view, the distribution derived from the histograms of hue and saturation is divided into multiple intervals, and the image is classified according to each interval. For example, histograms of multiple color components are generated for an image, and the image is divided according to the distribution range of the histograms. The image is classified within the regions belonging to the same combination of intervals, and the image region of the subject is recognized. By calculating an evaluation value for each recognized image region of the subject, the image region of the subject with the highest evaluation value can be determined as the main subject region. By the method described above, characteristic information about various subjects can be obtained from an image.

[0124] In step S902, the first control unit 223 calculates the image blur correction amount. The image blur correction amount is calculated by determining the absolute angle of camera shake based on the angular velocity and acceleration information detected by the camera shake detection unit 209, and then calculating the angle at which the camera is tilted / panned in an angular direction that cancels out the absolute angle to correct the image blur. Note that the calculation method for the image blur correction amount can be changed by the learning process described later.

[0125] In step S903, the first control unit 223 determines the state of the camera. Based on the camera's angle and movement detected from angular velocity information, acceleration information, GPS information, etc., the first control unit 223 determines the current state of vibration / movement of the camera. For example, consider the case where the camera 101 is mounted on a vehicle and taking pictures. In this case, subject information such as the surrounding scenery changes significantly depending on the distance the vehicle moves, so it is determined whether the camera 101 is in a "vehicle movement state" where it is mounted on a vehicle and moving at high speed, and the determination result is used in the subject search process described later. It is also determined whether the angle of the camera 101 has changed significantly. It is determined whether the camera 101 is in a "stationary shooting state" where there is almost no shaking, and if it is in a "stationary shooting state", it can be determined that there has been no change in the position of the camera 101. In this case, subject search processing for stationary shooting can be performed. If the angle change of the camera 101 is relatively large, it is determined to be in a "handheld state". In this case, subject search processing for handheld shooting can be performed.

[0126] In step S904, the first control unit 223 performs subject search processing. Subject search includes the following processing. (1) Area division (2) Calculation of importance for each area (3) Determining the area to be searched The following explains each of the processes (1) to (3) in order.

[0127] (1) Area division Refer to Figure 10 to explain area division. The origin O of the 3D Cartesian coordinate system is defined as the camera position. Figure 10(a) is a schematic diagram showing an example of dividing the spherical, omnidirectional shooting range centered on the camera position (origin O) into areas. In the example in Figure 10(a), the omnidirectional shooting range is divided into areas of 22.5 degrees each in the tilt direction and pan direction. In this type of area division, as the tilt angle moves away from 0 degrees, the horizontal circumference decreases and the area decreases. Figure 10(b) shows an example where the horizontal area division is set to a larger value than 22.5 degrees when the tilt angle is 45 degrees or more. Figures 10(c) and 10(d) illustrate the divided areas within the field of view. The axis 1001 shown in Figure 10(c) indicates the optical axis direction (shooting direction) of the camera 101, and area division is performed with the direction of axis 1001 as the reference direction. Area 1002 indicates the field of view (shooting range) of the image, and Figure 10(d) shows an example of an image corresponding to Area 1002. As shown in Figure 10(d), the image of Area 1002 is divided into multiple areas 1003 to 1018.

[0128] (2) Calculation of importance for each area For each divided area, an importance level is calculated that indicates the priority of the search, based on the situation of the subjects and the scene within the area. The importance level based on the situation of the subjects is calculated based on, for example, the number of people in the area, the size of their faces, the orientation of their faces, the certainty of face detection, their facial expressions, and the results of personal identification of the people. The importance level based on the scene is calculated based on, for example, the results of general object recognition, the results of scene discrimination (blue sky, backlight, sunset, etc.), the level of sound detected from the direction of the area, the results of speech recognition, and movement information within the area.

[0129] Furthermore, if camera vibration is detected in step S903 of Figure 9, the system may be configured to change the importance level depending on the vibration state. For example, consider the case where it is determined to be in a "placed-on shooting state." In this case, the system determines that the subject search will focus on subjects with high priority among those registered by face recognition (e.g., the camera owner). Also, in the case of automatic shooting described later, the camera will prioritize capturing the face of the camera owner, for example. This allows the camera owner to record many images of themselves even if they spend a long time carrying and shooting with the camera attached to their body, by removing the camera and placing it on a table or other surface. In this case, since face search is possible by pan and tilt movement, the owner does not need to consider the camera angle or other settings when setting it up; they can capture images of themselves or many faces simply by setting it up appropriately.

[0130] However, if the above conditions alone are not met, the area with the highest importance may remain the same unless there are changes in each area, in which case the area being searched would never change. Therefore, a process is performed to change the importance according to past shooting information. Specifically, a process is performed to lower the importance of areas that have been continuously designated as search targets for a predetermined period of time, and a process is performed to lower the importance of areas that were photographed in step S910, which will be described later, for a predetermined period of time.

[0131] (3) Determining the area to be searched Based on the importance of each area calculated as described above, a process is executed to determine the areas with high importance as the search target areas. Then, the target pan and tilt angles required to fit the search target areas within the field of view are calculated.

[0132] Returning to the explanation of Figure 9, in step S905, the first control unit 223 drives the pan and tilt. The first control unit 223 calculates the pan drive amount and tilt drive amount by adding the image blur correction amount and the angle based on the target angle of pan and tilt at a predetermined sampling period. The pan / tilt drive control unit 205 drives the tilt unit 104 and the pan unit 105 based on the pan drive amount and tilt drive amount.

[0133] In step S906, the first control unit 223 performs zoom drive by controlling the zoom unit 201. The first control unit 223 performs zoom drive according to the state of the subject to be searched determined in step S904. For example, suppose the subject to be searched is a person's face. In this case, if the face size in the image is too small, it may fall below the minimum detectable size and become undetectable, potentially causing the subject to be lost. In such cases, zoom control toward the telephoto side is performed to enlarge the face size in the image. On the other hand, if the face size in the image is too large, the subject may easily move out of the field of view due to the movement of the subject or the camera itself. In such cases, zoom control toward the wide-angle side is performed to reduce the size of the face on the screen. By performing zoom control in this way, a state suitable for tracking the subject can be maintained. Note that there are two types of zoom control: optical zoom control, which is performed by driving the lens, and electronic zoom control, which changes the field of view through image processing. Furthermore, there are methods that perform only one of these controls, and methods that combine both controls.

[0134] In step S907, the first control unit 223 determines whether or not it has received a manual shooting instruction from the user. Manual shooting instructions include instructions by pressing the shutter button, instructions by lightly tapping the camera housing with a finger, instructions by voice command input, and instructions from an external device. For example, a shooting instruction triggered by a tap is determined when the camera shake detection unit 209 detects a continuous high-frequency acceleration in a short period of time when the user taps the camera housing. Another method of shooting instructions by voice command is to have the voice processing unit 214 recognize the voice when the user utters a predetermined password to instruct shooting (for example, "Take a picture") and use that as a trigger to start shooting. An example of an instruction from an external device is to trigger the camera with a shooting instruction sent from a dedicated application such as a smartphone wirelessly connected to the camera.

[0135] If a manual shooting instruction is received in step S907, proceed to step S910. If it is determined that no manual shooting instruction was received in step S907, proceed to step S908.

[0136] In step S908, the first control unit 223 performs an automatic shooting determination process. The automatic shooting determination process determines whether or not to perform automatic shooting and determines the shooting method (which of still image shooting, video shooting, continuous shooting, panoramic shooting, etc., to perform).

[0137] Automatic shooting is a process that automatically records image data generated by the image sensor 206. The decision to perform automatic shooting is made in the following two cases. In the first case, the importance level for each area obtained in step S904 exceeds a predetermined value. In the second case, the determination result by the neural network (NN) is applied, and the second case will be described later. The recording of automatically captured images includes recording the image data in the work memory 215, recording the image data in the non-volatile memory 216, or automatically transferring the image data to the external device 301 and recording the image data in the external device 301.

[0138] In this embodiment, camera 101 automatically takes a picture based on an automatic shooting determination process using a neural network (NN). Depending on the shooting location and camera conditions, it may be better to change the parameters of the automatic shooting determination process. Unlike shooting at fixed time intervals, automatic shooting that responds to the camera conditions tends to be preferred in a form that meets the user's shooting intentions, such as the following. (1) I want to take a large number of pictures, including people and objects. (2) I don't want to miss capturing memorable moments. (3) I want to take photos in a power-saving manner, taking into consideration the remaining battery life and storage capacity. Automatic shooting is performed by calculating an evaluation value from the subject's state, comparing it to a threshold, and only executing if the evaluation value exceeds the threshold. The evaluation value for automatic shooting is determined by a judgment process using a neural network (NN).

[0139] Figure 11 illustrates the configuration of a neural network using a multilayer perceptron. The NN is used to predict output values ​​from input values. By pre-training the NN with input values ​​and exemplary output values ​​for those input values, it can estimate output values ​​that follow the learned examples for new input values. The training method will be described later.

[0140] Figure 11 shows that node 1101 and the multiple nodes indicated by the vertically aligned circles represent neurons in the input layer. Node 1103 and the multiple nodes indicated by the vertically aligned circles represent neurons in the hidden layer. Node 1104 represents a neuron in the output layer. Arrow 1102 indicates the connections between each neuron. In the NN-based decision processing, the neurons in the input layer are given feature quantities as input, based on the subject in the current field of view, the scene, and the state of the camera. The value output from the output layer is obtained after calculations based on the forward propagation rule of the multilayer perceptron. If the output value is above a threshold, it is determined that automatic shooting should be performed.

[0141] The following information, for example, is used as the feature vector for the subject: • Recognition results of general objects at the current zoom level and field of view. • Face detection results: number of faces in the current field of view, degree of smile, degree of eye closure, face angle, face recognition ID number, and angle of gaze of the person. • Scene detection results, elapsed time since last shooting, current time, GPS information, and change in location from the last shooting location. • Information on the current sound level, the person making the sound, and whether or not there is applause or cheering. • Vibration information (acceleration information, camera status), environmental information (temperature, atmospheric pressure, illuminance, humidity, UV radiation level), etc. Furthermore, if information is notified from an external device 501 or the like, the notification information (such as user movement information, arm action information, and biometric information such as heart rate) is also used as feature information. The feature information is converted into numerical values ​​within a predetermined range and provided as feature quantities to each neuron in the input layer. Therefore, each neuron in the input layer is required for each feature quantity used.

[0142] In automated image detection using neural networks (NNs), the output value can be changed by altering the weight coefficients that adjust the connectivity between each neuron during the learning process described later, thereby adapting the detection result to the learning result.

[0143] Furthermore, the automatic shooting determination also changes depending on the activation conditions of the first control unit 223 acquired in step S702 of Figure 7. For example, in the case of activation by tap detection or activation by a specific voice command, the activation conditions are set to increase the shooting frequency because it is highly likely that the user's intention is to instruct the camera to take a picture at that moment.

[0144] In determining the shooting method, automatic shooting is determined based on the state of camera 101 and the state of surrounding subjects detected in steps S901 to S904. For example, it is determined whether to take a still image, video, continuous shooting, or panoramic shot. For example, if the subject (a person) is stationary, still image shooting is selected and executed. If the subject is moving, video shooting or continuous shooting is executed. In addition, if multiple subjects are present surrounding the camera, or if it is determined to be a scenic spot based on GPS information, panoramic shooting is executed. Panoramic shooting is a process that generates a panoramic image by sequentially taking images while driving pan and tilt and then combining them. Similar to the automatic shooting determination method, it is also possible to input various information detected before shooting into the NN and determine the shooting method. Furthermore, in the case of determination using the NN, the determination conditions can be changed by the learning process described later.

[0145] Returning to the explanation in Figure 9, if it is determined in step S909 that automatic shooting should be performed, the process proceeds to step S910. If it is determined in step S909 that automatic shooting should not be performed, the shooting mode processing is terminated.

[0146] In step S910, the first control unit 223 starts automatic shooting. The first control unit 223 starts shooting using the shooting method determined in step S908. In this case, the focus drive control unit 204 performs autofocus control. In addition, exposure control is performed using an aperture control unit, a sensor gain control unit, and a shutter control unit (not shown) to adjust the brightness of the subject to an appropriate level. The image processing unit 207 performs various known image processing operations on the captured image, such as auto white balance processing, noise reduction processing, and gamma correction processing, to generate image data.

[0147] If predetermined conditions are met during automatic shooting in step S910, the camera 101 may notify the person being photographed that shooting is about to take place before proceeding with the shooting. The predetermined conditions are set based on, for example, the following information. • Number of faces within the field of view, degree of smile, degree of eye closure, angle of gaze and face, facial recognition ID number • Number of people registered through personal authentication, general object recognition results at the time of shooting • Scene classification results, elapsed time since the last photo was taken, time of shooting, and information on whether the current location is a scenic spot based on GPS data. • Information regarding the sound level during filming, whether or not there are people speaking, and whether or not there is applause or cheering. • Vibration information (acceleration information, camera status), environmental information (temperature, atmospheric pressure, illuminance, humidity, UV radiation level), etc. Notification methods include, for example, sound output from the audio output unit 218 or LED illumination by the light emission control unit 224. By performing shooting with notification based on these conditions, it is possible to record desirable images with the subject looking directly at the camera in high-priority scenes. For pre-shoot notification, the notification method and timing can be determined by inputting information from the captured image or various information detected before shooting into the neural network (NN). Furthermore, the determination conditions can be changed through the learning process described later.

[0148] In step S911, the first control unit 223 performs editing processing on the image data generated in step S910 using the image processing unit 207, such as processing the image data or adding it to the video. Image processing includes cropping based on people's faces and focus positions, image rotation, HDR (High Dynamic Range) effects, bokeh effects, and color conversion filter effects. Image processing may also generate multiple image data from the image data generated in step S910 by combining the above editing processes and save them separately from the captured image. For video processing, captured video or still images may be added to the generated edited video while applying special effects such as slide, zoom, and fade. For the editing processing in step S911, the image processing method can also be determined by inputting information from the captured image or various information detected before shooting into the NN. Furthermore, the determination conditions can be changed by the learning process described later.

[0149] In step S912, the first control unit 223 generates and records training data from the image data generated in step S910, which will be used for the training process described later. The training data includes, for example, the following information. • Data collected on the images taken: zoom magnification at the time of shooting, general object recognition results at the time of shooting, face detection results, number of faces in the image, degree of smile, degree of eye closure, face angle, face recognition ID number, and gaze angle of the person. • Scene classification results, elapsed time since last shooting, shooting time, GPS information, and change from the previous shooting location. • Information regarding the sound level during filming, the people making the noise, and whether or not there was applause or cheering. • Vibration information (acceleration information, camera status), environmental information (temperature, atmospheric pressure, illuminance, humidity, UV radiation level). • Information such as the video recording time and whether or not it was recorded manually. Furthermore, a score is calculated, which is the output of the neural network (NN) that quantifies the user's preferred images. This information is then generated and recorded as tag information in the captured image file. Alternatively, it can be stored in the non-volatile memory 216, or saved in the recording medium 221 in a format that lists the information of each captured image as so-called catalog data.

[0150] In step S913, the first control unit 223 updates past shooting information. The first control unit 223 updates the number of images taken per area, the number of images taken per person registered by personal authentication, the number of images taken per subject detected by general object recognition, and the number of images taken per scene for scene discrimination, as described in step S908. The first control unit 223 increases the number of images taken this time by one, and at the same time stores the time of this shooting and the evaluation value of automatic shooting, and retains them as shooting history information. After the processing in step S913 is completed, the processing shown in Figure 9 is terminated.

[0151] Next, we will explain the learning process tailored to the user's preferences.

[0152] In this embodiment, the learning processing unit 219 performs learning processing tailored to the user's preferences using machine learning with the neural network learning model shown in Figure 11. The NN is used for inference processing to predict output values ​​from input values, and by learning the actual input values ​​and actual output values ​​in advance, it can estimate output values ​​for new input values. By using the NN, the operation of the above-mentioned automatic shooting, automatic editing, and subject search can be learned to match the user's preferences. In addition, subject information (results such as face recognition and general object recognition), which also serves as feature data to be input to the NN, is registered, and shooting notification control, low power mode control, and automatic file deletion are also modified through the learning process.

[0153] In this embodiment, the operation to which the learning process is applied is illustrated below. (1) Automatic shooting (2) Automatic editing (3) Subject search (4) Subject registration (5) Shooting notification control (6) Low power mode control (7) Automatic file deletion (8) Image blur correction (9) Automatic image transfer The explanations for (2) automatic editing, (7) automatic file deletion, and (9) automatic image transfer, which are among the operations to which the learning process is applied, will be omitted.

[0154] <Automatic shooting> This section explains the learning process for automatic shooting. In automatic shooting, a learning process is performed to automatically capture images that match the user's preferences. After the shooting in step S910 in Figure 9, a learning information generation process (step S912) is performed. This is a process in which images for training are selected using a method described later, and learning is performed by changing the weight coefficients of the neural network (NN) based on the learning information contained in the images. The learning process is performed by changing the NN that determines the timing of automatic shooting and by changing the NN that determines the shooting method (still image shooting, video shooting, continuous shooting, panoramic shooting, etc.).

[0155] <Subject search> This section describes the learning process for subject search. In subject search, a learning process is performed to automatically search for subjects that match the user's preferences. In the subject search process shown in Figure 9 (step S904), the importance of each area is calculated, and subject search is performed by driving pan, tilt, and zoom. The learning process is performed based on the captured image and detection information during the search, and the learning results are reflected by changing the weight coefficients of the neural network (NN). By inputting various detection information during the search operation into the NN and determining its importance, subject search can be performed while reflecting the learning results. In addition to calculating importance, the search method (speed, frequency of movement) is also controlled by driving pan and tilt.

[0156] <Subject Registration> This section describes the learning process for subject registration. Subject registration involves a learning process to automatically register and rank subjects that match the user's preferences. The learning process includes, for example, facial recognition registration, general object recognition registration, gesture and voice recognition registration, and sound-based scene recognition registration. Authentication registration is performed for both people and objects, and ranking is set based on the number and frequency of image acquisition, the number and frequency of manual shooting, and the frequency of appearance of subjects during the search. All of this information is registered as input for judgment by the neural network.

[0157] <Shooting notification control> The learning process for the shooting notification will now be explained. Immediately before shooting in step S910 of Figure 9, if predetermined conditions are met, the camera 101 notifies the person to be photographed that it will take a picture, and then takes the picture. For example, processes are performed to visually guide the subject's gaze by driving the pan and tilt, or to attract the subject's attention using speaker sound emitted from the audio output unit 218 or LED light from the light emission control unit 224. Immediately after the notification, it is determined whether or not to use the detected information (e.g., degree of smile, eye contact detection, gesture) for the learning process based on whether or not the detected information has been acquired, and the learning process is performed by changing the weight coefficients of the neural network.

[0158] Each detection piece immediately before shooting is input into the neural network (NN) to determine whether or not to issue a notification. For a notification sound, the sound level, type and timing of the sound are determined, and for a notification light, the illumination time, speed, and camera orientation (pan and tilt) are determined.

[0159] <Low-power mode control> As explained using Figures 7 and 8, the power of the first control unit 223 (main processor) is turned on / off. A learning process is performed for the conditions for releasing the low-power mode and the conditions for transitioning to the low-power mode. First, the learning process for the conditions for releasing the low-power mode will be explained.

[0160] • Sound detection The learning process can be performed by manually setting the user's specific voice, specific sound scenes to be detected, and specific sound levels, for example, through communication using a dedicated application on an external device 301. Alternatively, multiple detection methods can be pre-configured in the voice processing unit, and the image to be trained can be selected using the method described later. The learning process can be performed by learning the information of the surrounding sounds contained in the selected image and setting the sound judgment to be used as the trigger (specific sound command or sound scenes such as "cheers" or "applause").

[0161] • Environmental information detection The system can be trained by manually setting environmental information changes that the user wants to use as activation conditions, for example, through communication using a dedicated application on an external device 301. For example, specific conditions such as the absolute amount or change in temperature, atmospheric pressure, illuminance, humidity, and ultraviolet radiation can be set, and the imaging device can be activated when these conditions are met. Judgment thresholds based on each environmental information can also be learned. If the camera detection information after activation based on the environmental information determines that the activation conditions are not met, the parameters of each judgment threshold are set to make it difficult to detect environmental changes.

[0162] Furthermore, each of the above parameters changes with the battery level. For example, when the battery level is low, it becomes more difficult to proceed to various judgments, and when the battery level is high, it becomes easier to proceed to various judgments. Specifically, even if the shaking state detection result or sound scene detection result is not a factor that the user intended to activate the camera, the camera may still be activated if the battery level is high.

[0163] Furthermore, the conditions for deactivating low-power mode can also be determined by inputting information such as vibration detection information, sound detection information, time elapsed detection information, various environmental information, and battery level into the neural network (NN). In this case, an image for training processing is selected using the method described later, and training processing is performed by changing the weight coefficients of the NN based on the training information contained in the image.

[0164] Next, we will explain the learning of the conditions for transitioning to low-power mode. In step S704 of Figure 7, the mode determination determines that the device is not in "automatic shooting mode," "automatic editing mode," "automatic image transfer mode," "learning mode," or "automatic file deletion mode," and the device transitions to low-power mode. The determination conditions for each mode also change through the learning process.

[0165] <Automatic shooting mode> The camera determines the importance of each area and automatically takes photos while searching for the subject by driving the pan and tilt. If it is determined that there is no subject to be photographed, the automatic shooting mode is canceled. For example, if the total importance of all areas, or the sum of the importance of each area, falls below a predetermined threshold, the automatic shooting mode is canceled. In this case, the predetermined threshold is lowered based on the elapsed time since the camera entered automatic shooting mode. As the elapsed time since the camera entered automatic shooting mode increases, it becomes easier to switch to low-power mode.

[0166] Furthermore, by changing a predetermined threshold based on the remaining battery level, low-power control that takes into account the usable battery life can be performed. For example, when the battery level is low, the threshold is increased to make it easier to switch to low-power mode, and when the battery level is high, the threshold is decreased to make it more difficult to switch to low-power mode. Here, the parameter for the next low-power mode deactivation condition (elapsed time threshold TimeC) is set for the second control unit 211 based on the elapsed time and the number of shots taken since the last time the automatic shooting mode was activated. Each of the above thresholds changes through a learning process. The learning process is performed, for example, by manually setting the shooting frequency and activation frequency through communication using a dedicated application on an external device 301.

[0167] Alternatively, the system may accumulate average time data for the period from when the camera 101's power button is turned on to when the power button is turned off, as well as distribution data for each time period, to learn each parameter. In this case, for users whose elapsed time from power-on to power-off is short, the time interval for returning from low-power mode or transitioning to low-power mode will be shortened through the learning process. Conversely, for users whose elapsed time from power-on to power-off is long, the time interval will be lengthened through the learning process.

[0168] Learning processing is also performed on detection information during subject search. While it is determined that there are many important subjects, the time interval for returning from low-power mode and transitioning to low-power mode is shortened by the learning process. Conversely, while it is determined that there are few important subjects, the time interval is lengthened by the learning process.

[0169] <Image blur correction> The learning process for image blur correction is described below. In step S902 of Figure 9, the amount of image blur correction is calculated, and in step S905, pan and tilt are driven based on the amount of image blur correction. Image blur correction involves a learning process to correct the image blur according to the characteristics of camera shake 101 caused by the user. For example, by using a PSF (Point Spread Function), it is possible to estimate the direction and magnitude of blur in the captured image. In the generation of learning information in step S912 of Figure 9, the estimated direction and magnitude of blur is added to the image data.

[0170] In the learning mode processing of step S716 in Figure 7, a process is performed to learn the weight coefficients of the neural network for image blur correction based on predetermined input information and output (estimated direction and magnitude of blur). The predetermined input information includes, for example, each detection information at the time of shooting (motion vector information of the image at a predetermined time before shooting, motion information of the detected subject (person or object), and vibration information (gyro output, acceleration output, camera status). Furthermore, environmental information (temperature, atmospheric pressure, illuminance, humidity), sound information (sound scene determination, specific sound detection, sound level change), time information (elapsed time since startup, elapsed time since the last shooting), and location information (GPS information, amount of positional movement) may also be added as input.

[0171] In step S902 of Figure 9, when calculating the image blur correction amount, the magnitude of blur during shooting can be estimated by inputting the above detection information into the neural network (NN). If the estimated magnitude of blur is greater than the threshold, it becomes possible to control the shutter speed, for example. Also, if the estimated magnitude of blur is greater than the threshold, there is a possibility that a blurred image will be acquired, so methods such as prohibiting shooting can be used.

[0172] Furthermore, since there are limitations on the pan and tilt drive angles, further image blur correction cannot be performed after reaching the drive end. In this embodiment, by estimating the magnitude and direction of blur during shooting, it is possible to estimate the range required for pan and tilt drive to correct image blur during exposure. Regarding the pan and tilt drive angles, if there is insufficient room in the movable range during exposure, the cutoff frequency of the filter that calculates the amount of image blur correction is increased to set the drive angle so that it does not exceed the movable range. This makes it possible to suppress large blurs. Also, if it is predicted that the drive angle will exceed the movable range, the drive angle is changed immediately before exposure, and the camera rotates in the opposite direction to the direction in which the drive angle will exceed the movable range, and exposure is started. This makes it possible to shoot with suppressed image blur while ensuring the movable range. By learning the user's shooting characteristics and usage, image blur correction can be suppressed or prevented in captured images.

[0173] In determining the shooting method of this embodiment, a panning shot determination process may be performed. In panning shots, images are taken so that there is no blur on the moving subject, and the image appears blurred against a stationary background. In the determination process for whether or not to perform a panning shot, the pan and tilt drive speeds required to photograph the subject without blur are estimated from the detection information up to the point before shooting, and image blur correction of the subject is performed. In this case, the drive speed can be estimated by inputting the information into a trained neural network using the above-mentioned detection information. By dividing the image into multiple blocks and estimating the PSF of each block, the direction and magnitude of blur in the block where the main subject is located are estimated. A learning process is performed based on this information.

[0174] Furthermore, the system can learn the amount of background blur from information in images selected by the user. In this case, the degree of blur in blocks (image areas) where the main subject is not located is estimated, and the system can learn the user's preferences based on this estimated information. By setting the shutter speed during shooting based on the learned preferred amount of background blur, the system can automatically take photos that produce a panning effect that matches the user's preferences.

[0175] Next, we will explain the learning methods. Learning methods include learning processes within the camera and learning processes in cooperation with external devices. First, we will explain the learning processes within the camera. The learning processes within the camera in this embodiment include the following methods. (1) Learning process based on detection information during manual shooting As explained in Figure 9, camera 101 can perform both manual and automatic shooting. If manual shooting is instructed in step S907, in step S912, information indicating that the image was taken manually is added to the captured image. Also, if automatic shooting is determined to be on in step S909 and shooting is performed, in step S912, information indicating that the image was taken automatically is added to the captured image. In the case of manual shooting, it is highly likely that the shooting was performed based on the user's preferred subject, preferred scene, preferred location, and preferred time interval. Therefore, learning processing is performed based on training data such as each feature data and captured image data obtained during manual shooting. Furthermore, learning processing is performed from the detection information during manual shooting regarding the extraction of feature quantities in the captured image, registration of personal authentication, registration of facial expressions for each individual, and registration of combinations of people. In addition, from the detection information during subject search, for example, learning processing is performed to change the importance of nearby people and objects based on the facial expressions of individually registered subjects. (2) Learning process using detection information during subject search During subject search, the system determines which people, objects, and scenes a registered subject is simultaneously in the frame, and calculates the time ratio in which the subjects are simultaneously within the frame. For example, the time ratio in which registered subject A is simultaneously in the frame with registered subject B is calculated. If both person A and person B are within the frame, various detection information is saved as training data to increase the automatic shooting score, and this data is used in the learning mode processing (step S716 in Figure 7). In another example, the time ratio in which registered subject A is simultaneously in the frame with a subject "cat" detected by general object recognition is calculated. If both person A and the "cat" are within the frame, various detection information is saved as training data to increase the automatic shooting score, and this data is used in the learning mode processing (step S716 in Figure 7).

[0176] Furthermore, if a high degree of smiling is detected in person A, a subject registered for personal authentication, or if expressions such as "joy" or "surprise" are detected, subjects simultaneously within the frame are learned to be important. Conversely, if expressions such as "anger" or "neutral" are detected in person A, subjects simultaneously within the frame are determined to be less likely to be important, and no learning process is performed.

[0177] Next, we will describe the learning process in cooperation with external devices. The learning process in cooperation with external devices in this embodiment can be performed in the following ways. (1) Learning process based on image acquisition by an external device (2) Learning process by assigning judgment values ​​to images using an external device (3) Learning process by analyzing images stored on an external device (4) Learning process from information uploaded to an SNS (Social Networking Service) server by an external device. (5) Learning process by changing camera parameters with an external device (6) Learning process from information where images have been manually edited using an external device. <Learning process based on image acquisition by an external device> As explained in Figure 3, the camera 101 and the external device 301 perform first communication 302 and second communication 303. Image data is transmitted and received via the first communication 302, and images from the camera 101 can be sent to the external device 301 using a dedicated application on the external device 301. In addition, thumbnails of the image data stored in the camera 101 can be viewed using a dedicated application on the external device 301. The user can select and review their favorite image from the thumbnails, or send image data to the external device 301 by issuing an image acquisition command. Images selected and acquired by the user are very likely to be images that the user likes. Therefore, the acquired images are determined to be images suitable for learning processing. Based on the learning information of the acquired images, learning processing tailored to the user's preferences can be performed.

[0178] An example of image selection operation will be explained with reference to Figure 12. Figure 12 illustrates an example of a user viewing images from camera 101 using a dedicated application on external device 301. Thumbnails 1204 to 1209 of image data stored in camera 01 are displayed on the display unit 407. The user can select and acquire their preferred image. Buttons 1201 to 1203 are operated to change the display method.

[0179] When the first button 1201 is pressed, the display mode is changed to date and time priority mode, and the images from camera 101 are displayed on the display unit 407 in order of the date and time they were taken. For example, the newer image is displayed at the position indicated by thumbnail 1204, and the older image is displayed at the position indicated by thumbnail 1209. When the second button 1202 is pressed, the display mode is changed to recommended image priority mode. Based on the score that determines the user's preference for each image, calculated in step S912 of Figure 9, the images from camera 101 are displayed on the display unit 407 in order of highest score. For example, the image with the highest score is displayed at the position indicated by thumbnail 1204, and the image with the lowest score is displayed at the position indicated by thumbnail 1209. Furthermore, when the user presses the third button 1203, they can specify a subject, such as a person or object, and if they subsequently specify a particular person or object, only that specific subject can be displayed.

[0180] Buttons 1201-1203 can also be used to turn on settings simultaneously. For example, if all settings are turned on, only the specified subject will be displayed, images with the most recent shooting date will be prioritized, and images with higher scores will be prioritized. In this way, the user's preferences are learned even for captured images, making it possible to extract only the user's preferred images from a large number of captured images with a simple review process.

[0181] <Learning process by assigning judgment values ​​to images using an external device> The user can view images stored in camera 101 and assign scores to them. Users can assign high scores (e.g., 5 points) to images they like and low scores (e.g., 1 point) to images they do not like. The camera learns image judgment values ​​in response to user input. The score for each image is used in the camera's retraining process along with training information. The output of the neural network, which takes feature data from the specified image information as input, is trained to approach the score specified by the user.

[0182] In addition to the configuration in which the user adds judgment values ​​to captured images using an external device 301, the user may also add judgment values ​​to images by operating the camera 101. In this case, the camera 101 is equipped with a touch panel display, and the user operates the GUI (Graphical User Interface) buttons on the touch panel display to set it to a mode that displays captured images. Then, the user can perform the same learning process as described above by setting judgment values ​​for each image while checking the captured images.

[0183] <Learning process by analyzing images stored on an external device> The storage unit 404 of the external device 301 also records images other than those captured by the camera 101. Images stored in the external device 301 are easily viewable by the user, and can also be easily uploaded to a shared server by the public wireless control unit 406, so there is a very high probability that they will contain many images of the user's preference.

[0184] The control unit 411 of the external device 301 can process images stored in the memory unit 404 with the same capabilities as the learning processing unit 219 of the camera 101 using a dedicated application. The learning process is performed by communicating the processed learning data to the camera 101. Alternatively, the system can be configured so that images or data to be used for learning are sent to the camera 101, and the camera 101 performs the learning process. Furthermore, the system can be configured so that the user can select images to be used for learning from the images stored in the memory unit 404 using a dedicated application.

[0185] <Learning process from information uploaded to the SNS server by an external device> This document describes how to use information from social networking services (SNS), which are services and websites that build social networks focusing on connections between people, for learning processing. One technique involves inputting tags related to an image from an external device 301 before sending the image along with the uploaded image to the SNS. Another technique involves inputting information about likes and dislikes regarding images uploaded by other users. It is also possible to determine whether an image uploaded by another user is to the liking of the user who owns the external device 301.

[0186] A dedicated SNS application downloaded to the external device 301 allows users to obtain images they have uploaded themselves and information about those images. Furthermore, by inputting whether users like or dislike images uploaded by other users, the system can obtain information about the user's preferred images and tags. The acquired images and tag information are then analyzed, and a learning process is performed on the camera 101.

[0187] The control unit 411 of the external device 301 acquires images uploaded by the user and images that the user has determined to like, and can process them with the same capabilities as the learning processing unit 219 of the camera 101. The learning process is performed by communicating the processed learning data to the camera 101. Alternatively, the system may be configured to send image data to be used for learning to the camera 101, and the camera 101 may perform the learning process.

[0188] Based on subject information set in the tag information (for example, object information such as dogs and cats, scene information such as beaches, and facial expression information such as smiles), it is possible to estimate subject information tailored to the user's preferences. In this case, the subject to be detected is input into the neural network (NN) for training. Furthermore, it is possible to estimate currently trending image information from statistical values ​​of tag information (image filter information and subject information) on social media and create a configuration that can be trained by camera 101.

[0189] <Learning process by changing camera parameters using an external device> The learning parameters currently set on camera 101 (such as the weight coefficients of the neural network and the selection of subjects to input to the neural network) can be transmitted to the external device 301 and stored in the memory unit 404 of the external device 301. Furthermore, using a dedicated application on the external device 301, learning parameters set on a dedicated server can be acquired by the public wireless control unit 406. These can also be set as the learning parameters for camera 101. The learning parameters can also be restored by saving the parameters at a certain point in time on the external device 301 and setting them on camera 101. Additionally, learning parameters held by other users can be acquired by the dedicated server and set on the owner's camera 101.

[0190] Furthermore, the system may be configured to allow users to register voice commands, authentication credentials, and gestures using a dedicated application on the external device 301, or to register important locations. This information is used as input data for the trigger to start shooting and for automatic shooting determination, as explained in the automatic shooting mode processing in Figure 9. The system may also be configured to allow users to set the shooting frequency, activation interval, ratio of still images to videos, and preferred images, and the activation interval and other settings described in the low-power mode control above may be configured accordingly.

[0191] <Learning process from information where images have been manually edited using an external device> A dedicated application for the external device 301 enables manual editing according to user operations, and the content of the editing work can also be fed back into the learning process. For example, it is possible to edit the application of image effects (cropping, rotation, slide, zoom, fade, color conversion filter effect, time, still image-to-video ratio, BGM). A learning process using an automated editing neural network is performed so that manually edited image effects are detected in the learning information of the image.

[0192] Next, the learning mode processing in step S716 of Figure 7 will be explained. In the mode determination in step S704 of Figure 7, it is determined whether or not to perform learning processing. If it is determined in step S715 that learning processing should be performed, the learning mode processing in step S716 is executed.

[0193] Now, referring to Figure 13, we will explain the process for determining whether or not to perform the learning process. The decision of whether or not to perform the learning process is made based on the elapsed time since the last learning process, the amount of information available for the learning process, and instructions for the learning process from an external device.

[0194] Figure 13 is a flowchart illustrating the determination process for whether or not to perform the learning process, which is performed in steps S704 (mode determination process) and S715 of Figure 7.

[0195] When the mode determination process starts in step S704, the process shown in Figure 13 begins.

[0196] In step S1301, the first control unit 223 determines whether or not there is a registration instruction from the external device 301. The registration instruction is for performing tasks such as <learning processing based on image acquisition by the external device>, <learning processing based on assigning judgment values ​​to images by the external device>, and <learning processing based on analysis of images stored in the external device>. If it is determined in step S1301 that there is a registration instruction from the external device 301, the process proceeds to step S1308.

[0197] In step S1308, the first control unit 223 sets the learning mode determination flag to TRUE, configures the unit to perform the processing in step S716, and then terminates the learning mode determination process. If it is determined in step S1301 that there is no registration instruction from the external device 301, the unit proceeds to the processing in step S1302.

[0198] In step S1302, the first control unit 223 determines whether or not there is a learning instruction from the external device 301. A learning instruction is an instruction to set learning parameters, such as <learning process by changing camera parameters with the external device>. If it is determined in step S1302 that there is a learning instruction from the external device 301, the process proceeds to step S1308. If it is determined in step S1302 that there is no learning instruction from the external device 301, the process proceeds to step S1303.

[0199] In step S1303, the first control unit 223 obtains the elapsed time TimeN since the previous learning process (recalculation of the NN weight coefficients) was performed.

[0200] In step S1304, the first control unit 223 obtains the number of new data points DN for training. The number of data points DN corresponds to the number of images specified to be trained during the elapsed time TimeN since the previous training process was performed.

[0201] In step S1305, the first control unit 223 calculates a threshold DT based on the elapsed time TimeN to determine whether or not to transition to learning mode. The smaller the value of the threshold DT, the easier it is to transition to learning mode. For example, the value of the threshold DT when the elapsed time TimeN is less than a predetermined value is denoted as DTa, and the value of the threshold DT when the elapsed time TimeN is greater than a predetermined value is denoted as DTb. DTa is set to be greater than DTb, and the threshold is set to decrease as time progresses. As a result, even with little learning data, if the elapsed time is long, it becomes easier to transition to learning mode, and the learning process is performed again. In other words, the ease or difficulty of the camera transitioning to learning mode is changed according to the usage time.

[0202] In step S1306, the first control unit 223 determines whether the number of training data DN is greater than the threshold DT. If it is determined that the number of data DN is greater than the threshold DT, the process proceeds to step S1307; if it is determined that the number of data DN is less than or equal to the threshold DT, the process proceeds to step S1309.

[0203] In step S1307, the first control unit 223 sets the data count DN to zero and proceeds to the process in step S1308.

[0204] In step S1309, the first control unit 223, since there are no registration or learning instructions from the external device 301 and the number of learning data DN is less than or equal to the threshold DT, sets the learning mode determination flag to FALSE, and after setting it not to perform the process in step S716, terminates the process.

[0205] Next, with reference to Figure 14, the learning mode processing in step S716 of Figure 7 will be described.

[0206] Figure 14 is a flowchart illustrating the learning mode processing in step S716 of Figure 7.

[0207] The process shown in Figure 14 is initiated when it is determined in step S715 of Figure 7 that the system is in learning mode.

[0208] In step S1401, the first control unit 223 determines whether or not there is a registration instruction from the external device 301. If it is determined in step S1401 that there is a registration instruction from the external device 301, the process proceeds to step S1402. If it is determined in step S1401 that there is no registration instruction from the external device 301, the process proceeds to step S1404.

[0209] In step S1402, the first control unit 223 performs various registration processes and proceeds to the process in step S1403. These registrations involve registering features to be input to the neural network (NN), such as facial recognition registration, general object recognition registration, sound information registration, and location information registration.

[0210] In step S1403, the first control unit 223 performs a process to change the information to be input to the NN from the feature information registered in step S1402, and then proceeds to the process in step S1407.

[0211] In step S1404, the first control unit 223 determines whether or not there is a learning instruction from the external device 301. If it is determined that there is a learning instruction from the external device 301, the process proceeds to step S1405; if it is determined that there is no learning instruction, the process proceeds to step S1406.

[0212] In step S1405, the first control unit 223 sets the learning parameters communicated from the external device 301 to each classifier (such as the weight coefficients of the neural network), and then proceeds to the process in step S1407.

[0213] In step S1406, the first control unit 223 performs learning (recalculation of the NN's weight coefficients) using the learning processing unit 219, and proceeds to step S1407. The transition to step S1406 occurs when the number of data points DN exceeds the threshold DT, as explained in Figure 13, and each classifier is retrained. Through retraining using methods such as backpropagation and gradient descent, the weight coefficients of the NN are recalculated, and the parameters of each classifier are changed.

[0214] In step S1407, the first control unit 223 performs a rescoring process for the image files stored on the recording medium 221. In this embodiment, scores are assigned to all image files stored on the recording medium 221 based on the learning results, and automatic editing or automatic file deletion is performed according to the scores. Therefore, if retraining or setting of learning parameters from an external device is performed, the scores of already captured images also need to be updated. In step S1407, after recalculation is performed to assign new scores to the image files stored on the recording medium 221, the process ends.

[0215] In this embodiment, we have described a method for acquiring images that the user prefers by learning the characteristics of scenes that are estimated to suit the user's preferences and reflecting the learning results in camera operations such as automatic shooting and automatic editing. However, this is not the only example. For example, it is also possible to apply this method to applications that intentionally use images that differ from the user's preferences, as shown below.

[0216] <A method using a neural network that has learned the user's preferences> The user's preferences are learned using the method described above, and the automatic shooting determination process is executed in step S908 of Figure 9. Automatic shooting is performed when the output value of the NN is different from the user's preferences, which are the training data. For example, suppose the user's preferred image is used as the training image, and the learning process is performed so that a high value is output when the image shows similar features to the training image. In this case, conversely, automatic shooting is performed on the condition that the output value is lower than a predetermined threshold. Similarly, in the subject search process and the automatic editing process, processing is performed so that the output value of the NN is different from the user's preferences, which are the training data.

[0217] <A method using a neural network that has been trained in situations different from the user's preferences> During the learning process, situations that differ from the user's preferences are used as training data. In this embodiment, images taken manually are assumed to be scenes that the user prefers to photograph, and a learning method using these as training data has been described. In contrast, manually taken images are not used as training data; instead, scenes that have not been manually photographed for a predetermined period of time are added as training data. Alternatively, if there is data in the training data that has similar features to manually taken images, this data is removed from the training data. Furthermore, images acquired by external devices that have different features from the training data are added to the training data, or images that have similar features to acquired images are removed from the training data. In this way, the training data accumulates data that differs from the user's preferences, so as a result of the learning process, the NN can distinguish situations that differ from the user's preferences. In automatic photography, photography is performed according to the output value of the NN, so scenes that differ from the user's preferences can be photographed.

[0218] By deliberately using images that differ from the user's preferences, scenes that the user would not manually capture are photographed, reducing the number of missed shots. Furthermore, by suggesting scenes that the user would not have thought of, it has the effect of encouraging user awareness and broadening their range of preferences.

[0219] By combining the methods described above, it is easy to suggest situations that are somewhat similar to the user's preferences but differ in some ways, and to adjust the degree of suitability to the user's preferences. The degree of suitability to the user's preferences can be changed depending on the setting mode, the status of various sensors, and the status of the detected information.

[0220] In this embodiment, a configuration in which the camera 101 performs the learning process has been described. On the other hand, if the external device 301 has a learning function, the same learning effect as described above can be achieved even in a configuration in which the data necessary for the learning process is sent to the external device 301 and the learning process is executed by the external device 301. For example, as described in <Learning process by changing camera parameters with an external device>, the external device 301 may perform the learning process by setting parameters such as the weight coefficients of the learned NN through communication with the camera 101.

[0221] Furthermore, there are configurations in which both the camera 101 and the external device 301 have learning functions. For example, when the camera 101 performs learning mode processing (step S716 in Figure 7), the learning information held by the external device 301 is transmitted to the camera 101, the learning parameters are merged, and the learning process is performed using the merged learning parameters.

[0222] <System Configuration> Next, referring to Figures 15 to 30, we will describe a system in which an external device controls multiple cameras to perform imaging.

[0223] In the following example, two cameras 101a and 101b installed in different locations are connected to the external device 301, but three or more cameras may be connected.

[0224] Furthermore, the configuration and functions of cameras 101a and 101b and the external device 301 are the same as those described in Figures 1 to 14.

[0225] Figure 15 illustrates a system that controls multiple cameras to perform photography.

[0226] In this embodiment, the system is connected to multiple cameras 101a and 101b and an external device 301, enabling wireless communication 1501 and 1502. In the example shown in Figure 15, two cameras 101a and 101b and the external device 301 are connected, enabling wireless communication 1501 and 1502, but three or more cameras may also be connected. Wireless communication 1501 and 1502 can be Wi-Fi or BLE. The external device 301 can connect to multiple cameras 101a and 101b and can identify multiple cameras based on unique information received from cameras 101a and 101b. The unique information used to identify multiple cameras can be MAC addresses or individual identification numbers assigned during manufacturing. The external device 301 transmits shooting control information to cameras 101a and 101b and receives omnidirectional images from cameras 101a and 101b.

[0227] Figure 16 illustrates a situation where multiple subjects are within the shooting range of multiple cameras.

[0228] In Figure 16, cameras A1602 and B1603, along with people 1610-1616, are located in a predetermined area 1601, with cameras A1602 and B1603 independently performing automatic shooting. The shooting range of camera A1602 is shown as field of view 1605, and the shooting range of camera B1603 is shown as field of view 1606.

[0229] In the example in Figure 16, person 1611 is included in the shooting range of both camera A1602 and camera B1603, meaning that camera A1602 and camera B1603 may both capture the same person. Furthermore, person 1611 is likely to produce similar images regardless of whether they are captured by camera A1602 or camera B1603.

[0230] Figure 17 shows examples of images captured in all directions by cameras A1602 and B1603.

[0231] Figure 17(c) illustrates a scenario in which two cameras and three people are present in a designated area 1701.

[0232] Camera A1702 and camera B1703 can capture images of all directions every 60 degrees starting from the reference angle of the camera by means of pan drive. Persons 1704 to 1706 are located in the middle of camera A1702 and camera B1703. The image of all directions captured by camera A1702 is shown in Fig. 17(a), and the image of all directions captured by camera B1703 is shown in Fig. 17(b).

[0233] Figs. 17(a) and 17(b) illustrate the states in which person 1704 is photographed frontally at 120 degrees in Fig. 17(a), person 1705 is photographed frontally at 240 degrees in Fig. 17(a), and person 1704 is photographed frontally at 120 degrees in Fig. 17(b).

[0234] Person 1706 at 300 degrees in Fig. 17(a), person 1706 at 0 degrees in Fig. 17(b), person 1705 at 60 degrees in Fig. 17(b), and person 1706 at 300 degrees in Fig. 17(b) all illustrate the states in which the subject is photographed from behind.

[0235] Here, image 1710 captured at an angle of 120 degrees in Fig. 17(a) and image 1711 captured at an angle of 120 degrees in Fig. 17(b) are similar. Also, image 1720 captured at an angle of 300 degrees in Fig. 17(a) and image 1721 captured at an angle of degrees in Fig. 17(b) are similar. The method for determining image similarity will be described later in Fig. 23.

[0236] Fig. 18 is a flowchart illustrating the control processing of cameras 101a and 101b of the system according to the present embodiment.

[0237] Each of cameras 101a and 101b according to the present embodiment can execute automatic shooting as described in Figs. 7 and 9. However, when connected to an external device 301 as shown in Fig. 15, each of cameras 101a and 101b executes the processing of Fig. 18 instead of the processing of Fig. 9.

[0238] The processing of Fig. 18 is realized by executing a program stored in the non-volatile memory 216 of the first control unit 223 of camera 101.

[0239] In step S1801, the first control unit 223 performs omnidirectional shooting processing. As shown in Figure 17, the first control unit 223 pans the cameras 101a and 101b to take pictures every 60 degrees within the 360-degree shooting range.

[0240] In step S1802, the first control unit 223 transmits the omnidirectional image captured in S1801 to the external device 301. In this case, the captured image and the shooting angle are added as supplementary information to the captured image so that they can be handled together.

[0241] In step S1803, the first control unit 223 determines whether the installation positions of cameras 101a and 101b have been changed. The installation positions of cameras 101a and 101b are detected using the gyro sensor and accelerometer sensor provided by cameras 101a and 101b. By acquiring movement information and motion detection information using the gyro sensor and accelerometer sensor, it is possible to determine whether the installation positions of cameras 101a and 101b have moved. If the installation positions of cameras 101a and 101b have been changed, the process returns to step S1801 in order to retake the omnidirectional image.

[0242] In step S1804, the first control unit 223 acquires imaging control information from the external device 301.

[0243] In step S1805, the first control unit 223 determines whether the shooting control information acquired from the external device 301 includes a shooting instruction. In this embodiment, if all the shooting permission settings described later in Figure 29 are turned off, it is determined that no shooting instruction is included. If the shooting control information acquired from the external device 301 includes a shooting instruction, the process proceeds to step S1806; otherwise, the process proceeds to step S1807.

[0244] In step S1806, the first control unit 223 performs a second automatic shooting based on a shooting instruction received from the external device 301. The shooting instruction is notified with the shooting permission set for each shooting angle, as shown in Figure 29. Cameras 101a and 101b are panned to the shooting angle for which the shooting permission is set to ON (shooting possible) based on the shooting instruction received from the external device 301, and the shooting process of step S910 in Figure 9 is performed on the main subject determined by the subject search process by the NN.

[0245] In step S1807, the first control unit 223 performs the first automatic shooting. Cameras 101a and 101b perform the automatic shooting shown in Figure 9 if the shooting control information acquired from the external device 301 does not include a shooting instruction, that is, if all the shooting permission settings described later in Figure 29 are turned off.

[0246] Next, with reference to Figure 19, the control process of the external device 301 of the system in this embodiment will be described.

[0247] The process shown in Figure 19 is achieved when the control unit 411 of the external device 301 executes a program stored in the storage unit 404.

[0248] In step S1901, the control unit 411 waits until it has received omnidirectional images from all cameras 101a and 101b in step S1802 of Figure 18. If it is determined that it has received omnidirectional images from all cameras 101a and 101b connected to the external device 301, it proceeds to step S1903.

[0249] In step S1903, the control unit 411 determines the shooting area for all images at all shooting angles received in step S1902. Figure 21 shows an example of the shooting area determination results. For each shooting angle, the shooting area determination result includes either an area not to be photographed, a shooting area (front view with a person), or a shooting area (other than front view with a person). The shooting area determination process will be described later in Figure 20.

[0250] In step S1904, the control unit 411 determines whether or not a target area exists based on the result of determining the shooting area in step S1903. If it is determined that a target area exists, the process proceeds to step S1905. If it is determined that a target area does not exist, that is, if all shooting angles in Figure 21 are outside the target area, the process proceeds to step S1908.

[0251] In step S1905, the control unit 411 determines whether or not to take a picture based on image similarity, if a target area exists. The control unit 411 evaluates the similarity of the images taken by cameras 101a and 101b, and controls the camera to prevent overlap of target areas for images with high similarity. Figure 25 shows an example of the results of the image similarity-based determination of whether or not to take a picture. The image similarity-based determination of whether or not to take a picture is saved in a table format and includes the determination results for whether or not to take a picture for each shooting angle. Figure 25(a) shows an example of the determination result for camera A1702 in Figure 17(c), and Figure 25(b) shows an example of the determination result for camera B1703 in Figure 17(c). The image similarity-based determination of whether or not to take a picture will be described later in Figure 22.

[0252] In step S1906, the control unit 411 determines whether or not to photograph based on the bias of the subject, if a target area exists. Based on the subject detection results for the images captured by cameras 101a and 101b, the control unit 411 controls the system to prevent overlapping target areas so that the same person is not photographed repeatedly. Figure 28 shows an example of the results of the shooting feasibility determination based on the bias of the subject. The results of the shooting feasibility determination based on the bias of the subject are saved in a table format and include the determination results of whether or not to photograph for each shooting angle. Figure 28(a) shows an example of the shooting feasibility determination result for camera A1702 in Figure 17(c), and Figure 28(b) shows an example of the shooting feasibility determination result for camera B1703 in Figure 17(c). The process of determining whether or not to photograph based on the bias of the subject will be described later in Figure 26.

[0253] In step S1907, the control unit 411 integrates the shooting permission determination results obtained in steps S1905 and S1906. FIG. 29 illustrates the final shooting permission determination result obtained by integrating the shooting permission determination result of step S1905 shown in FIG. 25 and the shooting permission determination result of step S1906 shown in FIG. 28. FIG. 29(a) illustrates the final shooting permission determination result of the camera A1702 in FIG. 17(c), and FIG. 29(b) illustrates the final shooting permission determination result of the camera B1703 in FIG. 17(c). The final shooting permission determination result is obtained by the logical sum (OR) of the shooting permission determination result based on image similarity and the shooting permission determination result based on the bias of the subject, and the final shooting permission determination result is stored in a table format.

[0254] In step S1908, since there is no shooting target area, the control unit 411 turns off all the final shooting permission determination results for each shooting angle shown in FIG. 29 so that the shooting instruction is not included in the shooting control information.

[0255] In step S1909, the control unit 411 transmits the shooting control information to the cameras 101a and 101b. The shooting permission determination result obtained in step S1907 or S1908 is transmitted to the cameras 101a and 101b.

[0256] Next, referring to FIG. 20, the shooting area determination process in step S1903 of FIG. 19 will be described.

[0257] In step S2001, the control unit 411 performs a subject detection process on the image obtained in step S1901 of FIG. 19. The subject detection process uses a known method such as a method using pattern matching or a method using a neural network (NN) to detect a person's face or body as a subject. Further, the control unit 411 obtains the orientation of the person's face or body detected by the subject detection process.

[0258] In step S2002, the control unit 411 determines whether the subject detection process has been completed for all areas (shooting angles) of the images from all cameras. If it is determined that the subject detection process has been completed for all areas (shooting angles) of the images from all cameras, the process proceeds to step S2003. If it is determined that the subject detection process has not been completed for all areas (shooting angles) of the images from all cameras, the process returns to step S2001 and continues.

[0259] In step S2003, the control unit 411 determines whether a person's face or body has been detected by the subject detection process in step S2001. If it is determined that a person's face or body has been detected, the process proceeds to step S2003; otherwise, the process proceeds to step S2007.

[0260] In step S2004, the control unit 411 determines whether the face or body of the person detected in step S2001 is facing forward. If multiple faces or bodies of people were detected in step S2001, the control unit 411 determines the orientation of the face or body of the main subject. If it is determined that the face or body of the person is facing forward, the process proceeds to step S2005; if the face or body of the person is facing sideways or backward, the process proceeds to step S2006.

[0261] In steps S2005 to S2007, the control unit 411 sets the target area for each camera based on the determination results from steps S2003 and S2004 (whether a person's face or body was detected, and the orientation of the person's face or body). In step SS2005, the shooting angle of images in which a person's face or body is detected and the person's face or body is facing forward is set as the target area, and the setting value is saved in the storage unit 404. In S2006, the shooting angle of images in which a person's face or body is detected and the person's face or body is not facing forward is set as the target area, and the setting value is saved in the storage unit 404. In S2007, the shooting angle of images in which a person's face or body is not detected is set as the non-target area, and the setting value is saved in the storage unit 404. Figure 21 illustrates the shooting area determination results for each camera and shooting angle. The shooting area determination results are saved in table format. Figure 21(a) shows the shooting area determination results based on the subject detection results for each image in Figure 17(a) taken by camera A1702. In the example in Figure 21(a), the shooting angles of 120 degrees and 240 degrees, where a person's face is detected and the person's face is not facing forward, and the shooting angle of 300 degrees, where a person's face is detected and the person's face is not facing forward, are set as the target shooting area, while the other shooting angles of 0 degrees, 60 degrees, and 180 degrees are set as areas not to be photographed. Figure 21(b) shows the shooting area determination results based on the subject detection results for each image in Figure 17(b) taken by camera B1703. In the example shown in Figure 21(b), the shooting area is set to include 120 degrees, where a person's face is detected in each image of Figure 17(b) and the person's face is not facing forward, as well as 0 degrees and 60 degrees, where a person's face is detected and the person's face is not facing forward. The other shooting angles of 180 degrees, 240 degrees, and 300 degrees are set to be excluded from shooting.

[0262] In step S2008, the control unit 411 determines whether or not a shooting area determination result has been obtained for images of all areas (shooting angles) from all cameras. If it is determined that a shooting area determination result has been obtained based on images of all areas (shooting angles) from all cameras, the process ends. If it is determined that a shooting area determination result has not been obtained based on images of all areas (shooting angles) from all cameras, the process returns to step S2003 and continues.

[0263] Next, referring to Figure 22, the process of determining whether or not to take a photograph based on image similarity in step S1905 of Figure 19 will be explained.

[0264] In step S2201, the control unit 411 determines whether there are any similar images for all images taken at all shooting angles from all cameras, and stores the determination result in the storage unit 404. Details of the similarity determination process will be described later in Figure 23. Figure 24 shows an example of the similarity determination results for all images taken at all shooting angles from all cameras. The image similarity determination results are stored in a table format. If there are similar images, a discriminable index is assigned to the shooting angle from which the similar image was taken as part of the similarity determination results for each camera. Figure 24(a) shows the similarity determination results for each image in Figure 17(a) taken by camera A1702 for each image in Figure 17(b) taken by camera B1703. In the example in Figure 24(a), the image in Figure 17(b) taken at a 120-degree angle is similar to the image in Figure 17(b) taken at a 120-degree angle, and an index indicating similarity (YES) and an index indicating similarity (Similarity 1) are set for the shooting angle of the image in Figure 17(a) taken at a 300-degree angle that is similar to the image in Figure 17(b) taken at a 0-degree angle, and an index indicating similarity (YES) and an index indicating similarity (Similarity 2) are set for the shooting angle of the image in Figure 17(b) taken at a 0-degree angle. Figure 24(b) shows the similarity determination results for each image in Figure 17(b) taken by camera B1703 for each image in Figure 17(a) taken by camera A1702. In the example in Figure 24(b), the image in Figure 17(b) taken at a 120-degree angle that is similar to the image in Figure 17(a) taken at a 120-degree angle is assigned an index (YES) indicating similarity and an index (Similar 1) indicating a similar image. Similarly, the image in Figure 17(b) taken at a 0-degree angle that is similar to the image in Figure 17(a) taken at a 300-degree angle is assigned an index (YES) indicating similarity and an index (Similar 2) indicating a similar image.

[0265] In step S2202, the control unit 411 refers to the shooting area determination result in Figure 21 and determines whether all shooting angles of all cameras are within the target shooting area. If the shooting angle of the camera being determined is determined to be within the target shooting area, the process proceeds to step S2203. If the shooting angle of the camera being determined is determined to be outside the target shooting area, the process proceeds to step S2212.

[0266] In step S2203, the control unit 411 refers to the similarity determination results in Figure 24 and determines whether or not there are similar images for all shooting angles of all cameras. If it is determined that there are similar images, the process proceeds to step S2204; if it is determined that there are no similar images, the process proceeds to step S2211.

[0267] In step S2204, the control unit 411 compares the number of target areas for each camera by referring to the shooting area determination result in Figure 21. In the example in Figure 21, the number of target areas for camera A1702 and the number of target areas for camera B1703 are the same (3).

[0268] In step S2205, the control unit 411 determines whether the number of target areas to be photographed by the cameras is the same, based on the result of the comparison of the number of target areas to be photographed by the cameras in step S2204. If it is determined that the number of target areas to be photographed by the cameras is the same, the process proceeds to step S2206; if it is determined that the number of target areas to be photographed by the cameras is different, the process proceeds to step S2208.

[0269] In step S2206, the control unit 411 refers to the subject detection results from step S2001 in Figure 20 for the images captured by each camera that were determined to be similar in step S2203, and determines which of the similar images captured by each camera has the main subject, the person, positioned closer to the center of the field of view.

[0270] In step S2207, the control unit 411, based on the result determined in step S2206, sets the shooting angle of the camera that captured the image in which the main subject, the person, is closer to the center of the field of view to enable shooting.

[0271] In step S2208, the control unit 411 sets the shooting angle of the camera with the smaller number of target areas to be photographed to enable shooting, based on the result determined in step S2205.

[0272] In step S2209, the control unit 411 sets the shooting feasibility determination result for the shooting angle of the camera that was set to be able to shoot in steps S2207 and S2208 to ON, and saves the setting value in the storage unit 404. In the example in Figure 21, since the number of target areas for cameras A1702 and B1703 is the same, among the similar images of a shooting angle of 120 degrees in Figure 17(a) and Figure 17(b), the shooting angle of the camera that took the image of Figure 17(a) in which the person is closer to the center of the field of view is set to be able to shoot in steps S2206 and S2207. Also in the example in Figure 21, among the similar images of a shooting angle of 300 degrees in Figure 17(a) and Figure 17(b) in which the shooting angle is 0 degrees, the shooting angle of the camera that took the image of Figure 17(a) in which the person is closer to the center of the field of view is set to be able to shoot in steps S2206 and S2207. In the example in Figure 24, the camera's shooting angle is set to enable shooting when it captures the image in Figure 24(a) where the person is closer to the center of the field of view, compared to the similar image in Figure 24(a) with a shooting angle of 102 degrees and the similar image in Figure 24(b) with a shooting angle of 102 degrees. Also in the example in Figure 24, the camera's shooting angle is set to enable shooting when it captures the image in Figure 24(a) where the person is closer to the center of the field of view, compared to the similar image in Figure 24(a) with a shooting angle of 300 degrees and the similar image in Figure 24(b) with a shooting angle of 0 degrees, compared to the similar image in Figure 24(a) with a shooting angle of 0 degrees, compared to the similar image in Figure 24(b) with a shooting angle of 0 degrees.

[0273] In step S2210, the control unit 411 sets the shooting angle of other cameras corresponding to the shooting angle of the camera that was set to be able to shoot in step S2209 to be unavailable, sets the shooting availability determination result to OFF, and saves the setting value in the storage unit 404.

[0274] Figure 25 illustrates the results of the image similarity-based shooting feasibility determination. The image similarity-based shooting feasibility determination results are saved in a table format. Figure 25(a) shows the shooting feasibility determination results for each shooting angle of camera A1702. In the example of Figure 25(a), the shooting feasibility determination results for camera A1702 at shooting angles of 120 degrees and 300 degrees in Figure 24(a) are set to ON. Figure 25(b) shows the shooting feasibility determination results for each shooting angle of camera B1703. In the example of Figure 25(b), the shooting feasibility for camera B1703 at shooting angles of 120 degrees and 300 degrees in Figure 24(b) is set to OFF (not possible to shoot).

[0275] In step S2211, the control unit 411 sets the shooting angle of the camera that has captured an image that is within the target area and is not similar to the image captured by the other camera to be able to shoot, sets the shooting feasibility determination result to ON, and saves the setting value in the storage unit 404. In the examples of Figures 21 and 24, the shooting angle of camera A1702 in Figure (a) at 240 degrees and the shooting angle of camera B1703 in Figure (b) at 60 degrees are both within the target area but are not similar images, so they are set to be able to shoot, and the shooting feasibility determination result for the shooting angle of camera A1702 in Figure 25(a) at 240 degrees and the shooting angle of camera B1703 in Figure 25(b) at 60 degrees is set to ON.

[0276] In step S2212, the control unit 411 sets the shooting angle of the camera that captured an image of an area not to be photographed to "cannot be photographed," sets the shooting permission / failure determination result to "off," and saves the setting value in the storage unit 404. In the examples of Figures 21 and 24, the shooting angles of camera A1702 in Figure (a) at 0 degrees, 60 degrees, and 180 degrees, and the shooting angles of camera B1703 in Figure (b) at 180 degrees, 240 degrees, and 300 degrees are all areas not to be photographed, so they are set to "cannot be photographed." The shooting permission / failure determination results for the shooting angles of camera A1702 in Figure 25(a) at 0 degrees, 60 degrees, and 180 degrees, and the shooting angles of camera B1703 in Figure 25(b) at 180 degrees, 240 degrees, and 300 degrees are set to "off."

[0277] In step S2213, the control unit 411 determines whether or not a shooting permission / failure judgment result has been set for all shooting angles of all cameras. If it is determined that a shooting permission / failure judgment result has been set for all shooting angles of all cameras, the process ends. If it is determined that a shooting permission / failure judgment result has not been set for all shooting angles of all cameras, the process returns to step S2202 and continues.

[0278] Next, referring to Figure 23, the image similarity determination process for all camera angles in step S2201 of Figure 22 will be explained.

[0279] In step S2301, the control unit 411 acquires image feature quantities from all cameras at all shooting angles and applies a feature detector to quantify the feature quantities of each image. Known methods such as AKAZE can be used to quantify the image feature quantities.

[0280] In step S2302, the control unit 411 determines whether or not it has acquired image feature quantities for all shooting angles from all cameras. If it is determined that it has acquired image feature quantities for all shooting angles from all cameras, it proceeds to step S2303. If it is determined that it has not acquired image feature quantities for all shooting angles from all cameras, it returns to step S2301 and continues processing until it has acquired image feature quantities for all shooting angles from all cameras.

[0281] In step S2303, the control unit 411 initializes the similarity determination results of images from all cameras at all shooting angles in the previous processing. In the example in Figure 24, the control unit sets NO data to indicate that no similar images exist for the similar image determination results for each shooting angle of camera A1702 in Figure 24(a) and for each shooting angle of camera B1703 in Figure 24(b).

[0282] In step S2304, the control unit 411 compares the feature quantities of the images for each shooting angle acquired by each camera in step S2301. In comparing the feature quantities, a known method called matching is used to calculate the distance between the feature quantities of the images for each shooting angle of each camera, and the average of the distances is calculated as the similarity score.

[0283] In step S2305, the control unit 411 compares the similarity calculated in step S2304 with a threshold and determines whether the similarity is equal to or greater than the threshold. If the similarity is equal to or greater than the threshold, the images are determined to be similar, and the process proceeds to step S2306. If the similarity is less than the threshold, the images are determined to be dissimilar, and the process proceeds to step S2307.

[0284] In step S2306, the control unit 411 determines that the images are similar, so it sets the similarity determination result for images at the shooting angles of all cameras that captured similar images to YES and saves the setting value to the storage unit 404. In the examples of Figures 17 and 24, the image taken at a shooting angle of 120 degrees by camera A1702 in Figure 24(a) and the image taken at a shooting angle of 120 degrees by camera B1703 in Figure 24(b) are determined to be similar, so an index indicating similarity (YES) and an index indicating similar images (Similar 1, Similar 2) are set for shooting angles of 120 degrees and 300 degrees in Figure 24(a) and shooting angles of 120 degrees and 0 degrees in Figure 24(b). In the example of Figure 24, a number numbered with a 1 origin is added as a suffix for Similar 1.

[0285] In step S2307, it is determined whether or not a similarity comparison has been performed for images from all cameras at all shooting angles. If it is determined that a similarity comparison has been performed for images from all cameras at all shooting angles, the process ends. If it is determined that a similarity comparison has not been performed for images from all cameras at all shooting angles, the process returns to step S2304 and continues until the similarity comparison for images from all cameras at all shooting angles is completed.

[0286] In step S2210 of Figure 22, if there are similar images for all shooting angles of all cameras, the shooting feasibility determination result for the shooting angle of other cameras corresponding to the shooting angle of the camera for which the shooting feasibility determination result was set to ON is turned OFF. However, other processing may be applied. For example, as shown in Figure 30, instead of setting the shooting feasibility determination result to OFF according to the similar image determination result, it may be set to ON, and the camera may be controlled to change the field of view by zooming and then take a picture. Methods for changing the field of view include enlarging the subject using the zoom unit 201 or cropping and enlarging using image processing.

[0287] Next, referring to Figure 26, we will explain the process of determining whether or not to photograph based on the bias of the subject in step S1906 of Figure 19.

[0288] In step S2601, the control unit 411 refers to the shooting area determination result in Figure 21 and determines whether the images from all cameras at all shooting angles are within the target shooting area and whether the person's face or body is facing forward. If the shooting angle of the camera being determined is determined to be within the target shooting area and the person's face or body is facing forward, the process proceeds to step S2602. If the shooting angle of the camera being determined is determined to be outside the target shooting area or the person's face or body is not facing forward, the process proceeds to step S2603.

[0289] In step S2602, the control unit 411 performs personal authentication processing on the person who was determined in step S2601 to be in the target area and whose face or body is facing forward, and stores the authentication result in the storage unit 404. For the personal authentication processing of a person, known methods can be applied to quantify the feature quantities of the entire face or each part of the face. Figure 27 illustrates the subject determination result by personal authentication. The subject determination result by personal authentication is stored in a table format. In Figure 27(a), person 1 (person's name) extracted from the image taken at a shooting angle of 120 degrees from the image of camera A1702 in Figure 17(a) is registered, and person 2 (person's name) extracted from the image taken at a shooting angle of 240 degrees is registered.

[0290] In step S2603, the control unit 411 sets the person information for camera shooting angles that are outside the target area in step S2601 or where the person's face or body is not facing forward to none, sets the shooting feasibility determination result to off, and saves the setting value to the storage unit 404. In the example in Figure 27, the person information for camera shooting angles that are outside the target area or where the person's face or body is not facing forward is set to none (unknown). Figure 28 illustrates the shooting feasibility determination result due to subject bias. In Figure 28(a), among the images from camera A1702 in Figure 17(a), the shooting feasibility determination results for shooting angles of 0 degrees, 60 degrees, 180 degrees, and 300 degrees, for which the person information was set to none (unknown) in Figure 27(a), are set to off. In Figure 28(b), the shooting feasibility determination results for shooting angles of 0 degrees, 60 degrees, 180 degrees, 240 degrees, and 300 degrees, for which the person information was set to none (unknown) in Figure 27(b), are set to off.

[0291] In step S2604, the control unit 411 determines whether or not personal authentication has been performed for images of the target area and all shooting angles in which the person's face or body is facing forward. If personal authentication has been performed for images of the target area and all shooting angles in which the person's face or body is facing forward, the process proceeds to step S2605. If it is determined that personal authentication has not been performed for images of the target area and all shooting angles in which the person's face or body is facing forward, the process returns to step S2601 and continues until personal authentication has been performed for images of the target area and all shooting angles in which the person's face or body is facing forward.

[0292] In step S2605, the control unit 411 compares the person's feature quantities obtained from images taken at each camera's shooting angle based on the personal authentication in step S2602. In comparing the person's feature quantities, a known method called matching is used to determine whether the same person is included in images taken at different camera angles. If the same person is included, person information with the same index assigned to each camera's shooting angle is input. If the same person is not included, person information with a non-repeating index is input. In the example in Figure 27, an index (Person 1) is set to indicate that the person in the image taken at a shooting angle of 120 degrees in Figure 27(a) and the person in the image taken at a shooting angle of 120 degrees in Figure 27(b) are the same person.

[0293] In step S2606, the control unit 411 refers to the subject determination result in Figure 27 and determines whether or not a person is present in the images from all cameras at all shooting angles. If it is determined that a person is present in the image to be determined, the process proceeds to step S2607; if it is determined that no person is present in the image to be determined, the process proceeds to step S2614.

[0294] In step S2607, the control unit 411 determines whether the person identified in step S2606 is present in all camera images. If it is determined that the same person is present in all camera images, the process proceeds to step S2610. If it is determined that the same person is not present in all camera images (i.e., present in only one camera image), the process proceeds to step S2608. In the example in Figure 27, it can be confirmed that the person in the image taken at a 120-degree angle by camera A1702 in Figure 17(a) is the same person as the person in the image taken at a 120-degree angle by camera B1703. Also, the person in the image taken at a 240-degree angle by camera A1702 in Figure 17(a) is not present in either of the images taken by camera B1703 in Figure 17(b), so it can be confirmed that the person is not present in all camera images.

[0295] In step S2608, the control unit 411 sets the shooting angle determination result for the camera capturing a person who is not present in the images of any of the cameras (i.e., a person who is present in the image of only one camera) to ON. In the examples of Figures 27 and 28, the shooting angle determination result for the 240-degree shooting angle in Figure 28(a), which corresponds to the 240-degree shooting angle of camera A1702 in Figure 17(a) capturing the image of person 2 in Figure 27(a), is set to ON.

[0296] Steps S2609 to S2612 are for processing when the same person is present in all camera images.

[0297] In step S2609, the control unit 411 compares the shooting conditions of all cameras that are photographing the same person. If the same person is being photographed by multiple cameras, the control unit compares the shooting conditions of each camera and controls the system to shoot with the camera that has the best shooting conditions. The shooting conditions are the position and size of the person, with the best shooting conditions being when the subject is large and positioned close to the center of the field of view. Note that the shooting conditions are not limited to the position and size of the subject, but may also include the brightness of the face and facial expressions such as a smile.

[0298] In step S2610, the control unit 411 determines the camera with the best shooting conditions based on the comparison results in step S2609.

[0299] In step S2611, the control unit 411 sets the camera shooting feasibility determination result, which was determined in step S2610, to ON.

[0300] In step S2612, the control unit 411 turns off the setting for enabling or disabling shooting for cameras other than the one determined in step S2610.

[0301] In the examples in Figures 27 and 28, the shooting conditions of camera A1702 in Figure 17(a), which is taking the image of person 1 in Figure 27(a), are compared with the shooting conditions of camera B1703 in Figure 17(b), which is taking the image of person 1 in Figure 27(b). Camera A1702 in Figure 17(a) is determined to have better shooting conditions because it captures person 1 closer to the center of the field of view. As a result, the shooting feasibility determination result for a 120-degree shooting angle in Figure 28(a), corresponding to camera A1702 in Figure 17(a), is set to ON, and the shooting feasibility determination result for a 120-degree shooting angle in Figure 28(b), corresponding to camera B1703 in Figure 17(b), is set to OFF.

[0302] In step S2613, the control unit 411 refers to the personal authentication result in step S2602 (subject determination result in Figure 27) and determines whether the processing in steps S2606 to S2612 has been performed for all subjects captured by all cameras. If it is determined that the processing has been performed for all subjects captured by all cameras, the process ends. If it is determined that the processing has not been performed for all subjects captured by all cameras, the process returns to step S2606 and continues until the processing for all subjects from all cameras is completed.

[0303] The control described above is an example of changing the shooting angle by panning the camera, but it is also possible to change the shooting angle by tilting the camera or by combining panning and tilting. In this case, a table of shooting feasibility determination results should be created for tilt shooting angles, similar to the panning method.

[0304] Additionally, the system may be controlled to increase the frequency of shooting pre-registered main subjects, or to perform control whenever the camera's position changes.

[0305] As described above, according to this embodiment, it is possible to take pictures without overlapping subjects or shooting ranges among multiple cameras.

[0306] Furthermore, the various controls described above, which are performed by the control unit, may be performed by a single piece of hardware, or multiple pieces of hardware may share the processing to control the entire device.

[0307] [Other embodiments] This embodiment can also be implemented by supplying a program that implements one or more of the functions of the above-described embodiment to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions.

[0308] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention.

[0309] The disclosures herein include the following information processing devices, control methods, and programs. [Item 1] A communication means for communicating with multiple imaging devices whose shooting direction can be changed, An acquisition means for acquiring images captured from the aforementioned plurality of imaging devices for each shooting direction, A comparison means for comparing images acquired from the aforementioned multiple imaging devices, An information processing apparatus characterized by having a control means for controlling the shooting directions of the plurality of imaging devices so as not to overlap, based on the results of the comparison. [Item 2] The information processing device according to item 1, characterized in that the control means determines the shooting direction of the target to be photographed by the plurality of imaging devices from images acquired from the plurality of imaging devices. [Item 3] The system includes a processing means for performing subject detection processing on images acquired from the aforementioned multiple imaging devices, The information processing apparatus according to item 2, characterized in that the control means determines the shooting direction of the target to be photographed by the plurality of imaging devices based on the result of the subject detection process. [Item 4] The information processing apparatus according to item 3, characterized in that the control means determines the shooting direction of the image in which the subject is detected by the subject detection process to be the shooting direction of the target to be photographed. [Item 5] The information processing apparatus according to item 3 or 4, characterized in that the control means determines that the shooting direction of an image in which no subject has been detected by the subject detection process is an outside shooting direction. [Item 6] The control means determines whether or not to take a photograph for each shooting direction of the target to be photographed by the plurality of imaging devices, The information processing device according to item 5, characterized in that the possibility of taking the aforementioned photograph includes both the possibility of taking the photograph and the possibility of not taking the photograph. [Item 7] The control means determines whether the images acquired from the plurality of imaging devices are similar or not. The number of shooting directions of the target object captured by the multiple imaging devices that have captured similar images is compared, The information processing device according to item 6, characterized in that, if the number of shooting directions for the target to be photographed by the multiple imaging devices is the same, the shooting direction of the target to be photographed by the imaging device that took the image in which the position of the subject is closer to the center of the field of view among the images acquired from the multiple imaging devices is set to be photographable, and the shooting direction of the target to be photographed by the imaging device that took the image in which the position of the subject is not close to the center of the field of view is set to be unphotographable. [Item 8] The control means determines whether the images acquired from the plurality of imaging devices are similar or not. The number of shooting directions of the target object captured by the multiple imaging devices that have captured similar images is compared, The information processing device according to item 6, characterized in that, if the number of shooting directions for the target to be photographed differs among the multiple imaging devices, the imaging device with fewer shooting directions for the target to be photographed is set to enable shooting, and the imaging device with more shooting directions for the target to be photographed is set to disable shooting. [Item 9] The control means determines whether the images acquired from the plurality of imaging devices are similar or not. The information processing apparatus according to item 6, characterized in that, if the images acquired from the plurality of imaging devices are not similar, the control means is set to enable the shooting direction of the target object of all imaging devices that captured the dissimilar images. [Item 10] The information processing apparatus according to any one of items 7 to 9, characterized in that the control means calculates the similarity of images acquired from the plurality of imaging devices and determines that the images acquired from the plurality of imaging devices are similar when the similarity is higher than a threshold. [Item 11] The information processing device according to any one of items 6 to 10, characterized in that the determination of whether or not to take the image is performed each time the position of any of the plurality of imaging devices is changed. [Item 12] The information processing device according to item 6, characterized in that the control means sets the shooting direction determined to be outside the target of shooting for the plurality of imaging devices to be unavailable for shooting. [Item 13] The control means determines whether the same subject exists in the images acquired from the plurality of imaging devices, The shooting conditions of the multiple imaging devices that photographed the same subject were compared, The information processing device according to any one of items 7 to 9, characterized in that it sets the shooting direction of the object to be photographed by the imaging device to be photographable when the aforementioned shooting conditions satisfy predetermined conditions, and sets the shooting direction of the object to be photographed by the imaging device to be photographable when the aforementioned conditions do not satisfy predetermined conditions. [Item 14] The aforementioned shooting conditions are the position and size of the same subject within the field of view. The information processing device according to item 13, characterized in that the predetermined conditions are that the same subject is close to the center of the field of view and the size of the same subject is large. [Item 15] The information processing apparatus according to item 13 or 14, characterized in that, if the same subject does not exist in the images acquired from the plurality of imaging devices, the control means sets the shooting direction of the target of the plurality of imaging devices that have taken images containing different subjects to be able to be photographed. [Item 16] The information processing device according to any one of items 13 to 15, characterized in that the control means integrates the result of determining whether or not to take the photograph based on the similarity of the images and the result of determining whether or not to take the photograph based on the presence of the same subject, and sets the final determination of whether or not to take the photograph for each shooting direction of the target to be photographed by the plurality of imaging devices. [Item 17] The information processing device according to any one of items 7 to 9, characterized in that the control means zooms in with respect to the shooting direction of the imaging device which is set to not be able to shoot. [Item 18] The plurality of imaging devices are rotatable in the pan and tilt directions. The information processing device according to any one of items 1 to 17, characterized in that the shooting direction is a shooting angle divided into predetermined angles in the pan direction and tilt direction. [Item 19] The information processing apparatus according to item 18, characterized in that the plurality of imaging devices are capable of capturing images in all directions. [Item 20] The plurality of imaging devices perform a first photograph based on a photographing instruction from the information processing device or a second photograph not based on the photographing instruction. The aforementioned shooting instructions include shooting instructions for each shooting angle. The information processing device according to item 18 or 19, characterized in that the plurality of imaging devices perform photography at a shooting angle set to be photographable based on the shooting instruction. [Item 21] The information processing device according to item 20, characterized in that the plurality of imaging devices automatically perform imaging at an imaging angle set to be possible in the second imaging. [Item 22] The information processing device according to any one of items 1 to 20, characterized in that the information processing device is an external device different from the plurality of imaging devices or one of the plurality of imaging devices. [Item 23] A control method for an information processing device that communicates with multiple imaging devices capable of changing the shooting direction, The steps include acquiring images captured from the plurality of imaging devices for each shooting direction, The steps include comparing images acquired from the aforementioned multiple imaging devices, A control method characterized by comprising the step of controlling the shooting directions of the plurality of imaging devices so as not to overlap, based on the results of the comparison. [Item 24] A program that causes a computer to function as an information processing device as described in any of items 1 through 22. [Explanation of Symbols]

[0310] 101, 101a, 101b... Imaging device, 222... Communication unit, 301... External device, 401... Wireless LAN control unit, 402... BLE control unit, 403... Packet transmission / reception unit, 411... Control unit

Claims

1. A communication means for communicating with multiple imaging devices whose shooting direction can be changed, An acquisition means for acquiring images captured from the aforementioned plurality of imaging devices for each shooting direction, A comparison means for comparing images acquired from the aforementioned multiple imaging devices, An information processing apparatus characterized by having a control means for controlling the shooting directions of the plurality of imaging devices so as not to overlap, based on the results of the comparison.

2. The information processing apparatus according to claim 1, characterized in that the control means determines the shooting direction of the target to be photographed by the plurality of imaging devices from images acquired from the plurality of imaging devices.

3. The system includes a processing means for performing subject detection processing on images acquired from the aforementioned multiple imaging devices, The information processing apparatus according to claim 2, characterized in that the control means determines the shooting direction of the target to be photographed by the plurality of imaging devices based on the result of the subject detection process.

4. The information processing apparatus according to claim 3, characterized in that the control means determines the shooting direction of the image in which the subject was detected by the subject detection process to be the shooting direction of the target to be photographed.

5. The information processing apparatus according to claim 3, characterized in that the control means determines that the shooting direction of an image in which no subject has been detected by the subject detection process is an outside shooting direction.

6. The control means determines whether or not to take a photograph for each shooting direction of the target to be photographed by the plurality of imaging devices, The information processing device according to claim 5, characterized in that the possibility of taking the photograph includes the possibility of taking the photograph and the possibility of not taking the photograph.

7. The control means determines whether the images acquired from the plurality of imaging devices are similar or not. The number of shooting directions of the target object captured by the multiple imaging devices that have captured similar images is compared, The information processing device according to claim 6, characterized in that, if the number of shooting directions for the target to be photographed by the multiple imaging devices is the same, the shooting direction of the target to be photographed by the imaging device that took the image in which the position of the subject is closer to the center of the field of view is set to be photographable, and the shooting direction of the target to be photographed by the imaging device that took the image in which the position of the subject is not close to the center of the field of view is set to be unphotographable.

8. The control means determines whether the images acquired from the plurality of imaging devices are similar or not. The number of shooting directions of the target object captured by the multiple imaging devices that have captured similar images is compared, The information processing apparatus according to claim 6, characterized in that, if the number of shooting directions for the target to be photographed differs among the multiple imaging devices, the imaging device with fewer shooting directions for the target to be photographed is set to enable shooting, and the imaging device with more shooting directions for the target to be photographed is set to disable shooting.

9. The control means determines whether the images acquired from the plurality of imaging devices are similar or not. The information processing apparatus according to claim 6, wherein the control means is configured to enable shooting in the shooting direction of the target object of all imaging devices that captured the dissimilar images if the images acquired from the plurality of imaging devices are not similar.

10. The information processing apparatus according to claim 7, characterized in that the control means calculates the similarity of images acquired from the plurality of imaging devices and determines that the images acquired from the plurality of imaging devices are similar when the similarity is higher than a threshold.

11. The information processing device according to claim 6, characterized in that the determination of whether or not to take the photograph is performed each time the position of any of the plurality of imaging devices is changed.

12. The information processing apparatus according to claim 6, characterized in that the control means sets the shooting direction determined to be outside the target of shooting for the plurality of imaging devices to be unavailable for shooting.

13. The control means determines whether the same subject exists in the images acquired from the plurality of imaging devices, The shooting conditions of the multiple imaging devices that photographed the same subject were compared, The information processing apparatus according to claim 7, characterized in that it sets the shooting direction of the object to be photographed by the imaging device to be photographable when the aforementioned shooting conditions satisfy predetermined conditions, and sets the shooting direction of the object to be photographed by the imaging device to be photographable when the aforementioned conditions do not satisfy predetermined conditions.

14. The aforementioned shooting conditions are the position and size of the same subject within the field of view. The information processing apparatus according to claim 13, characterized in that the predetermined conditions are that the same subject is close to the center of the field of view and the size of the same subject is large.

15. The information processing apparatus according to claim 13, wherein the control means, if there is no identical subject in the images acquired from the plurality of imaging devices, is set to enable the shooting direction of the target object captured by the plurality of imaging devices that have captured images containing different subjects.

16. The information processing apparatus according to claim 13, characterized in that the control means integrates the result of determining whether or not to take the photograph based on the similarity of the images and the result of determining whether or not to take the photograph based on the presence of the same subject, and sets the final determination of whether or not to take the photograph for each shooting direction of the target to be photographed by the plurality of imaging devices.

17. The information processing apparatus according to claim 7, characterized in that the control means zooms in with respect to the shooting direction of the imaging device which is set to not be able to shoot.

18. The plurality of imaging devices are rotatable in the pan and tilt directions. The information processing apparatus according to claim 1, characterized in that the shooting direction is a shooting angle divided into predetermined angles in the pan direction and tilt direction.

19. The information processing apparatus according to claim 18, characterized in that the plurality of imaging devices are capable of capturing images in all directions.

20. The plurality of imaging devices perform a first photograph based on a photographing instruction from the information processing device or a second photograph not based on the photographing instruction. The aforementioned shooting instructions include shooting instructions for each shooting angle. The information processing apparatus according to claim 18, characterized in that the plurality of imaging devices take photographs at shooting angles set to be photographable based on the shooting instruction.

21. The information processing apparatus according to claim 20, characterized in that the plurality of imaging devices automatically perform imaging at an imaging angle set to be capable of imaging during the second imaging.

22. The information processing device according to claim 1, characterized in that the information processing device is an external device different from the plurality of imaging devices or one of the plurality of imaging devices.

23. A control method for an information processing device that communicates with multiple imaging devices capable of changing the shooting direction, The steps include acquiring images captured from the plurality of imaging devices for each shooting direction, The steps include comparing images acquired from the aforementioned multiple imaging devices, A control method characterized by comprising the step of controlling the shooting directions of the plurality of imaging devices so as not to overlap, based on the results of the comparison.

24. A program for causing a computer to function as an information processing device as described in any one of claims 1 to 22.

Citation Information

Patent Citations

  • Camera and camera system

    JP2013223104A

  • Information processing apparatus and control method of the same

    JP2020022052A