Image pickup apparatus and control method thereof
The imaging device uses a prediction model to dynamically adjust shooting thresholds based on user preferences, optimizing shot frequency and interval for desired images.
Patent Information
- Application Number
- JP2024111705
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2026-01-23
AI Technical Summary
Existing automatic cameras fail to account for individual user preferences, leading to unnecessary shots of subjects that are not desired by the purchaser, potentially missing important subjects due to uniformly applied shooting thresholds.
An imaging device that utilizes a prediction model trained on user purchase history and image data to determine optimal shooting frequency and interval based on subject and scene information, adjusting thresholds dynamically to meet user preferences.
The system effectively predicts the optimal number of shots and intervals for each subject, ensuring that desired images are captured while minimizing unnecessary shots.
Smart Images

Figure 2026011250000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, and more particularly to an automatic photography technique. [Background technology]
[0002] When taking still or video images using an imaging device such as a camera, it is common for the photographer to determine the subject of the image through a viewfinder, check the shooting conditions themselves, adjust the framing of the image, and then take the image. Such imaging devices have traditionally been equipped with mechanisms for detecting user operation errors and the external environment, and notifying the user if the environment is not suitable for shooting, or for detecting the external environment and controlling the camera to bring the environment into a suitable state for shooting.
[0003] In contrast to such imaging devices that perform shooting by user operation, there is known an automatic camera (Patent Document 1) that takes pictures periodically and continuously without the user issuing a shooting instruction. The automatic camera sets a shooting threshold based on the shooting history and the evaluation value of the captured image, controls the shooting frequency, and is controlled to capture the subject evenly while minimizing the missed shots during automatic shooting.
[0004] Furthermore, various technologies have been proposed that allow purchasers to purchase, via the Internet, images taken by third parties that have been electronically published on websites on the Internet, etc. (for example, Patent Document 2). By using these technologies, users can view multiple images at home, select their favorite images, and obtain only the images they need.
[0005] In recent years, services have been provided that sell images taken at kindergartens and nursery schools to potential buyers such as parents. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Patent Publication No. 2021-57815 [Patent Document 2] Special Publication No. 2004-514976 Summary of the Invention [Problem to be solved by the invention]
[0007] The method described in Patent Document 1 uniquely calculates an image evaluation value from the orientation and size of the face and the degree of smile, and determines the number of images to be taken and the interval between shots.
[0008] However, since the number of images and intervals between images desired by each user differ, there is a problem in that a uniquely determined image evaluation value cannot reflect the user's preferences.
[0009] Furthermore, because automatic cameras are controlled to take pictures evenly depending on the subject, they may take pictures of multiple subjects evenly. This can result in taking pictures of subjects that are unnecessary for the person wanting to purchase the photo. Furthermore, if too many unnecessary subjects are taken within the limited number of shots, there is a risk that important subjects may be missed.
[0010] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to provide an imaging device that enables various automatic shooting according to user preferences, thereby enabling the user to take the photograph they desire. [Means for solving the problem]
[0011] In order to achieve the above object, the present invention provides a photograph purchasing system comprising an imaging means for imaging a subject and outputting image data, and a storage means for storing the image data output by the imaging means, the system comprising: a calculation means for calculating an evaluation value used to determine whether or not to perform a photographing operation with the imaging means; a first setting means for setting a first threshold used to determine whether or not to perform a photographing operation with the imaging means; a determination means for determining whether or not to perform a photographing operation using the evaluation value and the first threshold; a second threshold used to determine whether or not to control the first threshold; a processing means for referring to the purchase history of image data stored in the purchasing system and performing processing based on the number of purchases by the user, subject information of the stored photographs, and scene information; and a prediction model machine-learned using as training data the number of purchases by the user and the sales volume of the image data predicted from the subject information and scene information of the stored photographs stored in the purchasing system, to control the appropriate first threshold and photographing interval for each subject. [Effects of the Invention]
[0012] According to the present invention, by controlling the number of shots and the interval between shots according to the purchase history, it is possible to provide a photography system that can predict the optimal number of shots and the optimal interval between shots for each subject and provide the optimal number of shots. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram schematically illustrating a system configuration to which the present embodiment can be applied. [Figure 2] FIG. 1 is a diagram schematically illustrating an imaging device used in the present embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of an imaging device and a server according to an embodiment of the present invention. [Figure 4] FIG. 2 is a conceptual diagram illustrating the structure of a learning model in an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating an example of operation when a learning model is used in an embodiment of the present invention. [Figure 6]FIG. 1 is a diagram illustrating an example of the operation of a system according to an embodiment of the present invention. [Figure 7] FIG. 2 is a table illustrating an example of learning data according to an embodiment of the present invention. [Figure 8] FIG. 10 is a flowchart illustrating an example of a learning operation according to an embodiment of the present invention. [Figure 9] FIG. 10 is a flowchart illustrating an example of an estimation operation according to an embodiment of the present invention. [Figure 10] FIG. 3 is a flowchart illustrating an example of the operation of the imaging device according to the first embodiment of the present invention. [Figure 11] 10A and 10B are diagrams illustrating examples of the relationship between image evaluation values and threshold values in the present embodiment. [Figure 12] FIG. 10 is a flowchart illustrating an example of the operation of the imaging device according to the second embodiment of the present invention. [Figure 13] FIG. 11 is a flowchart illustrating an example of the operation of the imaging device according to the third embodiment of the present invention. [Figure 14] FIG. 10 is a flowchart illustrating an example of the operation of the imaging device according to the fourth embodiment of the present invention. [Figure 15] FIG. 1 is a diagram schematically illustrating a smart device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] [First embodiment] <System configuration> FIG. 1 is a diagram showing a schematic diagram of a system configuration to which the present invention can be applied.
[0015] 1, a client terminal 104, a content providing server 105, a data collection server 106, and a learning server 107 are each connected to the Internet 100. An imaging device 102 and a communication terminal 103 are each connected to a local network 101. The local network 101 is connected to the Internet 100, and the system of this embodiment is configured.
[0016] The local network 101 is a network in which an imaging device 102 and a communication terminal 103 are connected by cable or wirelessly. The imaging device 102 and the communication terminal 103 can exchange information with each other via the local network 101. The local network 101 is also connected to the Internet 100.
[0017] The imaging device 102 has a drive mechanism (described later) and is a camera capable of autonomously capturing images without user operation. The imaging device 102 can change camera control using learning information (described later) acquired from the learning server 107 via the local network 101 to create still images and videos. Still images and videos captured by the imaging device 102 are stored in the data collection server 106 from the communication terminal 103 via the local network 101. Note that this example is not particularly limited, and for example, still images and videos may be stored directly from the imaging device 102 in the data collection server 106 via the local network 101. Also, although one imaging device 102 is illustrated in FIG. 1, multiple imaging devices may be connected to the local network 101.
[0018] The communication terminal 103 is a communication terminal that has a communication function and can be connected to the imaging device 102. The communication function of the communication terminal 103 is realized by a wireless LAN communication module or a wired communication module. The communication terminal 103 has an operation mechanism that accepts user operations, and is capable of transmitting instructions based on information input to the operation mechanism to the imaging device 102 via the wireless LAN communication module or the wired communication module. Furthermore, the communication terminal 103 is also capable of obtaining information from the learning server 107 via the Internet 100 and transmitting it to the imaging device 102. It also obtains still images and videos captured by the imaging device 102 and transmits them to the data collection server 106 via the Internet 100.
[0019] The client terminal 104 is a computer that allows a user to select and purchase still images and videos provided by the content providing server 105 using a web browser.
[0020] The content providing server 105 is a computer that provides still images and videos stored in the data collection server 106 in a viewable and selectable format via the Internet 100. The content providing server 105 provides still images and videos that include the user himself / herself or the user's family or acquaintances as subjects to the client terminal 104 used by the user. The operator of the content providing server 105 according to this embodiment provides a service of providing and selling still images and videos captured by the imaging device 102 using an imaging sequence described below to users. The subjects of the still images and videos purchased by users from the operator are not necessarily the user himself / herself or family or acquaintances related to the user. However, for the sake of convenience, the following explanation will be given assuming that the still images and videos made available for sale by the operator to users are the user himself / herself or the user's family or acquaintances. The content providing server 105 also transmits the user's purchase information to the learning server 107.
[0021] The data collection server 106 is a computer that collects still images and moving images captured by the imaging device 102 and transmits still images or moving images requested by the content providing server 105 to the content providing server 105 .
[0022] The learning server 107 is a computer that generates a trained model using a learning method described below, using user purchase information transmitted from the content providing server 105. The learning server 107 transmits the trained model to the data collection server 106 via the Internet 100. The data collection server 106 can provide the trained model in response to a request from the imaging device 102 via the Internet 100.
[0023] <Configuration of imaging device> FIG. 2 is a diagram schematically showing an imaging device used in this embodiment.
[0024] The imaging device 102 shown in FIG. 2(a) is provided with an operating member for operating a power switch (hereinafter referred to as a power button, but operations such as tapping, flicking, or swiping on a touch panel may also be used). A lens barrel 202, which is a housing containing a group of photographic lenses and an image sensor for capturing images, is attached to the imaging device 102 and is provided with a rotation mechanism that can rotate the lens barrel 202 relative to a fixed part 203. The tilt rotation unit 204 is a motor-driven mechanism that can rotate the lens barrel 202 in the pitch direction shown in FIG. 2(b), and the pan rotation unit 205 is a motor-driven mechanism that can rotate the lens barrel 202 in the yaw direction. Thus, the lens barrel 202 can rotate in two or more axial directions. Note that FIG. 2(b) defines the axes of the fixed part 203 in the installation state shown in FIG. 2(a). Both the gyroscope 206 and the accelerometer 207 are mounted on the fixed part 203 of the imaging device 102. Then, a device vibration detection unit 309 (described later) detects vibrations of the imaging device 102 based on the angular velocity meter 206 and the accelerometer 207, and drives the tilt rotation unit 204 and the pan rotation unit 205 to rotate based on the detected vibration angle. This allows for the configuration to correct the vibration and tilt of the lens barrel 202, which is the movable part.
[0025] FIG. 3 is a block diagram showing the hardware configuration of each device constituting the system of FIG. 1 according to this embodiment.
[0026] 3, the first control unit 323 is made up of a processor (e.g., CPU, microprocessor, MPU, etc.) and memory (e.g., DRAM, SRAM, etc.). These execute various processes to control each block of the imaging device 102 and control data transfer between each block. The non-volatile memory (Flash ROM) 316 is an electrically erasable and recordable memory, and stores constants, programs, etc. for the operation of the first control unit 323.
[0027] 3, a zoom unit 301 includes a zoom lens that changes magnification. A zoom drive control unit 302 controls the drive of the zoom unit 301. A focus unit 303 includes a lens that adjusts focus. A focus drive control unit 304 controls the drive of the focus unit 303.
[0028] In the imaging unit 306, the imaging element receives light incident through each lens group, performs analog-to-digital conversion (A / D conversion) on the charge information corresponding to the amount of light, and outputs the result as a digital signal to the image processing unit 307. The image processing unit 307 generates digital image data from the received digital signal. Furthermore, the image processing unit 307 applies image processing such as distortion correction, white balance adjustment, and color interpolation processing to the digital image data, and outputs the digital image data after the processing. The digital image data output from the image processing unit 307 is converted into a recording format such as JPEG format by the image encoding unit 308, and is sent to the memory 315 or a video output unit 317, which will be described later.
[0029] The lens barrel rotation drive unit 305 drives the tilt rotation unit 204 and pan rotation unit 205 to drive the lens barrel 202 in the tilt direction and pan direction.
[0030] The device vibration detection unit 309 is equipped with, for example, an angular velocity meter 206 that detects the angular velocity in three axial directions of the imaging device 102, and an accelerometer 207 that detects the acceleration in three axial directions of the device. The device vibration detection unit 309 calculates the rotation angle of the device, the amount of shift of the device, etc. based on the detected signals.
[0031] The audio input unit 313 acquires audio signals from the surroundings of the imaging device 102 from a microphone provided in the imaging device 102, performs A / D conversion, and transmits the signals to an audio processing unit 314. The audio processing unit 314 performs audio-related processing such as optimization of the input digital audio signals. The audio signals processed by the audio processing unit 314 are then transmitted to a memory 315 by a first control unit 323. The memory 315 temporarily stores the image signals and audio signals obtained by the image processing unit 307 and audio processing unit 314.
[0032] The image processing unit 307 and the audio processing unit 314 read out the image signals and audio signals temporarily stored in the memory 315 and encode the image signals and audio signals to generate compressed image signals and compressed audio signals. The first control unit 323 transmits these compressed image signals and compressed audio signals to the recording and playback unit 320.
[0033] The recording / playback unit 320 records the compressed image signal, compressed audio signal, and other control data related to shooting generated by the image processing unit 307 and audio processing unit 314 on the recording medium 321. When the audio signal is not compression-encoded, the first control unit 323 transmits the audio signal generated by the audio processing unit 314 and the compressed image signal generated by the image processing unit 307 to the recording / playback unit 320, and causes them to be recorded on the recording medium 321.
[0034] The recording medium 321 may be a recording medium built into the imaging device 102 or a removable recording medium. The recording medium 321 can record various types of data such as compressed image signals, compressed audio signals, and audio signals generated by the imaging device 102, and generally uses a medium with a larger capacity than the nonvolatile memory 316. For example, the recording medium 321 includes any type of recording medium, such as a hard disk, optical disk, magneto-optical disk, CD-R, DVD-R, magnetic tape, nonvolatile semiconductor memory, flash memory, etc.
[0035] The recording / playback unit 320 reads (plays back) compressed image signals, compressed audio signals, audio signals, various data, and programs recorded on the recording medium 321. Then, the first control unit 323 transmits the read compressed image signals and compressed audio signals to the audio processing unit 314. The image processing unit 307 and the audio processing unit 314 temporarily store the compressed image signals and compressed audio signals in the memory 315, decode them in a predetermined procedure, and transmit the decoded signals to the video output unit 317 and the audio output unit 318.
[0036] The audio input unit 313 is equipped with multiple microphones, and the audio processing unit 314 can detect the direction of sound on a plane on which the multiple microphones are installed, and is used for searching and automatic photography, which will be described later. Furthermore, the audio processing unit 314 detects specific audio commands. The audio commands may be pre-registered commands, or the user may be able to register specific audio in the imaging device. The imaging device also performs audio scene recognition. In audio scene recognition, audio scenes are determined using a network trained using machine learning based on a large amount of pre-stored audio data. For example, a network for detecting specific scenes such as "cheers," "clapping," and "vocalization" is configured in the audio processing unit 314. When a specific audio scene or a specific audio command is detected, a detection trigger signal is output to the first control unit 323 or the second control unit 311.
[0037] The second control unit 311 is a control unit provided separately from the first control unit 323 and controls the power supply to the imaging device 102 .
[0038] The first power supply unit 310 and the second power supply unit 312 each supply power to operate the audio output unit 318 and the recording / playback unit 320. The second power supply unit 312 is always on when external power is supplied from a battery or an adapter or USB, and supplies power to the second control unit 311. The second control unit 311 is always in an operating state when power is supplied, and performs a determination process of whether to start the first control unit 323 based on pressing a power button provided on the first imaging device 102 or input information from the device vibration detection unit 309 or the audio processing unit 314. The first power supply unit 310 starts up based on the start-up determination result of the second control unit 311, and supplies power to the first control unit 323. When power is supplied, the first control unit 323 starts operating in accordance with the start-up factor of the second control unit 311.
[0039] The audio output unit 318 outputs a preset audio pattern from a speaker built into the image capture device 102, for example, during shooting.
[0040] The LED control unit 324 controls the LEDs provided in the image capturing device 102 so that they turn on and off in a predetermined pattern, for example, when capturing an image.
[0041] Video output unit 317 is, for example, a video output terminal, and transmits an image signal to display a video on a connected external display, etc. Audio output unit 318 and video output unit 317 may be combined into one terminal, such as an HDMI (registered trademark) (High-Definition Multimedia Interface) terminal.
[0042] The communication unit 322 communicates between the imaging device 102 and external devices, transmitting and receiving data such as audio signals and image signals. It also receives control signals related to imaging, such as commands to start and stop imaging and pan, tilt, and zoom, and drives the imaging device 102 based on instructions from external devices that can communicate with the imaging device 102. It also transmits and receives information between the imaging device 102 and external devices, such as various parameters related to learning, which are processed by the learning processing unit 319 (described later). The communication unit 322 is a wireless communication module, such as a Bluetooth communication module conforming to the Bluetooth standard or a wireless LAN communication module conforming to the IEEE 802.11 standard. The imaging device 102 can communicate with external devices via the local network 101 and the Internet 100 by using the communication unit 322. When multiple imaging devices 102 used by the same user are present on the same local network 101, they communicate with each other using the communication unit 322 to operate cooperatively.
[0043] When the image capture device 102 is running and the face authentication unit 325 finds a nearby human face, it performs face authentication processing on the found face and determines whether it matches the human face image received from the learning server 107. Types of face authentication include, for example, 2D authentication, which recognizes the positions of the eyes, nose, mouth, etc. on the face and compares them with a database for authentication. In addition to 2D authentication, there are also 3D authentication methods that use an infrared sensor or dot projector for authentication.
[0044] The setting unit 326 sets a shooting release condition according to the face information of the person authenticated by the face authentication unit 325. The shooting release condition is determined based on an image evaluation value and a shooting threshold value. The image evaluation value is a value that serves as an index of whether the scene captured by the image capture device 102 is suitable for shooting, and is determined based on face detection information, face authentication information, blink rate, facial expression of the subject, facial orientation, size of the subject, etc. The shooting threshold value is a criterion value that, when the image evaluation value exceeds this value, triggers shooting, and is determined based on the shooting frequency, the time elapsed since the previous shooting, etc.
[0045] The GPU 327 is a processor that performs arithmetic processing of video signals and is a calculation unit that can process input data in parallel. The GPU 327 is configured using, for example, a GPU (Graphical Processing Unit) or an FPGA (Field-Programmable Gate Array).
[0046] The GPU 327 can perform efficient calculations by processing a larger amount of data in parallel, and therefore, when calculations are performed multiple times using a learning model such as deep learning, it is effective to use the GPU 327 for processing. Therefore, in this embodiment, the GPU 327 is used in addition to the first control unit 323 for processing by the learning processing unit 319. Specifically, when a prediction program is executed, the first control unit 323 and the GPU 327 work together to perform calculations to make an estimation.
[0047] The scene discrimination unit 328 is a discrimination unit that discriminates the captured scene of an image from the image signal of the image processing unit 307. For example, the scene discrimination unit 328 inputs the image signal into a scene discrimination trained model (not shown) and discriminates the captured scene from the output result. Here, the scene discrimination trained model is a trained model that trains the image signal as input data using scenes such as sports days, play days, and entrance ceremonies as training data, and outputs a scene similar to the input image signal as output data. The scene discrimination unit 328 may also discriminate the captured scene by receiving scene information from the communication terminal 103 or by using scene designation information input by the user from the UI display unit 502.
[0048] In this embodiment, the information processing device capable of communicating with the imaging device 102 via a network is a device equipped with similar hardware, and the data collection server 106 will be described as a representative example. The CPU 332 is a central processing unit (CPU) capable of executing various processes, and is a processor capable of interpreting instructions from the user and instructions from other devices and executing program code. The system bus 331 is a bus for exchanging data between various pieces of hardware within the information processing device, including the CPU 332. The ROM 333 is a non-volatile memory that records various setting information for the information processing device, and is configured, for example, by an EEPROM or Flash ROM, and stores programs such as a BIOS (Basic Input Output System). The RAM 334 is the main storage device of the information processing device, and is used when the CPU 332 temporarily stores data to be processed and when the CPU 332 loads programs. The RAM 334 is configured, for example, by a dynamic random access memory (DRAM). The HDD 335 is an auxiliary storage device of the information processing device. The HDD 335 is configured using hardware such as a hard disk drive (HDD) or a solid state drive (SSD). The HDD 335 can be used to store the operating system (OS) of the information processing device, programs executed by the CPU 332, various data input to the programs, and program execution results.
[0049] The GPU 336 is a processor responsible for arithmetic processing of video signals and the like of the information processing device. The display unit 338 is a display unit capable of displaying video signals processed by the GPU 336, and is configured, for example, by an LCD (Liquid Crystal Display), and may be integrated with or separate from the information processing device. The input unit 337 is an input unit through which the user issues instructions to the information processing device, and may be, for example, a keyboard, mouse, or touch panel. The communication unit 339 is a communication unit used by the information processing device to communicate with external devices. The communication unit 339 is configured, for example, by a wireless LAN communication module conforming to the IEEE 802.11 standard or a wired LAN communication module conforming to the IEEE 802.3 standard. The communication terminal 103, the content providing server 105, and the learning server 107 have the same hardware configuration as the data collection server 106, and therefore, depending on the scale of the system, a single information processing device may share multiple functions. In this embodiment, the content providing server 105, the data collecting server 106, and the learning server 107 are independent devices, and are capable of communicating with each other via the communication unit 339 over a network.
[0050] FIG. 4 is a conceptual diagram showing an input / output structure using a trained model in this embodiment. Input data X401 is input data to be input to a neural network to be trained. Output data Y402 is output data obtained when input data X401 is input to the neural network. Trained model 403 is an example of a neural network that is a training model. Trained model 403 is made up of a network consisting of a multilayer perceptron, for example, in which a large number of neuron models called units are connected.
[0051] The trained model 403 is used to predict output data Y402 from input data X401. By training the trained model 403 in advance so as to output a teacher output value for input data, it is possible to predict output data for new input data that follows the trained teacher. Learning in this embodiment will be described later.
[0052] FIG. 5 is a diagram illustrating the software functions of each part of a system to which the present invention is applied, using the trained model shown in FIG.
[0053] The imaging device 102 includes a data transmission / reception unit 501, a UI display unit 502, and an estimation unit 503. The client terminal 104 includes a web browser 511.
[0054] The content providing server 105 includes a data presenting unit 521 , a content-related data managing unit 522 , and a data storing unit 523 .
[0055] The data collection server 106 includes a data receiving unit 531 , a data collecting and providing unit 532 , a data storage unit 533 , and an estimation model providing unit 534 .
[0056] The learning server 107 includes a learning unit 541 , a learning data generation unit 542 , a data storage unit 543 , and an estimation model output unit 544 .
[0057] The data transmission / reception unit 501 has the function of transmitting and receiving content data such as still images and videos, and transmitting and receiving estimation model data between the imaging device 102 and the data collection server 106, and executes this function using the communication unit 322. The data transmission / reception unit 501 only needs to be compatible with a communication protocol for transmitting and receiving files between a server and a client, such as an FTP (File Transfer Protocol) client, and can also be implemented using other protocols or proprietary protocols.
[0058] The estimation unit 503 has a function of inputting necessary input data to a trained model that has undergone training and performing predictive calculations to obtain output data. The estimation unit 503 applies the trained model received using the function of the data transmission / reception unit 501 using the communication unit 322, and performs predictive calculations through the cooperative operation of the first control unit 323 and the GPU 327.
[0059] The web browser 511 is software that a user uses to view and obtain content, such as still images and videos, stored in the content providing server 105 from the client terminal 104. The web browser 511 only needs to be able to send and receive data using a protocol based on HTTP (Hyper Text Transfer Protocol), for example, and display the results of the transmission and reception on the display unit 338 of the user's client terminal 104.
[0060] The data presentation unit 521 has a function for providing still images and videos stored in the content providing server 105 to the client terminal 104 used by the user. The data presentation unit 521 may have, for example, an HTTP server function, so long as it can receive requests from the client terminal 104 and transmit content and data in response to the received requests. The content-related data management unit 522 provides content stored in the data storage unit 523 to the data presentation unit 521 based on requests from the client terminal 104 received by the data presentation unit 521, and manages the provided content and content viewing information. The content-related data management unit 522 may manage and control content information by using, for example, a database function capable of operating a database written in a structured query language such as SQL.
[0061] The data storage unit 523 has a function of storing and providing content data that temporarily stores moving image and still image content to be provided to the client terminal 104 used by the user. The data storage unit 523 receives and stores still image and moving image content from the data collection server 106 so that it can provide content to the client terminal 104, and also has a function of generating and holding data for easy confirmation of content, such as thumbnail images.
[0062] The data receiving unit 531 has a data receiving function for receiving still image and video data captured by the imaging device 102. The data receiving unit 531 has, for example, an FTP server function and a function for storing still image and video data received using the communication unit 339 in the data storage unit 533.
[0063] The data collection and provision unit 532 manages the still image and video data stored in the data storage unit 533, and has the function of providing information on the still image and video data it manages in response to requests from the content providing server 105 and the learning server 107. The data collection and provision unit 532 may manage and control the information on the still image and video data it manages by using a database function that can operate a database written in a structured query language such as SQL, for example.
[0064] The data storage unit 533 has a function of storing and reading data in the auxiliary storage device shown as the HDD 335. The data storage unit 533 only needs to operate so as to be able to store still image and video data received by the data receiving unit 531 and to read out stored information based on instructions from the data collection and provision unit 532, for example.
[0065] The estimation model providing unit 534 is configured to provide a trained model to be used by the estimation unit 503 of the imaging device 102 in response to a request from the imaging device 102. The trained model provided from the estimation model providing unit 534 to the imaging device 102 can be obtained by acquiring the trained model output by the learning server 107 at a specific timing and making it possible to provide it to the imaging device 102.
[0066] The estimation model providing unit 534 may be configured to perform a function that allows the data of the trained model to be shared between the learning server 107 and the imaging device 102, for example, by using a function such as an FTP server.
[0067] 4, and operates to generate the trained model 403 using input data X401 and a teacher output value. The trained model 403 generated by the learning unit 541 is provided to the estimation model output unit 544, and then provided to the imaging device 102 via the data collection server 106. When learning in the learning unit 541, the learning unit 541 operates to execute learning by instructing the CPU 332 and GPU 336 to perform calculations.
[0068] The learning data generation unit 542 has a function of generating input data X401 to the learning unit 541 and output values that serve as teachers, requests information to be used for learning from the data collection server 106, and selects and processes the received information. The learning unit 541 performs learning using the input data X401 output by the learning data generation unit 542 and the output values that serve as teachers, and updates the trained model 403.
[0069] The data storage unit 543 has a function of storing and reading data from the auxiliary storage device shown in the HDD 335. The data storage unit 543 can store the learning input data X401 and the teacher output values generated by the learning data generation unit 542, and the trained model 403. The data storage unit 543 only needs to be able to operate, for example, to store the input data X401 and the teacher output values based on a request from the learning data generation unit 542, and to provide the stored data based on a request from the learning unit 541.
[0070] In response to a request from the data collection server 106, the estimation model output unit 544 provides the trained model 403 based on the request to the estimation model provision 534 of the data collection server 106. Furthermore, if the trained model 403 to be provided has not been updated since the last time it was provided, it may be deemed to have not been updated and its provision may be canceled.
[0071] As shown in Figure 5, the operation of this embodiment is realized by exchanging data between devices using software functions within each device. Note that while this embodiment illustrates a system in which each device exists independently and operates in coordination with others, it can also be realized by integrating the functions of the devices, for example by configuring the data collection server 106 and the learning server 107 as the same device. It can also be realized by making modifications, such as dividing the learning server 107 into multiple devices and allowing the learning calculations of a large-scale learning model to operate in parallel.
[0072] Fig. 6 is a diagram illustrating an example of the operation of the system in this embodiment. The operation of the system in this embodiment will be explained using Fig. 6. The following explanation illustrates the case where a user purchases still images and videos, which are content provided by the system, and the operation of each device in the present invention that operates in response to the purchase.
[0073] <Users can view and purchase still image and video content> The user operates the client terminal 104, browses the content providing server 105, selects and purchases the still image or video content they need. For example, thumbnails of a list of content provided by the content providing server 105 are displayed on the web browser on the client terminal 104, and the user selects each piece of content one by one to perform the purchase process.
[0074] The content providing server 105 receives a purchase process from the client terminal 104 and transmits to the data collection server 106, as purchase information, information indicating the purchased still image and video data and information indicating the user who made the purchase. The data collection server 106 may also obtain information on usage from the content providing server 105 regardless of purchase, for example, information on the history of content selected and viewed by the user. The data collection server 106 transmits the still image and video data to the content providing server 105 based on the information received from the content providing server 105. The content providing server 105 stores the still image and video data received from the data collection server 106 so that it can be provided to the user's client terminal 104. The content providing server 105 then transmits the still image and video data, which are the purchased content, in response to a request from the client terminal 104. Once the client terminal 104 has received all of the purchased content, it transmits a notification of reception completion to the content providing server 105, and the process for purchasing the content required by the user is completed. After the purchase is complete, the content providing server 105 discards the still image and video data that has already been provided, and if there is a request from the user again, the content providing server 105 makes a request to the data collecting server 106 in the same manner as described above and provides the data in the same manner.
[0075] <Learning model based on purchase information> The data collection server 106 transmits still image and video data related to the user and its purchase information to the learning server 107. The learning server 107 stores purchase information and image data of users who have previously made purchases, selected and processed for learning purposes. When new purchase information and related still image and video data are received, the learning server 107 stores the new purchase information and related still image and video data as learning data. The learning server 107 processes the purchase information and related still image and video data received from the data collection server 106 as learning data through processing by the learning data generation unit 542. In the case of still image data, the learning data generation unit 542 processes and selects the data by resizing it to the required size or selecting the subjects included in the data by linking it to the purchase information, and then passes the processed and selected data to the data storage unit 543 for storage as learning data. In addition, in the case of video data, the learning data generation unit 542 of the learning server 107 extracts video frames, processes and selects them in the same way as for still images, and passes them to the data storage unit 543 for storage as learning data. Here, while a video is composed of multiple frames, learning data is not acquired from all frames; only representative frames are used. The representative frame selected here is, for example, the first frame of the video, or the index frame if the video is in IPB format. In addition to the image data processed and selected here, the learning server 107 associates the user who purchased the data and the purchase price of each still image and video data purchased by the user using the learning data generation unit 542. Furthermore, the input data X401 used for training the trained model 403 and the training data are stored in the HDD 335 in accordance with the data storage unit 543.
[0076] The input data and training data used for learning in this embodiment are exemplified in FIG. 7 and will be described later.
[0077] After preparing input data X401 and training data, the learning server 107 uses these to train the trained model 403. The input data X401 to the trained model 403 uses image data processed by the training data generation unit 542, a user ID, and scene information at the time of shooting, and the training data uses the purchase prices of videos and still images, respectively, to train the trained model 403. The trained model generates a trained model 403 that predicts the rating values of still images and videos preferred by a specific user from images acquired by an image sensor.
[0078] <Application of learning model to imaging devices> The data collection server 106 periodically queries the learning server 107 to check whether the trained model 403 has been updated. For example, the data collection server 106 calculates a hash value of the stored trained model, and the learning server 107 similarly calculates a hash value of the trained model 403 and passes the hash value of the trained model 403 upon request from the data collection server 106. The data collection server 106 compares the hash values of the trained model 403 on the data collection server 106 and the learning server 107 to check whether they match.
[0079] If the hash values of the respective trained models 403 do not match, the data collection server 106 determines that the trained model 403 has been updated. Then, the data collection server 106 requests the trained model from the learning server 107, and the learning server 107 transmits the trained model to the data collection server 106.
[0080] When the imaging device 102 starts operation, it inquires of the data collection server 106 whether the learning model has been updated, and if an update has been found, it requests the learning model from the data collection server 106 .
[0081] The data collection server 106 transmits the learning model based on a request from the image capture device 102, and the image capture device 102 updates the learning model in its own device with the received learning model.
[0082] <Automatic photography and post-photography actions based on learning models> After completing the learning model update, the imaging device 102 receives an automatic shooting instruction and an instruction for shooting scene information and starts automatic shooting operation. When not shooting during automatic shooting, the imaging device 102 periodically drives the imaging unit 306 to acquire images and uses them to determine whether to shoot automatic images. The imaging device 102 inputs the acquired image data and scene information as input data into the learning model, and acquires predicted evaluation values for still images and videos for a specific user as output from the learning model. The first control unit 323 of the imaging device 102 compares the predicted evaluation values for the acquired still images and videos, the history since receiving the shooting instruction, and the state of the imaging device 102 itself, such as the remaining battery capacity of the first power supply unit 310, against a lookup table to determine whether to shoot still images or videos.
[0083] After receiving the instruction to end shooting, if the image capturing device 102 is in a state where it can communicate with the data collection server 106 , it reads the recording medium 321 and transmits the captured still image and video data to the data collection server 106 .
[0084] FIG. 7 is a table showing an example of learning data indicating input data and teacher data used for learning in the learning server 107 in this embodiment.
[0085] The learning data is generated from information acquired by the learning server 107 from the data collection server 106. The learning data is generated on the learning server 107 based on still image and video content purchased by the user, and the learning data ID, user ID, image identifier, subject information, scene information, shooting time, and the images associated with each are used as input data. In addition, from the learning data, whether the image data is a video or a still image and the purchase price are held as expected values for each learning data ID, and are stored as still image expected values and video expected values.
[0086] The table shown in Figure 7 is stored as table data for each specific community. In this case, using the learning method described below, it is possible to train the learning model to predict the expected values of still images and videos so as to satisfy the preferences of each user in the specific community.
[0087] In the example of Figure 7, only still images and videos with a purchase history are shown as examples of learning data, but non-purchased still images and videos may also be used for further training of the learning model as images with low expected values.
[0088] After receiving the instruction to end shooting, if the image capturing device 102 is in a state where it can communicate with the data collection server 106 , it reads the recording medium 321 and transmits the captured still image and video data to the data collection server 106 .
[0089] FIG. 8 is a flowchart illustrating the operation of learning the trained model 403 executed by the learning unit 541 of the learning server 107 in this embodiment.
[0090] FIG. 8(a) is a flowchart illustrating the operation of the content providing server 105 related to the learning operation of this embodiment.
[0091] In step (hereinafter abbreviated as S) 801, the content providing server 105 checks whether the user's login information entered into the data presentation unit 521 via the web browser 511 exists. Here, the login information is, for example, a user ID required for the user to purchase still images and videos via the web browser 511. If the login information does not exist, S801 is repeated until the user enters the login information (NO in S801). If the login information exists, the process proceeds to S802 (YES in S801).
[0092] In S802, the content providing server 105 requests the data collection server 106 for an image / video list corresponding to the user ID, and in S803, it checks whether the requested image / video list has been received. If the image / video list has been received, the process proceeds to S804 (YES in S803).
[0093] In S804, the content providing server 105 displays the still images and videos included in the image / video list on the user's web browser 511 via the data presentation unit 521. Note that in a flow not shown, the viewing time during which the user views the still images and videos is measured between S804 and S808.
[0094] In S805, the content providing server 105 determines whether or not it has received an instruction to select a displayed still image or video from the client terminal 104. If it has received an instruction to select, it transitions to S806 (YES in S805) and causes the web browser 511 to display a purchase screen. In S807, the content providing server 105 determines whether or not a purchase request for the displayed still image or video has been received from the client terminal 104. If a purchase request has been received, the process proceeds to S808 (YES in S807) and payment processing is performed. In S808, the content providing server 105 performs payment processing using a payment processing unit (not shown).
[0095] In S809, the content providing server 105 transmits purchase information to the data collecting server 106. The purchase information includes information such as the user ID of the user who purchased each still image or video, the viewing time, the favorite rating, and the purchase amount. The purchase information is selected from the learning data described below and is used as input data or training data for the trained model.
[0096] FIG. 8(b) is a flowchart illustrating the operation of the data collection server 106 related to the learning operation of this embodiment.
[0097] In S821, the data collection server 106 acquires images and videos from the imaging device 102 and stores them in the data storage unit 533. In S822, the acquired images and videos are stored in the data storage unit 533. The acquired images and videos are assigned shooting information by the imaging device 102 at the time of capture. The shooting information includes, as subject information, general object recognition results for the current angle of view, face detection results, the number of faces captured in the current angle of view, the degree of smile and eye closure, face angle, face recognition ID number, gaze angle of the subject, and scene determination results. The shooting information also includes, as setting information, zoom magnification, pan / tilt motion method (movement pattern and driving speed), shutter speed, white balance, audio level, LED lighting method (color, flashing duration, etc.). Furthermore, as environmental information, the current time, height information from GPS location information and a barometer, vibration information from acceleration information, and tilt information from a gyro sensor are also included. This information attached to images and videos is selected as training data, which will be described later, and used as input data or training data for the trained model.
[0098] In S823, the data collection server 106 determines whether or not there is a request for an image / video list from the content providing server 105. If there is a request (YES in S823), the process proceeds to S824, where the data collection / providing unit 532 reads images / videos corresponding to the user ID from the data storage unit 533, creates an image / video list, and transmits it to the content providing server 105.
[0099] In S825, the data collection server 106 determines whether or not purchase information has been received from the content providing server 105. If the purchase information has been received (YES in S825), the process proceeds to S826.
[0100] In S826, the data collection server 106 combines the purchase information acquired by the data collection / provision unit 532 with the image and video shooting information to generate an image attribute list. This image attribute list is a list that associates shooting information such as the learning data ID, image identifier, subject information, scene information, shooting time, location information, altitude, and success / failure, with purchase information such as the user ID, image type, photo sales amount, and video sales amount, as shown in Figure 7. The generated image attribute list is sent to the learning server 107.
[0101] In S827, the data collection server 106 determines whether a request to send the trained model has been sent from the imaging device 102. If a request has been sent (YES in S827), the process proceeds to S828. Note that in this embodiment, the request to send the trained model is sent from the imaging device 102, but it may also be sent from the communication terminal 103.
[0102] In S828, the data collection server 106 requests the learning server 107 to send the latest trained model, and in S829, repeats confirmation of reception until the trained model is received. Once reception of the trained model is complete, the process proceeds to S830 (YES in S829). Note that the data collection server 106 requests the trained model in accordance with a request from the imaging device 102, but the data collection server 106 may periodically request transmission from the learning server 107 based on updates to the purchase information from the content providing server 105.
[0103] In S830, the data collection server 106 transmits the latest trained model to the imaging device 102.
[0104] FIG. 8(c) is a flowchart illustrating the operation of the learning server 107 related to the learning operation of this embodiment.
[0105] In S841, the learning server 107 determines whether or not it has received an image attribute list from the data collection server 106. If it has not received an image attribute list, the learning server 107 waits until it receives an image attribute list from the data collection server 106 (NO in S841). If it has received an image attribute list, it transitions to S842.
[0106] In S842, the learning data generation unit 542 of the learning server 107 selects input data X401 and training data to be input to the trained model from the image attribute list, and transmits them to the learning unit 541. In this embodiment, for example, the subject, event name, and image evaluation value are selected as the input data X401, and the sales amount of content such as photos and videos is selected as the training data.
[0107] In S843, S844, and S845, the learning unit 541 uses a machine learning algorithm to generate the trained model 403. Specific examples of machine learning algorithms include nearest neighbor algorithms, naive Bayes algorithms, decision trees, and support vector machines. Deep learning, which uses a neural network to generate features and connection weighting coefficients for learning, is also an example. Any available algorithm among the above may be used as appropriate and applied to this embodiment. For example, in this embodiment, the learning unit 541 generates the trained model 403 using a neural network. The learning unit 541 may include an error detection unit and an update unit. The error detection unit obtains an error between the training data and output data output from the neural network in response to input data input to the input layer. The error detection unit may use a loss function to calculate the error between the training data and output data from the neural network.
[0108] The update unit updates the connection weighting coefficients between the nodes of the neural network based on the error obtained by the error detection unit so as to reduce the error. This update unit updates the connection weighting coefficients, for example, using backpropagation. Backpropagation is a technique for adjusting the connection weighting coefficients between the nodes of each neural network so as to reduce the error. In this embodiment, in S843, subject information and shooting condition information are used as input data X401, and the weights of the connection weighting coefficients are adjusted based on the sales of the purchase information as training information. In S844, a trained model 403 is trained, which outputs the expected value of sales for each subject as output data Y402. Then, after it is confirmed that all the data selected in S845 has been input, the process proceeds to S846 (YES in S845).
[0109] In S846, the learning server 107 determines whether or not there is a request for the trained model 403 from the data collection server 106. If there is a request, the process proceeds to S847 (YES in S846).
[0110] In S847, the learning server 107 transmits the trained model 403 to the estimation model providing unit 534 of the data collection server 106 via the estimation model output unit 544.
[0111] FIG. 9 is a flowchart illustrating the prediction operation executed by the estimation unit 503 of the image capturing device 102 in this embodiment.
[0112] In S901, the estimation unit 503 of the imaging device 102 determines whether input information has been acquired from the first control unit 323. If input information has been acquired, the process proceeds to S902. The first control unit 323 acquires information from each function as shown in FIG. 3 to operate the imaging device 102. In this embodiment, the shooting condition information is set using the communication terminal 103 via the local network 101. Alternatively, the input unit 337 may be used to directly select from a list of shooting conditions (not shown) prepared in advance. Furthermore, it may be confirmed whether subject information acquired by the face authentication unit 325 and shooting condition information acquired by a shooting condition identification unit (not shown) have been acquired as input information. Here, the shooting condition identification unit (not shown) identifies the shooting conditions from the video obtained from the image processing unit 307, for example, using a trained model (not shown) that identifies the shooting conditions.
[0113] In S902, the estimation unit 503 selects input data X401 from the input information acquired in S901. In this embodiment, subject information and shooting condition information are selected as the input data X401.
[0114] In S903, the estimation unit 503 inputs the input data X401 to the trained model 403.
[0115] In S904, the estimation unit 503 outputs the output data Y402 from the trained model 403. In this embodiment, an expected value for each subject is obtained as the output data Y402.
[0116] In S905, the estimation unit 503 transmits the output data Y402 acquired in S904 to the first control unit 323. Based on the output data Y402, the first control unit 323 controls the imaging device 102 so as to perform an operation sequence described below.
[0117] FIG. 10 is a flowchart illustrating the operations executed by the image capture device 102 in this embodiment.
[0118] The operation of the image capture device 102 in this embodiment will be described using the flowchart in Fig. 10. Before the start of the flowchart shown in Fig. 10, the image capture device 102 is in a power-off state, and begins operation upon start.
[0119] In S1001, the imaging device 102 communicates with the data collection server 106 to check whether the currently applied learning model needs to be updated.
[0120] In S1002, the imaging device 102 compares the learning model applied to the imaging device 102 with the learning model on the data collection server 106. If it is determined in S1002 that the learning model needs to be updated, in S1003 the learning model is updated via the communication unit 322 (YES in S1002). For example, when a DNN (Deep Neural Network) model is used as the learning model, if there is a change in the number of layers or the number of units, the entire DNN network is updated, and if the number of layers or the number of units remains the same, parameters such as coefficients and connection weights are updated. If there is no update to the learning model in S1002 or if the learning model update process in S1003 has finished, the process transitions to S1004, where shooting condition information is acquired.
[0121] In S1005, as described above with reference to FIG. 9, the shooting condition information acquired in S1004 is input into the trained model, and an expected value according to the camera settings is predicted.
[0122] In S1006, the initial value of TH_SHOT is set to the shooting threshold value obtained from the predicted expected value in S1005. TH_SHOT will be described later in S1011.
[0123] In S1007, it is confirmed whether or not an automatic shooting instruction has been issued by the operator of the imaging device 102, and if an automatic shooting instruction has been issued, the process proceeds to S1008, where the imaging device 102 starts automatic shooting (YES in S1007). If an automatic shooting instruction has not been issued, the imaging device 102 waits until an instruction is issued (NO in S1007).
[0124] When automatic shooting starts, in S1009, the image processing unit 307 performs image processing on the signal captured by the imaging unit 306 to generate an image for subject detection. The generated image is subjected to subject detection processing to detect people, objects, etc. After subject detection processing is performed, the process transitions to S1010.
[0125] In S1010, an image evaluation value is calculated from the subject information acquired in S1009 according to the subject's situation and shooting conditions. The image evaluation value is calculated based on, for example, the number of people, the size and orientation of the people's faces, the people's facial expressions, the results of personal authentication of the people, etc.
[0126] FIG. 11 is a diagram showing the relationship between the image evaluation value and the threshold value in this embodiment, with the horizontal axis representing the elapsed time and the vertical axis representing the image evaluation value.
[0127] From S1011, automatic photography determination processing is performed.
[0128] In this embodiment, the photographing frequency is changed by updating the photographing threshold and the photographing interval. The method for updating the photographing threshold is as follows, and updating the photographing interval will be described later in S1014.
[0129] Automatic photography calculates an evaluation value from the state of the subject, compares the evaluation value with a threshold, and performs automatic photography if the evaluation value exceeds the threshold. The threshold at this time is TH_SHOT in FIG. 11; the higher TH_SHOT, the less likely automatic photography is performed, and the lower TH_SHOT, the more likely it is performed. Note that TH_SHOT is an indefinite value, and is updated when it exceeds a threshold representing the image evaluation value. The threshold at this time is TH_TR, and TH_TR is a value greater than TH_SHOT. If the image evaluation value calculated in S1010 exceeds TH_TR, the process transitions to S1012 (YES in S1011).
[0130] In S1012, as described above in FIG. 9, the image data acquired in S1009 and the shooting condition information acquired in S1004 are input into the trained model, and an expected value according to the camera settings is predicted.
[0131] In S1013, TH_SHOT is updated to the shooting threshold value obtained from the predicted expected value in S1012.
[0132] Similarly, in S1014, the shooting interval is updated to the shooting interval obtained from the predicted expected value in S1012, and the process proceeds to S1016 to execute shooting. The shooting interval at this time is Δt in FIG. 11, and the longer Δt is, the more single-shot automatic shooting is performed, and the shorter Δt is, the more continuous automatic shooting is performed. The updated shooting threshold and shooting interval are applied from the next automatic shooting determination process. If the image evaluation value is equal to or less than TH_TR in S1011, the process proceeds to S1015.
[0133] In S1015, it is determined whether the image evaluation value exceeds TH_SHOT, and if it exceeds TH_SHOT, the process transitions to S1016 and photography is performed (YES in S1015). If it is equal to or less than TH_SHOT, the process of S1016 is skipped and control is exercised so that photography is not performed (NO in S1015). If photography is performed in S1016, or if the image evaluation value is equal to or less than TH_SHOT in S1015, the process transitions to S1017 and waits for the photography interval time updated in S1014.
[0134] Thereafter, if an end instruction is received in S1018, it is determined that the image capture is to be ended, a predetermined power-off process is performed for the image capture device 102, and the operation of the image capture device 102 is ended (YES in S1018). Unless an end instruction is received in S1018, S1009 to S1018 are repeated.
[0135] As described above, by controlling the number of shots and the shooting interval according to the purchase history, it is possible to provide a photography system that can predict the optimal number of shots and shooting interval for each subject and provide the optimal number of shots. [Second embodiment]
[0136] In the second embodiment, these are the same as 100 to 107 shown in FIG. 1, so a description thereof will be omitted.
[0137] 15 shows an example of reinstalling an imaging device via a dedicated application on the communication terminal 103 according to the second embodiment of the present invention. The communication terminal 103 comprises a display unit 1501 and operation units 1502 to 1504. Thumbnail images (1505 to 1510) of image data stored in the imaging device are displayed on the display unit 1501.
[0138] The user can select a displayed thumbnail image to check the image data in detail, or switch the display on an application in response to communication from the imaging device 102. For operation, the communication terminal 103 has operation units 1502, 1503, and 1504 on the surface of the display unit 1501. The operation units 1502, 1503, and 1504 are configured, for example, as a touch panel, and can be implemented by the user designating a specific area on the surface of the display unit 1501 using a finger or a pen-like operation member. Note that while the illustrated example in FIG. 15 shows the above-mentioned three operation units as an example, when implemented on a touch panel, various operations can be performed by moving, deforming, or increasing or decreasing the area of the operation unit within the area of the display unit 1501 to suit the operation.
[0139] The operation of the image capture device 102 in this embodiment will be described using the flowchart in Fig. 12. Steps S1201 to S1204 are the same as steps S1001 to S1004 in the first embodiment shown in Fig. 10.
[0140] In S1205, it is confirmed whether or not an automatic shooting instruction has been issued by the user of the imaging device 102, and if an automatic shooting instruction has been issued, the process transitions to S1206, where the imaging device 102 starts automatic shooting (YES in S1205). In S1207, it is determined whether or not a time equal to or greater than the set shooting interval, which is one of the shooting conditions set in S1204, has elapsed, and if the set time has elapsed, the process transitions to S1209 (YES in S1207). If the set time has not elapsed but an end instruction has been issued, control ends (S1208).
[0141] In S1209, the imaging unit 306 is driven to acquire image data, and the process proceeds to S1210. If it is detected in S1210 that a subject is present in the live view image data, the process proceeds to S1211, where the face authentication unit 325 performs personal authentication on the main subject of the image data to identify the individual (YES in S1210). For example, if three subjects are detected in the image data, personal authentication is performed on each person to identify the individual in the image data.
[0142] In S1212, the information on individuals identified from the image data acquired in S1209, the number of people, and the combination of people, as well as the shooting conditions set in S1204, are input into the trained model. After that, whether the combination of subjects is suitable is predicted by calculating an expected value that indicates the relationship between the subjects.
[0143] In S1213, the number of times of shooting is updated based on the predicted expected value obtained in S1212 and the past shooting history information, and the process proceeds to S1214 to perform shooting, and then to S1215 to update the shooting history.
[0144] In S1216, it is determined whether or not the number of times of shooting set in S1213 has been reached based on the shooting history updated in S1215.
[0145] Steps S1214 to S1216 are repeated until the number of shots set in step S1213 is reached, and shooting is repeatedly performed. When the set number of shots is reached (YES in step S1216), the process transitions to step S1207, and thereafter, shooting is similarly repeated at predetermined intervals until an instruction to end the process is received.
[0146] As described above, in the second embodiment, it is possible to increase the frequency of capturing images with a suitable combination in automatic capture. [Third embodiment]
[0147] The operation of the image capture device 102 according to the third embodiment of the present invention will be described with reference to the flowchart of FIG.
[0148] Before the flowchart shown in FIG. 13 starts, the imaging device 102 is in a power-off state, and starts operating upon start.
[0149] Operations S1301 to S1309 are similar to operations S1201 to S1209 of the second embodiment shown in FIG.
[0150] In S1310, the face authentication unit 325 determines whether or not a subject is included in the image based on the image data obtained from the image processing unit 307. If a subject is included in the image, the process proceeds to S1311 (YES in S1310). The face authentication unit 325 also identifies the individual of the subject and transmits the subject information to the first control unit 323.
[0151] In S1311 , the scene determination unit 328 determines the scene of the image from the image data obtained by the image processing unit 307 , and transmits the scene information to the first control unit 323 .
[0152] In S1312, the first control unit 323 transmits the subject information and the scene information to the learning processing unit 319. The learning processing unit 319 inputs the received subject information and scene information as input data X401 to the trained model 403. Here, the trained model 403 of the third embodiment is a trained model that has trained using the sales of subjects in specific scenes as training data, and outputs sales forecasts for each subject as expected sales values according to the input subject and scene information. Furthermore, if the input subject cannot be identified as an individual, the expected sales value is set lower than that of a subject with the highest expected sales value and higher than that of a subject with the lowest expected sales value. The learning processing unit 319 transmits the expected sales value for each subject obtained from the trained model 403 to the first control unit 323.
[0153] In S1313, the first control unit 323 checks the shooting history since the automatic shooting started in S1306.
[0154] In S1314, the first control unit 323 checks whether there are any unphotographed subjects among the subjects checked in S1310 and S1311 in the photography history checked in S1313. If there are any unphotographed subjects, the process proceeds to S1315 (YES in S1314). If there are no unphotographed subjects, the process proceeds to S1320.
[0155] In S1315, the first control unit 323 sets the order of photographing the subjects based on the expected sales value obtained in S1312. Specifically, the order is set so that the subjects are photographed in order starting from the subject with the highest expected sales value.
[0156] In S1316, the first control unit 323 checks whether a full-body photo showing all of the subjects checked in S1310 has been taken, based on the shooting history checked in S1313. If a full-body photo has not been taken, the process proceeds to S1317 (YES in S1316). If a full-body photo has been taken, the process proceeds to S1319.
[0157] In S1317, the first control unit 323 controls the lens barrel rotation drive unit 305 to perform photography so that the entire photograph can be taken.
[0158] In S1318, the first control unit 323 checks whether or not the entire photograph has been taken. If the entire photograph has not been taken, the process proceeds to S1319 (NO in S1318). If the entire photograph has been taken, the process proceeds to S1321 (YES in S1318).
[0159] In S1319, the first control unit 323 adds 1 to the whole imaging incompletion count N.
[0160] In S1320, the first control unit 323 checks the count number of the whole image capturing incomplete count N. If the whole image capturing incomplete count N is greater than 3, the process proceeds to S1321 (YES in S1320). If the whole image capturing incomplete count N is 3 or less, the process proceeds to S1317 (NO in S1320). Note that in this embodiment, the number determined in S1320 is 3, but it may be any number greater than 0, for example.
[0161] In S1321, the first control unit 323 sets the whole image capturing incomplete count N to 0, and updates the history of the subject captured in S1322.
[0162] In S1323, the first control unit 323 controls the lens barrel rotation drive unit 305 to adjust the angle of view to the unphotographed subject confirmed in S1314, according to the photographing order set in S1315, and photographs the individual subject.
[0163] In S1324, the first control unit 323 checks whether all of the subjects acquired in S1310 have been photographed. If it is confirmed that all of the subjects have been photographed, the process proceeds to S1307 (YES in S1324). If there are subjects that have not yet been photographed, the process proceeds to S1314 (NO in S1324).
[0164] This control process provides a photo that shows the shooting scene in a full-frame shot, and by changing the shooting order for each subject, it is possible to provide a subject photo that the user likes while preventing the user from missing a subject.
[0165] In the third embodiment of the present invention, the order in which subjects are photographed is set based on the expected sales value in S1315. However, it is also possible to set a tracking time for each subject. In this case, for example, the tracking time is set longer for subjects with a high expected sales value. A subject for which a long tracking time is set is photographed more often, resulting in a larger number of photographs being taken. The number of photographs to be taken for each subject may also be set based on the expected sales value. In this case, for example, a larger number of photographs is set for subjects with a high expected sales value. When a predetermined number of photographs is to be taken, the likelihood that a photograph will be selected after shooting can be increased by photographing subjects with a high expected sales value more than subjects with a low expected sales value. [Fourth embodiment]
[0166] The operation of the image capture device 102 in this embodiment will be described using the flowchart in Fig. 14. Before the start of the flowchart shown in Fig. 14, the image capture device 102 is in a power-off state, and begins operation upon start.
[0167] Steps S1401 to S1409 are the same as steps S1201 to S1209 of the second embodiment shown in FIG.
[0168] In S1410, the image data acquired in S1409 and the shooting conditions set in S1404 are input into the trained model, and the expected values for each of the still images and video images are predicted.
[0169] In S1411, the shooting history since the start of automatic shooting in S1406 is checked, and in S1412, it is determined whether or not to execute video recording based on the predicted expected value in S1410 and the history in S1411.
[0170] In S1412, for example, if there is no video shooting history and only still image shooting history and the predicted expected value of the video exceeds the predicted expected value of the still image, it is determined that video shooting will be performed, and the process transitions to S1413, where video shooting is performed (YES in S1412). Details of the determination method in S1412 will be described later. In S1413, when performing autonomous shooting using multiple image capture devices 102, if the multiple image capture devices 102 are in a communication state, they notify each other of the shooting state using the communication units 322 of the image capture devices 102. It is sufficient if multiple image capture devices 102 can shoot videos simultaneously so that an opportunity to capture a still image is not missed.
[0171] After the predetermined video is shot in S1413, the video shooting history is updated in S1414, and the process transitions to S1407 to wait for the next predetermined interval to elapse.
[0172] On the other hand, if it is determined in S1412 that video shooting will not be performed, for example, if there is no history of still image shooting but there is a history of video shooting, the process transitions to S1415 to determine whether to perform still image shooting (NO in S1412).
[0173] In S1415, it is determined whether or not to perform still image shooting. For example, if the predicted expected value of the still image exceeds a preset threshold, it is determined to perform still image shooting. If it is determined to perform still image shooting, a predetermined still image is shot in S1416, the process transitions to S1417 to update the still image shooting history, and the process transitions to S1407 to wait for the next predetermined interval to elapse (YES in S1415). The determination method in S1415 will be described in detail later.
[0174] On the other hand, if it is determined in S1415 that still images will not be taken, for example, if the predicted expected values for both moving images and still images are low and it is determined that shooting is not necessary, the process returns to S1407 without taking any still images or moving images, and waits for the next shooting interval to elapse (NO in S1415).
[0175] When the imaging device 102 determines whether to shoot a video, shoot a still image, or not shoot at all based on the expected value predicted from the learning model in S1410 and the shooting history of automatic shooting in S1411, the determination may be made based on a predetermined rule that is set in advance. For example, history information may be acquired from the data collection server 106 and the video / still image ratio in the past purchase history may be referenced, or if an individual can be identified by personal authentication of the main subject of the image data in S1409, a determination criterion for each individual may be set by comparing the individual with the past history.
[0176] While waiting for the next shooting interval to elapse (NO in S1407), NO in S1408 and NO in S1407 are repeated, and when the predetermined interval elapses, the next shooting begins. On the other hand, if an end instruction is received in S1408, it is determined that shooting is to end, a predetermined power-off process for the image capture device 102 is performed, and the operation of the image capture device 102 is ended (YES in S1408).
[0177] The determination of whether to perform video recording in S1412 described above is made based on the expected values of still images and videos predicted in S1410 and the shooting history confirmed in S1411. Specifically, when setting the shooting conditions in S1404, a determination value can be instructed to the image capture device 102, and the determination can be made by setting conditions such as the predicted expected value of video shooting being equal to or greater than a predetermined value and the number of video shots being less than a predetermined number. Also, if personal information can be detected in the image data using the face authentication unit 325, it is possible to calculate the predicted expected value for each user and use the total value obtained by adding them up for the determination.
[0178] In the fourth embodiment of the present invention, the decision on whether to shoot a video is made in S1412 based on the predicted expected value of the video. However, the decision on whether to shoot a still image may be made first. In this case, since the still image has a higher priority in the operation sequence of the still image, adjustments can be made by weighting the predicted expected value of the video. Furthermore, when updating the trained model in S1403, past shooting history information may be acquired from the data collection server 106 and set as the shooting conditions in S1404, thereby determining the shooting conditions. Based on the past shooting history, it is only necessary to set the amount of content, such as the shooting time for one video, to be suitable for the user. As described above, in autonomous shooting, it is possible to shoot videos and still images according to the user's preferences, thereby making it possible to provide content that the user desires.
[0179] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.
[0180] The present invention can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having the computer of the system or device read and execute the program. The computer has one or more processors or circuits, and may include multiple separate computers or a network of multiple separate processors or circuits to read and execute computer-executable instructions.
[0181] The processor or circuitry may include a central processing unit (CPU), a microprocessing unit (MPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a data flow processor (DFP), or a neural processing unit (NPU). [Explanation of symbols]
[0182] 100 Internet 101 Local Network 102 Imaging device 103 Communication terminal 104 client terminals 105 Video Server 105 Content Provider Server 106 Data Collection Server 107 Learning Server
Claims
1. an imaging device that performs pan-tilt control to perform automatic photography; a mobile terminal device that controls the imaging device; an automatic photography system comprising a storage means for storing electronic photographs and moving images photographed by the imaging device; a providing system that provides the content recorded by the storage means, The imaging device is A means for acquiring information about the detected subject; a calculation means for calculating an evaluation value used to determine whether or not to perform a photographing operation in the automatic photographing system; a setting means for setting a first threshold value used to determine whether or not to perform a photographing operation in the automatic photographing system; a determination means for determining whether or not a photographing operation is to be performed using the evaluation value and the first threshold value; a second threshold value that triggers updating of the first threshold value; means for updating the first threshold when the evaluation value exceeds the second threshold; a processing means for referencing the purchase history of image data recorded by the storage means and acquiring the number of images purchased by the user, information on the subjects of the stored photographs, and information on the photographing conditions; A photographing device characterized by controlling the first threshold value and photographing interval appropriate for each subject using a predictive model that has been machine-learned using training data such as the number of purchases by the user, the subject information of the stored photographs, and the sales performance of the image data relative to the photographing condition information, all of which are recorded by the storage means.
2. 2. The imaging apparatus according to claim 1, further comprising a detection unit that detects information about a subject, wherein the calculation unit determines the evaluation value using the information about the subject detected by the detection unit.
3. 3. The imaging device according to claim 1, wherein the information about the subject is associated with user information recorded by the storage means.
4. 4. The imaging device according to claim 1, wherein the imaging condition information input to the prediction model is input for each imaging condition by an operator of the imaging device.
5. 5. The imaging device according to claim 1, wherein the initial value of the first threshold is set based on the prediction model using shooting condition information input for each shooting condition by an operator of the imaging device.
6. an imaging device that performs pan-tilt control to perform automatic photography; a mobile terminal device that controls the imaging device; an automatic photography system comprising a storage means for storing electronic photographs and moving images photographed by the imaging device; a providing system that provides content recorded in the storage means, The imaging device is A means for acquiring information about the detected subject; a means for determining, from the subject information, a subject to be included within the photographed area; means for calculating a pan / tilt control amount for the imaging device and controlling the pan / tilt so that the target subject is included; The system is characterized in that the storage means comprises a means for storing photographs taken by the photographing device, a processing means for acquiring subject, scene information, shooting time, and sales figures from the purchase history of the photographs, and an imaging device that uses a machine-learned prediction model using the subject information of the photograph and the actual sales figures of the photograph for the shooting information as training data to predict the optimal tracking order of the subject to be photographed by the photographing device and then takes the photograph.
7. 7. The imaging device according to claim 6, wherein the imaging device sets a subject tracking time, which is one of various settings of the imaging device, in accordance with the result of the prediction model.
8. 7. The imaging device according to claim 6, wherein the imaging device sets the number of images of the subject, which is one of various settings of the imaging device, in accordance with the result of the prediction model.
9. The imaging device according to claim 6, characterized in that, if a subject is not included in the results of the prediction model, the imaging device sets the order of the subject not included in the results of the prediction model above the bottom of the predicted preferred tracking order.
10. The imaging device according to claim 6, characterized in that, if a subject is not included in the results of the prediction model, the imaging device sets the tracking time of the subject not included in the results of the prediction model above the bottom of the predicted preferred tracking order.
11. The imaging device according to claim 6, characterized in that, if a subject is not included in the results of the prediction model, the imaging device sets the number of shots of the subject not included in the results of the prediction model above the lowest predicted preferred tracking order.
12. The imaging device according to claim 6, wherein the imaging device is controlled to take an image by pan-tilting to obtain a composition that includes a large number of subjects within the captured image plane, and then to photograph the subjects individually in order according to a prediction model.
13. an imaging device that performs pan-tilt control to perform automatic photography; a mobile terminal device that controls the imaging device; an automatic photography system comprising a storage means for storing electronic photographs and moving images photographed by the imaging device; a providing system that provides content recorded in the storage means, The imaging device comprises a means for acquiring detected subject information, a means for determining a subject to be included in an image plane from the subject information, and a means for calculating a pan / tilt control amount of the imaging device and controlling the pan / tilt so that the subject is included, the storage means includes a means for storing the photographs or videos taken by the imaging device, The provision system uses at least a portion of the video and still images as input image data, and purchase information and browsing history of the provision system as training data to generate a machine-learned prediction model that predicts expected values at the time of shooting each of the video and still images from the input image data; From the image data obtained at the time of shooting, we predict the expected values for both still images and videos, The imaging device determines whether to capture a still image or a video based on at least the output of the prediction model, and then performs imaging.
14. The imaging device according to claim 13, characterized in that the imaging device is configured to be able to acquire at least a portion of the provision history of the content held by the provision system, and sets the shooting time of the video based at least on the provision history.
15. The imaging device according to claim 13, characterized in that the prediction model is configured to allow subject information to be input, and based on the prediction model to which the subject information is input in addition to the image data, the imaging device predicts expected values for still images and videos for each subject and controls the imaging device to take at least still images or videos.
16. an imaging means for imaging a subject and outputting image data; A photograph purchasing system comprising a storage means for storing image data output by the photographing means, a calculation means for calculating an evaluation value used to determine whether or not to perform a photographing operation by the imaging means; a first setting means for setting a first threshold value used to determine whether or not to perform a photographing operation by the imaging means; a determination means for determining whether or not a photographing operation is to be performed using the evaluation value and the first threshold value; a second threshold value used to determine whether to control the first threshold value; a processing means for referring to the purchase history of image data stored in the purchasing system and performing processing based on the number of images purchased by the user, the subject information of the stored photographs, and the scene information; This photographing device is characterized by using a machine-learned prediction model that uses the number of photos purchased by the user, the subject information of the stored photos, and the sales performance of the image data for each combination of subjects stored in the purchasing system as training data to predict and photograph the optimal combination for each subject.
17. 17. The imaging apparatus according to claim 16, further comprising a detection unit that detects information about a subject, wherein the calculation unit determines the evaluation value using the information about the subject detected by the detection unit.
18. 17. The imaging device according to claim 16, wherein the information about the subject is associated with user information held by the purchasing system.
Citation Information
Patent Citations
Methods for storing, searching, editing and outputting photographic images
JP2004514976A
Imaging apparatus, control method of the same, program, and storage medium
JP2021057815A