Image pickup apparatus and control method thereof
The imaging device uses machine-learning prediction models to adjust settings based on purchase history and location for optimal photography, addressing the issue of missed shots in automatic cameras by ensuring desired environments and times are captured.
Patent Information
- Application Number
- JP2024111706
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2026-01-23
AI Technical Summary
Automatic cameras capture images evenly across multiple subjects, often missing important shots in desired environments, and user preferences are not adequately reflected by a single image evaluation value.
An imaging device with machine-learning prediction models that analyze purchase history and indoor location information to adjust shooting frequency and settings, ensuring optimal photography environments and times.
Provides preferred photography environments and times based on user preferences, enhancing the likelihood of capturing desired images.
Smart Images

Figure 2026011251000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, and more particularly to an automatic photography technique. [Background technology]
[0002] When taking still or video images using an imaging device such as a camera, it is common for the photographer to determine the subject of the image through a viewfinder, check the shooting conditions themselves, adjust the framing of the image, and then take the image. Such imaging devices have traditionally been equipped with mechanisms for detecting user operation errors and the external environment, and notifying the user if the environment is not suitable for shooting, or for detecting the external environment and controlling the camera to bring the environment into a suitable state for shooting.
[0003] In contrast to such imaging devices that perform shooting by user operation, there is known an automatic camera (Patent Document 1) that takes pictures periodically and continuously without the user issuing a shooting instruction. The automatic camera sets a shooting threshold based on the shooting history and the evaluation value of the captured image, controls the shooting frequency, and is controlled to capture the subject evenly while minimizing the missed shots during automatic shooting.
[0004] Furthermore, various technologies have been proposed that allow purchasers to purchase, via the Internet, images taken by third parties that have been electronically published on websites on the Internet, etc. (for example, Patent Document 2). By using these technologies, users can view multiple images at home, select their favorite images, and obtain only the images they need.
[0005] In recent years, services have been provided that sell images taken at kindergartens and nursery schools to potential buyers such as parents. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Patent Publication No. 2021-57815 [Patent Document 2] Special Publication No. 2004-514976 Summary of the Invention [Problem to be solved by the invention]
[0007] However, automatic cameras are controlled to capture images evenly depending on the subject, so they capture images of multiple subjects evenly. As a result, photographs are taken in locations that are unnecessary for those who wish to purchase the photographs, and it is easy to miss capturing a subject in a suitable photographic environment.
[0008] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to provide an imaging device that can adjust the shooting frequency required for each subject while reducing the risk of missing a shot in a suitable shooting environment.
[0009] The method described in Patent Document 1 uniquely calculates an image evaluation value from the orientation and size of the face and the degree of smile, and determines the number of images to be taken and the interval between shots.
[0010] However, since the desired shooting environment differs from user to user, a uniquely determined image evaluation value alone cannot reflect the user's preferences.
[0011] The present invention has been made in view of the above-mentioned problems, and its object is to predict a suitable photographing environment common to users and provide the user with a photograph that meets their desire. [Means for solving the problem]
[0012] In order to achieve the above-mentioned object, the photography system of the present invention is a photograph purchase system comprising an imaging device that performs automatic photography, a smartphone that also functions as the imaging device, and a storage means for saving photographs taken by the imaging device, wherein the imaging device comprises a means for acquiring information on a subject detected within the screen, a means for determining a target subject to be included within the screen when photographing from the subject information, and a means for calculating a pan-tilt control amount for the camera so that the target subject is included, and the storage means comprises a means for storing photographs taken by the imaging device, a processing means for acquiring location information of the room where the photograph was taken, the time of photographing, and the number of sales from the purchase history of the photograph, and a prediction model trained by machine learning using training data of the time of photographing the photograph and the number of sales of the photograph predicted from the indoor location information, to predict various settings of the imaging device and estimate indoor location information suitable for photographing. [Effects of the Invention]
[0013] According to the present invention, preferred photography environments and times are analyzed from purchase history, and a machine-learned prediction model is used as training data, based on the time the photo was taken and the number of sales of the photo predicted from indoor location information, to predict an ideal installation location and encourage the user to install the device, thereby providing photos in an ideal photography environment. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram schematically illustrating a system configuration to which the present embodiment can be applied. [Figure 2] FIG. 1 is a diagram schematically illustrating an imaging device used in the present embodiment. [Figure 3] FIG. 2 is a block diagram showing the configuration of an imaging device and a server according to an embodiment of the present invention. [Figure 4] FIG. 2 is a conceptual diagram illustrating the structure of a learning model in an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating an example of operation when a learning model is used in an embodiment of the present invention. [Figure 6]FIG. 1 is a diagram illustrating an example of the operation of a system according to an embodiment of the present invention. [Figure 7] FIG. 2 is a table illustrating an example of learning data according to an embodiment of the present invention. [Figure 8] FIG. 10 is a flowchart illustrating an example of a learning operation according to an embodiment of the present invention. [Figure 9] FIG. 10 is a flowchart illustrating an example of an estimation operation according to an embodiment of the present invention. [Figure 10] FIG. 3 is a flowchart illustrating an example of the operation of the imaging device according to the first embodiment of the present invention. [Figure 11] 10A and 10B are diagrams illustrating examples of the relationship between image evaluation values and threshold values in the present embodiment. [Figure 12] FIG. 10 is a flowchart illustrating an example of the operation of the imaging device according to the second embodiment of the present invention. [Figure 13] 1 is a schematic diagram illustrating an operation method using a smart device according to an embodiment of the present invention. [Figure 14] 10A and 10B are conceptual diagrams illustrating an example of an operation for guiding the gaze in the second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0015] [First embodiment] <System configuration> FIG. 1 is a diagram showing a schematic diagram of a system configuration to which the present invention can be applied.
[0016] 1, a client terminal 104, a content providing server 105, a data collection server 106, and a learning server 107 are each connected to the Internet 100. An imaging device 102 and a communication terminal 103 are each connected to a local network 101. The local network 101 is connected to the Internet 100, and the system of this embodiment is configured.
[0017] The local network 101 is a network in which an imaging device 102 and a communication terminal 103 are connected by cable or wirelessly. The imaging device 102 and the communication terminal 103 can exchange information with each other via the local network 101. The local network 101 is also connected to the Internet 100.
[0018] The imaging device 102 has a drive mechanism (described later) and is a camera capable of autonomously capturing images without user operation. The imaging device 102 can change camera control using learning information (described later) acquired from the learning server 107 via the local network 101 to create still images and videos. Still images and videos captured by the imaging device 102 are stored in the data collection server 106 from the communication terminal 103 via the local network 101. Note that this example is not particularly limited, and for example, still images and videos may be stored directly from the imaging device 102 in the data collection server 106 via the local network 101. Also, although one imaging device 102 is illustrated in FIG. 1, multiple imaging devices may be connected to the local network 101.
[0019] The communication terminal 103 is a communication terminal that has a communication function and can be connected to the imaging device 102. The communication function of the communication terminal 103 is realized by a wireless LAN communication module or a wired communication module. The communication terminal 103 has an operation mechanism that accepts user operations, and is capable of transmitting instructions based on information input to the operation mechanism to the imaging device 102 via the wireless LAN communication module or the wired communication module. Furthermore, the communication terminal 103 is also capable of obtaining information from the learning server 107 via the Internet 100 and transmitting it to the imaging device 102. It also obtains still images and videos captured by the imaging device 102 and transmits them to the data collection server 106 via the Internet 100.
[0020] The client terminal 104 is a computer that allows a user to select and purchase still images and videos provided by the content providing server 105 using a web browser.
[0021] The content providing server 105 is a computer that provides still images and videos stored in the data collection server 106 in a viewable and selectable format via the Internet 100. The content providing server 105 provides still images and videos that include the user himself / herself or the user's family or acquaintances as subjects to the client terminal 104 used by the user. The operator of the content providing server 105 according to an embodiment of the present invention provides a service of providing and selling still images and videos captured by the imaging device 102 using an imaging sequence described below to users. The subjects of the still images and videos purchased by users from the operator are not necessarily the user himself / herself or family or acquaintances related to the user. However, for the sake of convenience, the following description will be given assuming that the still images and videos made available for sale by the operator to users are the user himself / herself or the user's family or acquaintances. The content providing server 105 also transmits the user's purchase information to the learning server 107.
[0022] The data collection server 106 is a computer that collects still images and moving images captured by the imaging device 102 and transmits still images or moving images requested by the content providing server 105 to the content providing server 105 .
[0023] The learning server 107 is a computer that generates a learning model using a learning method described below, using the user's purchase information transmitted from the content providing server 105. The learning server 107 transmits the learning model to the data collection server 106 via the Internet 100. The data collection server 106 can provide the learning model in response to a request from the imaging device 102 via the Internet 100.
[0024] <Configuration of imaging device> FIG. 2 is a diagram schematically showing an imaging device used in this embodiment.
[0025] The imaging device 102 shown in FIG. 2(a) is provided with an operating member for operating a power switch (hereinafter referred to as a power button, but operations such as tapping, flicking, or swiping on a touch panel may also be used). A lens barrel 202, which is a housing containing a group of photographic lenses and an image sensor for capturing images, is attached to the imaging device 102 and is provided with a rotation mechanism that can rotate the lens barrel 202 relative to a fixed part 203. The tilt rotation unit 204 is a motor-driven mechanism that can rotate the lens barrel 202 in the pitch direction shown in FIG. 2(b), and the pan rotation unit 205 is a motor-driven mechanism that can rotate the lens barrel 202 in the yaw direction. Therefore, the lens barrel 202 can rotate in one or more directions. Note that FIG. 2(b) defines the axes at the position of the fixed part 203. Both the gyroscope 206 and the accelerometer 207 are mounted on the fixed part 203 of the imaging device 102. Vibrations of the imaging device 102 are detected based on the angular velocity meter 206 and the accelerometer 207, and the tilt rotation unit and pan rotation unit are driven to rotate based on the detected shaking angle, thereby correcting the vibration and tilt of the lens barrel 202, which is the movable part.
[0026] FIG. 3 is a block diagram showing the hardware configuration of each device constituting the system of FIG. 1 according to this embodiment.
[0027] 3, the first control unit 323 is made up of a processor (e.g., CPU, microprocessor, MPU, etc.) and memory (e.g., DRAM, SRAM, etc.). These execute various processes to control each block of the imaging device 102 and control data transfer between each block. The non-volatile memory (EEPROM) 316 is an electrically erasable and recordable memory, and stores constants, programs, etc. for the operation of the first control unit 323.
[0028] 3, a zoom unit 301 includes a zoom lens that changes magnification. A zoom drive control unit 302 controls the drive of the zoom unit 301. A focus unit 303 includes a lens that adjusts focus. A focus drive control unit 304 controls the drive of the focus unit 303.
[0029] In the imaging unit 306, the imaging elements receive light incident through each lens group, and output charge information corresponding to the amount of light as analog image data to an image processing unit 307. The image processing unit 307 applies image processing such as distortion correction, white balance adjustment, and color interpolation to the digital image data output by A / D conversion, and outputs the digital image data after the processing. The digital image data output from the image processing unit 307 is converted into a recording format such as JPEG format by an image recording unit 308, and sent to a memory 315 or a video output unit 317, which will be described later.
[0030] The lens barrel rotation drive unit 305 drives the tilt rotation unit 204 and pan rotation unit 205 to drive the lens barrel 202 in the tilt direction and pan direction.
[0031] The device vibration detection unit 309 is equipped with, for example, an angular velocity meter 206 that detects the angular velocity in three axial directions of the imaging device 102, and an accelerometer 207 that detects the acceleration in three axial directions of the device. The device vibration detection unit 309 calculates the rotation angle of the device, the amount of shift of the device, etc. based on the detected signals.
[0032] The audio input unit 313 acquires audio signals from the surroundings of the imaging device 102 from a microphone provided in the imaging device 102, performs analog-to-digital conversion, and transmits the signals to the audio processing unit 314. The audio processing unit 314 performs audio-related processing such as optimization of the input digital audio signals. The audio signals processed by the audio processing unit 314 are then transmitted to the memory 315 by the first control unit 323. The memory 315 temporarily stores the image signals and audio signals obtained by the image processing unit 307 and audio processing unit 314.
[0033] The image processing unit 307 and the audio processing unit 314 read out the image signals and audio signals temporarily stored in the memory 315 and encode the image signals and audio signals to generate compressed image signals and compressed audio signals. The first control unit 323 transmits these compressed image signals and compressed audio signals to the recording and playback unit 320.
[0034] The recording / playback unit 320 records the compressed image signal, compressed audio signal, and other control data related to shooting generated by the image processing unit 307 and audio processing unit 314 on the recording medium 321. When the audio signal is not compression-encoded, the first control unit 323 transmits the audio signal generated by the audio processing unit 314 and the compressed image signal generated by the image processing unit 307 to the recording / playback unit 320, and causes them to be recorded on the recording medium 321.
[0035] The recording medium 321 may be a recording medium built into the imaging device 102 or a removable recording medium. The recording medium 321 can record various types of data such as compressed image signals, compressed audio signals, and audio signals generated by the imaging device 102, and generally uses a medium with a larger capacity than the nonvolatile memory 316. For example, the recording medium 321 includes any type of recording medium, such as a hard disk, optical disk, magneto-optical disk, CD-R, DVD-R, magnetic tape, nonvolatile semiconductor memory, flash memory, etc.
[0036] The recording / playback unit 320 reads (plays back) compressed image signals, compressed audio signals, audio signals, various data, and programs recorded on the recording medium 321. The first control unit 323 then transmits the read compressed image signals and compressed audio signals to the image processing unit 307 and audio processing unit 314. The image processing unit 307 and audio processing unit 314 temporarily store the compressed image signals and compressed audio signals in memory 315, decode them in a predetermined procedure, and transmit the decoded signals to the video output unit 317 and audio output unit 318.
[0037] The audio input unit 313 is equipped with multiple microphones in the imaging device 102, and the audio processing unit 314 can detect the direction of sound on a plane on which the multiple microphones are installed, and is used for searching and automatic shooting, which will be described later. Furthermore, the audio processing unit 314 detects specific audio commands. The audio commands may be pre-registered commands, or the user may be able to register specific audio in the imaging device. The audio processing unit 314 also performs audio scene recognition. In audio scene recognition, audio scenes are determined using a network trained by machine learning based on a large amount of audio data in advance. For example, a network for detecting specific scenes such as "cheers," "clapping," and "vocalization" is set in the audio processing unit 314. When a specific audio scene or a specific audio command is detected, a detection trigger signal is output to the first control unit 323 or the second control unit 311.
[0038] A second control unit 311 , which is provided separately from a first control unit 323 that controls the entire main system of the imaging device 102 , controls the power supply to the first control unit 323 .
[0039] The first power supply unit 310 and the second power supply unit 312 supply power to operate the first control unit 323 and the second control unit 311, respectively. When a power button provided on the imaging device 102 is pressed, power is first supplied to both the first control unit 323 and the second control unit 311, but as will be described later, the first control unit 323 is controlled to turn off its own power supply to the first power supply unit 310. Even while the first control unit 323 is not operating, the second control unit 311 continues to operate, and information is input from the device vibration detection unit 309 and the audio processing unit 314. The second control unit is configured to determine whether or not to activate the first control unit 323 based on various input information, and if it is determined that the first control unit 323 should be activated, it instructs the first power supply unit to supply power.
[0040] The audio output unit 318 outputs a preset audio pattern from a speaker built into the image capture device 102, for example, during shooting.
[0041] The LED control unit 324 controls the LEDs provided in the image capturing device 102 to light up and flash in a preset pattern, for example, during shooting.
[0042] Video output unit 317 is, for example, a video output terminal, and transmits an image signal to display a video on a connected external display, etc. Audio output unit 318 and video output unit 317 may be combined into one terminal, such as an HDMI (registered trademark) (High-Definition Multimedia Interface) terminal.
[0043] The communication unit 322 communicates between the imaging device 102 and external devices, transmitting and receiving data such as audio signals, image signals, compressed audio signals, and compressed image signals. It also receives control signals related to imaging, such as commands to start and stop imaging and pan, tilt, and zoom, and drives the imaging device 102 based on instructions from external devices that can communicate with the imaging device 102. It also transmits and receives information such as various parameters related to learning, which are processed by the learning processing unit 319 (described later), between the imaging device 102 and external devices. The communication unit 322 is a wireless communication module, such as a Bluetooth communication module conforming to the Bluetooth standard or a wireless LAN communication module conforming to the IEEE 802.11 standard. The imaging device 102 can communicate with external devices via the local network 101 and the Internet 100 by using the communication unit 322. When multiple imaging devices 102 used by the same user are present on the same local network 101, they communicate with each other using the communication unit 322 to operate cooperatively.
[0044] When the image capture device 102 is running and the face authentication unit 325 finds the face of a person in the vicinity, it performs face authentication processing on the found face and determines whether it matches the face image of the person received from the learning server 107. Types of face authentication processing include, for example, a 2D authentication method that recognizes the positions of the eyes, nose, mouth, etc. on the face and compares them with a database for authentication, and a 3D authentication method that uses an infrared sensor in addition to the 2D method for authentication.
[0045] The setting unit 326 sets a shooting release condition according to the face information of the person authenticated by the face authentication unit 325. The shooting release condition is determined based on an image evaluation value and a shooting threshold value. The image evaluation value is a value that serves as an index of whether the scene captured by the image capture device 102 is suitable for shooting, and is determined based on face detection information, face authentication information, blink rate, facial expression of the subject, facial orientation, size of the subject, etc. The shooting threshold value is a criterion value that, when the image evaluation value exceeds this value, triggers shooting, and is determined based on the shooting frequency, the time elapsed since the previous shooting, etc.
[0046] The GPU 327 is a processor that handles arithmetic processing of video signals, and is configured using, for example, a GPU (Graphical Processing Unit) or an FPGA (Field-Programmable Gate Array), and is a calculation unit that can process input data in parallel.
[0047] The GPU 327 can perform efficient calculations by processing a larger amount of data in parallel, and therefore, when calculations are performed multiple times using a learning model such as deep learning, it is effective to use the GPU 327 for processing. Therefore, in the first embodiment, the GPU 327 is used in addition to the first control unit 323 for processing by the learning processing unit 319. Specifically, when a prediction program is executed, the first control unit 323 and the GPU 327 work together to perform calculations to make an estimation.
[0048] The scene discrimination unit 328 is a discrimination unit that discriminates the captured scene of an image from the image signal of the image processing unit 307. For example, the scene discrimination unit 328 inputs the image signal into a scene discrimination trained model (not shown) and discriminates the captured scene from the output result. Here, the scene discrimination trained model is a trained model that trains the image signal as input data using scenes such as sports days, play days, and entrance ceremonies as training data, and outputs a scene similar to the input image signal as output data. The scene discrimination unit 328 may also discriminate the captured scene by receiving scene information from the communication terminal 103 or by using scene designation information input by the user from the UI display unit 502.
[0049] <Indoor location information measurement> The image capture device 102 performs indoor positioning using the communication unit 332.
[0050] The communication unit 332 included in the image capture device 102 receives a beacon signal (MAC address, SSID, radio wave intensity) transmitted from an access point and estimates the distance to the access point based on the radio wave intensity to perform positioning. Three or more access points are used, distances from the three access points are calculated, and the intersection of these distances is determined by three-point positioning to derive the indoor installation position, i.e., two-dimensional coordinates. The two-dimensional coordinates are associated with the MAC address and SSID and recorded in the memory 315 of the image capture device 102.
[0051] The more access points there are, the more accurate the position measurement becomes. Access points already installed indoors are used, and measurements are performed repeatedly as the imaging device moves. If sufficient radio wave strength is not available, the signal arrival time can be used instead.
[0052] Furthermore, the height of three-dimensional coordinates is determined by three-point positioning, as described above, using three or more access points positioned at different heights. As with two-dimensional coordinates, the height is associated with the MAC address and SSID and recorded in the memory of the imaging device. As with two-dimensional coordinates, the more access points there are, the more accurate the position measurement becomes.
[0053] If the number of access points available for positioning in the height direction is less than three, positioning in the height direction may be performed using the image shake detection unit 309 inside the image capture device 102. In this case, the relative movement direction and movement speed are measured using an acceleration sensor or gyro included in the device shake detection unit 309, and the displacement from the initial installation position, i.e., the relative height, is derived to measure the position.
[0054] In this embodiment, the information processing device capable of communicating with the imaging device 102 via a network is a device equipped with similar hardware, and the data collection server 106 will be described as a representative example. The CPU 332 is a central processing unit (CPU) capable of executing various processes, and is a processor capable of interpreting user instructions and instructions from other devices and executing program code. The system bus 331 is a bus for exchanging data between various pieces of hardware within the information processing device, including the CPU 332. The ROM 333 is a non-volatile memory that records various setting information for the information processing device, and is configured, for example, by an EEPROM or Flash ROM, and stores programs such as a BIOS (Basic Input Output System). The RAM 334 is the main storage device of the information processing device, and is used when the CPU 332 temporarily stores data to be processed and when programs are loaded. The RAM 334 is configured, for example, by a dynamic random access memory (DRAM). The HDD 335 is an auxiliary storage device of the information processing device. The HDD 335 is configured using hardware such as a hard disk drive (HDD) or a solid state drive (SSD). The HDD 335 can be used to store the operating system (OS) of the information processing device, programs executed by the CPU 332, various data input to the programs, and program execution results.
[0055] The GPU 336 is a processor responsible for arithmetic processing of video signals and the like of the information processing device. The display unit 338 is a display unit capable of displaying video signals processed by the GPU 336, and is configured, for example, by an LCD (Liquid Crystal Display), and may be integrated with or separate from the information processing device. The input unit 337 is an input unit through which the user issues instructions to the information processing device, and may be, for example, a keyboard, mouse, or touch panel. The communication unit 339 is a communication unit used by the information processing device to communicate with external devices. The communication unit 339 is configured, for example, by a wireless LAN communication module conforming to the IEEE 802.11 standard or a wired LAN communication module conforming to the IEEE 802.3 standard. The communication terminal 103, the content providing server 105, and the learning server 107 have the same hardware configuration as the data collection server 106, and therefore, depending on the scale of the system, a single information processing device may share multiple functions. In this embodiment, the content providing server 105, the data collecting server 106, and the learning server 107 are independent devices, and are capable of communicating with each other via a network via the communication unit 339.
[0056] FIG. 4 is a conceptual diagram showing the input / output structure using the learning model in this embodiment.
[0057] Input data X401 is input data to be input to the neural network to be trained. Output data Y402 is output data obtained when input data X401 is input to the neural network. Trained model 403 is an example of a neural network that is a training model. Trained model 403 is made up of a network consisting of a multilayer perceptron, for example, in which a large number of neuron models called units are connected.
[0058] The trained model 403 is used to predict output data Y 402 from input data X 401, and by training in advance to output a teacher output value for input data, it is possible to predict output data for new input data that follows the trained teacher. The training in this embodiment will be described later.
[0059] FIG. 5 is a diagram for explaining the software functions of each part of a system to which the present invention is applied, using the learning model shown in FIG.
[0060] The imaging device 102 includes a data transmission / reception unit 501 , a UI display unit 502 , and an estimation unit 503 .
[0061] The client terminal 104 includes a web browser 511 .
[0062] The content providing server 105 includes a data presenting unit 521 , a content-related data managing unit 522 , and a data storing unit 523 .
[0063] The data collection server 106 includes a data receiving unit 531 , a data collecting and providing unit 532 , a data storage unit 533 , and an estimation model providing unit 534 .
[0064] The learning server 107 includes a learning unit 541 , a learning data generation unit 542 , a data storage unit 543 , and an estimation model output unit 544 .
[0065] The data transmission / reception unit 501 has the function of transmitting and receiving content data such as still images and videos, and transmitting and receiving estimation model data between the imaging device 102 and the data collection server 106, and executes this function using the communication unit 322. The data transmission / reception unit 501 only needs to be compatible with a communication protocol for transmitting and receiving files between a server and a client, such as an FTP (File Transfer Protocol) client, and can also be implemented using other protocols or proprietary protocols.
[0066] The estimation unit 503 has a function of inputting necessary input data to the learned estimation model and performing predictive calculations to obtain output data. The estimation unit 503 applies the estimation model received using the function of the data transmission / reception unit 501 using the communication unit 322, and performs predictive calculations through the cooperative operation of the first control unit 323 and the GPU 327.
[0067] The web browser 511 is software that a user uses to view and obtain content, such as still images and videos, stored in the content providing server 105 from the client terminal 104. The web browser 511 only needs to be able to send and receive data using a protocol based on HTTP (Hyper Text Transfer Protocol), for example, and display the results of the transmission and reception on the display unit 338 of the user's client terminal 104.
[0068] The data presentation unit 521 has a function for providing still images and videos stored in the content providing server 105 to the client terminal 104 used by the user. The data presentation unit 521 may have, for example, an HTTP server function, so long as it can receive requests from the client terminal 104 and transmit content and data in response to the received requests. The content-related data management unit 522 provides content stored in the data storage unit 523 to the data presentation unit 521 based on requests from the client terminal 104 received by the data presentation unit 521, and manages the provided content and content viewing information. The content-related data management unit 522 may manage and control content information by using, for example, a database function capable of operating a database written in a structured query language such as SQL.
[0069] The data storage unit 523 has a function of storing and providing content data that temporarily stores moving image and still image content to be provided to the client terminal 104 used by the user. The data storage unit 523 receives and stores still image and moving image content from the data collection server 106 so that it can provide content to the client terminal 104, and also has a function of generating and holding data for easy confirmation of content, such as thumbnail images.
[0070] The data receiving unit 531 has a data receiving function for receiving still image and video data captured by the imaging device 102. The data receiving unit 531 has, for example, an FTP server function and a function for storing still image and video data received using the communication unit 339 in the data storage unit 533.
[0071] The data collection and provision unit 532 manages the still image and video data stored in the data storage unit 533, and has the function of providing information on the still image and video data it manages in response to requests from the content providing server 105 and the learning server 107. The data collection and provision unit 532 may manage and control the information on the still image and video data it manages by using a database function that can operate a database written in a structured query language such as SQL, for example.
[0072] The data storage unit 533 has a function of storing and reading data in the auxiliary storage device shown as the HDD 335. The data storage unit 533 only needs to operate so as to be able to store still image and video data received by the data receiving unit 531 and to read out stored information based on instructions from the data collection and provision unit 532, for example.
[0073] The estimation model providing unit 534 is configured to provide an estimation model to be used by the estimation unit 503 of the imaging device 102 in response to a request from the imaging device 102. The estimation model provided from the estimation model providing unit 534 to the imaging device 102 can be obtained by acquiring the estimation model output by the learning server 107 at a specific timing and making it possible to provide the estimation model to the imaging device 102.
[0074] The estimation model providing unit 534 may be configured to execute a function that allows the learning server 107 and the image capture device 102 to share estimation model data using a function such as an FTP server.
[0075] 4, and operates to generate the learning model 403 using input data X401 and a teacher output value. The learning model 403 generated by the learning unit 541 is provided to the estimation model output unit 544, and then provided to the imaging device 102 via the data collection server 106. When learning in the learning unit 541, the learning unit 541 operates to execute learning by instructing the CPU 332 and GPU 336 to perform calculations.
[0076] The learning data generation unit 542 has a function of generating input data X401 to the learning unit 541 and output values that serve as teachers, requests information to be used for learning from the data collection server 106, and selects and processes the received information. The learning unit 541 performs learning using the input data X401 output by the learning data generation unit 542 and the output values that serve as teachers, and updates the learning model 403.
[0077] Data storage unit 543 has a function for storing and reading data from an auxiliary storage device such as HDD 335. Data storage unit 543 can store learning input data X401 and teacher output values generated by learning data generation unit 542, and learning model 403. Data storage unit 543 only needs to be able to store input data X401 and teacher output values based on a request from learning data generation unit 542, and provide the stored data based on a request from learning unit 541, for example.
[0078] In response to a request from the data collection server 106, the estimation model output unit 544 provides the learning model 403 based on the request to the estimation model provision 534 of the data collection server 106. Furthermore, if the learning model 403 to be provided has not been updated since the last time it was provided, it may be deemed to have not been updated and its provision may be canceled.
[0079] As shown in Figure 5, the operation of this embodiment is realized by exchanging data between devices using software functions within each device. Note that while this embodiment illustrates a system in which each device exists independently and operates in coordination with others, it can also be realized by integrating the functions of the devices, for example by configuring the data collection server 106 and the learning server 107 as the same device. It can also be realized by making modifications, such as dividing the learning server 107 into multiple devices and allowing the learning calculations of a large-scale learning model to operate in parallel.
[0080] Fig. 6 is a diagram illustrating an example of the operation of the system in this embodiment. The operation of the system in this embodiment will be explained using Fig. 6. The following explanation illustrates the case where a user purchases still images and videos, which are content provided by the system, and the operation of each device of the present invention that operates in response to the purchase.
[0081] <Users can view and purchase still image and video content> The user operates the client terminal 104, browses the content providing server 105, selects and purchases the still image or video content they need. For example, thumbnails of a list of content provided by the content providing server 105 are displayed on the web browser on the client terminal 104, and the user selects each piece of content one by one to perform the purchase process.
[0082] The content providing server 105 receives the purchase process from the client terminal 103 and transmits to the data collection server 106, as purchase information, information indicating the purchased still image and video data and information indicating the user who made the purchase. The data collection server 106 transmits the still image and video data to the content providing server 105 based on the information received from the content providing server 104. The content providing server 105 stores the still image and video data received from the data collection server 106 so that it can be provided to the user's client terminal 104, and transmits the still image and video data that are the purchased content in response to a request from the client terminal 104. Once the client terminal 104 has received all the purchased content, it transmits a notification of reception completion to the content providing server 105, completing the process for purchasing the content required by the user. After the purchase is complete, the content providing server 105 discards the provided still image and video data, and if the user requests it again, the content providing server 105 makes a request to the data collection server 106 in the same manner as described above and provides the data in the same manner.
[0083] <Learning model based on purchase information> The data collection server 106 transmits still image and video data related to the user and its purchase information to the learning server 107. The learning server 107 stores purchase information and image data of users who have previously made purchases, selected and processed for learning purposes, and stores additional learning data when it receives new purchase information and related still image and video data. The learning server 107 processes the purchase information and related still image and video data received from the data collection server 106 into learning data through processing by the learning data generation unit 542. In the case of still image data, the learning data generation unit 542 processes and selects the data by resizing it to the required size or selecting the subjects in the image by linking it to the purchase information, and then passes the processed and selected data to the data storage 543 for storage as learning data. In addition, in the case of video data, the learning data generation unit 542 of the learning server 107 cuts out video frames, processes and selects them in the same way as in the case of still images, and passes them to the data storage unit 543 for storage as learning data.
[0084] Here, while video is composed of multiple frames, learning data is not acquired from all frames; only representative frames are used. The representative frame shown here is selected, for example, as the first frame of the video, or the index frame if the video is in IPB format. In addition to the image data processed and selected here, the learning server 107 associates the user who purchased the data and the purchase price of each still image and video data purchased by the user using the learning data generation unit 542. Furthermore, the input data X401 used for learning the learning model 403 and the training data are stored in the HDD 335 according to the data storage unit 543.
[0085] The input data and training data used for learning in this embodiment are exemplified in FIG. 7 and will be described later.
[0086] After preparing input data X401 and training data, the learning server 107 uses these to train the learning model 403. The input data X401 to the learning model 403 uses image data processed by the learning data generation unit 542, a user ID, and scene information at the time of shooting, and the training data uses the purchase prices of videos and still images, respectively, to train the learning model 403. The trained model generates a learning model 403 that predicts the rating values of still images and videos preferred by a specific user from images acquired by an imaging element.
[0087] <Application of learning model to imaging devices> The data collection server 106 periodically queries the learning server 107 to check whether the trained learning model 403 has been updated. For example, the data collection server 106 calculates a hash value of the stored learning model, and the learning server 107 similarly calculates a hash value of the learning model 403 and passes the hash value of the learning model 403 upon request from the data collection server 106. The data collection server 106 compares the hash values of the learning model 403 on the data collection server 106 and the learning server 107 to check whether they match.
[0088] If the hash values of the respective learning models 403 do not match, the data collection server 106 determines that the learning model 403 has been updated and requests the learned learning model from the learning server 107, and the learning server 107 sends the learning model to the data collection server 106.
[0089] When the imaging device 102 starts operation, it inquires of the data collection server 106 whether the learning model has been updated, and if an update has been found, it requests the learning model from the data collection server 106 .
[0090] The data collection server 106 transmits the learning model based on a request from the image capture device 102, and the image capture device 102 updates the learning model in its own device with the received learning model.
[0091] <Automatic photography and post-photography actions based on learning models> After completing the update of the learning model, the imaging device 102 receives an automatic shooting instruction and an instruction for shooting scene information and starts automatic shooting operation. When not shooting during automatic shooting, the imaging device 102 periodically drives the imaging unit 306 to acquire images, acquires image data, and uses it to determine whether to shoot automatic shooting. The imaging device 102 inputs the acquired image data and scene information as input data into the learning model, and acquires predicted evaluation values for still images and videos for a specific user as output of the learning model. The first control unit 323 of the imaging device 102 compares the predicted evaluation values for the acquired still images and videos, the history since receiving the shooting instruction, and the state of the imaging device 102 itself, such as the remaining battery capacity of the first power supply unit 310, against a lookup table to determine whether to shoot still images or videos.
[0092] After receiving the instruction to end shooting, if the image capturing device 102 is in a state where it can communicate with the data collection server 106 , it reads the recording medium 321 and transmits the captured still image and video data to the data collection server 106 .
[0093] FIG. 7 is a table showing an example of learning data indicating input data and teacher data used for learning in the learning server 107 in this embodiment.
[0094] The learning data is generated from information acquired by the learning server 107 from the data collection server 106. The learning data is generated on the learning server 107 based on still image and video content purchased by the user, and the learning data ID, user ID, image identifier, subject information, scene information, shooting time, and the images associated with each are used as input data. In addition, from the learning data, whether the image data is a video or a still image and the purchase price are held as expected values for each learning data ID, and are stored as still image expected values and video expected values.
[0095] For example, by storing the table diagram shown in Figure 7 as table data for each specific community, it is possible to train a learning model to predict the expected values of still images and videos so as to satisfy the preferences of each user within the specific community.
[0096] In the example of Figure 7, only still images and videos with a purchase history are shown as examples of learning data, but non-purchased still images and videos may also be used for further training of the learning model as images with low expected values.
[0097] A lookup table is consulted to determine whether to take a still or video image.
[0098] After receiving the instruction to end shooting, if the image capturing device 102 is in a state where it can communicate with the data collection server 106 , it reads the recording medium 321 and transmits the captured still image and video data to the data collection server 106 .
[0099] FIG. 7 is a table showing an example of learning data indicating input data and teacher data used for learning in the learning server 107 in this embodiment.
[0100] The learning data is generated from information acquired by the learning server 107 from the data collection server 106. The learning data is generated on the learning server 107 based on still image and video content purchased by the user, and the learning data ID, user ID, image identifier, subject information, scene information, shooting time, and the images associated with each are used as input data. In addition, from the learning data, whether the image data is a video or a still image and the purchase price are held as expected values for each learning data ID, and are stored as still image expected values and video expected values.
[0101] For example, by storing the table diagram shown in Figure 7 as table data for each specific community, it is possible to train a learning model to predict the expected values of still images and videos so as to satisfy the preferences of each user within the specific community.
[0102] In the example of Figure 7, only still images and videos with a purchase history are shown as examples of learning data, but non-purchased still images and videos may also be used for further training of the learning model as images with low expected values.
[0103] FIG. 8 is a flowchart illustrating the operation of learning the trained model 403 executed by the learning unit 541 of the learning server 107 in this embodiment.
[0104] FIG. 8(a) is a flowchart illustrating the operation of the content providing server 105 related to the learning operation of this embodiment.
[0105] In step (hereinafter abbreviated as S) 801, the content providing server 105 checks whether the user's login information entered into the data presentation unit 521 via the web browser 511 exists. Here, the login information is, for example, a user ID required for the user to purchase still images and videos via the web browser 511. If the login information does not exist, S801 is repeated until the user enters the login information (NO in S801). If the login information exists, the process proceeds to S802 (YES in S801).
[0106] In S802, the content providing server 105 requests the data collection server 106 for an image / video list corresponding to the user ID, and in S803, it checks whether the requested image / video list has been received. If the image / video list has been received, the process proceeds to S804 (YES in S803).
[0107] In S804, the content providing server 105 displays the still images and videos included in the image / video list on the user's web browser 511 via the data presentation unit 521. Note that in a flow not shown, the viewing time during which the user views the still images and videos is measured between S804 and S808.
[0108] In S805, the content providing server 105 determines whether or not it has received an instruction to select a displayed still image or video from the client terminal 104. If it has received an instruction to select, it transitions to S806 (YES in S805) and causes the web browser 511 to display a purchase screen.
[0109] In S807, the content providing server 105 determines whether or not a purchase request for the displayed still image or video has been received from the client terminal 104. If a purchase request has been received, the process proceeds to S808 (YES in S807) and payment processing is performed. In S808, the content providing server 105 performs payment processing using a payment processing unit (not shown).
[0110] In S809, the content providing server 105 transmits purchase information to the data collecting server 106. The purchase information includes information such as the user ID of the user who purchased each still image or video, the viewing time, the favorite rating, and the purchase amount. The purchase information is selected from the learning data described below and is used as input data or training data for the trained model.
[0111] FIG. 8(b) is a flowchart illustrating the operation of the data collection server 106 related to the learning operation of this embodiment.
[0112] In S821, the data collection server 106 acquires images and videos from the imaging device 102 and stores them in the data storage unit 533. In S822, the acquired images and videos are stored in the data storage unit 533. The acquired images and videos are assigned shooting information by the imaging device 102 at the time of capture. The shooting information includes, as subject information, general object recognition results for the current angle of view, face detection results, the number of faces captured in the current angle of view, the degree of smile and eye closure, face angle, face recognition ID number, gaze angle of the subject, and scene determination results. The shooting information also includes, as setting information, zoom magnification, pan / tilt motion method (movement pattern and driving speed), shutter speed, white balance, audio level, LED lighting method (color, flashing duration, etc.). Furthermore, as environmental information, the current time, height information from GPS location information and a barometer, vibration information from acceleration information, and tilt information from a gyro sensor are also included. This information attached to images and videos is selected as training data, which will be described later, and used as input data or training data for the trained model.
[0113] In S823, the data collection server 106 determines whether or not there is a request for an image / video list from the content providing server 105. If there is a request (YES in S823), the process proceeds to S824, where the data collection / providing unit 532 reads images / videos corresponding to the user ID from the data storage unit 533, creates an image / video list, and transmits it to the content providing server 105.
[0114] In S825, the data collection server 106 determines whether or not purchase information has been received from the content providing server 105. If the purchase information has been received (YES in S825), the process proceeds to S826.
[0115] In S826, the data collection server 106 integrates the purchase information and the image and video shooting information acquired by the data collection / provision unit 532, and generates an image attribute list, which is learning data for learning (described later), shown in FIG. 7. The generated image attribute list is sent to the learning server 107.
[0116] In S827, the data collection server 106 determines whether a request to send the trained model has been sent from the imaging device 102. If a request has been sent (YES in S827), the process proceeds to S828. Note that in this embodiment, the request to send the trained model is sent from the imaging device 102, but it may also be sent from the communication terminal 103.
[0117] In S828, the data collection server 106 requests the learning server 107 to send the latest trained model, and in S829, repeats confirmation of reception until the trained model is received. Once reception of the trained model is complete, the process proceeds to S830 (YES in S829). Note that the data collection server 106 requests the trained model in accordance with a request from the imaging device 102, but the data collection server 106 may periodically request transmission from the learning server 107 based on updates to the purchase information from the content providing server 105.
[0118] In S830, the data collection server 106 transmits the latest trained model to the imaging device 102.
[0119] FIG. 8(c) is a flowchart illustrating the operation of the learning server 107 related to the learning operation of this embodiment.
[0120] In S841, the learning server 107 determines whether or not it has received an image attribute list from the data collection server 106. If it has not received an image attribute list, the learning server 107 waits until it receives an image attribute list from the data collection server 106 (NO in S841). If it has received an image attribute list, it transitions to S842.
[0121] In S842, the learning data generation unit 542 of the learning server 107 selects input data X401 and training data to be input to the trained model from the image attribute list, and transmits them to the learning unit 541. In the first embodiment, for example, the subject, event name, and image evaluation value are selected as the input data X401, and the photo sales amount is selected as the training data.
[0122] In S843, S844, and S845, the learning unit 541 uses a machine learning algorithm to generate the trained model 403. Specific examples of machine learning algorithms include nearest neighbor algorithms, naive Bayes algorithms, decision trees, and support vector machines. Deep learning, which uses a neural network to generate features and connection weighting coefficients for learning, can also be used. Any available algorithm from the above can be used as appropriate and applied to this embodiment. For example, in this embodiment, the learning unit 541 generates the trained model 403 using a neural network. The learning unit 541 may include an error detection unit and an update unit. The error detection unit obtains an error between the training data and output data output from the neural network in response to input data input to the input layer. The error detection unit may use a loss function to calculate the error between the training data and output data from the neural network.
[0123] The update unit updates the connection weighting coefficients between the nodes of the neural network based on the error obtained by the error detection unit so as to reduce the error. This update unit updates the connection weighting coefficients, for example, using the backpropagation algorithm. The backpropagation algorithm is a method for adjusting the connection weighting coefficients between the nodes of each neural network so as to reduce the error. In the first embodiment of this embodiment, subject information and scene information are used as input data X401, and the weights of the connection weighting coefficients are adjusted based on the sales of the purchase information as training information. Then, the trained model 403, which outputs the expected sales value for each subject in the scene as output data Y402, is trained. When all the data selected in S842 has been input, the process proceeds to S846 (YES in S845).
[0124] In S846, the learning server 107 determines whether or not there is a request for the trained model 403 from the data collection server 106. If there is a request, the process proceeds to S847 (YES in S846).
[0125] In S847, the learning server 107 transmits the trained model 403 to the estimation model providing unit 534 of the data collection server 106 via the estimation model output unit 544.
[0126] FIG. 9 is a flowchart illustrating the prediction operation executed by the estimation unit 503 of the image capturing device 102 in this embodiment.
[0127] In S901, the estimation unit 503 of the imaging device 102 determines whether input information has been acquired from the first control unit 323. If input information has been acquired, the process proceeds to S902. The first control unit 323 acquires information from each function as shown in FIG. 3 to operate the imaging device 102. In the first embodiment, the first control unit 323 confirms whether subject information acquired by the face authentication unit 325 and scene information acquired by a scene identification unit (not shown) have been acquired as input information. Here, the scene identification unit (not shown) identifies a scene from, for example, a video obtained from the image processing unit 307 using a trained model (not shown) that identifies the scene. The scene information may also be set from the communication terminal 103 via the local network 101. Furthermore, the input unit 337 may be used to directly select from a pre-prepared scene list (not shown).
[0128] In S902, the estimation unit 503 selects input data X401 from the input information acquired in S901. In the first embodiment, subject information and scene information are selected as the input data X401.
[0129] In S903, the estimation unit 503 inputs the input data X401 to the trained model 403.
[0130] In S904, the estimation unit 503 outputs the output data Y402 from the trained model 403. In the first embodiment, an expected value for each subject is obtained as the output data Y402.
[0131] In S905, the estimation unit 503 transmits the output data Y402 acquired in S904 to the first control unit 323. Based on the output data Y402, the first control unit 323 controls the imaging device 102 so as to perform an operation sequence described below.
[0132] FIG. 10 is a flowchart illustrating the operations executed by the image capture device 102 in this embodiment.
[0133] The operation of the image capture device 102 in this embodiment will be described using the flowchart in Fig. 10. Before the start of the flowchart shown in Fig. 10, the image capture device 102 is in a power-off state, and begins operation upon start.
[0134] In S1001, the imaging device 102 communicates with the data collection server 106 to check whether the currently applied learning model needs to be updated.
[0135] In S1002, the imaging device 102 compares the learning model applied to the imaging device 102 with the learning model on the data collection server 106. If it is determined in S1002 that the learning model needs to be updated, in S1003 the learning model is updated via the communication unit 322 (YES in S1002). For example, when a DNN (Deep Neural Network) model is used as the learning model, if there is a change in the number of layers or the number of units, the entire DNN network is updated, and if the number of layers or the number of units remains the same, parameters such as coefficients and connection weights are updated. If there is no update to the learning model in S1002 or if the learning model update process in S1003 has finished, the process transitions to S1004, where shooting condition information is acquired.
[0136] In the shooting condition setting, the user who installed the image capturing device 102 sets the conditions related to shooting, and may also set rules such as the interval between automatic shooting and shooting conditions for moving images and still images.
[0137] In addition, in S1004, the imaging device 102 uses a smart device display to prompt the user to reinstall the imaging device in the recommended placement position.
[0138] First, the smart device display displays an indoor floor plan, the current installation position of the imaging device, and a recommended installation position. Next, a UI message (not shown) or the like is displayed, urging the user to change the imaging device from its current installation position to the recommended installation position. After that, when it is detected that the user has changed the installation position to the recommended installation position, the UI message stops displaying and the reinstallation is completed. The recommended installation position is displayed as a point or concentric circle, and the reinstallation is completed when the installation position is at the same coordinate as the point or within the concentric circle.
[0139] In S1005, as described above with reference to FIG. 9, the shooting condition information acquired in S1004 is input into the trained model, and an expected value according to the camera settings is predicted.
[0140] In S1006, the initial value of TH_SHOT is set to the shooting threshold value obtained from the predicted expected value in S1005. TH_SHOT will be described later in S1011.
[0141] In S1007, it is confirmed whether or not an automatic shooting instruction has been issued by the operator of the imaging device 102, and if an automatic shooting instruction has been issued, the process proceeds to S1008, where the imaging device 102 starts automatic shooting (YES in S1007). If an automatic shooting instruction has not been issued, the imaging device 102 waits until an instruction is issued (NO in S1007).
[0142] When automatic shooting starts, in S1009, the image processing unit 307 performs image processing on the signal captured by the imaging unit 306 to generate an image for subject detection. The generated image is subjected to subject detection processing to detect people, objects, etc. After subject detection processing is performed, the process transitions to S1010.
[0143] In S1010, an image evaluation value is calculated from the subject information acquired in S1009 according to the subject's situation and shooting conditions. The image evaluation value is calculated based on, for example, the number of people, the size and orientation of the people's faces, the people's facial expressions, the results of personal authentication of the people, etc.
[0144] FIG. 11 is a diagram showing the relationship between the image evaluation value and the threshold value in this embodiment, with the horizontal axis representing the elapsed time and the vertical axis representing the image evaluation value.
[0145] From S1011, automatic photography determination processing is performed.
[0146] In this embodiment, the photographing frequency is changed by updating the photographing threshold and the photographing interval. The method for updating the photographing threshold is as follows, and updating the photographing interval will be described later in S1014.
[0147] Automatic photography calculates an evaluation value from the state of the subject, compares the evaluation value with a threshold, and performs automatic photography if the evaluation value exceeds the threshold. The threshold at this time is TH_SHOT in FIG. 11; the higher TH_SHOT, the less likely automatic photography is performed, and the lower TH_SHOT, the more likely it is performed. Note that TH_SHOT is an indefinite value, and is updated when it exceeds a threshold representing the image evaluation value. The threshold at this time is TH_TR, and TH_TR is a value greater than TH_SHOT. If the image evaluation value calculated in S1010 exceeds TH_TR, the process transitions to S1012 (YES in S1011).
[0148] In S1012, as described above in FIG. 9, the image data acquired in S1009 and the shooting condition information acquired in S1004 are input into the trained model, and an expected value according to the camera settings is predicted.
[0149] In S1013, TH_SHOT is updated to the shooting threshold value obtained from the predicted expected value in S1012.
[0150] Similarly, in S1014, the shooting interval is updated to the shooting interval obtained from the predicted expected value in S1012, and the process proceeds to S1016 to execute shooting. The shooting interval at this time is Δt in FIG. 11, and the longer Δt is, the more single-shot automatic shooting is performed, and the shorter Δt is, the more continuous automatic shooting is performed. The updated shooting threshold and shooting interval are applied from the next automatic shooting determination process. If the image evaluation value is equal to or less than TH_TR in S1011, the process proceeds to S1015.
[0151] In S1015, it is determined whether the image evaluation value exceeds TH_SHOT, and if it exceeds TH_SHOT, the process transitions to S1016 and photography is performed (YES in S1015). If it is equal to or less than TH_SHOT, the process of S1016 is skipped and control is exercised so that photography is not performed (NO in S1015). If photography is performed in S1016, or if the image evaluation value is equal to or less than TH_SHOT in S1015, the process transitions to S1017 and waits for the photography interval time updated in S1014.
[0152] Thereafter, if an end instruction is received in S1018, it is determined that the image capture is to be ended, a predetermined power-off process is performed for the image capture device 102, and the operation of the image capture device 102 is ended (YES in S1018). Unless an end instruction is received in S1018, S1009 to S1018 are repeated.
[0153] As a result, it is possible to set up the camera in a suitable shooting position and perform automatic shooting.
[0154] <Smart device display> FIG. 12 shows an example of reinstalling an imaging device using a dedicated application of an external device 1201, which is a smart device according to the first embodiment of the present invention, using FIG. 12 1).
[0155] The smart device display unit 1202 displays an indoor floor plan, the current installation position of the imaging device, and a recommended installation position (1203).
[0156] The indoor floor plan suitable for the installation location is loaded in advance by the user operating a UI button (not shown) or the like.
[0157] The coordinates of the current installation position of the imaging device are derived by indoor position measurement by the imaging device, and the coordinates are transferred to the smart device through communication between the imaging device and the smart device, and are reflected on the display.
[0158] The recommended installation location is output to the imaging device in S905 of Figure 9 and transferred to the smart device through communication between the imaging device and the smart device, where it is reflected on the display. The floor plan and location information are associated, and the recommended installation location is displayed as a point on the floor plan. Alternatively, the recommended installation location may be displayed as concentric circles with an arbitrary radius and centered on the point. A UI message is displayed encouraging the user to move, and when it is detected that the user has subsequently changed to the recommended installation location, the UI message stops displaying and the reinstallation is completed.
[0159] Second Embodiment In the second embodiment, steps 100 to 107 are the same as those shown in Fig. 1, and therefore a description thereof will be omitted. The operation of the imaging device 102 in this embodiment will be described using the flowchart in Fig. 12. Steps S1201 to S1203 are the same as steps S1001 to S1003 in the first embodiment shown in Fig. 10.
[0160] In step S1204, the imaging device 102 prompts the user to reset the imaging device to a recommended placement position using a smart device display. First, the user is prompted to reset the imaging device to a suitable shooting position, and then, if there is a recommended placement position in the height direction, the user is prompted to change the height of the imaging device by, for example, displaying a UI message encouraging the use of a tripod (S1205).
[0161] This makes it possible to set up the camera at an appropriate shooting position and height for guiding the viewer's line of sight.
[0162] Operations from S1206 to S1219 are similar to those from S1005 to S1018 of the first embodiment shown in FIG.
[0163] As described above, in the second embodiment, in automatic photography, it is possible to set the camera at a suitable height, guide the camera's line of sight, and take a photograph at the camera's line of sight.
[0164] <Smart device display> An example of resetting the imaging device in the height direction via a dedicated application of the external device 1201, which is a smart device according to the second embodiment of the present invention, will be described using FIG. 13(2).
[0165] The recommended height placement is performed after the recommended installation position is completed. The recommended height placement position is output to the image capture device in S905 in Figure 9, and is transferred to the smart device through communication between the image capture device and the smart device, and is reflected on the display.
[0166] A UI message (1304) is displayed to prompt the user to relocate the imaging device by using a tripod or moving it to a higher position. When it is detected that the user has changed the installation position in the vertical direction, the display of the UI message is stopped and the relocation is completed.
[0167] It is desirable to display UI messages in stages according to the recommended vertical placement position, such as: 1) encouraging the use of accessories that can increase height, such as a tripod; 2) encouraging the user to find a higher location within the recommended installation range; or 3) encouraging the user to reconsider the installation range itself in order to take eye-level photos.
[0168] <Gaze guidance configuration> An example of eye guidance will be described with reference to Fig. 14. Eye guidance is performed using the audio output unit 318, lens barrel rotation driver 305, image processor 307, and first controller 323 of the imaging device.
[0169] First, an imaging device installed indoors uses input to the image processing unit 307, and when the face of the subject is detected by a face recognition module (not shown) contained within the image processing unit 307, the audio output unit 318 is activated and at the same time the lens barrel rotation drive unit 305 is caused to repeatedly move back and forth to attract the subject's gaze (1 in Figure 14).
[0170] The above-mentioned operation of the lens barrel is stored as several fixed patterns in the memory 315 of the imaging device, and is read out randomly by the first control unit. The above-mentioned repetitive operation is configured from patterns that are likely to attract the subject's attention, such as high-speed reciprocating movement of the lens barrel in conjunction with panning and tilting, unlike the operation during photography.
[0171] The sound output unit has several fixed patterns stored in the memory 315 of the imaging device, which are read out randomly by the first control unit. The sound patterns are different from the warning sounds and shutter sounds when taking a picture and are composed of fixed sound patterns for guiding the viewer's line of sight.
[0172] Next, after a series of lens barrel operations, if the face of the subject is detected again using the input to the image processing unit 307 and a face recognition module (not shown), it is determined that the specific subject has turned their face toward the imaging device due to the sound of the imaging device, and the eye guidance is successful. After that, an image is immediately taken (2) in FIG. 14).
[0173] If the face detection of the subject fails, it is determined that the gaze guidance has failed.
[0174] Thereafter, the success or failure of the eye guidance is recorded in the memory 315 of the image capturing device together with the image data, image evaluation value, and subject information.
[0175] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.
[0176] The present invention can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having the computer of the system or device read and execute the program. The computer has one or more processors or circuits, and may include multiple separate computers or a network of multiple separate processors or circuits to read and execute computer-executable instructions.
[0177] The processor or circuitry may include a central processing unit (CPU), a microprocessing unit (MPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a data flow processor (DFP), or a neural processing unit (NPU). [Explanation of symbols]
[0178] 100 Internet 101 Local Network 102 Imaging device 103 Communication terminal 104 client terminals 105 Video Server 105 Content Provider Server 106 Data Collection Server 107 Learning Server
Claims
1. A photograph purchasing system comprising an imaging means for imaging a subject and outputting image data, and a storage means for storing the image data output by the imaging means, a calculation means for calculating an evaluation value used to determine whether or not to perform a photographing operation by the imaging means; a first setting means for setting a first threshold value used to determine whether or not to perform a photographing operation by the imaging means; a determination means for determining whether or not a photographing operation is to be performed using the evaluation value and the first threshold value; a second threshold value used to determine whether to control the first threshold value; a processing means for referencing the purchase history of image data stored in the purchasing system and processing the number of purchases by the user, the subject information, scene information, location information, and shooting time of the stored photographs; Using a prediction model that has been machine-learned using training data as training data, the sales volume of the image data predicted from the number of purchases by users stored in the purchasing system, An imaging device characterized in that an installation position suitable for imaging of said imaging means is predicted and imaging is performed.
2. 2. The imaging apparatus according to claim 1, further comprising a detection unit that detects information about a subject, wherein the calculation unit determines the evaluation value using the information about the subject detected by the detection unit.
3. 3. The imaging device according to claim 1, wherein the information about the subject is associated with user information held by the purchasing system.
4. 4. The imaging device according to claim 1, wherein the scene information input to the prediction model is input for each scene by an operator of the imaging device.
5. 5. The imaging device according to claim 1, wherein the initial value of the first threshold is set based on the prediction model using scene information input for each scene by an operator of the imaging device.
6. 6. The imaging device according to claim 1, wherein the first setting unit determines whether to change the first threshold value using the evaluation value and the second threshold value.
7. A photograph purchasing system comprising an imaging means for imaging a subject and outputting image data, and a storage means for storing the image data output by the imaging means, a calculation means for calculating an evaluation value used to determine whether or not to perform a photographing operation by the imaging means; a first setting means for setting a first threshold value used to determine whether or not to perform a photographing operation by the imaging means; a determination means for determining whether or not a photographing operation is to be performed using the evaluation value and the first threshold value; a second threshold value used to determine whether to control the first threshold value; A processing means for referencing the purchase history of image data stored in the purchasing system and processing the number of purchases by the user, the subject information of the stored photographs, scene information, location information, and shooting time; and a gaze guidance means for guiding the gaze of the user by sound and lens barrel operation in the imaging means. Using a prediction model that has been machine-learned using training data, the number of purchases by the user stored in the purchasing system, the shooting time of the stored photograph, the shooting location, and the sales volume of the image data predicted from the success or failure of gaze guidance, An imaging device characterized in that an installation position suitable for guiding the line of sight of the imaging means is predicted and an image is taken.
8. 8. The imaging apparatus according to claim 7, further comprising a detection unit that detects information about a subject, wherein the calculation unit determines the evaluation value using the information about the subject detected by the detection unit.
9. 9. The imaging device according to claim 7, wherein the information about the subject is associated with user information held by the purchasing system.
10. 10. The imaging device according to claim 7, wherein the scene information input to the prediction model is input for each scene by an operator of the imaging device.
11. 11. The imaging device according to claim 7, wherein the initial value of the first threshold is set based on the prediction model using scene information input for each scene by an operator of the imaging device.
12. 12. The imaging device according to claim 7, wherein the first setting unit determines whether to change the first threshold value using the evaluation value and the second threshold value.
Citation Information
Patent Citations
Methods for storing, searching, editing and outputting photographic images
JP2004514976A
Imaging apparatus, control method of the same, program, and storage medium
JP2021057815A