Image capture device, image capture device control method and program
The imaging device uses a line-of-sight sensor and learning model to analyze eye movements, addressing the time lag and accuracy issues in conventional devices, ensuring precise scene capture.
Patent Information
- Application Number
- JP2021162763
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-01
- Publication Date
- 2025-11-17
- Estimated Expiration
- 2041-10-01
AI Technical Summary
Conventional imaging devices suffer from a time lag between the user's decision to shoot and the actual shooting operation, leading to missed scenes, and existing gaze-based shooting instructions have low detection accuracy.
An imaging device with a line-of-sight sensor, analysis unit, and learning model that detects eye movements, analyzes gaze patterns, and estimates the optimal shooting time based on learned user behavior.
Prevents missing desired scenes by accurately timing the shooting operation based on user eye movements, enhancing capture precision.
Smart Images

Figure 0007770844000001 
Figure 0007770844000002 
Figure 0007770844000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, a control method for an imaging device, and a program. [Background technology]
[0002] In recent years, digital imaging devices have undergone remarkable evolution. To achieve further evolution in the future, it is expected that they will be integrated with other devices to achieve higher performance and more functionality. In this case, it is conceivable that information input by eye gaze will be utilized as input information. Patent Document 1 describes a technology that utilizes eye gaze information, in which a change signal is output to an image switching control device to instruct switching between an electronic image and an external world image, provided that the state of the convergence angle θ between the left and right eyeballs changes and the same state continues for a long period of time. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 8-160345 Summary of the Invention [Problem to be solved by the invention]
[0004] In conventional imaging devices, a shooting instruction is input by pressing the release button, but there is a time lag between when the user decides to shoot and when the shooting operation actually starts, which can result in the user missing the scene they wanted to capture.Even if a shooting instruction is input based on information about the convergence angle between the left and right eyes, as in the technology described in Patent Document 1, the detection accuracy of the shooting instruction is low, and the above problem has not been solved.
[0005] The present invention aims to prevent a user from missing a scene that he or she wants to capture. [Means for solving the problem]
[0006] The imaging device of the present invention includes a detection means for detecting the state of the user's eyes, an analysis means for analyzing eye movements based on the detection result by the detection means, and an image capturing device obtained by inputting movement information representing the eye movements analyzed by the analysis means into a learning model. Indicate whether this is the scene you want to capture. and an estimation means for estimating the timing to start the release operation based on the output data. [Effects of the Invention]
[0007] According to the present invention, it is possible to prevent missing a scene that you want to capture. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an imaging apparatus. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of an imaging apparatus. [Figure 3] FIG. 10 is a diagram illustrating an example of a functional configuration in a learning phase. [Figure 4] FIG. 10 is a diagram illustrating an example of information used to generate learning data. [Figure 5] FIG. 1 is a conceptual diagram showing an input / output structure using a learning model. [Figure 6] FIG. 1 is a diagram for explaining a method for analyzing eye movements. [Figure 7] 10 is a flowchart showing a process executed in a learning phase. [Figure 8] FIG. 10 is a diagram illustrating an example of a functional configuration in an estimation phase. [Figure 9] 10 is a flowchart showing a process executed in an estimation phase. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings.
[0010] <Overall configuration of the imaging device> FIG. 1 is a diagram showing an example of the overall configuration of an imaging device 100 according to this embodiment. The imaging device 100 is a digital still camera and includes an imaging lens 101 and an imaging element 102. The imaging lens 101 is a lens group including a zoom lens and a focus lens. The imaging element 102 forms an optical image guided by the imaging lens 101 on an imaging plane and converts the image into an electrical signal.
[0011] The imaging device 100 also includes a CPU 103, a memory 104, a GPU (Graphics Processing Unit) 105, and an FPGA (Field Programmable Gate Array) 106. The CPU 103 controls the entire imaging device 100. The memory 104 is a RAM, a ROM, a HDD, or the like, and stores programs. The CPU 103 executes the programs stored in the memory 104 to realize processing of the flowcharts described below. The GPU 105 and FPGA 106 can perform efficient calculations by processing a larger amount of data in parallel. Therefore, when learning is performed multiple times using a learning model such as deep learning, it is effective to perform processing using the GPU 105 and FPGA 106. In this embodiment, when analysis and learning are performed, the GPU 105 and FPGA 106 work together with the CPU 103 to perform processing.
[0012] The imaging device 100 also includes a line-of-sight sensor 107 , a display element 108 , a display element drive circuit 109 , an eyepiece 110 , and a release button 111 . The line-of-sight detection sensor 107 is a sensor for detecting that the user is looking through the finder, and outputs the detection result to the CPU 103. The display element 108 is configured with a liquid crystal screen or the like, and is provided inside the viewfinder. A display element drive circuit 109 drives the display element 108 under the control of the CPU 103 to display an image captured by the image capture element 102 on the screen of the display element 108. The display element 108 is an example of a display unit. The eyepiece 110 is used to enlarge and observe the image displayed on the display element 108. The release button 111 is a button operated by the user when taking a picture, and an operation signal is input to the signal input circuit 203 (FIG. 2).
[0013] Furthermore, imaging device 100 includes illumination light sources 112a-112b, a light splitter 114, a light receiving lens 115, and an eye image sensor 116, and has a function of detecting the state of the eyes of a user looking through the viewfinder. The state of the eyes here includes the direction of gaze, the degree to which the eyes are open, the degree to which the eyes are narrowed, the state of the eyes being open or closed, etc. Illumination light sources 112a-112b are light sources that illuminate user's eyeball 113 in order to detect the user's gaze direction from the relationship between the pupil and the image resulting from the corneal reflection of the light source. Illumination light sources 112a-112b are composed of infrared light-emitting diodes and are arranged around eyepiece 110. The illuminated eyeball image and the image resulting from the corneal reflection of illumination light sources 112a-112b pass through eyepiece 110, are reflected by light splitter 114, and are then focused by light receiving lens 115 on eyeball image sensor 116, which is a two-dimensional array of photoelectric elements such as CCDs. Light receiving lens 115 positions the pupil of user's eyeball 113 and eyeball image sensor 116 in a conjugate imaging relationship. From the extent of the white of the eye in the eyeball image focused on eyeball image sensor 116, it is possible to detect the degree to which the eyes are wide open, narrowed, or open. Furthermore, the line of sight direction can be detected from the positional relationship between the eyeball image formed on the eyeball image pickup device 116 and the image of the illumination light sources 112a to 112b due to the corneal reflection.
[0014] <Hardware configuration of imaging device> Fig. 2 is a diagram showing an example of the hardware configuration of the imaging device 100 of Fig. 1. The same components as those in Fig. 1 are assigned the same numbers.
[0015] As shown in FIG. 2, the CPU 103 is connected to the image sensor 102, memory 104, display device drive circuit 109, line of sight detection circuit 201, photometry circuit 202, signal input circuit 203, illumination light source drive circuit 205, and GPU 105, and controls these devices. The image sensor 102 outputs an electrical signal to the CPU 103 as a captured image. The memory 104 records the captured image output from the image sensor 102, the image information output from the eye image sensor 116, data necessary for learning, and the like. The display element drive circuit 109 executes processing for displaying on the display element 108 under the control of the CPU 103. In this embodiment, live view display is performed by sequentially displaying captured images output from the image sensor 102 on the display element 108. Hereinafter, images displayed in live view will be referred to as through images.
[0016] The gaze detection circuit 201 A / D converts the output resulting from the formation of an eyeball image from the eyeball imaging element 116 and outputs this image information to the CPU 103. The CPU 103 acquires the image information from the gaze detection circuit 201, extracts each feature point of the eyeball image required for detecting the state of the eye from the acquired image information according to a predetermined algorithm, and detects the position of each extracted feature point. The detection results are output to the GPU 105 and FPGA 106. The photometry circuit 202 calculates luminance information of the object scene based on an electrical signal obtained from the image sensor 102 , which also functions as a photometry sensor, and outputs the information to the CPU 103 . Signal input circuit 203 is connected to SW1, which is a switch that is turned on by the first stroke (first operation) of release button 111 and that instructs the start of preparations for shooting, such as photometry and distance measurement, and outputs information that SW1 has been turned on to CPU 103. Signal input circuit 203 is also connected to SW2, which is a switch that is turned on by the second stroke (second operation) of release button 111 and that instructs the start of a release operation, and outputs information that SW2 has been turned on to CPU 103. The release operation here refers to an operation for storing a captured image output from image sensor 102 at the timing that SW2 is turned on in memory 104 as still image data. Under the control of the CPU 103, the illumination light source drive circuit 205 executes a process for driving the illumination light sources 112a to 112b used when detecting the direction of the user's line of sight.
[0017] The CPU 103 analyzes eye movement based on the information output from the gaze detection circuit 201. Here, the eye movement refers to changes in the degree of eye opening, changes in the degree of eye squinting, gaze movement, blinking cycles, etc. In the learning phase, the CPU 103 generates learning data based on the eye movement information obtained by the analysis and the output information from the signal input circuit 203, and uses the generated learning data to train a learning model. In the estimation phase, the CPU 103 inputs the eye movement information obtained by analyzing the information output from the gaze detection circuit 201 into the learning model, thereby estimating the timing to start a release operation. In this embodiment, the CPU 103 executes the above-described processing in cooperation with the GPU 105 and FPGA 106.
[0018] <Functional configuration in the learning phase> Fig. 3 is a diagram showing an example of the functional configuration of the imaging device 100 according to this embodiment in the learning phase. The CPU 103 executes a program stored in the memory 104 to control each device connected to the CPU 103 and function as a state detection unit 301, a data storage unit 302, an analysis unit 303, a shooting request detection unit 304, and a learning unit 305 shown in Fig. 3. Note that the imaging device 100 can switch between activating and deactivating functions in the learning phase according to user settings.
[0019] The state detection unit 301 detects the state of the user's eyes based on the information output from the line-of-sight detection circuit 201. The state detection unit 301 (CPU 103) is an example of a detection means. The data storage unit 302 stores information about the eye condition detected by the condition detection unit 301 in the memory 104 .
[0020] The analysis unit 303 reads information about the eye state for a predetermined time period stored in the memory 104 and analyzes eye movements based on the read information. The analysis unit 303 (CPU 103) is an example of an analysis means. The analysis of eye movements may be performed by the GPU 105 and FPGA 106. In this embodiment, the analysis unit 303 reads information about the eye state for a predetermined time period from the memory 104 and determines, for each time period, each item, such as whether the eyes are wide open, narrowed, moving, or closed. Then, the analysis unit 303 determines each item for the information about the eye state for the predetermined time period and analyzes changes in the degree of eye opening, changes in the degree of eye narrowing, gaze movements, blinking cycles, and the like. The analysis method is not limited to the above. For example, a method of patterning eye movements or a method of classifying eye movements into predetermined patterns may also be used. The analysis unit 303 provides the analysis results to the learning unit 305 as eye movement information.
[0021] The imaging request detection unit 304 detects the timing at which SW1 and SW2 are turned on when it receives an input that SW1 and SW2 are turned on from the signal input circuit 203. The imaging request detection unit 304 provides the learning unit 305 with the timing at which SW1 and SW2 are turned on. The learning unit 305 generates learning data by associating the eye movement information with the timing when SW1 and SW2 are turned on. The learning unit 305 generates multiple pieces of learning data and stores them in the memory 104. The learning unit 305 also learns a learning model using the multiple pieces of learning data stored in the memory 104. The learning unit (CPU 103) 305 is an example of a learning means.
[0022] <Explanation about the training data> Fig. 4 is a diagram showing an example of information used to generate learning data. The information shown in Fig. 4 is composed of accumulated data of record 400, and is stored in memory 104. Record 400 includes information such as time 401, elapsed time 402, coordinate a 403, coordinate b 404, output a 405, output b 406, ON of SW1 407, ON of SW2 408, and analysis result 409. Time 401 is the time when record 400 is stored. Elapsed time 402 is the time that has elapsed since the start of eye state detection until the time 401 in the same record 400. In this embodiment, eye state detection starts when the line-of-sight detection sensor 107 detects that the user is looking into the viewfinder. The coordinate a403, the coordinate b404, the output a405, and the output b406 store information about the state of the eye at the time 401 in the record 400. The coordinates a403 and b404 represent position coordinates (reflection image coordinates 631a and 631b in FIG. 6) on the eye image sensor 116 (CCD) corresponding to the reflection images from the illumination light sources 112a and 112b. The outputs a405 and b405 represent the output intensities of the CCD at the coordinate positions of the coordinates a403 and b404. Details will be described later with reference to FIG. 6. Note that, as the coordinate a403, the pupil edge coordinate 641a or the iris edge coordinate 651a in FIG. 6 may be used together with the reflection image coordinate 631a in FIG. 6. Furthermore, as the coordinate b404, the pupil edge coordinate 641b or the iris edge coordinate 651b in FIG. 6 may be used together with the reflection image coordinate 631b in FIG. 6.
[0023] In the record stored at time 401 when SW1 was turned ON, flag information indicating that SW1 was turned ON is stored in ON 407 for SW1. The same is true for ON 408 for SW2. In the example of Fig. 4, time 401 indicates that SW1 was turned ON at "9:34:00", and time 401 indicates that SW2 was turned ON at "9:34:22". The analysis result 409 holds the results obtained by analyzing the information about the eye condition at time 401 in the record 400. The analysis result 409 is a description for explaining the analysis result in an easy-to-understand manner. The CPU 103 analyzes the eye movement from the time series transition of the information about the eye condition. In the example shown in FIG. 4, the following information is obtained as a result of analyzing the eye movement. After the gaze moved slightly to the left from the center, the release button 111 was pressed and SW1 was turned on. After SW1 is turned on, the eye blinks, the iris shrinks little by little, and then the release button is pressed to turn SW2 on.
[0024] <Explanation on learning methods> FIG. 5 is a conceptual diagram showing the input / output structure using the learning model of this embodiment. Eye movement information 501 is input data to be input to a learning model 503. As the learning model 503, for example, a neural network is used. The eye movement information 501 is movement information that represents eye movement, and is information obtained by analyzing, for example, changes in the degree of eye opening, changes in the degree of eye squinting, gaze movement, and blinking cycles. In this embodiment, the information is obtained by analyzing information relating to the eye state for a predetermined period of time (for example, accumulated data of coordinates a403, coordinates b404, output a405, and output b406 in FIG. 4). The eye movement information 501 includes time information. Note that the information relating to the eye state for a predetermined period of time may be used as is as the eye movement information 501.
[0025] The learning unit 305 associates time information of the eye movement information 501 with timing 502 when the release button 111 is operated (timing when SW1 is turned ON, timing when SW2 is turned ON). The learning unit 305 then uses the eye movement information 501 as input data, and learns a learning model 503 using information from the eye movement information 501 that is from when detection of the eye state begins to when the release button 111 is operated to timing 502 as correct answer data. Information 504 indicating whether the scene is one that the user wishes to photograph is output from the learning model 503. The learning unit 305 trains the learning model 503 so that, when an eye movement similar to one that is easily detected at or immediately before the timing when the release button 111 is operated is input, information indicating that the scene is one that the user wishes to photograph is output as output information 504.
[0026] Specific examples of machine learning algorithms include nearest neighbor algorithms, naive Bayes algorithms, decision trees, and support vector machines. Deep learning, which uses a neural network to generate features and connection weighting coefficients for learning, is also an example. Any of the above algorithms that can be used can be used as appropriate and applied to this embodiment.
[0027] The learning unit 305 may also include an error detection unit and an update unit. The error detection unit obtains the error between the training data and output data output from the output layer of the neural network in response to input data input to the input layer. The error detection unit may use a loss function to calculate the error between the output data from the neural network and the training data. The update unit updates the connection weighting coefficients between the nodes of the neural network based on the error obtained by the error detection unit so as to reduce the error. This update unit updates the connection weighting coefficients, for example, using an error backpropagation method. The error backpropagation method is a technique for adjusting the connection weighting coefficients between the nodes of each neural network so as to reduce the error.
[0028] <Method for analyzing eye movements> Next, a method for analyzing eye movement will be specifically described using Fig. 6. The upper part of each of Figs. 6(i) to (v) shows a schematic diagram of an eyeball image projected onto the eyeball image sensor 116. In each schematic diagram of the eyeball image, regions A, B, and C are provided corresponding to the upper, middle, and lower parts of the eyeball image. The lower parts of Figs. 6(i) to (v) show graphs representing the output intensity of the CCD corresponding to regions A, B, and C of the eyeball image, respectively. The horizontal axis of each graph represents the X-axis of the CCD. The vertical axis of each graph represents the output intensity of the CCD at that X-coordinate. The pupil 610 represents a projected image of the pupil of the user's eyeball 113. The iris 620 represents a projected image of the iris of the user's eyeball 113. The reflected images 630a and 630b represent reflected images from the illumination light sources 112a to 112b irradiating the pupil of the user's eyeball 113. The upper eyelid 660 represents the position of the upper eyelid in the eyeball image. The lower eyelid 670 represents the position of the lower eyelid in the eyeball image. The white of the eye 680 represents the range of the white of the eye in the eyeball image.
[0029] Reflected image coordinates 631a and 631b represent CCD coordinate positions corresponding to reflected images 630a and 630b. Pupil edge coordinates 641a and 641b represent CCD coordinate positions corresponding to pupil edges 640a and 640b representing the left and right ends of pupil 610. Iris edge coordinates 651a and 651b represent CCD coordinate positions corresponding to iris edges 650a and 650b representing the left and right ends of iris 620.
[0030] In this embodiment, the analysis unit 303 analyzes the movement of the user's eyes based on the reflected images by the illumination light sources 112a-112b, the coordinate positions of the pupil edge and the iris edge, and the output intensity of the CCD in the eyeball image projected on the eyeball imaging element 116. As a specific example, the state in Fig. 6(i) is defined as a reference state, and each of the states in Fig. 6(ii) to Fig. 6(v) will be described below.
[0031] (ii) Eye movement Compared to the reference state shown in FIG. 6(i), in FIG. 6(ii), the pupil edge coordinates 641a, 641b and reflected image coordinates 631a, 631b in area B are shifted to the left as one faces. At this time, it can be seen that the user is moving their gaze to the right relative to the center. Furthermore, although not shown, when the pupil edge coordinates 641a, 641b and reflected image coordinates 631a, 631b are shifted to the right as one faces, it can be seen that the user is moving their gaze to the left relative to the center. By using the accumulated data of position information such as the pupil edge coordinates 641a, 641b and reflected image coordinates 631a, 631b, the analysis unit 303 analyzes how the user's gaze is moving.
[0032] (iii) blinking Compared to the reference state shown in FIG. 6(i), FIG. 6(iii) shows a state in which the CCD output intensity at reflected image coordinates 631a, 631b, etc. in region B is lower. At this time, the user is looking through the viewfinder, and it is understood that the pupil 610 is blocked by the upper eyelid 660, and the user has their eyes closed. Furthermore, if it is detected that the CCD output intensity has returned to its original state after being lowered, it is understood that the user is blinking. By using such accumulated data of the CCD output intensity at reflected image coordinates 631a, 631b, etc., the analysis unit 303 analyzes whether the user is blinking, blinking quickly, blinking slowly, etc.
[0033] (iv) Eyes wide open Compared to the reference state shown in Figure 6(i), Figure 6(iv) shows a state in which the output is increasing at the same level in the range from iris edge coordinate 651a in region A and region C toward the starting point, X coordinate 0, and in the range in which the coordinate increases from iris edge coordinate 651b. At this time, the range of the white of the eye 680 in region A and region C increases, indicating that the user's eyes are wide open. By using this accumulated data of the output intensity of the CCD outside the iris edge coordinates 651a and 651b, the analysis unit 303 analyzes how the degree to which the user's eyes are wide open is changing.
[0034] (v) Squinting Compared to the reference state shown in Figure 6(i), Figure 6(v) shows a state in which the output is decreasing at the same level in the range from iris edge coordinate 651a in region A and region C toward the starting point, X coordinate 0, and in the range where the coordinate increases from iris edge coordinate 651b. At this time, the area of the white of the eye 680 in region A and region C is decreasing, indicating that the user is squinting. By using this accumulated data of the output intensity of the CCD outside the iris edge coordinates 651a and 651b, the analysis unit 303 analyzes how the degree to which the user is squinting is changing.
[0035] As described above, the state of the user's eyes is detected using output information of reflected images 630a, 630b, pupil edges 640a, 640b, iris edges 650a, 650b, reflected image coordinates 631a, 631b, pupil edge coordinates 641a, 641b, iris edge coordinates 651a, 651b, etc. Furthermore, by using this accumulated data, changes in the state of the user's eyes are analyzed.
[0036] <Processing performed during the learning phase> 7 is a flowchart showing the processing executed in the learning phase by the imaging device 100 according to this embodiment. The processing shown in this flowchart is realized by the CPU 103 executing a program stored in the memory 104. The processing shown in this flowchart starts when the imaging device 100 is powered on. In step S701, CPU 103 determines whether or not line-of-sight detection sensor 107 has detected that the user has looked into the viewfinder. The process of step S701 is repeated until CPU 103 determines that line-of-sight detection sensor 107 has detected that the user has looked into the viewfinder. If CPU 103 determines that line-of-sight detection sensor 107 has detected that the user has looked into the viewfinder, the process proceeds to step S702. In step S702, CPU 103 starts acquiring a through image from image sensor 102 and displays the acquired through image on display device 108. This allows the user to visually recognize the subject by looking at the through image displayed on display device 108 in the viewfinder. At this time, CPU 103 starts acquiring data from eye image sensor 116. This starts detection of the state of the eye.
[0037] Next, in step S703, the CPU 103 detects information about the state of the user's eyes from the data acquired from the eye image sensor 116, and stores the detected information in the memory 104. In this embodiment, one row of records in FIG. 4 is stored. Next, in step S704, CPU 103 determines whether or not SW1 of release button 111 has been turned on by a user operation. If CPU 103 determines that SW1 has been turned on, the process proceeds to S705, and if CPU 103 determines that SW1 has not been turned on, the process returns to S703, and the accumulation of information regarding the state of the eye continues. In step S705, the CPU 103 stores the timing when SW1 of the release button 111 was turned ON in the memory 104. For example, when the CPU 103 receives a notification from the signal input circuit 203 that SW1 has been turned ON, the CPU 103 stores time information such as the current time and the time elapsed since the line of sight detection sensor 107 detected the line of sight in step S701 in the memory 104.
[0038] Next, in step S706, the CPU 103 stops the acquisition of data from the eye image pickup device 116 that began in step S702. That is, the CPU 103 stops the acquisition of data because SW1 has been turned ON. For example, flag information or the like is added to the record last stored in the memory 104 so that it can be identified as the record when SW1 was turned ON. Next, in step S707, CPU 103 determines whether or not SW2 of release button 111 has been turned on by a user operation. If CPU 103 determines that SW2 has not been turned on, the process proceeds to step S709, and if CPU 103 determines that SW2 has been turned on, the process proceeds to step S708.
[0039] Next, a series of processes from step S709 to step S712 will be described. If a long time has passed since SW1 was turned ON until SW2 was turned ON, it is assumed that a major change, such as a movement of the subject, has occurred. Therefore, it is highly likely that the information accumulated up to SW1 does not correspond to the scene actually captured when SW2 was turned ON. Therefore, if SW2 is not turned ON within a predetermined time after SW1 was turned ON, in step S709, CPU 103 resumes the acquisition of data stopped in step S706. Next, in S710, the CPU 103 detects information about the state of the user's eyes from the data acquired from the eye image sensor 116, as in S703, and stores the detected information in the memory 104.
[0040] Next, in step S711, CPU 103 determines whether or not SW2 of release button 111 has been turned on by a user operation. If CPU 103 determines that SW2 has been turned on, the process proceeds to S712, and if CPU 103 determines that SW1 has not been turned on, the process returns to S710, and the accumulation of information regarding the state of the eye continues. In S712, the CPU 103 stops the data acquisition that was resumed in step S709. Next, in step S708, the CPU 103 stores the timing at which SW2 of the release button 111 was turned ON in the memory 104. The processing content of step S708 overlaps with that of step S705, and therefore a description thereof will be omitted.
[0041] Next, in step S713, the CPU 103 controls the GPU 105 and FPGA 106 to read information relating to the eye state from the memory 104 and analyze the eye movement. Next, in step S714, CPU 103 controls GPU 105 and FPGA 106 to associate the eye movement information obtained by the analysis in S713 with the timing at which SW1 and SW2 were pressed. At this time, CPU 103 associates time information, such as the current time stored in S705 and S708 and the time elapsed since detection by gaze detection sensor 107, with the eye movement information from detection by gaze detection sensor 107 until SW1 is turned ON and until SW2 is turned ON. After that, the series of processes in the flowchart ends.
[0042] The eye movement information obtained as described above is used as learning data to be input to the learning model 503. In this embodiment, if the time from when SW1 is turned on to when SW2 is turned on is less than a predetermined time, the CPU 103 uses the eye movement information from when the line-of-sight detection sensor 107 detects SW1 to when SW1 is turned on as correct data. If the time from when SW1 is turned on to when SW2 is turned on is equal to or longer than a predetermined time, the CPU 103 uses the eye movement information from when the line-of-sight detection sensor 107 detects SW1 to when SW2 is turned on as correct data.
[0043] In addition, the CPU 103 may exclude eye movement information detected in a specific state that is assumed to be the preparation stage for shooting, such as a state where the zoom ratio is changing or a state where the imaging device 100 is moving, from the target of correct answer data. Furthermore, the CPU 103 may use the eye movement information detected when the time interval at which SW2 is turned on is short as learning data for continuous shooting scenes. When training the learning model using the learning data for continuous shooting scenes, the learning model is trained so as to output information indicating whether the scene is one that should be captured by continuous shooting.
[0044] 7 is repeatedly executed, whereby a plurality of pieces of training data are stored in the memory 104. Thereafter, the CPU 103 controls the GPU 105 and the FPGA 106 to train a training model using the plurality of pieces of training data. The trained training model is stored in the memory 104 and is used in the estimation phase, which will be described next.
[0045] <Function configuration in the estimation phase> 8 is a diagram showing an example of the functional configuration in the estimation phase of the imaging device 100 according to this embodiment. The CPU 103 executes a program stored in the memory 104 to control each device connected to the CPU 103, and functions as a state detection unit 301, a data storage unit 302, an analysis unit 303, an estimation unit 801, and an imaging control unit 802 shown in FIG. In this embodiment, the imaging device 100 can switch between activating and deactivating the function in the estimation phase according to a user setting. When the function in the estimation phase is activated, a release operation is executed by operating the release button 111 or at the timing when it is estimated that the scene has been captured. When the function in the estimation phase is deactivated, a release operation is executed only when the release button 111 is operated.
[0046] The functions of the state detection unit 301, data storage unit 302, and analysis unit 303 are the same as those described in the learning phase, and therefore a description thereof will be omitted. The estimation unit 801 uses the eye movement information analyzed by the analysis unit 303 as estimation data, and inputs the estimation data into a learning model stored in the memory 104. Indicates whether the scene currently being captured is the scene the user wants to capture. The estimation unit 801 then estimates, based on the output data obtained, whether the scene currently being shot is the scene that the user wants to shoot. If the output data indicates that the scene currently being captured is the scene the user wants to capture, it is inferred that the scene currently being captured is the scene the user wants to capture. If the output data does not indicate that the scene currently being captured is the scene the user wants to capture, it is inferred that the scene currently being captured is not the scene the user wants to capture. When the estimation unit 801 estimates that the scene is one that the photographer wants to capture, it outputs an instruction to start a release operation to the shooting control unit 802. In other words, the estimation unit 801 estimates the timing to execute the release operation. The estimation unit 801 (CPU 103) is an example of an estimation means. The photographing control unit 802 controls each device to execute the release operation when an instruction to start the release operation is input from the estimation unit 801. The photographing control unit 802 (CPU 103) is an example of a control means.
[0047] <Processing performed in the estimation phase> 9 is a flowchart showing the processing executed in the estimation phase by the imaging device 100 according to this embodiment. The processing shown in this flowchart is realized by the CPU 103 executing a program stored in the memory 104. The processing shown in this flowchart starts when the imaging device 100 is powered on.
[0048] The processing in steps S901 to S904 is the same as the processing in steps S701 to S703 and step S713 in the learning phase described with reference to FIG. 7, and therefore a description thereof will be omitted here. In step S905, the CPU 103 controls the GPU 105 and FPGA 106 to read out the learning model stored in the memory 104, and inputs the eye movement information obtained by the analysis in S904 into the read out learning model. Next, in step S906, the CPU 103 controls the GPU 105 and FPGA 106 to obtain, as output data from the learning model, information indicating whether the scene is one that the user wants to capture.
[0049] Next, in step S907, CPU 103 controls GPU 105 and FPGA 106 to estimate whether the scene currently being captured is one that the user wants to capture, based on the information indicating whether the scene is one that the user wants to capture, acquired in S906. If CPU 103 estimates that the scene is one that the user wants to capture, the process proceeds to S908. If CPU 103 estimates that the scene is not one that the user wants to capture, the process proceeds to S901.
[0050] In step S908, CPU 103 controls each device to execute a release operation. At this time, CPU 103 outputs a shutter sound from an alarm sound output unit (not shown) in synchronization with the release operation. In this case, the shutter sound is output even though the release button 111 is not operated, which may cause discomfort to the user. Therefore, CPU 103 differentiates the shutter sound output due to the estimation in S907 from the normal shutter sound output due to operation of the release button 111. For example, the CPU 103 outputs a sound higher in pitch than the normal shutter sound. After that, the series of processes in the flowchart ends.
[0051] The CPU 103 displays the still image data obtained by the release operation in S908 on a rear monitor (not shown) of the imaging device 100 to present it to the user. Note that, depending on the accuracy of the estimation, it is conceivable that the release operation may be performed at a timing different from the scene that the user actually wants to capture. Therefore, after performing the release operation in S908, the CPU 103 determines whether the release button 111 has been operated at a timing that would be linked to the data used for the analysis in step S904. For example, after performing the release operation in S908, it determines whether the release button 111 has been operated within a predetermined time. If the release button 111 has been operated within the predetermined time, the CPU 103 displays the still image data obtained by the release operation in S908 on a rear monitor of the imaging device 100 to present it to the user. If the release button 111 is not operated within a predetermined time, the CPU 103 discards the still image data obtained by the release operation in S908 from the memory 104 so that it is not presented to the user.
[0052] Note that instead of a configuration that performs processing using a learning model, the imaging device 100 may be configured to perform processing using eye movement patterns extracted from multiple pieces of learning data obtained in the learning phase. In that case, for example, the eye movement patterns are stored in memory 104 or the like. The imaging device 100 calculates a similarity by comparing the results of the eye movement analysis with the stored eye movement patterns, and estimates whether the scene is a captured scene based on the calculated similarity. In other words, the imaging device 100 performs processing equivalent to that of the estimation unit 801 described above, using the eye movement patterns extracted from multiple pieces of learning data.
[0053] Furthermore, the image capturing device 100 may be configured to perform rule-based processing such as a lookup table (LUT). In this case, for example, a relationship between pattern data indicating eye movement patterns and whether the scene is one that the user wants to capture is created in advance as an LUT, and the created LUT is stored in the memory 104 or the like. The image capturing device 100 refers to the stored LUT and estimates whether the scene is one that the user wants to capture based on the analysis results of the eye movement. In other words, the image capturing device 100 uses the LUT to perform processing equivalent to that of the estimation unit 801 described above.
[0054] As described above, the imaging device 100 of this embodiment learns the characteristics of the user's eye movement at or immediately before the release button 111 is operated. Using the results of this learning, it becomes possible to estimate the timing to perform the release operation from the user's eye movement, such as gaze movement, changes in eye opening and closing, changes in eye squinting, and blinking cycles. This makes it possible to immediately start the release operation at the estimated timing when the user wants to capture a scene, thereby preventing the user from missing a desired scene. Furthermore, when capturing a still image during video capture, it is also possible to prevent image misalignment that occurs when the user operates the release button 111.
[0055] Although the present invention has been described above with reference to the embodiments, the above embodiments are merely illustrative of specific examples of how the present invention can be implemented, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.
[0056] The present invention can also be realized by providing a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having the computer in the system or device read and execute the program. The computer may have one or more processors or circuits, and may include multiple separate computers or a network of multiple separate processors or circuits to read and execute computer-executable instructions. The processor or circuit may include a central processing unit (CPU), a microprocessing unit (MPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a field-programmable gateway (FPGA). The processor or circuit may also include a digital signal processor (DSP), a data flow processor (DFP), or a neural processing unit (NPU).
[0057] As a first modification of this embodiment, the CPU 103 may analyze image information obtained from the eyeball image sensor 116 to acquire eyeball information such as features of the eyeball 113 and features of the eyelashes, and identify the user based on the acquired eyeball information. In this case, in the learning phase, the CPU 103 learns a learning model for each identified user. The memory 104 stores the learning model learned for each user. In this case, in the estimation phase, the CPU 103 performs estimation using the learning model learned for each user.
[0058] As a second modification of this embodiment, CPU 103 may detect the user's gaze position on display element 108 and analyze the relationship between the user's gaze position on display element 108 and the subject position on the through image. In this case, CPU 103 acquires information on whether the user's gaze position is in the subject area and whether the user's gaze movement follows the subject's movement. In the learning phase, CPU 103 may exclude from the target of correct answer data eye movement information that is detected when the user's gaze position is not in the subject area or when the gaze movement does not follow the subject's movement. In other words, eye movement information that is detected when the user's gaze position and the subject position on the through image are not linked is excluded from the target of correct answer data. [Explanation of symbols]
[0059] 100: imaging device, 102: imaging element, 103: CPU, 104: memory, 107: line-of-sight detection sensor, 116: eye imaging element
Claims
1. a detection means for detecting the state of the user's eyes; analysis means for analyzing eye movements based on the detection result by the detection means; an estimation means for estimating the timing to start a shutter release operation based on output data indicating whether the scene is a desired scene, which is obtained by inputting the movement information representing the eye movement analyzed by the analysis means into a learning model; and An imaging device comprising:
2. The imaging device according to claim 1, further comprising a learning means for learning the learning model using learning data in which the movement information is input data and the movement information when the release button is operated is used as correct data.
3. 3. The imaging device according to claim 2, wherein the learning means uses the movement information from when the detection means starts detection until the release button is operated as correct data.
4. The imaging device described in claim 3, characterized in that if the time between when preparation for shooting is instructed by a first operation of the release button and when shooting is instructed by a second operation of the release button is less than a predetermined time, the learning means uses the movement information from when detection by the detection means begins to when the first operation is performed as correct data.
5. The imaging device described in claim 3 or 4, characterized in that if the time between when preparation for shooting is instructed by a first operation of the release button and when shooting is instructed by a second operation of the release button is longer than a predetermined time, the learning means uses the movement information from when detection by the detection means begins to when the second operation is performed as correct data.
6. 6. The imaging device according to claim 2, wherein whether or not the learning by the learning means is performed can be switched according to a setting.
7. The imaging device described in any one of claims 1 to 6, characterized in that the analysis means analyzes at least one of the following items based on a predetermined period of data regarding the eye condition detected by the detection means: movement of the gaze, changes in the degree of eye opening, changes in the degree of eye squinting, and blinking cycle.
8. 8. The imaging device according to claim 1, further comprising a control means for controlling the release operation to be performed at the estimated timing.
9. 9. The imaging device according to claim 8, wherein a shutter sound emitted when the control means executes a release operation is made different from a shutter sound emitted when the release button is operated.
10. The control means If the release button is operated within a predetermined time after the control means has performed the release operation, still image data obtained by the release operation is presented to the user; If the release button is not operated within a predetermined time after the control means executes the release operation, the still image data obtained by the release operation is discarded.
10. The imaging device according to claim 8, wherein the imaging device is a lens.
11. 11. The imaging device according to claim 1, wherein the learning model is configured as a neural network.
12. the detection means detects a line of sight position relative to a display unit on which a captured image is displayed, the analyzing means analyzes a relationship between the gaze position and the position of the subject on the captured image, The imaging device according to claim 3, wherein the learning means excludes the movement information obtained when the gaze position and the position of the subject in the captured image are not linked from the target of correct data.
13. The imaging device according to claim 3, wherein the learning means excludes the motion information obtained when the imaging device is moving and / or when the zoom ratio is changing from the target of correct answer data.
14. a detecting step of detecting an eye condition of a user; an analysis step of analyzing eye movements based on the detection result of the detection step; an estimation step of estimating the timing to start a release operation based on output data indicating whether the scene is a desired scene to be photographed, which is obtained by inputting the movement information representing the eye movement analyzed in the analysis step into a learning model; 11. A method for controlling an imaging device, comprising:
15. A program for causing a computer to function as each of the means of the imaging device according to any one of claims 1 to 13.
Citation Information
Patent Citations
Head mounted display device
JP1996160345A
Method and device for image pickup
JP2001183735A
Imaging apparatus
JP2006345276A
Control method based on voluntary eye signals, especially for imaging.
JP2010518666A
Information processing device, processing method thereof, program and imaging apparatus
JP2012226665A