Imaging device and its control method, program, and storage medium

The imaging device uses a learning model to adjust shooting parameters based on user feedback, ensuring settings align with individual user preferences, enhancing user experience.

JP2026103340APending Publication Date: 2026-06-24CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2024-12-12
Publication Date
2026-06-24

AI Technical Summary

Technical Problem

Existing imaging devices struggle to set shooting parameters to values preferred by the user, as they do not reflect individual user preferences.

Method used

An imaging device equipped with a learning model that generates shooting parameters based on user instructions and preferences, using a neural network to adjust settings dynamically.

Benefits of technology

Enables setting of shooting parameters to values preferred by the user, improving the user experience by capturing images that better match their preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026103340000001_ABST
    Figure 2026103340000001_ABST
Patent Text Reader

Abstract

The present invention provides an imaging device that allows users to set shooting parameters to their preferred values. [Solution] The system comprises an imaging unit that images a subject, a recognition unit that recognizes the shooting scene and the characteristics of the object being photographed, a generation unit that generates shooting parameters corresponding to the composite information, using a learning model that has learned the relationship between the composite information, which associates the shooting scene and the characteristics of the object being photographed, and shooting parameters that can obtain appropriate shooting results, a display unit that displays the image captured based on the generated shooting parameters, an acquisition unit that obtains instructions from the user to modify the displayed image, and a control unit that repeats a series of operations, which involves taking images based on the generated shooting parameters, displaying the captured image, obtaining instructions from the user to modify the displayed image, and generating modified shooting parameters based on the composite information and the user's modification instructions, thereby training the learning model to obtain the shooting parameters desired by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for assisting shooting in an imaging device.

Background Art

[0002] In imaging devices such as digital cameras, it is not user-friendly for users to manually set many shooting parameters such as shutter speed, aperture, and ISO sensitivity, and it is also difficult to quickly set them for moving subjects. Therefore, in recent years, most imaging devices have an auto shooting mode that automatically sets shooting parameters according to the shooting scene.

[0003] In Patent Document 1, a technique is disclosed in which a shooting scene is determined based on a live view image, and shooting parameters suitable for the shooting scene are displayed together with a setting recommendation range. According to Patent Document 1, it is possible to assist shooting by visually showing the user basic setting items suitable for the user and prompting the user to change the setting content.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the technique disclosed in Patent Document 1, the shooting parameters are displayed by the camera determining the shooting scene. However, since the shooting parameters do not reflect the user's intention, they may not be set to appropriate values preferred by the user.

[0006] The present invention has been made in view of the above problems, and an object thereof is to provide an imaging device capable of setting shooting parameters to appropriate values preferred by the user. [Means for solving the problem]

[0007] The imaging device according to the present invention is characterized by comprising: imaging means for capturing an image of a subject and acquiring an image signal; recognition means for recognizing a shooting scene and the characteristics of an object to be photographed based on the image signal; generation means for generating shooting parameters for the imaging means corresponding to the composite information, using a learning model that has learned the relationship between composite information relating the shooting scene and the characteristics of the object to be photographed and shooting parameters that can obtain an appropriate shooting result for the composite information; display means for displaying an image captured by the imaging means based on the shooting parameters generated by the generation means; acquisition means for acquiring instructions from the user to modify the image displayed on the display means; and control means for learning the learning model so that the shooting parameters desired by the user can be obtained by repeating a series of operations: capturing an image with the imaging means based on the shooting parameters generated by the generation means, displaying the captured image on the display means, acquiring instructions from the user to modify the image displayed on the display means, and generating modified shooting parameters for the imaging means with the generation means based on the composite information and the instructions from the user. [Effects of the Invention]

[0008] According to the present invention, it becomes possible to set the shooting parameters to appropriate values ​​preferred by the user. [Brief explanation of the drawing]

[0009] [Figure 1] A system configuration diagram of a digital camera, which is a first embodiment of the imaging device of the present invention. [Figure 2] External view of a digital camera. [Figure 3A] A flowchart illustrating the operation of a digital camera. [Figure 3B] A flowchart illustrating the operation of a digital camera. [Figure 4] A conceptual diagram illustrating the operation of the digital camera, as explained in Figures 3A and 3B. [Figure 5] A diagram showing a configuration example of a generation unit composed of one learning model. [Figure 6] A diagram illustrating the data array of the input data of the generation unit. [Figure 7] A diagram showing the configuration of the digital camera according to the second embodiment. [Figure 8] A diagram for explaining the configuration of the gaze detection unit. [Figure 9] A diagram for explaining the principle of the gaze detection method. [Figure 10] A schematic diagram of the eyeball image projected onto the imaging element for the eyeball. [Figure 11] A schematic flowchart of the gaze detection process. [Figure 12A] A flowchart showing the operation of the digital camera in the second embodiment. [Figure 12B] A flowchart showing the operation of the digital camera in the second embodiment. [Figure 13] A conceptual diagram showing the operation of the digital camera described in FIGS. 12A and 12B. [Figure 14] A diagram showing a configuration example of a generation unit composed of one learning model in the second embodiment. [[ID=...]]

Embodiments for Carrying Out the Invention

[0010] [[ID=...]] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential for the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant explanations are omitted. [[ID=4...]]

[0011] [[ID=...]] (First Embodiment) Hereinafter, the first embodiment of the present invention will be described with reference to FIGS. 1 to 6.

[0012] <System Configuration of Digital Camera> It should be noted that the "..." in the translation indicates that the original text has some consecutive tags that are not numbered in a sequential way in the given list. The translation is done as accurately as possible while following the rules. If there are any specific requirements or corrections regarding these unnumbered tags, please let me know.FIG. 1 is a diagram showing the system configuration of a digital camera 100, which is a first embodiment of the imaging apparatus of the present invention.

[0013] FIG. 1(a) is a diagram showing the mechanical configuration of the digital camera 100. The photographing optical system 110 includes an aperture 11, a shake correction lens group 12, and a focus lens group 13, and can guide subject light to the camera body 130. The camera body 130 includes an imaging element 21 that photoelectrically converts the optical image formed by the photographing optical system 110 and a mechanical shutter 22 that adjusts the exposure time.

[0014] The camera body 130 includes a rear liquid crystal 23 on its back surface, and a small liquid crystal 24 and an optical system 25 in the finder unit 40, and can display the image captured by the imaging element 21. If the imaging element has an electronic shutter function, the mechanical shutter is not necessary. Even when the mechanical shutter is provided, when the exposure time is adjusted by the electronic shutter, the mechanical shutter remains fully open. At the time of shooting, by gently pressing the shutter button (not shown) to the first stage, so-called "half-press" (hereinafter referred to as SW1 press), automatic focus adjustment is performed, and shooting parameters such as shutter speed and aperture value by the automatic exposure mechanism are set. Further, by deeply pressing the shutter button from the half-press to the second stage, so-called "full-press" (hereinafter referred to as SW2 press), the mechanical shutter 22 or the electronic shutter function of the imaging element 21 operates to perform imaging.

[0015] FIG. 1(b) is a diagram showing the electrical configuration of the digital camera 100. The camera body 130 includes an electric circuit 30, and a CPU 31, an image processing unit 32, a control unit 33, a generation unit 34, an audio acquisition unit 35, etc. are arranged in the electric circuit 30. In addition, a memory 36 for storing image data, programs, etc. is connected to the electric circuit 30. The aperture 11, the shake correction lens group 12, the focus lens group 13, and the mechanical shutter 22 are driven and controlled by the control unit 33 via driving means not shown respectively.

[0016] The image signal generated by photoelectric conversion in the image sensor 21 is output as digital data from the image processing unit 32 and stored on a recording medium (not shown). The image processing unit 32 also performs processing to recognize the shooting scene and the characteristics of the object being photographed from the captured image data.

[0017] The viewfinder unit 40 is further equipped with an eyepiece sensor 26, which can detect whether or not the photographer is looking through the viewfinder unit 40. The camera body 130 is equipped with an audio acquisition unit 35, which can process audio input from the microphone 27. The generation unit 34 consists of a neural network and the like, and is a processing unit that generates a new image, a corrected image, and shooting parameters for capturing that corrected image from image data based on learning. Further details will be described later.

[0018] The CPU31 is a processing unit that can electrically control all of the above elements. In Figure 1, the control signal lines are omitted, and only the flow of information between each element is shown by arrows.

[0019] Figure 2 is an external view of the digital camera 100. The same reference numerals are used for the same blocks as in Figure 1.

[0020] The user changes the shooting parameters and takes a picture using the control unit 202 or electronic dial 204, which are setting change means such as buttons attached to the imaging device. The rear LCD 23 may also have a touch panel function, allowing the user to change the shooting parameters using the rear LCD 23. The user can understand the current shooting parameter settings through the display output on the rear LCD 23 or the small LCD 24.

[0021] <Generation part 34> The shooting parameters in the Digital Camera 100 are extensive. Typical shooting parameters include ISO sensitivity, aperture value, shutter speed, exposure compensation value, white balance setting, and contrast. Of course, other parameters that affect the captured image can also be included. For example, parameters related to image processing such as saturation, sharpness, lighting correction, tone priority, shooting style, noise reduction, and color correction may also be included.

[0022] Furthermore, since users have different preferences when it comes to photography, and some may want to focus on multiple subjects or on a specific subject, information about the focus position may also be included in the shooting parameters.

[0023] The generation unit 34 is a processing unit consisting of a learning model (in this embodiment, a large-scale language model) that can present modified images and shooting parameters that match the user's preferences based on information about the shooting conditions and user instructions. The learning model is composed of, for example, a multi-layer neural network.

[0024] In this embodiment, the generation unit 34 performs inference processing using the shooting scene, the characteristics of the object to be photographed, and the user instruction information (prompt) via voice as input to generate shooting parameters. The learning method will be described later.

[0025] <Camera operation flow> Figures 3A and 3B are flowcharts showing the operation of the digital camera 100 shown in Figure 1. Each process in Figures 3A and 3B is realized when the CPU 31 executes a program stored in memory 36, sending commands to the control unit 33 and controlling each part of the device.

[0026] The process shown in Figure 3A is initiated when the user turns on the power of the digital camera 100, and in step S301, the CPU 31 performs the startup process for the digital camera 100.

[0027] In step S302, the CPU 31 stores the live view image in memory 36.

[0028] In step S303, the CPU 31 uses the image processing unit 32 to recognize the shooting scene based on the live view image stored in the memory 36.

[0029] In step S304, the CPU 31 uses the image processing unit 32 to recognize the features of the object being photographed based on the live view image stored in the memory 36.

[0030] In step S305, the CPU 31 associates the characteristics of the shooting scene and the object being photographed and inputs them as composite information to the generation unit 34.

[0031] In step S306, the CPU 31 sets the shooting parameters based on the output from the generation unit 34.

[0032] In step S307, the CPU 31 determines whether SW1 is pressed or not. If SW1 is pressed, the CPU 31 proceeds to step S308; otherwise, it returns to step S302.

[0033] In step S308, the CPU 31 confirms the shooting parameters set in step S306.

[0034] In step S309, the CPU 31 determines whether SW2 is pressed or not. If SW2 is pressed, the CPU 31 proceeds to step S310; otherwise, it returns to step S307.

[0035] In step S310, the CPU 31 performs shooting with the shooting parameter settings confirmed in step S308.

[0036] In step S311, the CPU 31 displays the captured image on the rear LCD 23 or the small LCD 24.

[0037] In step S312, the CPU 31 uses the voice acquisition unit 35 to determine whether or not there is a user voice instruction regarding the captured image input to the microphone 27. If there is a voice instruction, the CPU 31 proceeds to step S313; otherwise, it proceeds to step S317.

[0038] In step S313, the CPU 31 converts the user's voice instructions input to the voice acquisition unit 35 into a prompt. A prompt is, for example, a verbal expression of how the user wants to modify the image that has been captured and displayed.

[0039] In step S314, the CPU 31 associates a prompt with the composite information, which links the shooting scene and the characteristics of the object being photographed, and inputs it to the generation unit 34. Based on this input, the generation unit 34 generates a corrected image by modifying the captured image based on the voice instructions (prompts) from the user.

[0040] In step S315, the CPU 31 displays the corrected image based on the output from the generation unit 34.

[0041] In step S316, the CPU 31 sets the shooting parameters based on the output from the generation unit 34 so that a captured image similar to the corrected image is obtained.

[0042] In step S317, the CPU 31 determines whether the display mode has been canceled. If the display mode has been canceled, the CPU 31 proceeds to step S318; otherwise, it returns to step S312.

[0043] In step S318, the CPU 31 determines whether the power to the digital camera 100 has been turned off. If the power is off, the CPU 31 terminates the operation of this flow; otherwise, it returns to step S302.

[0044] As described above, in this embodiment, imaging is performed by the image sensor 21 based on the shooting parameters generated by the generation unit 34 having a learning model, the captured image is displayed on the small LCD 24, and user instructions for corrections to the image displayed on the small LCD 24 are obtained. Then, the generation unit 34 generates corrected shooting parameters based on composite information relating the shooting scene and the characteristics of the object to be photographed, and the user's correction instructions. This series of operations is repeated to train the learning model so that the shooting parameters desired by the user can be obtained.

[0045] Figure 4 is a conceptual diagram illustrating the camera operation described in Figures 3A and 3B. The time-series update of images and the setting of shooting parameters are explained using the conceptual diagram in Figure 4.

[0046] Figure 4 shows an example of a user photographing a soccer player. It shows the user checking the live view image, edited image, and captured image through the viewfinder unit 40, and the diagram shown along the timeline shows the content displayed on the small LCD 24 that the user is looking at through the viewfinder.

[0047] First, at time t0, the shooting scene A and the characteristics A of the object being photographed, recognized from the live view image, are input to the generation unit 34.

[0048] At time t1, the shooting parameter setting A is output from the generation unit 34 and set on the digital camera 100.

[0049] At time t2, pressing SW2 initiates shooting with shooting parameter setting A, and the captured image A is displayed.

[0050] At time t3, the user looks at the captured image A and gives a voice prompt I saying "Make it brighter!", and this user prompt I is associated with the shooting scene A and the object A being photographed and input to the generation unit 34.

[0051] At time t4, the user instruction I is reflected, and the brightened corrected image AI and the shooting parameter setting AI for brighter shooting are output from the generation unit 34. At this point, the user may give further voice instructions regarding the corrected image and input this prompt to the generation unit 34.

[0052] At time t5, the display mode is deactivated, and the Live View Image AI is displayed. The shooting scene A and the characteristics A of the object being photographed, recognized from the Live View Image AI, are input to the generation unit 34.

[0053] At time t6, the shooting parameter setting AI is output from the generation unit 34 and set to the digital camera 100.

[0054] At time t7, pressing SW2 initiates shooting using the AI-controlled shooting parameter settings intended by the user, and the captured image AI is displayed.

[0055] In actual shooting, as shown in Figure 4, the process of generating shooting parameters based on inputs of the shooting scene, the characteristics of the object being photographed, and user instructions is repeated along the time axis. As a result, each time the user gives instructions on the captured image, the system can capture images that better suit the user's preferences.

[0056] <Learning method of this embodiment> The base learning model used in the generation unit 34 is a large-scale language model (LLM) capable of inference processing, taking the shooting scene, the characteristics of the object being photographed, and user instruction information (prompts) as input data. The initial learning model is a learning model that has been repeatedly trained using the shooting scene and the characteristics of the object being photographed as input data, and shooting results (shooting parameters at that time) that are generally considered appropriate for that input data as training data. In the process of taking many shots from this initial state, the learning of the generation unit 34 is further advanced by using images (shooting parameters at that time) that have been modified by user instruction information (prompts) as training data for the shooting scene and the characteristics of the object being photographed input to the generation unit 34. In this way, a learning model that is more suited to the user's preferences can be constructed with each shot.

[0057] <Model configuration of the first embodiment> In this embodiment, the generation unit 34 generates two things: a corrected image and shooting parameters. Figure 5 shows an example of the configuration of the generation unit 34, which consists of one learning model.

[0058] The live view image output from the image processing unit 32, the shooting scene, the characteristics of the object being photographed, and the user instruction information (prompt) output from the audio acquisition unit 35 are input to the learning model (LLM), and the corrected image and shooting parameters are output. Here, a configuration is used in which both the corrected image and shooting parameters are output from a single learning model, but in order to reduce the model size and shorten the processing time, a configuration may be adopted in which the corrected image and shooting parameters are output from two separate learning models.

[0059] <Selectivity of input data for generation unit 34> Users do not necessarily speak. Therefore, the generation unit 34 is configured to generate corrected images and shooting parameters if at least the characteristics of the shooting scene and the object being photographed are provided as input data.

[0060] Figure 6 is a diagram illustrating the data array of the input data for the generation unit 34 in this embodiment.

[0061] In Figure 6, the input data consists of a header (1 or 0) indicating whether or not there is significant input information, and a payload in which the data is arranged in the following order: shooting scene, characteristics of the object being photographed, user instruction (prompt), live view image, and shooting parameters associated with the live view image. In this embodiment, each input piece of information is of a fixed length, but is not limited to that; it can be of a variable length, and the data size may also be added as a header. In this way, in this embodiment, the data is notified to the generation unit 34 as a pair of input data.

[0062] The generation unit 34 accepts the notified significant data as input data and provides a 0 as the input signal for non-significant input information. If some of the input data is set to 0 during the training of the generation unit 34's learning model, the generation unit 34 will be able to generate corrected images and shooting parameters even when no input data is available.

[0063] As described above, by inputting user instructions to modify the displayed image into the digital camera via voice or other means, and by inputting these user instructions along with the characteristics of the shooting scene and subject into a learning model, the learning model can learn to capture images that match the user's preferences. This makes it possible to set parameters to capture images that better suit the user's preferences and to provide appropriate shooting assistance functions.

[0064] (Second embodiment) The digital camera 700 of the second embodiment will be described below with reference to Figures 7 to 12.

[0065] Figure 7 shows the configuration of the digital camera 700. Here, only the differences from the block diagram of the digital camera 100 shown in Figure 1(b), which describes the first embodiment, will be explained.

[0066] In Figure 7, a gaze detection unit 732 is further provided within the viewfinder unit 40 of the camera body 730 to detect the user's gaze position relative to the small LCD 24.

[0067] Figure 8 is a diagram illustrating the configuration of the gaze detection unit 732.

[0068] The illumination light source 801 is a light source that projects infrared light onto the eyeball 804 for gaze detection, and is composed of, for example, multiple infrared light-emitting diodes. The illuminated eyeball image and the image from the corneal reflection of the light source are formed by a light-receiving lens 802 onto an image sensor 803 for the eyeball, which has a two-dimensional arrangement of photoelectric conversion elements such as CMOS.

[0069] The light-receiving lens 802 positions the pupil of the user's eyeball 804 and the eyeball image sensor 803 in a complementary imaging relationship. The direction of the line of sight is detected by a predetermined algorithm, described later, based on the positional relationship between the image of the eyeball formed on the eyeball image sensor 803 and the image due to corneal reflection from the light source 801. The illumination light source 801, the light-receiving lens 802, and the eyeball image sensor 803 constitute the line of sight detection unit 732.

[0070] The memory 36 has a function to store imaging signals from the image sensor 21 and the eyeball image sensor 803, as well as a function to store gaze correction data and eye characteristic information.

[0071] Figure 9 is a diagram illustrating the principle of the gaze detection method, and corresponds to a summary diagram of the optical system for gaze detection shown in Figure 8 above.

[0072] In Figure 9, 801a and 801b are light sources such as light-emitting diodes that emit infrared light insensitive to the user. Each light source is positioned approximately symmetrically with respect to the optical axis of the light-receiving lens 802 and illuminates the user's eyeball 901. A portion of the illumination light reflected by the eyeball 901 is focused by the light-receiving lens 802 onto the image sensor 803 for the eyeball.

[0073] Figure 10(a) is a schematic diagram of the eyeball image projected onto the eyeball image sensor 803, and Figure 10(b) is an output intensity diagram of the eyeball image sensor 803. Figure 11 is a flowchart illustrating the schematic operation of the gaze detection process.

[0074] The method for detecting eye movements will be explained below using Figures 9 to 11.

[0075] <Explanation of eye-tracking operation> In Figure 11, when the gaze detection routine is started, in step S1101, the CPU 31 uses light sources 801a and 801b to emit infrared light towards the user's eyeball 804. The image of the user's eyeball illuminated by the infrared light is formed on the eyeball image sensor 803 through the light-receiving lens 802, and the eyeball image sensor 803 performs photoelectric conversion, making it possible to process it as an electrical signal.

[0076] In step S1102, the CPU 31 acquires an eyeball image signal from the eyeball image sensor 803.

[0077] In step S1103, the CPU 31 uses the information from the eyeball image signal obtained in step S1102 to determine the coordinates of the points corresponding to the corneal reflection images Pd and Pe of the light sources 801a and 801b and the pupil center c, as shown in Figure 9. The infrared light emitted from the light sources 801a and 801b illuminates the cornea 903 of the user's eyeball 804. At this time, the corneal reflection images Pd and Pe, formed by a portion of the infrared light reflected from the surface of the cornea 903, are focused by the light-receiving lens 802 and imaged onto the eyeball image sensor 803 (points Pd' and Pe' shown in the figure). Similarly, the light beams from the ends a and b of the pupil 902 are also imaged onto the eyeball image sensor 803.

[0078] Figure 10(a) shows an example of a reflected image obtained from the ophthalmic image sensor 803, and Figure 10(b) shows an example of luminance information obtained from the ophthalmic image sensor 803 in region α of the above image example. As shown in the figure, the horizontal direction is the X-axis and the vertical direction is the Y-axis. In this case, the coordinates in the X-axis direction (horizontal direction) of the images Pd' and Pe' formed by the corneal reflected images of light sources 801a and 801b are denoted as Xd and Xe. Also, the coordinates in the X-axis direction of the images a' and b' formed by the light beams from the ends a and b of the pupil 902 are denoted as Xa and Xb.

[0079] In the example of luminance information in Figure 10(b), extremely high levels of luminance are obtained at positions Xd and Xe, which correspond to the images Pd' and Pe' formed by the corneal reflection images of light sources 801a and 801b. In the region between coordinates Xa and Xb, which corresponds to the area of ​​the pupil 902, extremely low levels of luminance are obtained, except at the positions Xd and Xe mentioned above. In contrast, in the region of the iris 1001 outside the pupil 902, which has an X-coordinate value lower than Xa and an X-coordinate value higher than Xb, intermediate values ​​between the two types of luminance levels are obtained.

[0080] From the information on the variation in brightness level relative to the above X coordinate position, the X coordinates Xd and Xe of the images Pd' and Pe' formed by the corneal reflection images of light sources 801a and 801b, and the X coordinates Xa and Xb of the images a' and b' at the pupillary tip can be obtained. Furthermore, when the rotation angle θx of the optical axis of the eyeball 901 with respect to the optical axis of the light-receiving lens 802 is small, the coordinate Xc of the point corresponding to the pupillary center c (let's call it c') that is imaged on the eyeball image sensor 803 can be expressed as Xc ≈ (Xa + Xb) / 2. From the above, the X coordinate of c' corresponding to the pupillary center that is imaged on the eyeball image sensor 803, and the coordinates of the corneal reflection images Pd' and Pe' of light sources 801a and 801b can be estimated.

[0081] Returning to the explanation of Figure 11, in step S1104, the CPU 31 calculates the imaging magnification β of the eyeball image. β is a magnification determined by the position of the eyeball 901 relative to the light-receiving lens 802, and can essentially be obtained as a function of the interval (Xd-Xe) between the corneal reflection images Pd',Pe'.

[0082] In step S1105, the CPU 31 calculates the rotation angle of the eyeball 901. Since the X-coordinate of the midpoint of the corneal reflection images Pd' and Pe' is approximately the same as the X-coordinate of the center of curvature O of the cornea 903, if we let Oc be the standard distance between the center of curvature O of the cornea 903 and the center c of the pupil 902, then the rotation angle θX of the optical axis of the eyeball 901 in the ZX plane is: β*Oc*SINθX≒{(Xd+Xe) / 2}-Xc This can be determined from the given relationship. Furthermore, while Figures 9 and 10 show examples of calculating the rotation angle θX when the user's eyeball rotates in a plane perpendicular to the Y-axis, the method for calculating the rotation angle θy when the user's eyeball rotates in a plane perpendicular to the X-axis is similar.

[0083] In step S1105, when the rotation angles θx and θy of the optical axis of the user's eyeball 901 are calculated, in steps S1106 and S1107, the CPU 31 determines the position of the user's line of sight. Specifically, using θx and θy, the position of the user's line of sight (point of fixation) is determined on the small LCD 24. Assuming that the point of fixation is the coordinate (Hx, Hy) corresponding to the center c of the pupil 902 on the small LCD 24, Hx = m × (Ax × θx + Bx) Hy = m × (Ay × θy + By) The following can be calculated. The coefficient m is a constant determined by the optical system configuration, and is a conversion coefficient that converts the rotation angles θx and θy into position coordinates corresponding to the center c of the pupil 902 on the small LCD 24, and is assumed to be predetermined and stored in memory 36. In addition, Ax, Bx, Ay, and By are gaze correction coefficients that correct for individual differences in the user's gaze, and are obtained by performing a calibration operation and are assumed to be stored in memory 36 before the gaze detection routine is started.

[0084] After calculating the coordinates (Hx, Hy) of the center c of the pupil 902 on the small LCD 24 as described above, the coordinates are stored in the memory 36 in step S1108, and the gaze detection routine is terminated.

[0085] The above describes a method for acquiring the coordinates of the point of gaze on a small liquid crystal 24 using corneal reflection images from light sources 801a and 801b. However, this is not the only method; any method that can acquire the eyeball rotation angle from the captured eyeball image can be applied to this embodiment.

[0086] <Camera operation flow> Next, with reference to Figures 12A and 12B, the operation of the digital camera 700 in the second embodiment will be described. Here, steps that perform the same operations as the digital camera 100 shown in Figure 3, which was described in the first embodiment, will be numbered the same way, and only the parts that differ from Figure 3 will be described.

[0087] In step S1213, the CPU 31 obtains the user's gaze position on the small LCD 24 from the gaze detection unit 732. In step S1214, the CPU 31 inputs composite information, which associates the shooting scene and the characteristics of the object being shot, with the user's gaze position and prompts (user instructions), into the generation unit 34.

[0088] Figure 13 is a conceptual diagram illustrating the camera operation described in Figures 12A and 12B. The time-series update of images and the setting of shooting parameters are explained using the conceptual diagram in Figure 13.

[0089] Figure 13 shows an example of a user taking a photo of two children running towards the finish line at a school sports day. It shows the user checking the live view image, edited image, and captured image through the viewfinder unit 40, and the diagram shown along the timeline shows the content displayed on the small LCD 24 that the user is looking at through the viewfinder.

[0090] First, the shooting scene B and the characteristics B of the object being photographed, recognized from the live view image at time t0, are input to the generation unit 34.

[0091] At time t1, the generation unit 34 outputs shooting parameter setting B such that the player holding the ball is the main subject, and this is set on the digital camera 700.

[0092] When SW2 is pressed in Jikoku t2, shooting is performed with shooting parameter setting B, and the captured image B is displayed on the small LCD 24.

[0093] At time t3, when the user views the captured image B, the user's gaze position on the small LCD 24 is calculated. At the same time, the user looks at the captured image and gives a voice prompt, "Make it stand out!", and the user's gaze position and the user prompt II are associated with the captured scene B and the object B being captured and input to the generation unit 34.

[0094] At time t4, the user instruction II is reflected, and the generation unit 34 outputs a corrected image BII, which has been modified to make the child at the gaze position more prominent, and a shooting parameter setting BII for taking a picture in a way that makes the child at the gaze position more prominent.

[0095] At time t5, the display mode is deactivated, and the live view image BII is displayed. The shooting scene B and the characteristics B of the object being photographed, recognized from the live view image BII, are input to the generation unit 34.

[0096] At time t6, the shooting parameter setting BII is output from the generation unit 34 and set to the digital camera 700.

[0097] At time t7, pressing SW2 initiates a capture using the user's intended shooting parameter settings BII, and the captured image BII is displayed.

[0098] While voice input alone doesn't specify which area of ​​the captured image the user is directing to, as described above, inputting the user's gaze position to the generation unit 34 allows the camera to reflect the user's intentions. This enables the camera to capture images with more appropriate shooting parameters for the user.

[0099] <Learning model configuration of the second embodiment> In the second embodiment as well, the generation unit 34 generates two things: a corrected image and shooting parameters. Figure 14 shows an example of the configuration of the generation unit 34, which consists of one learning model.

[0100] In the configuration shown in Figure 5 of the first embodiment, information on the user's gaze position on the small LCD 24 from the gaze detection unit 732 is further input to the learning model. The user's gaze position is input to the learning model in association with the shooting scene, the characteristics of the object being shot, and user instructions (prompts), and a corrected image and shooting parameters are output.

[0101] In the first embodiment, it may be unclear which area of ​​the live view image a user instruction (prompt) refers to. In contrast, in this embodiment, information about the user's gaze position on the small LCD 24 is added, so that the user's intentions are better reflected in the corrected image and shooting parameters.

[0102] In the first and second embodiments, examples were described in which a large-scale language model is used as the learning model, but other AI learning models may also be used.

[0103] (Other embodiments) Furthermore, the present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by a process in which one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0104] The disclosures herein include the following imaging devices and their control methods, programs, and storage media.

[0105] (Item 1) An imaging means for capturing an image of a subject and acquiring an image signal, A recognition means that recognizes the shooting scene and the characteristics of the object being photographed based on the aforementioned image signal, A generation means that generates shooting parameters for the imaging means corresponding to the composite information, using a learning model that has learned the relationship between composite information relating the shooting scene and the characteristics of the object to be photographed, and shooting parameters that can obtain appropriate shooting results for the composite information. A display means for displaying an image captured by the imaging means based on the shooting parameters generated by the generation means, An acquisition means for acquiring instructions from the user to modify an image displayed on the display means, A control means that performs a series of operations, including taking images with the imaging means based on the shooting parameters generated by the generation means, displaying the captured images on the display means, obtaining correction instructions from the user for the images displayed on the display means, and repeating the operations, including generating corrected shooting parameters for the imaging means with the generation means based on the composite information and the correction instructions from the user, in order to learn the learning model so that the shooting parameters desired by the user can be obtained. An imaging device characterized by comprising:

[0106] (Item 2) The imaging device according to item 1, characterized in that the learning model takes the composite information as input and learns shooting parameters that yield appropriate shooting results for the composite information as training data.

[0107] (Item 3) The imaging device according to item 1 or 2, characterized in that the learning model is further trained using the modified shooting parameters as training data.

[0108] (Item 4) The imaging device according to any one of items 1 to 3, characterized in that the learning model is a large-scale language model.

[0109] (Item 5) The imaging device according to any one of items 1 to 4, characterized in that the user's instructions for correction are given by the user's voice.

[0110] (Item 6) The imaging device according to any one of items 1 to 5, further comprising a detection means for detecting the user's gaze position on the image displayed on the display means.

[0111] (Item 7) The imaging device according to item 6, characterized in that the learning model learns the relationship between the composite information and the shooting parameters based on the information from the detection means.

[0112] (Item 8) The imaging apparatus according to any one of items 1 to 7, characterized in that the generation means further generates a modified image obtained by modifying the image displayed on the display means based on a modification instruction from the user.

[0113] (Item 9) The imaging device according to item 8, characterized in that the display means further displays the modified image.

[0114] (Item 10) A method for controlling an imaging device equipped with imaging means for capturing an image of a subject and acquiring an image signal, A recognition process that recognizes the shooting scene and the characteristics of the object being photographed based on the aforementioned image signal, A generation step of generating shooting parameters for the imaging means corresponding to the composite information, using a learning model that has learned the relationship between composite information relating the shooting scene and the characteristics of the object to be photographed, and shooting parameters that can obtain appropriate shooting results for the composite information. A display step, in which an image captured by the imaging means is displayed based on the shooting parameters generated in the generation step, The acquisition step involves obtaining user instructions for modifications to the image displayed in the aforementioned display step, A control step which involves repeating a series of operations, including: taking images with the imaging means based on the shooting parameters generated in the generation step; displaying the captured images in the display step; obtaining correction instructions from the user for the images displayed in the display step; and generating corrected shooting parameters for the imaging means in the generation step based on the composite information and the correction instructions from the user, in order to train the learning model so that the shooting parameters desired by the user can be obtained. A control method for an imaging device, characterized by having the following features.

[0115] (Item 11) A program for causing a computer to function as one of the means of an imaging apparatus described in any one of items 1 to 9.

[0116] (Item 12) A computer-readable storage medium storing a program for causing a computer to function as one of the means of an imaging device described in any one of items 1 to 9.

[0117] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]

[0118] 11: Aperture, 12: Image stabilization lens group, 13: Focus lens group, 21: Image sensor, 22: Mechanical shutter, 23: Rear LCD, 24: Small LCD, 30: Electrical circuit, 31: CPU, 32: Image processing unit, 33: Control unit, 34: Generation unit, 35: Audio acquisition unit, 36: Memory, 100: Digital camera, 110: Shooting optical system, 130: Camera body

Claims

1. An imaging means for capturing an image of a subject and acquiring an image signal, A recognition means that recognizes the shooting scene and the characteristics of the object being photographed based on the aforementioned image signal, A generation means generates shooting parameters for the imaging means corresponding to the composite information, using a learning model that has learned the relationship between composite information relating the shooting scene and the characteristics of the object to be photographed, and shooting parameters that yield appropriate shooting results for the composite information. A display means for displaying an image captured by the imaging means based on the shooting parameters generated by the generation means, An acquisition means for acquiring instructions from the user to modify an image displayed on the display means, A control means that performs a series of operations, including taking images with the imaging means based on the shooting parameters generated by the generation means, displaying the captured images on the display means, obtaining correction instructions from the user for the images displayed on the display means, and repeating the operations, including generating corrected shooting parameters for the imaging means with the generation means based on the composite information and the correction instructions from the user, in order to learn the learning model so that the shooting parameters desired by the user can be obtained. An imaging device characterized by comprising:

2. The imaging apparatus according to claim 1, characterized in that the learning model takes the composite information as input and learns shooting parameters that yield appropriate shooting results for the composite information as training data.

3. The imaging apparatus according to claim 1, characterized in that the learning model is further trained using the modified shooting parameters as training data.

4. The imaging apparatus according to claim 1, characterized in that the learning model is a large-scale language model.

5. The imaging apparatus according to claim 1, characterized in that the user's instruction for correction is given by the user's voice.

6. The imaging apparatus according to claim 1, further comprising a detection means for detecting the user's gaze position relative to the image displayed on the display means.

7. The imaging apparatus according to claim 6, characterized in that the learning model learns the relationship between the composite information and the shooting parameters based on the information from the detection means.

8. The imaging apparatus according to claim 1, characterized in that the generation means further generates a modified image obtained by modifying the image displayed on the display means based on a modification instruction from the user.

9. The imaging apparatus according to claim 8, wherein the display means further displays the modified image.

10. A method for controlling an imaging device equipped with imaging means for capturing an image of a subject and acquiring an image signal, A recognition process that recognizes the shooting scene and the characteristics of the object being photographed based on the aforementioned image signal, A generation step of generating shooting parameters for the imaging means corresponding to the composite information, using a learning model that has learned the relationship between composite information relating the shooting scene and the characteristics of the object to be photographed, and shooting parameters that can obtain appropriate shooting results for the composite information. A display step, in which an image captured by the imaging means is displayed based on the shooting parameters generated in the generation step, The acquisition step involves obtaining user instructions for modifications to the image displayed in the aforementioned display step, A control step which involves repeating a series of operations, including: taking images with the imaging means based on the shooting parameters generated in the generation step; displaying the captured images in the display step; obtaining correction instructions from the user for the images displayed in the display step; and generating corrected shooting parameters for the imaging means in the generation step based on the composite information and the correction instructions from the user, in order to train the learning model so that the shooting parameters desired by the user can be obtained. A control method for an imaging device, characterized by having the following features.

11. A program for causing a computer to function as one of the means of an imaging apparatus according to any one of claims 1 to 9.

12. A computer-readable storage medium storing a program for causing a computer to function as one of the means of an imaging apparatus according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Imaging apparatus, learning device, imaging method, learning method, image data creation method, and program

    JP2019213130A