Imaging apparatus, image processing device, and method

JP2023065313A5Pending Publication Date: 2025-10-17CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022165023
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-27
Filing Date
2022-10-13
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing imaging devices struggle to quickly align a user's line of sight with an intended position in a displayed image, as they do not efficiently emphasize the characteristic region based on device settings.

Method used

An imaging apparatus and method that detects a user's gaze position and generates image data to visually emphasize a characteristic region determined by device settings, using processing techniques such as edge enhancement, luminance reduction, or pseudo-color conversion to highlight the intended subject.

Benefits of technology

This approach assists users in quickly gazing at the intended position or object by visually distinguishing it from other regions, thereby reducing the time required to align their line of sight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0001_ABST
    Figure 00000000_0001_ABST
  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an imaging apparatus and a method for generating image data for displaying to assist a user to quickly gaze at an intended location or subject.SOLUTION: An imaging apparatus is capable of detecting a user's gazing location in an image being displayed. The imaging apparatus applies a processing process that visually emphasizes a feature area over other areas when generating image data to be displayed when the detection of the gazing location is valid. The feature area is an area of a subject of a type determined on the basis of the setting of the imaging apparatus.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an imaging device, an image processing device, and a method.

Background Art

[0002] Patent Document 1 discloses an imaging device that detects a user's gaze position in a display image and enlarges and displays a region including the gaze position.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to the technique described in Patent Document 1, it becomes easier for the user to confirm whether or not they are gazing at the intended position in the display image. However, it is not possible to shorten the time required for the user to direct their line of sight to the intended position (subject) in the display image.

[0005] The present invention has been made in view of such problems of the prior art. In one aspect of the present invention, there is provided an imaging device and method for generating display image data for assisting a user to quickly gaze at an intended position or subject.

Means for Solving the Problems

[0006] The above objective is achieved by an imaging device having a detection means capable of detecting the user's gaze position in an image displayed by the imaging device, and a generation means for generating image data for display, wherein the generation means applies a processing treatment to the image data generated when the detection means is effective, which visually emphasizes a feature region compared to other regions, and the feature region is the region of a subject of a type determined based on the settings of the imaging device. [Effects of the Invention]

[0007] According to one aspect of the present invention, an imaging apparatus and method can be provided that generates display image data to assist a user in quickly focusing on an intended location or subject. [Brief explanation of the drawing]

[0008] [Figure 1] Block diagram showing the configuration of an imaging device according to an embodiment of the present invention. [Figure 2] This figure shows the correspondence between the pupil surface of a pixel and the photoelectric conversion unit of an imaging device according to an embodiment of the present invention. [Figure 3] This figure shows the configuration of the eye-tracking input unit according to an embodiment of the present invention. [Figure 4] Flowchart of the first embodiment of the present invention [Figure 5] Method for setting the shooting mode according to an embodiment of the present invention [Figure 6] Image processing example 1 of the first embodiment of the present invention [Figure 7] Image processing example 2 of the first embodiment of the present invention [Figure 8] Imaging device configuration of the second embodiment of the present invention [Figure 9] Flowchart of a second embodiment of the present invention [Figure 10] Image processing example 3 of the second embodiment of the present invention [Figure 11] Image processing example 4 of the second embodiment of the present invention [Figure 12] This figure shows an example of a calibration screen presented by the imaging device according to the third embodiment. [Figure 13] Figure showing an example of a scene suitable for performing processing according to visual characteristics and an example of the processing [Figure 14] Figure showing an example of a scene suitable for performing processing according to visual characteristics and an example of the processing [Figure 15] Figure showing an example of a scene suitable for performing processing according to visual characteristics and an example of the processing [Figure 16] Figure showing an example of a scene suitable for performing processing according to visual characteristics and an example of the processing [Figure 17] Flowchart relating to the operation of generating display image data in the third embodiment [Figure 18] Figure showing an example of the virtual space presented in the fourth embodiment [Figure 19] Figure showing an example of highlighting in the fourth embodiment [Figure 20] Figure showing an example of the relationship between the type of virtual space and the type of subject that can be highlighted in the fourth embodiment [Figure 21] Figure showing an example of the GUI for selecting the main subject type in the fourth embodiment [Figure 22] Figure showing an example of using metadata as image data in the fourth embodiment [Figure 23] Figure showing an example of an indicator to be displayed together with the virtual space image in the fourth embodiment [Figure 24] Figure showing an example of the display system according to the fifth embodiment and its operation [Figure 25] Block diagram showing an example of the functional configuration of a computer device that can be used as a server in the fifth embodiment [Figure 26] Figure showing an example of the display area in the fifth embodiment [Figure 27] Figure showing another example of the display system according to the fifth embodiment and its operation [Figure 28] Figure showing an example of the configuration of a camera used in the fifth embodiment

MODE FOR CARRYING OUT THE INVENTION

[0009] The present invention will be described in detail below with reference to the attached drawings, based on exemplary embodiments thereof. Note that the following embodiments do not limit the invention to the claims. Furthermore, while multiple features are described in the embodiments, not all of them are essential to the invention, and the multiple features may be combined arbitrarily. In addition, in the attached drawings, the same or similar configurations are given the same reference numeral, and redundant descriptions are omitted.

[0010] The following description will focus on the implementation of the present invention using an imaging device such as a digital camera. However, the present invention can be implemented using any electronic device capable of detecting the gaze position on a display screen. Such electronic devices include, in addition to imaging devices, computer equipment (personal computers, tablet computers, media players, PDAs, etc.), mobile phones, smartphones, game consoles, robots, and in-vehicle equipment. These are examples, and the present invention can be implemented using other electronic devices as well.

[0011] ●(First Embodiment) [Description of the imaging device configuration] Figure 1 is a block diagram showing an example of the functional configuration of an imaging device 1, which is an example of an image processing apparatus according to an embodiment. The imaging device 1 has a main body 100 and a lens unit 150. Here, the lens unit 150 is a replaceable lens unit that can be attached to and removed from the main body 100, but it may also be a lens unit integrated with the main body 100.

[0012] The lens unit 150 and the main body 100 are mechanically and electrically connected via a lens mount. Communication terminals 6 and 10 provided on the lens mount are contacts that electrically connect the lens unit 150 and the main body 100. The lens unit control circuit 4 and the system control circuit 50 can communicate through communication terminals 6 and 10. Power necessary for the operation of the lens unit 150 is also supplied from the main body 100 to the lens unit 150 via communication terminals 6 and 10.

[0013] The lens unit 150 constitutes a photographic optical system that forms an optical image of the subject on the imaging surface of the imaging unit 22. The lens unit 150 has an aperture 102 and a plurality of lenses 103, including a focusing lens. The aperture 102 is driven by the aperture drive circuit 2, and the focusing lens is driven by the AF drive circuit 3. The operation of the aperture drive circuit 2 and the AF drive circuit 3 is controlled by the lens system control circuit 4 according to instructions from the system control circuit 50.

[0014] The focal-plane shutter 101 (hereinafter simply referred to as shutter 101) is driven by the control of the system control circuit 50. When taking still images, the system control circuit 50 controls the operation of the shutter 101 to expose the imaging unit 22 according to the shooting conditions.

[0015] The imaging unit 22 is an image sensor having multiple pixels arranged in a two-dimensional array. The imaging unit 22 converts the optical image formed on the imaging surface into a group of pixel signals (analog image signals) by the photoelectric conversion unit of each pixel. The imaging unit 22 may be, for example, a CCD image sensor or a CMOS image sensor.

[0016] The imaging unit 22 of this embodiment is capable of generating a pair of image signals used for phase-difference detection autofocus (hereinafter referred to as phase-difference AF). Figure 2 shows the correspondence between the pupil plane of the lens unit 150 and the photoelectric conversion units of the pixels in the imaging unit 22. Figure 2(a) shows an example of a configuration in which the pixels have multiple (in this case, two) photoelectric conversion units 201a, 201b, and Figure 2(b) shows an example of a configuration in which the pixels have one photoelectric conversion unit 201.

[0017] Each pixel is equipped with one microlens 251 and one color filter 252. The color of the color filter 252 differs for each pixel, and the colors are arranged in a predetermined pattern. Here, as an example, we assume that the color filters 252 are arranged according to a primary color Bayer pattern. In this case, the color of the color filter 252 that each pixel possesses is either red (R), green (G), or blue (B).

[0018] In the configuration shown in Figure 2(a), light input to the pixel from region 253a of the pupil plane 253 is incident on the photoelectric conversion unit 201a, and light input to the pixel from region 253b is incident on the photoelectric conversion unit 201b. Phase-difference autofocus can be performed by using the signal group obtained from the photoelectric conversion unit 201a and the signal group obtained from the photoelectric conversion unit 201b as a pair of image signals for multiple pixels.

[0019] When the signals obtained by the photoelectric conversion units 201a and 201b are handled individually, each signal functions as a focus detection signal. On the other hand, when the signals obtained by the photoelectric conversion units 201a and 201b of the same pixel are handled together (added), the countable signal functions as a pixel signal. Therefore, a pixel having the configuration shown in Figure 2(a) functions as both a focus detection pixel and a capture pixel. The imaging unit 22 is assumed to have all pixels having the configuration shown in Figure 2(a).

[0020] On the other hand, Figure 2(b) shows an example configuration of a dedicated focus detection pixel. In the pixel shown in Figure 2(b), a light-shielding mask 254 is provided between the color filter 252 and the photoelectric conversion unit 201 to restrict the light incident on the photoelectric conversion unit 201. Here, the light-shielding mask 254 has an opening that allows only light from the pupil surface 253 region 253b to be incident on the photoelectric conversion unit 201. As a result, the pixel becomes substantially the same as the state in Figure 2(a) which has only the photoelectric conversion unit 201b. Similarly, by configuring the opening of the light-shielding mask 254 so that only light from the pupil surface 253 region 253a is incident on the photoelectric conversion unit 201, the pixel can be made substantially the same as the state in Figure 2(a) which has only the photoelectric conversion unit 201a. Even if multiple pairs of these two types of pixels are arranged in the imaging unit 22, signal pairs for phase-difference AF can be generated.

[0021] Alternatively, contrast-detection autofocus (hereinafter referred to as contrast AF) may be implemented instead of, or in combination with, phase-detection AF. When only contrast AF is implemented, the pixels can be configured as shown in Figure 2(b) without the light-shielding mask 254.

[0022] The A / D converter 23 converts the analog image signal output from the imaging unit 22 into a digital image signal. If the imaging unit 22 is capable of outputting a digital image signal, the A / D converter 23 can be omitted.

[0023] The image processing unit 24 applies predetermined image processing to the digital image signal from the A / D converter 23 or the memory control unit 15 to generate signals and image data according to the application, and to acquire and / or generate various types of information. The image processing unit 24 may be a dedicated hardware circuit such as an ASIC designed to realize a specific function, or it may be a configuration in which a programmable processor such as a DSP executes software to realize a specific function.

[0024] Here, the image processing applied by the image processing unit 24 includes preprocessing, color interpolation, correction, detection, data processing, evaluation value calculation, and special effects processing. Preprocessing includes signal amplification, reference level adjustment, and defective pixel correction. Color interpolation is the process of interpolating the values ​​of color components that cannot be obtained at the time of shooting, and is also called demosaicing or simulcasting. Correction processing includes white balance adjustment, gradation correction (gamma processing), processing to correct the effects of optical aberrations and vignetting of the lens 103, and color correction. Detection processing includes detection of feature areas (e.g., face areas and human body areas) and their movement, and person recognition processing. Data processing includes synthesis processing, scaling processing, encoding and decoding processing, and header information generation processing. Evaluation value calculation processing includes the generation of signals and evaluation values ​​used for autofocus detection (AF), and the calculation of evaluation values ​​used for automatic exposure control (AE). Special effects processing includes blurring, color tone changes, relighting, and processing applied when the gaze position detection described later is enabled. These are examples of image processing that the image processing unit 24 can apply, and do not limit the image processing that the image processing unit 24 can apply.

[0025] A specific example of the feature region detection process is described below. The image processing unit 24 applies horizontal and vertical bandpass filters to the image data to be detected (for example, data from a live view image) to extract edge components. Then, the image processing unit 24 applies a matching process to the edge components using a pre-prepared template according to the type of feature region to be detected, and detects image regions similar to the template. For example, when detecting a human face region as a feature region, the image processing unit 24 applies the matching process using templates of facial features (for example, eyes, nose, mouth, and ears).

[0026] Through matching processing, candidate regions for eyes, nose, mouth, and ears are detected. The image processing unit 24 narrows down the candidate eye region to those that satisfy pre-set conditions (e.g., distance between two eyes, tilt, etc.) with other eye candidates. The image processing unit 24 then associates other parts (nose, mouth, ears) that satisfy the positional relationship with the narrowed-down candidate eye region. Furthermore, the image processing unit 24 detects face regions by applying a pre-set non-face condition filter and excluding combinations of parts that do not constitute a face. The image processing unit 24 outputs the total number of detected face regions and information about each face region (position, size, detection confidence, etc.) to the system control circuit 50. The system control circuit 50 stores the feature region information obtained from the image processing unit 24 in the system memory 52.

[0027] The human face region detection method described here is illustrative, and any other known method, such as machine learning, can be used. Furthermore, the detection method may not be limited to human faces, but may also detect other types of feature regions, such as human torsos, limbs, animal faces, landmarks, characters, automobiles, airplanes, and railway vehicles.

[0028] The detected feature regions can be used, for example, to set the focus detection region. For instance, a primary face region can be determined from the detected face regions, and the focus detection region can be set to this primary face region. This allows autofocus (AF) to focus on face regions within the shooting range. The primary face region may also be selected by the user.

[0029] The output data from the A / D converter 23 is stored in the memory 32 via the image processing unit 24 and the memory control unit 15, or via the memory control unit 15 alone. The memory 32 is used as a buffer memory for still image data and video data, as working memory for the image processing unit 24, as video memory for the display unit 28, and so on.

[0030] The D / A converter 19 converts the image data for display stored in the video memory area of ​​the memory 32 into an analog signal and supplies it to the display unit 28. The display unit 28 displays the image on a display device such as a liquid crystal display according to the analog signal from the D / A converter 19.

[0031] By continuously generating and displaying image data while shooting video, the display unit 28 can function as an electronic viewfinder (EVF). The image displayed to enable the display unit 28 to function as an EVF is called a through image or live view image. The display unit 28 may be located inside the main body 100 so as to be observed through the eyepiece, or it may be located on the surface of the housing of the main body 100 (for example, the back), or it may be provided in both locations.

[0032] In this embodiment, the display unit 28 is assumed to be located at least inside the main body 100 in order to detect the user's gaze position.

[0033] The non-volatile memory 56 is electrically rewritable, such as an EEPROM. The non-volatile memory 56 stores programs that the system control circuit 50 can execute, various settings, GUI data, and the like.

[0034] The system control circuit 50 has one or more processors (also called CPUs, MPUs, etc.) capable of executing programs. The system control circuit 50 realizes the functions of the imaging device 1 by loading programs recorded in the non-volatile memory 56 into the system memory 52 and executing them using the processors.

[0035] System memory 52 is used to store programs executed by the system control circuit 50, as well as constants, variables, and other data used during program execution. The system timer 53 measures the time used for various controls and the time of the built-in clock.

[0036] The power switch 72 is an operating component that switches the power of the imaging device 1 ON and OFF. The mode selector switch 60, the first shutter switch 62, the second shutter switch 64, and the operation unit 70 are operating components for inputting instructions to the system control circuit 50.

[0037] The mode switch 60 switches the operating mode of the system control circuit 50 to one of the following: still image recording mode, video recording mode, playback mode, etc. Modes included in the still image recording mode include auto shooting mode, auto scene detection mode, manual mode, aperture priority mode (Av mode), and shutter speed priority mode (Tv mode). There are also various scene modes, program AE mode, custom mode, etc., which are shooting settings for different shooting scenes. The mode switch 60 allows direct switching to any of these modes included in the menu buttons. Alternatively, the mode switch 60 can be used to switch to the menu buttons first, and then another operating element can be used to switch to any of these modes included in the menu buttons. Similarly, the video recording mode may also include multiple modes.

[0038] The first shutter switch 62 turns ON when the shutter button 61 is half-pressed, generating the first shutter switch signal SW1. The system control circuit 50 recognizes the first shutter switch signal SW1 as an instruction to prepare for still image shooting and starts the shooting preparation operation. The shooting preparation operation includes, for example, AF processing, automatic exposure control (AE) processing, auto white balance (AWB) processing, and EF (flash pre-flash) processing, but these are not mandatory, and other processing may also be included.

[0039] The second shutter switch 64 is turned ON when the shutter button 61 is fully pressed, generating the second shutter switch signal SW2. The system control circuit 50 recognizes the second shutter switch signal SW2 as an instruction to take a still image and executes the shooting process and recording process.

[0040] The operation unit 70 is a general term for all operating components other than the shutter button 61, mode selector switch 60, and power switch 72. The operation unit 70 includes, for example, directional keys, a set (execute) button, a menu button, and a video recording button. If the display unit 28 is a touch display, software keys that are realized through display and touch operation also constitute the operation unit 70. When the menu button is operated, the system control circuit 50 displays a menu screen on the display unit 28 that can be operated using the directional keys and the set button. The user can change the settings of the imaging device 1 through the operation of software keys and the menu screen.

[0041] Figure 3(a) is a schematic side view showing an example configuration of the eye-tracking input unit 701. The eye-tracking input unit 701 is a unit that acquires an image (eye-tracking detection image) for detecting the rotation angle of the optical axis of the user's eyeball 501a as they look through the eyepiece at the display unit 28 located inside the main unit 100.

[0042] The image processing unit 24 processes the image for gaze detection and detects the rotation angle of the optical axis of the eyeball 501a. Since the rotation angle represents the direction of the gaze, the gaze position on the display unit 28 can be estimated based on the rotation angle and a preset distance from the eyeball 501a to the display unit 28. In addition, user-specific information obtained by a pre-performed calibration operation may be taken into consideration when estimating the gaze position. The gaze position estimation may be performed by either the image processing unit 24 or the system control circuit 50. The gaze input unit 701 and the image processing unit 24 (or system control circuit 50) constitute a detection means capable of detecting the user's gaze position in the image displayed on the display unit 28 by the imaging device 1.

[0043] The image displayed on the display unit 28 is visible to the user through the eyepiece 701d and the dichroic mirror 701c. The illumination light source 701e emits infrared light outwards from the housing through the eyepiece. The infrared light reflected by the eyeball 501a is incident on the dichroic mirror 701c. The dichroic mirror 701c reflects the incident infrared light upwards. A light-receiving lens 701b and an image sensor 701a are positioned above the dichroic mirror 701c. The image sensor 701a captures the image of the infrared light formed by the light-receiving lens 701b. The image sensor 701a may be a monochrome image sensor.

[0044] The image sensor 701a outputs the analog image signal obtained by imaging to the A / D converter 23. The A / D converter 23 outputs the obtained digital image signal to the image processing unit 24. The image processing unit 24 detects the eyeball image from the image data and further detects the pupil region within the eyeball image. The image processing unit 24 calculates the rotation angle of the eyeball (direction of gaze) from the position of the pupil region in the eyeball image. The process of detecting the direction of gaze from an image including the eyeball image can be carried out by known methods.

[0045] Figure 3(b) is a schematic side view showing an example of the configuration of the gaze input unit 701 when the display unit 28 is located on the back of the imaging device 1. In this case as well, infrared light is shone in the direction in which the user's face 500, which is observing the display unit 28, is likely to be located. Then, an infrared image of the user's face 500 is acquired by taking a picture with the camera 701f located on the back of the imaging device 1, and the gaze direction is detected by detecting the pupil region from the images of the eyeballs 501a and / or 501b.

[0046] Furthermore, as long as the gaze position on the display unit 28 can ultimately be detected, there are no particular restrictions on the configuration of the gaze input unit 701 and the processing of the image processing unit 24 (or system control circuit 50), and any other configuration and processing can be adopted.

[0047] Returning to Figure 1, the power control unit 80 is composed of a battery detection circuit, a DC-DC converter, a switch circuit for switching which blocks are energized, and the like. If the power supply unit 30 is a battery, the power control unit 80 detects whether it is installed, its type, and its remaining charge. The power control unit 80 also controls the DC-DC converter based on these detection results and instructions from the system control circuit 50, supplying the necessary voltage to each part, including the recording medium 200, for the required period of time.

[0048] The power supply unit 30 can supply one or more primary batteries such as alkaline batteries or lithium batteries, secondary batteries such as NiCd batteries, NiMH batteries, or Li batteries, and / or an AC adapter.

[0049] The recording medium I / F18 is an interface to a recording medium 200, such as a memory card or hard disk. The recording medium 200 may or may not be removable. The recording medium 200 is the destination for recording image data obtained through shooting.

[0050] The communication unit 54 transmits and receives image signals and audio signals to and from external devices connected wirelessly or via wired connections. The communication unit 54 supports one or more communication standards, such as wireless LAN (Local Area Network) and USB (Universal Serial Bus). The system control circuit 50 can transmit image data (including through images) obtained by the imaging unit 22 and image data recorded on the recording medium 200 to external devices via the communication unit 54. The system control circuit 50 can also receive image data and other various information from external devices via the communication unit 54.

[0051] The attitude detection unit 55 detects the orientation of the imaging device 1 relative to the direction of gravity. Based on the orientation detected by the attitude detection unit 55, it is possible to determine whether the imaging device 1 was oriented horizontally or vertically at the time of shooting. The system control circuit 50 can add the orientation of the imaging device 1 at the time of shooting to the image data file, or align the orientation of the images before recording. An acceleration sensor or a gyroscope sensor can be used as the attitude detection unit 55.

[0052] [Gaze position detection operation] Figure 4 is a flowchart relating to the gaze position detection operation of the imaging device 1. The gaze position detection operation is performed when the gaze detection function is enabled. Furthermore, the gaze position detection operation can be performed in parallel with the live view display operation.

[0053] In S2, the system control circuit 50 acquires the currently set shooting mode. The shooting mode can be set using the mode switch 60. If the scene selection mode is set using the mode switch 60, the scene type set within the scene selection mode is also treated as a shooting mode.

[0054] Figure 5 shows an example of the external appearance of the imaging device 1. Figure 5(a) shows an example of the arrangement of the mode selector switch 60. Figure 5(b) is a top view of the mode selector switch 60 and shows examples of selectable shooting modes. For example, Tv is shutter speed priority mode, Av is aperture priority mode, M is manual setting mode, P is program mode, and SCN is scene selection mode. The desired shooting mode can be set by rotating the mode selector switch 60 so that the letter indicating the desired shooting mode is in the position of mark 63. Figure 5(b) shows the state when scene selection mode is set.

[0055] Scene selection mode is a shooting mode for capturing specific scenes or subjects. Therefore, in scene selection mode, the type of scene or subject must be set. The system control circuit 50 sets the shooting conditions (shutter speed, aperture value, sensitivity, etc.) and AF mode appropriate for the set scene or subject type.

[0056] In this embodiment, the scene and subject type in scene selection mode can be set by operating the menu screen displayed on the display unit 28, as shown in Figure 5(c). Here, as an example, one of portrait, landscape, kids, or sports can be set, but there may be more options. As described above, in scene selection mode, the set scene and subject type are treated as the shooting mode.

[0057] In S3, the system control circuit 50 acquires image data for display. The system control circuit 50 reads the image data for the live view display that will be displayed, which is stored in the video memory area of ​​the memory 32, and supplies it to the image processing unit 24.

[0058] In S4, the image processing unit 24, as a generation means, applies processing to the display image data supplied from the system control circuit 50 to generate display image data for gaze position detection. The image processing unit 24 then stores the generated display image data back into the video memory area of ​​memory 32. In this example, the display image data acquired from the video memory area in S3 is processed, but the image processing unit 24 could also apply processing when generating the display image data to generate display image data for gaze position detection from the beginning.

[0059] Here, we will describe some examples of processing applied by the image processing unit 24 for gaze position detection. The processing applied for gaze position detection is a process in which feature regions determined based on the setting information of the imaging device 1 (here, a shooting mode as an example) are visually emphasized compared to other regions. This processing makes it easier for the user to quickly focus their gaze on the desired subject. What feature information should be detected in accordance with the setting information, and the parameters necessary for detecting the feature information, can be stored in advance for each piece of setting information, for example, in the non-volatile memory 56. For example, in association with a shooting mode for capturing a specific scene, templates and parameters for detecting the type of main subject corresponding to that specific scene and the feature regions of that main subject can be stored in advance.

[0060] (Example 1) Figure 6 schematically shows an example of processing that can be applied when the scene is set to "Sports" in scene selection mode. Figure 6(a) shows the image represented by the display image data before processing, and Figures 6(b) to 6(d) show the images represented by the display image data after processing, respectively.

[0061] If "Sports" is selected in the scene selection mode, it can be inferred that the user intends to photograph a sports scene. In this case, the image processing unit 24 determines that the area of ​​a moving person should be emphasized as a feature area, and applies a processing step to emphasize the feature area.

[0062] Specifically, the image processing unit 24 detects the human body region as a feature region and identifies the moving person region by comparing it with the previous detection result (for example, the detection result in the live view image from the previous frame). Then, the image processing unit 24 applies a processing step to the current frame's live view image to emphasize the moving person region.

[0063] Here, we assume that in the image of the current frame shown in Figure 6(a), moving person regions P1, P2, and P3 have been detected. Then, Figure 6(b) shows an example in which a processing method is applied to emphasize the person regions P1 to P3 by superimposing frames A1 to A3 that surround the person regions.

[0064] Furthermore, Figure 6(c) shows an example of a processing technique to emphasize the person area P1-P3, where the display of the area surrounding the person area P1-P3 is not changed, and the brightness of other areas A4 is reduced. Furthermore, Figure 6(d) shows an example of a processing technique to emphasize the person area P1-P3, where the display of the rectangular area A5 surrounding all of the person area P1-P3 is not changed, and the brightness of other areas is reduced.

[0065] In this way, by detecting feature regions according to the set scene and the type of main subject, and applying processing to emphasize the detected feature regions, it is expected that the user will find it easier to locate the intended main subject. By making it easier for the user to find the intended main subject, it is expected that the time it takes for the user's gaze to focus on the main subject will be shortened.

[0066] The processing techniques used to enhance feature regions are not limited to the examples described above. For example, the processing techniques could enhance the edges of the human body regions P1, P2, and P3 detected as feature regions. Frames A1 to A3 could also be made to blink or displayed in a specific color. In addition, instead of reducing the brightness in Figures 6(c) and 6(d), monochrome display could be used. Furthermore, if the feature regions are areas of humans or animals, the entire image could be converted into a thermographic-style pseudo-color image to enhance the areas of humans or animals.

[0067] (Example 2) Figures 7(a) and 7(b) schematically show examples of processing that can be applied when the main subject is set to "Kids" in scene selection mode. Figure 7(a) shows the image represented by the display image data before processing, and Figure 7(b) shows the image represented by the display image data after processing.

[0068] If "Kids" is selected in the scene selection mode, it can be inferred that the user intends to photograph children as the main subject. In this case, the image processing unit 24 determines that the area of ​​the person subject, which is presumed to be a child, is a feature area that should be emphasized, and applies a processing operation to emphasize the feature area.

[0069] Whether a person's facial region, detected as a feature area, is an adult or a child can be determined, for example, by checking if the torso length or the ratio of head length to height is below a threshold, or by using machine learning, but is not limited to these methods. Alternatively, only individuals pre-registered as children may be detected using facial recognition.

[0070] Here, in the image of the current frame shown in Figure 7(a), human regions P1, K1, and K2 are detected, and regions K1 and K2 are determined to be children. Figure 7(b) shows an example of a processing method applied to emphasize the child regions K1 and K2, which involves enhancing the edges of the child regions K1 and K2 and further reducing the gradation of regions other than the child regions K1 and K2. The reduction in gradation may be, but is not limited to, a reduction in maximum brightness (brightness compression) or a reduction in the number of brightness gradations (from 256 gradations to 16 gradations). Reduction in brightness or monochrome display, as shown in Example 1, may also be applied.

[0071] (Example 3) Figures 7(a) and 7(c) schematically show examples of processing that can be applied when the main subject is set to "text" in scene selection mode. Figure 7(a) shows the image represented by the display image data before processing, and Figure 7(c) shows the image represented by the display image data after processing.

[0072] If "Text" is selected in the scene selection mode, it can be inferred that the user intends to focus on and photograph text present in the scene. In this case, the image processing unit 24 determines that the area estimated to be text is a feature area that should be emphasized, and applies a processing step to emphasize the feature area.

[0073] Here, we assume that the character region MO has been detected in the image of the current frame shown in Figure 7(a). Figure 7(c) shows an example of a processing method that enhances the character region MO, which involves enhancing the edges of the character region MO and reducing the gradation of areas other than the character region MO. The method for reducing gradation may be the same as in Example 2.

[0074] As explained in Examples 2 and 3, different processing can be applied to the same original image (Figure 7(a)) depending on the set shooting mode. Note that the same processing as in Example 1 may also be applied to Examples 2 and 3. Furthermore, edge enhancement for areas to be emphasized and reduction of brightness or gradation for other areas may be applied only to one of these.

[0075] In this embodiment, the processing that can be applied to the image data for gaze position detection is a processing that visually emphasizes the region to be emphasized (feature region), which is determined based on the settings information of the imaging device, compared to other regions. The processing may be any of the following four types, for example. (1) Processing that leaves areas that should be emphasized untouched, while processing other areas to make them less noticeable (by reducing brightness or gradation, etc.). (2) A process that emphasizes areas that should be emphasized (such as edge highlighting) while leaving other areas unprocessed. (3) Processing that emphasizes areas that should be emphasized (such as edge enhancement) and further processes other areas to make them less prominent (such as by reducing brightness or gradation). (4) Processing the entire image to highlight areas that need to be emphasized (e.g., conversion to a pseudo-color image) These are merely examples, and any processing can be applied to make the areas to be emphasized visually stand out (more prominent) compared to other areas.

[0076] Returning to Figure 4, in S5, the system control circuit 50 displays the image data for display generated by the image processing unit 24 in S4 on the display unit 28. The system control circuit 50 also obtains from the image processing unit 24 the rotation angle of the optical axis of the eyeball, which the image processing unit 24 detected based on the gaze detection image from the gaze input unit 701. Based on the obtained rotation angle, the system control circuit 50 determines the coordinates (gaze position) within the image displayed on the display unit 28 where the user is fixated. The system control circuit 50 may also notify or provide feedback to the user of the gaze position by superimposing a mark or other indicator showing the obtained gaze position onto the live view image.

[0077] This concludes the gaze position detection operation. The gaze position obtained through the gaze position detection operation can be used, but is not limited to, setting the focus detection area or selecting the main subject. When recording image data obtained during shooting, the gaze position information detected at the time of shooting may be recorded in association with the image data. For example, the gaze position information at the time of shooting can be recorded as supplementary information in the header of the data file that stores the image data. The gaze position information recorded in association with the image data can be used in application programs that handle the image data to identify the main subject, etc.

[0078] If the eye-tracking input function is not enabled, the image processing unit 24 will not apply any processing to the display image data to support eye-tracking input, but it may apply processing for other purposes.

[0079] As described above, in this embodiment, when eye-tracking input is enabled, a processing treatment is applied to the displayed image in which the feature region determined based on the settings information of the imaging device is visually emphasized compared to other regions. This makes it easier to see the region that the user is likely to intend as the main subject, and is expected to shorten the time it takes for the user to focus on the main subject.

[0080] In this embodiment, the area to be emphasized was determined based on the shooting mode setting. However, other settings may be used if it is possible to determine the type of main subject that the user is likely intending.

[0081] ●(Second Embodiment) Next, a second embodiment will be described. The second embodiment is an embodiment in which an XR goggle (a head-mounted display device or HMD) is used as the display unit 28 in the first embodiment. Note that XR is a general term for VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality).

[0082] The left image in Figure 8(a) is a perspective view showing an example of the appearance of the XR Goggles 800. The XR Goggles 800 is typically worn on the face area SO shown in the right image in Figure 8(a). Figure 8(b) is a schematic diagram showing the mounting surface (the surface that comes into contact with the face) of the XR Goggles 800. Figure 8(c) is a schematic top view showing the positional relationship between the eyepiece lens 701d, the display units 28A and 28B, and the user's right eye 501a and left eye 501b when the XR Goggles 800 is worn.

[0083] The XR goggles 800 have a display unit 28A for the right eye 501a and a display unit 28B for the left eye 501b. Stereoscopic vision is enabled by displaying the right-eye image, which constitutes a parallax image pair, on display unit 28A and the left-eye image on display unit 28B. For this reason, the eyepiece lens 701d described in the first embodiment is provided for each of the display units 28A and 28B.

[0084] In this embodiment, the imaging unit 22 is assumed to have pixels with the configuration shown in Figure 2(a). In this case, the right-eye image can be generated from the pixel signal group obtained from the photoelectric conversion unit 201a, and the left-eye image can be generated from the pixel signal group obtained from the photoelectric conversion unit 201b. The right-eye and left-eye images may also be generated using other configurations, such as making the lens unit 150 a lens capable of capturing stereo images. Furthermore, the gaze input unit 701 is provided in the eyepiece of the XR goggles and generates a gaze detection image for either the right or left eye.

[0085] Since the other configurations can be implemented using the same configuration as the imaging device 1 shown in Figure 1, the following explanation will use the components of the imaging device 1. In this embodiment, the display image is generated using the right-eye image and the left-eye image that have been pre-recorded on the recording medium 200, rather than the live view image.

[0086] Figure 9 is a flowchart relating to the gaze position detection operation in this embodiment. Steps that perform the same processing as in the first embodiment are given the same reference numerals as in Figure 4, thereby omitting redundant explanations.

[0087] In S91, the system control circuit 50 acquires the currently set experience mode. In this embodiment, since no shooting is performed, the system acquires an experience mode related to XR. The experience mode is, for example, the type of virtual environment in which the XR experience is performed, and options such as "art museum," "museum," "zoo," and "diving" are available. The experience mode can be set using a remote controller, an input device provided on the XR goggles, or by displaying a menu screen and selecting it with eye movements. The recording medium 200 is assumed to store display image data corresponding to each of the virtual environments that can be selected as an experience mode.

[0088] In S3, the system control circuit 50 acquires display image data corresponding to the experience mode selected in S91 by reading it from the recording medium 200 and supplies it to the image processing unit 24.

[0089] In S92, the image processing unit 24 applies processing to the display image data supplied from the system control circuit 50 to generate display image data for gaze position detection. In this embodiment, since the display image data is stereo image data including an image for the right eye and an image for the left eye, the image processing unit 24 applies processing to both the image for the right eye and the image for the left eye.

[0090] The processing applied by the image processing unit 24 in S92 is a processing process in which feature regions determined based on the setting information (in this case, the experience mode) of the device providing the XR experience (in this case, the imaging device 1) are visually emphasized compared to other regions. This processing is expected to enhance the immersive feeling of the XR experience.

[0091] An example of the processing applied by the image processing unit 24 in S92 will be described. (Example 4) Figure 10 schematically shows an example of processing that can be applied when the experience mode is set to "diving". Figure 10(a) shows the image represented by the display image data before processing.

[0092] If "Diving" is selected in the experience mode, it can be inferred that the user is interested in marine life. In this case, the image processing unit 24 determines that the area of ​​moving marine life is a feature area that should be emphasized, and applies processing to emphasize the feature area.

[0093] Specifically, the image processing unit 24 detects areas of fish and marine mammals as feature regions and identifies moving feature regions by comparing them with past detection results. Then, the image processing unit 24 applies a processing operation to the frame image to be processed to emphasize the moving feature regions.

[0094] Here, we assume that feature regions f1 to f4, which represent the areas of moving fish and humans, are detected in the frame image to be processed shown in Figure 10(a). In this case, the image processing unit 24 applies a processing process to enhance feature regions f1 to f4, which maintains the display of feature regions f1 to f4 while reducing the number of colors in other areas (for example, making them monochrome). Note that the processing process to enhance feature regions may include other processing processes, including those described in the first embodiment.

[0095] Returning to Figure 9, in S5, the system control circuit 50 obtains from the image processing unit 24 the rotation angle of the optical axis of the eyeball detected by the image processing unit 24 based on the gaze detection image from the gaze input unit 701. Based on the obtained rotation angle, the system control circuit 50 determines the coordinates (gaze position) in the image displayed on the display unit 28A or 28B where the user is fixated. Then, in S92, the system control circuit 50 superimposes marks indicating the gaze position onto the right eye and left eye image data generated by the image processing unit 24 and displays them on the display units 28A and 28B.

[0096] In S93, the system control circuit 50 determines whether or not to apply further processing to the display image using the gaze position information detected in S5. This determination can be performed based on arbitrary determination conditions, such as based on user settings regarding the use of gaze position information.

[0097] The system control circuit 50 terminates the gaze position detection operation if it determines that gaze position information is not to be used. On the other hand, if it determines that line of sight position information is to be used, the system control circuit 50 executes S94.

[0098] In S94, the system control circuit 50 reads the display image data stored in the video memory area of ​​the memory 32 and supplies it to the image processing unit 24. The image processing unit 24 uses the gaze position detected in S5 to apply further processing to the display image data.

[0099] Examples of processing using gaze position information performed in S94 are shown in Figures 10(b) and 10(c). The display image data has a marker P1 indicating the gaze position detected in S5 superimposed on it. Here, since the detected gaze position p1 is within feature region f1, it is highly likely that the user is interested in feature region f1. Therefore, processing is applied in S92 to make feature region f1 more visually emphasized than the other feature regions f2 to f4 among the feature regions f1 to f4 that have been emphasized by applying processing.

[0100] For example, in S92, feature regions f1 to f4 are emphasized by maintaining color display and displaying other regions in monochrome. In this case, in S94, the image processing unit 24 also displays feature regions f2 to f4 in monochrome, while maintaining color display for feature region f1, or the region encompassing the gaze position and the feature region closest to the gaze position (in this case, feature region f1). Figure 10(b) schematically shows a state where the region C1 encompassing the gaze position p1 and feature region f1 is displayed in color, while other regions including feature regions f2 to f4 are displayed in monochrome. Here, feature regions f2 to f4 are changed to the same display format as regions other than the feature region, but feature regions f2 to f4 may be displayed in a format that is less conspicuous than feature region f1 but more conspicuous than regions other than the feature region.

[0101] In this way, by using the gaze position to narrow down the feature regions that should be emphasized, it becomes possible to more accurately identify the feature regions that the user is interested in and apply processing to emphasize them, compared to when gaze position is not used. Therefore, an even greater effect of increasing immersion during XR experiences can be expected. In addition, in applications that utilize gaze position, it becomes easier for the user to confirm that they are looking at the intended subject.

[0102] The time change of the detected gaze position can also be used. In Figure 10(c), suppose the gaze position detected at time T=0 was P1, and the gaze position detected at T=1 (units are arbitrary) was P2. In this case, since the gaze position moved from P1 to P2 between time T=0 and T=1, it can be seen that the user shifted their gaze to the left.

[0103] In this case, highlighting feature region f4, which lies in the direction of the gaze position's movement, is expected to make it easier for the user to focus on the new feature region. Here, we show an example where C2, which maintains color display, is extended to include feature region f4, which is the shortest distance in the direction of the gaze position's movement. While this example shows extending the highlighted region in the direction of the gaze position's movement, the region can also be moved in accordance with the gaze position's movement without being extended.

[0104] By determining the areas to emphasize while considering changes in the user's gaze position over time, it is possible to highlight subjects that the user is likely to focus on, and this is expected to help the user easily focus on the desired subject.

[0105] In S92, another example of the processing applied by the image processing unit 24 will be described. (Example 5) Figure 11 schematically shows an example of processing that can be applied when the experience mode is set to "Art Museum". Figure 11(a) shows the image represented by the display image data before processing.

[0106] If "Art Museum" is selected in the experience mode, it can be inferred that the user is interested in works of art such as paintings and sculptures. In this case, the image processing unit 24 determines that the area of ​​the artwork is a feature area that should be emphasized, and applies a processing operation to emphasize the feature area.

[0107] Here, we assume that feature regions B1 to B5 are detected as the area of ​​the artwork in the frame image to be processed shown in Figure 11(a). In this case, at S92, the image processing unit 24 applies a processing process to enhance feature regions B1 to B5, for example as shown in Figure 11(b), which maintains the display of feature regions B1 to B5 and reduces the brightness of other areas. Note that the processing process to enhance feature regions may be any other processing process, including the one described in the first embodiment.

[0108] When applying further processing to the display image using gaze position information, in S94, the image processing unit 24 can superimpose pre-stored associated information CM1 onto the feature region B2 (indicated by marker p3), which includes the gaze position, as shown in Figure 11(c). There are no particular restrictions on the associated information CM1; for example, if it is a painting, it may be bibliographic information such as the name of the painting, the artist, and the year of creation, or other information appropriate to the type of feature region. In this embodiment, since the display image data is prepared in advance, information regarding the position of the artwork in the image and associated information about the artwork can also be prepared in advance. Therefore, the image processing unit 24 can identify the artwork present at the gaze position and acquire its associated information.

[0109] Here, we've added supplementary information about the artwork at the point of focus to further emphasize it, but you could also emphasize it in other ways, such as overlaying a magnified image of the artwork at the point of focus.

[0110] As described above, in this embodiment, in addition to the processing described in the first embodiment, processing that takes into account the gaze position can be applied to more effectively highlight feature regions that are likely to be of interest to the user. Therefore, it becomes possible to help the user quickly focus on the desired subject and to provide a more immersive XR experience.

[0111] ●(Third embodiment) Next, a third embodiment will be described. The eye-tracking input function is a function that uses the user's vision, but there are individual differences in the user's visual characteristics. Therefore, in this embodiment, the usability of the eye-tracking input function is improved by applying processing to the display image data that takes into account the user's visual characteristics.

[0112] Examples of individual differences in visual characteristics include: (1) Individual differences in the range of brightness in which differences in brightness can be distinguished (dynamic range) (2) Individual differences in central vision (1-2° around the point of fixation) and effective field of view (4-20° around central vision) (3) Individual differences in the ability to perceive hue differences These are some examples. These individual differences can arise both congenitally and acquiredly (typically due to aging).

[0113] Therefore, in this embodiment, visual information reflecting the individual differences described in (1) to (3) is registered for each user, and processing that reflects the visual information is applied to the display image data, thereby providing an eye-tracking input function that is easy for each individual user to use.

[0114] The following describes specific examples of calibration functions for acquiring visual information. The calibration function can be executed by the system control circuit 50, for example, when instructed to do so by the user through a menu screen, or when the user's visual characteristics have not been registered.

[0115] (1) The luminance dynamic range can be set to the range of maximum and minimum luminance that the user does not find uncomfortable. For example, the system control circuit 50 displays a grayscale gradient chart on the display unit 28, which represents the range from maximum to minimum luminance with a predetermined number of gradations, as shown in Figure 12(a). The user then selects a luminance range that they do not find uncomfortable, for example, by operating the operation unit 70. The user can adjust the position of the upper and lower ends of the bar 1201 using, for example, the up and down keys of the four-way key, to set the maximum luminance that does not feel dazzling and the minimum luminance that allows the user to distinguish the difference between adjacent gradations (or does not feel too dark).

[0116] The system control circuit 50 registers brightness ranges KH and KL that are preferable for the user to not use, based on, for example, the positions of the upper and lower ends of the bar 1201 when the set (confirm) button is pressed. Alternatively, the brightness corresponding to the positions of the upper and lower ends of the bar 1201 may be registered.

[0117] Alternatively, the system control circuit 50 may, for example, increase the overall brightness of the screen in response to the press of the up key on the four-way key, and decrease the overall brightness of the screen in response to the press of the down key, thereby allowing the user to set the maximum and minimum brightness levels. The system control circuit 50 then prompts the user to press the set button when the screen is displayed at the maximum brightness level that is not perceived as dazzling. The system control circuit 50 then registers the display brightness level at the time the set button is pressed as the maximum brightness level. The system control circuit 50 also prompts the user to press the set button when the screen is displayed at the minimum brightness level at which the difference between adjacent gradations can be distinguished (or when the screen is not perceived as too dark). The system control circuit 50 then registers the display brightness level at the time the set button is pressed as the minimum brightness level. In this case as well, instead of maximum and minimum brightness levels, the system control circuit 50 may register a high-brightness range KH and a low-brightness range KL that are preferable not to use for the user.

[0118] User visual characteristics related to luminance dynamic range can be used to determine whether or not luminance adjustment is necessary and to decide on the parameters to use when adjusting luminance.

[0119] (2) The effective field of view is the range in which information can be identified, including central vision. The effective field of view may be, for example, a field of view called the Useful Field of View (UFOV). The system control circuit 50 displays an image on the display unit 28 in which a variable-size circle 1202 is displayed against a background of a relatively fine pattern, as shown in Figure 12(b). The system control circuit 50 then prompts the user to adjust the size of the circle 1202 so that the background pattern can be clearly distinguished while fixating on the center of the circle 1202. The user can set the size of the effective field of view by adjusting the size of the circle 1202 using, for example, the up and down keys of the four-way key to correspond to the largest range in which the background pattern can be clearly recognized, and then pressing the set key. When the system control circuit 50 detects that the up and down keys have been pressed, it changes the size of the circle 1202, and when it detects that the set key has been pressed, it registers the range of the effective field of view corresponding to the size of the circle 1202 at that time.

[0120] User visual characteristics related to the effective field of view can be used to extract the gaze range.

[0121] (3) is the magnitude of the hue difference in which differences can be recognized within the same color family. The system control circuit 50 displays on the display unit 28 an image in which multiple color samples of the same color family with gradually changing hues are selected and arranged, for example as shown in Figure 12(c). The color samples displayed here can be colors that occupy a large area in the background of a subject, such as green, yellow, and blue. Information may also be acquired for multiple color families, such as green families, yellow families, and blue families.

[0122] Figure 12(c) shows a color chart using colored pencils, but strip-shaped color charts or other methods may also be used. The leftmost colored pencil is the reference color, and color charts with varying hues are arranged to the right. The system control circuit 50 prompts the user to select the leftmost colored pencil that can be recognized as having a different color from the leftmost colored pencil. The user selects the corresponding colored pencil using, for example, the left or right arrow keys of a 4-way keypad and presses the set key. When the system control circuit 50 detects the press of the left or right arrow keys, it moves the selected colored pencil, and when it detects the press of the set key, it registers the difference between the hue corresponding to the currently selected colored pencil and the hue of the reference color as the smallest hue difference that the user can recognize. When registering information for multiple color systems, the same operation is repeated for each color system.

[0123] A user's visual characteristics regarding their ability to perceive hue differences can be used to determine whether hue adjustment is necessary and to decide on the parameters to use when adjusting hue.

[0124] The methods described above for obtaining individual differences in visual characteristics (1) to (3) and user-specific information regarding visual characteristics (1) to (3) are merely examples. User information regarding other visual characteristics and / or information regarding visual characteristics (1) to (3) can be registered by other means.

[0125] Next, we will explain specific examples of processing using the registered user's visual characteristics (1) to (3). Note that if visual characteristics can be registered for multiple users, the visual characteristics of the user selected through the settings screen will be used, for example.

[0126] Figure 13(a) shows a scene with multiple aircraft E1 against a high-brightness sky background in a backlit condition. When the background is this bright, depending on the user's visual characteristics, the background may be dazzling, making it difficult to focus on the aircraft E1.

[0127] To address this situation, when generating display image data while the eye-tracking function is enabled, the image processing unit 24 can determine whether the background luminance value (e.g., average luminance value) is appropriate for the user's visual characteristics (luminance dynamic range). If the background luminance value is outside the user's luminance dynamic range (i.e., included in the luminance range KH in Figure 12), the image processing unit 24 determines that the luminance is not appropriate for the user's visual characteristics. The image processing unit 24 then applies a processing step to the display image data to reduce the luminance so that the luminance value of the background area of ​​the image is included within the user's luminance dynamic range (the luminance range represented by bar 1201 in Figure 12).

[0128] Figure 13(b) schematically shows the state after applying a processing treatment to reduce the brightness of the background area. M1 is the main subject area. The area of ​​the image excluding the main subject area M1 is defined as the background area. Here, the image processing unit 24 defines the main subject area M1 as an area large enough to encompass a feature area (in this case, an airplane) that is within a certain range from the user's gaze position, and separates it from the background area. The size of the main subject area may also be the size of the user's effective field of view. Furthermore, the determination of the main subject area based on the user's gaze position may be based on other methods.

[0129] When applying processing to adjust the brightness to suit the user's brightness dynamic range, the target brightness value can be appropriately determined within the brightness dynamic range. For example, it may be set to the median value of the brightness dynamic range. Although only processing to adjust (correct) the brightness value of the background area has been explained here, the brightness value of the main subject area can be adjusted in the same way. When adjusting the brightness of both the background and main subject areas, the visibility of the main subject area can be improved by setting the target brightness of the main subject area higher than that of the background area.

[0130] Figure 14(a) shows an example of a scene where it is easy to lose sight of the main subject, such as a team sport or game where many similar subjects move in various directions. In Figure 14(a), let's assume that the main subject intended by the user is E2. If the user loses sight of the main subject E2 and the main subject moves out of the user's effective field of view, the main subject will appear blurred like other subjects, making it even more difficult to distinguish.

[0131] To address this situation, the image processing unit 24 applies a processing step that reduces (blurs) the resolution of areas other than the main subject area M2 (background area), as shown in Figure 14(b). This relatively increases the sharpness of the main subject area M2, so that even if the user loses sight of the main subject E2, it can be easily found. The main subject area M2 can be determined in the same way as described for brightness adjustment.

[0132] Furthermore, if the size of the main subject area is larger than the central field of view, the area of ​​the main subject area that is outside the central field of view may also be processed as a background area. In this way, by relatively increasing the sharpness of the main subject, the user's attention is naturally directed to the main subject, and as a result, the effect of supporting subject tracking based on the gaze position can also be achieved.

[0133] Figure 15(a) shows an example of a scene where the main subject is an animal moving in a dark place, as the main subject has low brightness and is difficult for the user to recognize. In Figure 15(a), assume that the main subject intended by the user is E3.

[0134] To address this situation, when generating display image data while the gaze input function is enabled, the image processing unit 24 can determine whether the luminance value (e.g., average luminance value) of the area surrounding the gaze position is appropriate for the user's visual characteristics (luminance dynamic range). If the luminance value of the area surrounding the gaze position is outside the user's luminance dynamic range (i.e., included in the luminance range KL in Figure 12), the image processing unit 24 determines that the luminance is not appropriate for the user's visual characteristics. The image processing unit 24 then applies a processing step to the display image data to increase the luminance so that the luminance value of the area surrounding the gaze position is within the user's luminance dynamic range (the luminance range represented by bar 1201 in Figure 12). Figure 15(b) schematically shows the state after applying a processing step to increase the luminance of the area surrounding the gaze position M3. The area surrounding the gaze position may be, for example, the area corresponding to the effective field of view, a feature area including the gaze position, or an area used as a tracking template.

[0135] The reason for adjusting (increasing) the brightness only in the area surrounding the point of focus, rather than the entire image, is that increasing the brightness of a dark scene image reduces its visibility due to noise. Increasing the brightness of the entire screen tends to reduce the accuracy of detecting moving subjects between frames due to the effects of noise. Furthermore, if noise becomes visible across the entire screen, the flickering of the noise can easily cause eye fatigue in the user.

[0136] Furthermore, if the scene is dark, it is quite possible that the main subject is not present at the point of focus. For this reason, the brightness of the entire screen may be increased until the point of focus stabilizes, and once the point of focus stabilizes, the brightness of areas other than the area surrounding the point of focus may be returned to its original level (no processing applied). The system control circuit 50 can determine that the point of focus has stabilized, for example, if the amount of movement of the point of focus remains below a threshold for a certain period of time.

[0137] Figure 16(a) shows an example of a scene where the main subject and background colors are similar, making it easy to lose sight of the main subject. The scene depicts a bird E4 moving against a background of grass, with similar colors. In Figure 16(a), the user intends the main subject to be bird E4. If the user loses sight of bird E4, it is difficult to find bird E4 because the background and bird E4 are similar in color.

[0138] Therefore, the image processing unit 24 can determine whether the difference between the hue of the main subject area (the area of ​​bird E4) and at least the hue of the surrounding background area is appropriate in light of the user's ability to recognize hue differences, which is one of the user's visual characteristics. If the difference between the hue of the main subject area and the background area is less than or equal to the hue difference that the user can recognize, the image processing unit 24 determines that it is inappropriate. In this case, the image processing unit 24 applies a processing operation to the displayed image data to change the hue of the main subject area so that the difference between the hue of the main subject area and the surrounding background area becomes greater than the hue difference that the user can recognize. Figure 16(b) schematically shows the state after applying the processing operation to change the hue of the main subject area M4.

[0139] It should be noted that, in addition to the processing methods exemplified here, it is possible to apply processing methods that utilize the user's visual characteristics. Furthermore, multiple processing methods can be combined and applied according to the brightness and hue of the main subject area and background area.

[0140] Figure 17 is a flowchart illustrating the operation for generating display image data according to this embodiment. This operation can be performed in parallel with the detection of the gaze position when the eye-tracking input function is enabled. In S1701, the system control circuit 50 captures one frame of image using the imaging unit 22 and supplies a digital image signal to the image processing unit 24 via the A / D converter 23.

[0141] In S1702, the image processing unit 24 detects a feature region to be designated as the main subject region based on the most recently detected gaze position. Here, the image processing unit 24 may, after detecting a feature region of the type determined from the shooting mode as described in the first embodiment, designate the feature region containing the gaze position, or the feature region closest in distance from the gaze position, as the main subject region.

[0142] In S1703, the image processing unit extracts the feature region (main subject region) detected in S1702. This separates the main subject region from other regions (background region).

[0143] In S1704, the image processing unit 24 acquires information about the user's visual characteristics, for example, stored in the non-volatile memory 56.

[0144] In S1705, the image processing unit 24 calculates the difference in average brightness and hue between the main subject area and the background area. The image processing unit 24 then compares the calculated difference in average brightness and hue with the user's visual characteristics to determine whether or not it is necessary to apply processing to the main subject area. As described above, the image processing unit 24 determines that it is necessary to apply processing to the main subject area if the brightness of the main subject or the difference in hue between the main subject area and the background area is not appropriate for the user's visual characteristics. If it is determined that it is necessary to apply processing to the main subject area, the image processing unit 24 executes S1706; otherwise, it executes S1707.

[0145] In S1706, the image processing unit 24 applies processing to the main subject area according to the content that was determined to be inappropriate, and then executes S1707.

[0146] In S1707, the image processing unit 24 determines, in the same manner as in S1705, whether or not it is necessary to apply processing to other areas (background areas). If the image processing unit 24 determines that it is necessary to apply processing to the background areas, it executes S1708. If it does not determine that it is necessary to apply processing to the background areas, the image processing unit 24 executes S1701 and starts the operation for the next frame.

[0147] In S1708, the image processing unit 24 applies processing to the background area according to the content that was determined to be inappropriate, and then executes S1701.

[0148] Furthermore, the type of processing to be applied to the main subject area and the background area can be predetermined, depending on what is appropriate for the user's visual characteristics. Therefore, depending on the judgment result in S1705, the content of the processing to be applied is specified: whether to apply the processing only to the main subject area, only to the background area, or to both the main subject area and the background area.

[0149] As described above, according to this embodiment, when the eye-tracking input function is enabled, processing that takes into account the user's visual characteristics is applied to generate display image data. Therefore, it is possible to generate display image data appropriate for the individual user's visual characteristics, and to provide an eye-tracking input function that is easier for the user to use.

[0150] Furthermore, the processing described in the first embodiment to facilitate the selection of the main subject by gaze, and the processing described in this embodiment to create an image suitable for the user's visual characteristics, can be applied in combination.

[0151] ●(Fourth embodiment) Next, a fourth embodiment will be described. This embodiment relates to improving the visibility of a virtual space experienced using an XR goggle (head-mounted display device or HMD) that incorporates the components of the imaging device 1. The image of the virtual space viewed through the XR goggle is generated by drawing display image data, which is prepared in advance for each virtual space, according to the orientation and posture of the XR goggle. The display image data may be stored in advance on the recording medium 200, or it may be acquired from an external device.

[0152] Here, as an example, we assume that the display data for providing the "Diving" and "Art Museum" experience modes in the virtual space is stored on the recording medium 200. However, there are no particular restrictions on the types and number of virtual spaces that can be provided.

[0153] Examples of virtual space images for providing the "Diving" and "Art Museum" experience modes are schematically shown in Figures 18(a) and (b). For the sake of clarity and ease of explanation, the entire virtual space is assumed to be represented by a computer-generated image (CG). Therefore, the main subject to be highlighted is a part of the CG image. By highlighting the main subject included in the virtual space image, the visibility of the main subject can be improved. The main subject is set by the imaging device 1 (system control circuit 50) at least in the initial state. The user may change the main subject set by the imaging device 1.

[0154] Furthermore, in cases where a composite image is displayed, such as with a video see-through type HMD, which overlays computer graphics (CG) as a virtual image onto an image of real space, the main subject area (feature area) to be highlighted may be included in the real-world image portion or in the CG portion.

[0155] Figures 19(a) and (b) schematically illustrate examples of applying image processing to the scenes shown in Figures 18(a) and (b), respectively, to emphasize the main subject. Here, the main subject is emphasized by reducing the saturation of elements other than the main subject, thereby improving its visibility. Note that other methods may be used to emphasize the main subject.

[0156] The example shown in Figure 19 is an image processing technique that leaves the main subject area untouched while making other areas less prominent. Other possible processing techniques include emphasizing the main subject while leaving other areas untouched, or emphasizing the main subject while making other areas less prominent. Alternatively, the entire image may be processed to emphasize the subject, or the main subject area may be emphasized using other methods.

[0157] When experiencing diving in a virtual space, the primary subject can be considered to be "living creatures." When experiencing an art museum in a virtual space, the primary subject can be considered to be "exhibits" (paintings, sculptures, etc.) or objects with distinctive colors (here, we'll consider them to be vividly colored). In other words, the primary subject that should be emphasized may differ depending on the type of virtual space or experience being presented.

[0158] Figure 20 shows the relationship between the type of virtual space (or experience) provided and the types of subjects (types of feature regions) that can be highlighted. Here, the types of subjects that can be highlighted are associated with the type of virtual space as metadata. The type of primary subject that is highlighted by default is also associated with the type of virtual space. The types of subjects listed as metadata here correspond to the types of subjects that the image processing unit 24 can detect. Furthermore, for each type of virtual space, the types of subjects that can be set as the primary subject are indicated with a circle (○), and the types of subjects that are selected as the primary subject by default are indicated with a double circle (◎). Therefore, the user can select a new primary subject from the subjects indicated with a circle (○).

[0159] There are no particular restrictions on how the user can change the type of primary subject. For example, the system control circuit 50 displays a GUI for changing the primary subject on the display unit 28 of the imaging device 1 or on the display unit of the XR goggles in response to operations on the menu screen via the operation unit 70. The system control circuit 50 can then change the primary subject setting for the currently provided virtual space type in response to operations on this GUI via the operation unit 70.

[0160] Figure 21 shows an example of a GUI displayed to change the main subject. Figure 21(a) is a GUI that mimics a mode dial, and by operating the dial included in the operation unit 70, one of the options can be set as the main subject. Figure 21(a) shows the state where the landscape is set as the main subject. The options displayed in the GUI for changing the main subject correspond to the metadata types marked with a circle in Figure 20. In the example of Figure 21(a), in addition to the metadata types, "OFF" is included as an option to disable highlighting. Figure 21(b) shows another example of a GUI for changing the type of main subject. It is the same GUI as shown in Figure 21(a), except that it is displayed in a list format instead of a dial format. The user can change the main subject to be highlighted (and turn off highlighting) by selecting the desired option using the operation unit 70. It may also be possible to select options using gaze.

[0161] Figure 22 is an image illustrating examples of metadata for the virtual space types "diving," "museum," and "safari." For example, the image processing unit 24 can extract the subject area detected by the image processing unit 24 for each subject type as metadata and store it in memory 32. This makes it easy to respond to changes in the main subject to be highlighted. If it is possible to pre-generate the image of the virtual space to be displayed on the XR goggles, the metadata can also be recorded in advance. Furthermore, the metadata may be numerical information representing the subject area (for example, the center position and size, and coordinate data of the outer edge).

[0162] Furthermore, the gaze position information described in the second embodiment may be used to identify the type of subject the user is interested in, and the identified type of subject area may be highlighted. In this case, since the type of main subject highlighted changes according to the gaze position, the user can change the main subject without explicitly changing the settings.

[0163] Furthermore, if the main subject is not present in the current field of view, or if the number or size of the main subject area is below a threshold, an indicator showing the direction in which more main subjects will be within the field of view may be superimposed on the virtual space image.

[0164] Figure 23(a) shows an example of a virtual space image currently displayed on the XR goggles in the "Diving" experience mode. The displayed virtual space image does not contain the area of ​​the fish, which is the main subject. In this case, the system control circuit 50 can superimpose an indicator P1 indicating the direction in which the main subject exists onto the virtual space image. The system control circuit 50 can determine the direction in which the fish enters the field of view of the XR goggles, for example, based on the position information of the fish object in the virtual space data used to generate the display image data.

[0165] The user can spot the fish as shown in Figure 23(b) by turning their head to look in the direction indicated by indicator P1. Multiple indicators showing the direction of the main subject may be superimposed. In this case, the system control circuit 50 can display the indicator showing the direction requiring the shortest line of sight to include the main subject in the field of view, or the indicator showing the direction that allows the most main subjects to be included in the field of view, in the most prominent (e.g., large) position.

[0166] According to this embodiment, subject areas of a type corresponding to the provided virtual space are highlighted. As a result, the area that the user is most likely to intend as the main subject in the virtual space image becomes easier to see, and it is expected that the time it takes for the user to focus on the main subject will be shortened.

[0167] ●(Fifth embodiment) Next, a fifth embodiment will be described. This embodiment relates to a display system that acquires virtual space images to be displayed on the XR goggles in the fourth embodiment from an external device of the XR goggles, such as a server.

[0168] Figure 24(a) is a schematic diagram of a display system in which the XR goggles DP1 and the server SV1 are connected in a communicative manner. A network such as a LAN or the Internet may exist between the XR goggles DP1 and the server SV1.

[0169] Generally, generating virtual space images requires a large amount of virtual space data and the processing power to generate (render) virtual space images from that data. Therefore, the XR goggles output information necessary for generating virtual space images, such as the posture information detected by the posture detection unit 55, to the server. The server then generates the virtual space image to be displayed on the XR goggles and sends it to the XR goggles.

[0170] By storing virtual space data (3D data) on server SV1, it becomes possible to share the same virtual space with multiple XR goggles connected to the server.

[0171] Figure 25 is a block diagram showing an example configuration of a computer device that can be used as server SV1. In the figure, display 2501 displays information about data being processed by an application program, various message menus, etc., and is composed of an LCD (Liquid Crystal Display), etc. CRTC2502, acting as a video RAM (VRAM) display controller, controls the screen display on display 2501. Keyboard 2503 and pointing device 2504 are used for inputting characters, operating icons and buttons in a GUI (Graphical User Interface), etc. CPU 2505 controls the entire computer device.

[0172] ROM (Read Only Memory) 2506 stores programs and parameters executed by the CPU 2505. RAM (Random Access Memory) 2507 is used as a work area and data buffer when the CPU 2505 executes various programs.

[0173] The hard disk drive (HDD) 2508 and the removable media drive (RMD) 2509 function as external storage devices. The removable media drive is a device that reads or writes to removable recording media and may be an optical disc drive, magneto-optical disc drive, memory card reader, etc.

[0174] Furthermore, the programs that implement the various functions of server SV1, as well as the OS, application programs such as browsers, data, libraries, etc., are stored in one or more of the following storage media: ROM2506, HDD2508, or RMD2509, depending on their purpose.

[0175] Expansion slot 2510 is a slot for installing expansion cards that comply with standards such as the PCI (Peripheral Component Interconnect) bus. Various expansion boards, such as video capture boards and sound boards, can be installed in expansion slot 2510.

[0176] Network interface 2511 is an interface for connecting server SV1 to local and external networks. In addition to network interface 2511, server device SV1 also has one or more communication interfaces for external devices that comply with standards. Examples of standards include USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface) (registered trademark), wireless LAN, and Bluetooth (registered trademark).

[0177] Bus 2512 consists of an address bus, a data bus, and a control bus, and connects the aforementioned blocks.

[0178] Next, the operation of server SV1 and XR goggles DP1 will be explained using the flowchart shown in Figure 24(b). The operation of server SV1 is achieved by CPU 2501 executing a predetermined application.

[0179] In S2402, the XR goggles DP1 specify the type of virtual space (Figure 20) to the server SV1. The system control circuit 50 displays a GUI for specifying the type of virtual space on the display unit 28 of the XR goggles DP1, for example. When the system control circuit 50 detects a selection operation via the operation unit 70, it sends data indicating the selected type to the server SV1 via the communication unit 54.

[0180] Here, we assume that the range of the virtual space displayed on the XR goggles DP1 is fixed. Therefore, server SV1 sends image data (virtual space image data) of a specific scene in the specified type of virtual space, along with the accompanying metadata, to the XR goggles DP1.

[0181] In S2403, the system control circuit 50 receives virtual space image data and associated metadata from server SV1.

[0182] In S2404, the system control circuit 50 stores the virtual space image data and metadata received from the server SV1 in the memory 32.

[0183] In step S2405, the system control circuit 50 uses the image processing unit 24 to apply enhancement processing to the virtual space image data, as explained in Figure 19, for the main subject area. The enhanced virtual space image data is then displayed on the display unit 28. If the virtual space image consists of an image for the right eye and an image for the left eye, the enhancement processing is applied to each individual image.

[0184] Figure 24(c) is a flowchart illustrating the operation of server SV1 when generating virtual space image data according to the orientation (gaze direction) of the XR goggles DP1 and applying enhancement processing to the virtual space data. The operation of server SV1 is realized by CPU 2501 executing a predetermined application.

[0185] In S2411, server SV1 receives data from XR goggles DP1 specifying the type of virtual space.

[0186] The operations from S2412 onwards are performed for each frame of the video displayed on the XR goggles DP1. In S2412, server SV1 receives attitude information from XR goggles DP1. In S2413, server SV1 generates virtual space image data corresponding to the orientation of the XR goggles DP1. The virtual space data can be generated by any known method, such as rendering 3D data or cropping from a 360-degree image. For example, as shown in Figure 26, server SV1 can determine the display area of ​​the XR goggles DP1 from the virtual space image based on the orientation information of the XR goggles DP1 and crop the area corresponding to the display area. Alternatively, the XR goggles DP1 may transmit information that identifies the display area (e.g., center coordinates) instead of orientation information.

[0187] In S2415, server SV1 receives the type of main subject from XR goggles DP1. Note that receiving the type of main subject in S2415 is performed when the type of main subject changes in XR goggles DP1, and is skipped if there is no change.

[0188] In S2416, server SV1 applies enhancement processing to the main subject area of ​​the virtual space image data generated in S2413. If there is no change in the type of main subject, server SV1 applies enhancement processing to the default main subject area corresponding to the type of virtual space.

[0189] In S2417, server SV1 transmits the enhanced virtual space image data to XR goggles DP1. XR goggles DP1 displays the received virtual space image data on display unit 28.

[0190] Figure 27(a) is a schematic diagram of a display system in which a camera CA capable of generating VR images is added to the configuration of Figure 26(a). Here, the type of virtual space is assumed to be the experience sharing example given in Figure 20. By displaying the image recorded with added XR information by the camera CA on the XR goggles, the wearer of the XR goggles DP1 can also virtually experience the experience of the camera CA user.

[0191] Figure 28 is a block diagram showing an example configuration of camera CA. Camera CA has a main body 100' and a lens unit 300 mounted on the main body 100'. The lens unit 300 and the main body 100' are detachable via lens mounts 304 and 305. Furthermore, the lens system control circuit 303 of the lens unit 300 and the system control circuit 50 (not shown) of the main body 100' can communicate with each other via communication terminals 6 and 10 provided on the lens mounts 304 and 305.

[0192] The lens unit 300 is a stereo fisheye lens, and the camera CA can capture a stereo circular fisheye image with a field of view of 180°. Specifically, the two optical systems 301L and 301R of the lens unit 300 each generate a circular fisheye image by projecting a field of view of 180 degrees in the left-right direction (horizontal angle, azimuth angle, yaw angle) and 180 degrees in the up-down direction (vertical angle, elevation angle, pitch angle) onto a circular two-dimensional plane.

[0193] Although only a portion of its configuration is shown, the main unit 100' is assumed to have the same configuration as the main unit 100 of the imaging device 1 shown in Figure 1. Images captured by a camera CA with such a configuration (for example, moving images conforming to the VR180 standard) are recorded as XR images on the recording medium 200.

[0194] The operation of the display system shown in Figure 27(a) will be explained using the flowchart shown in Figure 27(b). It is assumed that server SV1 is in a state where it can communicate with XR goggles DP1 and camera CA.

[0195] S2602 transmits image data from camera CA to server SV1. The image data includes additional information such as Exif information including the shooting date and shooting conditions, the photographer's gaze information recorded at the time of shooting, and the main subject information detected at the time of shooting. Alternatively, instead of communication between camera CA and server SV1, the image data may be read by inserting the recording medium 200 of camera CA into server SV1.

[0196] In S2603, server SV1 generates image data and metadata to be displayed on XR goggles DP1 from image data received from camera CA. In this embodiment, since camera CA records a stereo circular fisheye image, the display image data is generated by cropping the display range using a known method and converting it into a rectangular image. Server SV1 also detects a predetermined type of subject area from the display image data and generates metadata based on the detected subject area. Server SV1 transmits the generated display image data and metadata to XR goggles DP1. Server SV1 also transmits additional information obtained from camera CA, such as main subject information and gaze information, to XR goggles DP1.

[0197] The operations performed by the system control circuit 50 of the XR goggles DP1 in S2604 and S2605 are the same as those in S2404 and S2405, so a detailed explanation is omitted. The system control circuit 50 can determine the type of main subject to which enhancement processing is applied in S2605 based on the main subject information received from the server SV1. The system control circuit 50 may also apply enhancement processing to the main subject area identified based on the photographer's gaze information. In this case, the subject that the photographer was focusing on at the time of shooting is highlighted, allowing for a greater sharing of the photographer's experience.

[0198] Figure 27(c) is a flowchart showing the operation of server SV1 when the highlighting process is performed on server SV1 in the same way as in Figure 24(c) in the display system shown in Figure 27(a).

[0199] Since S2612 is the same as S2602, the explanation will be omitted. Furthermore, since S2613~S2617 are the same as S2412, S2413, and S2415~S2417 respectively, their explanations will be omitted. Note that the type of main subject to which highlighting is applied will be the type specified by the XR Goggles DP1 if specified, otherwise it will be determined based on the main subject information at the time of shooting.

[0200] According to this embodiment, it becomes possible to apply appropriate enhancement processing to virtual space images and VR images. Furthermore, by performing computationally intensive processing on external devices such as servers, the resources required for XR goggles can be reduced, and it becomes easier for multiple users to share the same virtual space.

[0201] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0202] This embodiment includes the following imaging apparatus, method, image processing apparatus, image processing method, and program. (Item 1) An imaging device, A detection means capable of detecting the user's gaze position in the image displayed by the imaging device, It has a generation means for generating image data for the aforementioned display, The generation means applies a processing operation to the image data generated when the detection means is effective, which visually emphasizes the feature region compared to other regions. The characteristic region is the region of a subject of a type determined based on the settings of the imaging device. An imaging device characterized by the following features. (Item 2) The aforementioned settings are for photographing a specific scene or a specific subject. The imaging device according to item 1, characterized in that the feature region is a region of a type of subject corresponding to the specific scene, or a region of the specific subject. (Item 3) The aforementioned processing is The aforementioned feature region is left untouched, while other regions are processed to make them less prominent. A process that emphasizes the aforementioned feature region while leaving other regions unprocessed. A process that emphasizes the aforementioned feature region while making other regions less prominent. A process of processing the entire image including the aforementioned feature region to emphasize the aforementioned feature region. An imaging device according to item 1 or 2, characterized in that it is one of the following. (Item 4) The imaging apparatus according to any one of items 1 to 3, characterized in that the generation means generates the image data as image data for live view display. (Item 5) Furthermore, the imaging device according to any one of items 1 to 4 is characterized by having a setting means for setting a focus detection area based on the gaze position detected by the detection means. (Item 6) A method performed by an imaging device having detection means capable of detecting the user's gaze position in the displayed image, The process includes a generation step for generating image data for the aforementioned display, In the above generation process, When the detection means is effective, the image data generated is subjected to a processing that visually emphasizes the feature region compared to other regions. When the aforementioned detection means is ineffective, the image data generated is not subjected to any processing that visually emphasizes the feature region compared to other regions. The characteristic region is the region of a subject of a type determined based on the settings of the imaging device. A method characterized by the following: (Item 7) A program to cause the computer of the imaging device to function as one of the means of the imaging device described in any one of items 1 to 5. (Item 8) It has a generation means for generating image data to be displayed on a head-mounted display device, The generation means generates the image data by applying a processing operation that visually emphasizes a feature region corresponding to the type of virtual environment provided to the user through the display device compared to other regions. An image processing apparatus characterized by the following: (Item 9) The aforementioned processing is The aforementioned feature region is left untouched, while other regions are processed to make them less prominent. A process that emphasizes the aforementioned feature region while leaving other regions unprocessed. A process that emphasizes the aforementioned feature region while making other regions less prominent. A process of processing the entire image including the aforementioned feature region to emphasize the aforementioned feature region. The image processing apparatus according to item 8, characterized in that it is one of the following. (Item 10) Furthermore, the display device has detection means capable of detecting the user's gaze position in the image it is displaying, The image processing apparatus according to item 8 or 9, characterized in that the generating means generates the image data by applying the processing and then applying further processing based on the gaze position detected by the detection means. (Item 11) The image processing apparatus according to item 10, characterized in that the further processing is a processing that visually emphasizes the feature region including the gaze position among the feature regions more than other feature regions. (Item 12) The image processing apparatus according to item 10, characterized in that the further processing is a processing that superimposes and displays supplementary information relating to the feature region including the gaze position among the feature region. (Item 13) The image processing apparatus according to item 10, characterized in that the further processing is a processing that visually emphasizes a feature region located in the direction of movement of the gaze position. (Item 14) The image processing apparatus according to item 8, characterized in that, for each type of virtual environment, the type of feature region to which the processing can be applied and the type of feature region to which the processing is applied by default are associated. (Item 15) The image processing apparatus according to item 14, characterized in that the generation means applies the processing based on a type specified by the user from among the types of feature regions associated with the virtual environment provided to the user. (Item 16) The image processing apparatus according to item 14 or 15, characterized in that, if the user does not specify otherwise, the generation means applies the processing based on the type of feature region to which the processing is applied by default, which is associated with the virtual environment provided to the user. (Item 17) Furthermore, the display device has detection means capable of detecting the user's gaze position in the image it is displaying, The image processing apparatus according to any one of items 14 to 16, characterized in that the generation means applies the processing to the feature region based on the gaze position detected by the detection means. (Item 18) The image processing apparatus according to any one of items 14 to 17, characterized in that, if the generated image data does not contain the feature region, the generated means includes an index in the image data indicating the direction in which the feature region exists. (Item 19) The image processing apparatus according to any one of items 14 to 18, characterized in that the head-mounted display device is an external device capable of communicating with the image processing apparatus. (Item 20) The image processing apparatus according to any one of items 14 to 18, characterized in that the image processing apparatus is part of the head-mounted display device. (Item 21) The system further includes an acquisition means for acquiring data of a VR image representing the virtual environment, The generation means generates the image data from the VR image. An image processing apparatus according to any one of items 14 to 20, characterized in that (Item 22) The acquisition means further acquires the main subject information and / or gaze information obtained when the VR image was captured, The image processing apparatus according to item 21, characterized in that the generation means determines the feature region to which the processing is applied based on the main subject information or the gaze information. (Item 23) An image processing method performed by an image processing device, It has a generation process for generating image data to be displayed on a head-mounted display device, In the generation step, the image data is generated by applying a processing treatment that visually emphasizes a feature region corresponding to the type of virtual environment provided to the user through the display device compared to other regions. An image processing method characterized by the following: (Item 24) A program for causing a computer to function as one of the means of an image processing apparatus as described in any one of items 8 through 22.

[0203] The present invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]

[0204] 1…Imaging device, 22…Imaging unit, 24…Image processing unit, 28…Display unit, 50…System control circuit, 70…Operation unit, 100…Main unit, 150…Lens unit

Claims

1. An imaging device, a detection means for detecting a user's gaze position in an image displayed by the imaging device; generating means for generating image data for the display; The generating means generates the image data when the detecting means is active. determining a type of subject to be defined as a feature region based on settings of the imaging device; detecting a subject of the determined type from the image and determining a region of the subject of the determined type as the characteristic region; An imaging device characterized in that a processing process is applied to the characteristic region to visually emphasize it more than other regions.

2. the settings are for photographing a specific scene or a specific subject, 2. The imaging device according to claim 1, wherein the characteristic region is a region of a subject of a type corresponding to the specific scene, or a region of the specific subject.

3. The processing step is A process of not processing the characteristic region and processing other regions so that they are not noticeable; A process of emphasizing the characteristic region and leaving other regions untouched; A process of emphasizing the characteristic region and making other regions less noticeable; a process of processing the entire image including the characteristic region to emphasize the characteristic region; 2. The imaging device according to claim 1, wherein the imaging device is one of the following:

4. 2. The imaging device according to claim 1, wherein the generating means generates the image data as image data for live view display.

5. 2. The image pickup apparatus according to claim 1, further comprising setting means for setting a focus detection area based on the gaze position detected by said detecting means.

6. The imaging device of claim 1, wherein the type of subject is determined from a plurality of subject types including at least one of a human face, a human torso, a human limb, an animal face, a landmark, a character, a car, an airplane, and a railroad vehicle.

7. The setting is a scene, a storage means for storing a set scene and a type of subject to be set as a feature region in association with the scene; 2. The imaging device according to claim 1, wherein the generating means determines the type of subject to be the feature region corresponding to the set scene by referring to the information stored in the storage means.

8. A method executed by an imaging device having a detection means capable of detecting a user's gaze position in a displayed image, a generating step of generating image data for the display, In the generating step, determining a type of subject to be defined as a feature region based on settings of the imaging device; The image data generated when the detecting means is active is: detecting a subject of the determined type from the image and determining a region of the subject of the determined type as the characteristic region; applying a processing process to visually emphasize the characteristic region more than other regions; A method characterized in that the image data generated when the detection means is not effective is not subjected to processing that visually emphasizes the characteristic region more than other regions.

9. A program for causing a computer included in an imaging apparatus to function as each of the means included in the imaging apparatus according to any one of claims 1 to 7.

10. a generating means for generating image data to be displayed on a head-mounted display device; the generating means generates the image data by applying a processing process to visually emphasize a characteristic area corresponding to the type of virtual environment to be provided to the user through the display device.

1. An image processing device comprising:

11. The processing step is A process of not processing the characteristic region and processing other regions so that they are not noticeable; A process of emphasizing the characteristic region and leaving other regions untouched; A process of emphasizing the characteristic region and making other regions less noticeable; a process of processing the entire image including the characteristic region to emphasize the characteristic region; 11. The image processing device according to claim 10, wherein the image processing device is one of the above.

12. Further, the display device has a detection means for detecting a user's gaze position in an image displayed by the display device, 11. The image processing device according to claim 10, wherein the generating means generates the image data by applying further processing based on the gaze position detected by the detecting means after applying the processing.

13. 13. The image processing device according to claim 12, wherein the further processing is processing for visually emphasizing, among the characteristic regions, a characteristic region including the gaze position more than other characteristic regions.

14. 13. The image processing apparatus according to claim 12, wherein the further processing is processing for superimposing and displaying accompanying information relating to a feature region that includes the gaze position among the feature regions.

15. 13. The image processing device according to claim 12, wherein the further processing is processing for visually emphasizing a characteristic area that exists in the moving direction of the gaze position.

16. 11. The image processing device according to claim 10, wherein for each type of virtual environment, a type of feature region to which the processing can be applied and a type of feature region to which the processing is applied by default are associated.

17. 17. The image processing apparatus according to claim 16, wherein the generating means applies the processing based on a type designated by the user from types of feature regions associated with the virtual environment being provided to the user.

18. The image processing device according to claim 16, characterized in that, when no user designation is made, the generation means applies the processing based on the type of feature region to which the processing is applied by default, which is associated with the virtual environment being provided to the user.

19. Further, the display device has a detection means for detecting a user's gaze position in an image displayed by the display device, 17. The image processing apparatus according to claim 16, wherein the generating means applies the processing to a feature region based on the gaze position detected by the detecting means.

20. 17. The image processing apparatus according to claim 16, wherein said generating means, when the generated image data does not include the characteristic region, includes in the image data an index indicating a direction in which the characteristic region exists.

21. 17. The image processing device according to claim 16, wherein the head-mounted display device is an external device capable of communicating with the image processing device.

22. 17. The image processing device of claim 16, wherein the image processing device is part of the head-mounted display device.

23. The system further includes an acquisition means for acquiring data of a VR image representing the virtual environment, The generating means generates the image data from the VR image.

17. The image processing device according to claim 16,

24. The acquisition means further acquires main subject information and / or line of sight information obtained when capturing the VR image, 24. The image processing apparatus according to claim 23, wherein said generating means determines said feature region to which said processing is to be applied based on said main subject information or said line of sight information.

25. An image processing method executed by an image processing device, a generating step of generating image data to be displayed on a head-mounted display device; In the generating step, the image data is generated by applying a processing process that visually emphasizes a feature area corresponding to the type of virtual environment provided to the user through the display device more than other areas. An image processing method comprising:

26. A program for causing a computer to function as each of the means included in the image processing device according to any one of claims 10 to 24.